This is a remedial run for missed papers from 08/21/2026 to 08/23/2026.
Results generated on 09/14/2026.
Personalized Daily ArXiv Papers 2026-08-24
| Model | Metric | Usage | Papers | ||||
|---|---|---|---|---|---|---|---|
| Prompt | Completion | Total | Total arXiv | Scanned | Relevant | ||
local agent CLI |
Tokens | not reported | not reported | not reported | 883 | 883 | 56 |
| Cost | not reported | not reported | not reported | ||||
Token counts are not reported for this run. 39 of 44 model calls succeeded, 12,052s of model wall clock.
Topic Coverage:
| Topic | Papers |
|---|---|
| MoE Training | 1 |
| Large-Scale Training Systems and Efficiency | 5 |
| Architecture and Training Dynamics | 24 |
| Efficiency, Compression, and Large-Scale Training | 26 |
Table of contents by topic:
MoE Training (1)
- More Experts, Worse Dynamics: Inverse Scaling and Spectral Bias in Mixture-of-Experts State-Space Models Authors: Chandresh Pandey
Large-Scale Training Systems and Efficiency (5)
-
Unveiling the Depth-Performance Dilemma in Split-Federated Fine-tuning of LLMs Authors: Hariharan Ramesh, Someshwaran Murugaiyan, Jyotikrishna Dass
-
FeLoG: Scalable and Efficient Distributed Graph Embedding with Feedback Loop Mechanism Authors: Peng Fang, Arijit Khan, Ziqiang Wu, Zhenli Li, Yibo Zhou, Fang Wang, Dan Feng
-
Two-level domain-decomposition AdaGrad method for scalable training of graph neural networks Authors: Laurynas Varnas, Julien Herrmann, Alexander Heinlein, Serge Gratton, Alena KopaniÄáková
-
Thermo-FL: Thermal-Aware Robust Federated Fine-Tuning of Large Language Models for Edge AI Authors: Shiva Shrestha, Kazi Shaharair Sharif, Zongxing Xie, Jiajing Huang, Anhao Xiang, Honghui Xu
-
Model-Consistent Byzantine-Resilient Decentralized Federated Learning for Collaborative Missions Authors: Yue Li, Sudip Bhujel, Cameron Lira, Ning Wang, Yang Xiao
Architecture and Training Dynamics (24)
-
Improving Few-Step Language Flows with Untied Self-Conditioning Authors: Bocheng Li, Linli Xu
-
Variational Structure at the Edge of Stability Authors: Eric Regis
-
Dynamic Multi-Byte Prediction With Hierarchical Language Models Authors: Abraham Toluwase Owodunni, Chibuzor Okocha, Christan Grant, Tomasz Limisiewicz, Sachin Kumar
-
Closing the Curvature Gap: Full Transformer Hessians Authors: Egor Petrov, Nikita Kiselev, Vladislav Meshkov, Andrey Grabovoy
-
TANGO: Token-Aggregated Nonlinear Gating Operators for Natural and Formal Language Modeling Authors: Joshua Nunley
-
Provable Edge-of-Stability for Adam on a One-Dimensional Quadratic Authors: Yiman Fong, Heng Yang
-
From Attention Masks to Inert Zero-Vector Tokens: OAttention and O-Closure for Token Dynamics Authors: Heyang Gong
-
Minimax Optimality of Score-Entropy Discrete Diffusion Authors: Cholyeon Cho, Yuchen Wu
-
Finite-Nudge Equilibrium Propagation in Thermal Ensembles Authors: Elon Litman
-
Learning with Boolean threshold functions Authors: Veit Elser, Manish Krishan Lal
-
Control Allocation in Neural Network Optimization: Joint Affine Control of Weight and Bias Updates Authors: Zhang Gongyue, Sheng Yixuan, Wang Zhiyong, Liu Donghan, Ren Weihong, Liu Honghai
-
MentorPulse: Refreshing Cross-Model Latent Guidance for Long-Form Generation Authors: Ziwu Liu, Guozhong Li, Chen Qiu, Weiyang Kong, Panos Kalnis
-
Free-Probability Kernels for Zero-Rollout Hyperparameter Selection in Reservoir Computing Authors: Sara Malacarne, Andrea Ceni, Claudio Gallicchio
-
Read, Write, Relax: Why Neural PDE Surrogates Need Both Global and Local Processing Authors: Anuj Kumar, Heiko Zimmermann, Josiah Bjorgaard, Jacan Chaplais, Nikolaos Bouklas, Matteo Salvador, Alexander Lavin
-
Divisive Normalization Shapes Low-Rank Slow Manifolds for Continuous Working Memory Authors: Zhaotian Gu, Jie Su, Weiwei Wang, Chang Liu, Tianyi Qian, Dahui Wang
-
Continuous-Time Quantum Walks based Graph Neural Network Authors: Yuliang Zhan, Zefeng Gao, Jian Li, Yang Liu, Hao sun
-
Where Cognition Lives: Dissecting Emergent from Computed Function in a Minimal Complete Cognitive Architecture Authors: Francisco M. Arrabal-Campos, Francisco G. Montoya, Alfredo Alcayde, Ignacio Fernández
-
Loss-Parameterized Fisher Width Along Learning Trajectories Authors: Vu Khac Ky
-
Dataset Complexity Shapes Finite-Distance Loss Geometry in Neural Networks Authors: Jaeyong Bae, Hawoong Jeong
-
Training, learning and inference: unified dynamics of neural systems Authors: Mian Wang
-
SPARCL: Spectral Partitioned Analytic Continual Learning Authors: James Hartley, Zeropy Surio, Daniel Whitmore, Hannah Clarke, Thomas Reed
-
Learning in PINNs: Phase transition, diffusion equilibrium, and generalization Authors: Sokratis J. Anagnostopoulos, Juan Diego Toscano, Nikolaos Stergiopulos, George Em Karniadakis
-
Bounded Precision-Geometry Scaling for Robust Multi-Task Learning under Loss Scale Mismatch Authors: Krishna Subedi
-
Socialized Division and Collaboration: Rethinking Class-Incremental Learning under Optimization Conflicts Authors: Xinjie Yao, Zhihe Fan, Yunqi Zhu, Jiaqi Zhou, Dengyu Zhao, Zhoupeng Guo, Yan Fan, Guosong Jiang, Pengfei Zhu
Efficiency, Compression, and Large-Scale Training (26)
-
Beyond Sparse Weights: When Is Attention Compressible? Authors: Chiwun Yang, Xiaoyu Li
-
A JoLT for the KV cache: Near-lossless KV cache compression via joint Lagrangian allocation of Tucker ranks and a rotated residual for llms Authors: Rahul Krishnan, Volker Schulz
-
Jacobian-guided Noise Injection for Quantization Robustness in Large Language Models Authors: Deepanshu Pandey, Arnav Chavan, Nahush Lele, Sankalp Dayal, Deepak Gupta
-
NeuroPrefetcher: Storage-Aware Sparse LLM Inference via Delta Prefetching Authors: Nobel Dhar, Md Romyull Islam, Xuechen Zhang, Gongjin Sun, Sahidul Islam, Bobin Deng, Kun Suo
-
HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models Authors: Zhinan Xie, Peisong Wang, Shuang Qiu, Jian Cheng
-
SSDi8: Accurate and Efficient 8-bit Quantization for State Space Duality Authors: Hyunwoo Kim, Byoungchan Ko, Minseok Kang, Minwoo Kim, Dongjin Lee, Jaehoon Lee, Sungroh Yoon, Dahuin Jung
-
COEC: Calibrated Orthogonal-Equivalence Compensation for Structured Pruning of Large Language Models Authors: Peiqi Yu, Nam Ling, Wei Wang, Wei Jiang
-
Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs Authors: Luka Ribar, Jeevan Bhoot, Douglas Orr
-
Width-Independent Compressibility of Deep Neural Networks Authors: Hong-Yi Wang, Mingze Wang, Liu Ziyin
-
Tensor Seeks Layout: Formalizing Layout Selection for ML Compilers Authors: Clemens Eisenhofer, Yuwen Jia, Daniel Kroening, Sergey Pupyrev
-
Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs Authors: Bakbergen Ryskulov, Iker GarcÃa-Ferrero, David Montero, David Jansen, Ali Hashemi, Jezabel R. Garcia, Antonio Tiene, Román Orús
-
Scaling Laws for Task-Specific LLM Distillation Authors: Lavinia Ghita, Dhruv Desai, Ioana Boier
-
GraniKV: Asymmetric Granularity KV-Cache Paging for Multi-Agent Systems with Long Shared Prefix Authors: Jinhyun Jeon, Sungjoo Yoo
-
When Less Latent Leads to Better Relay: Information-Preserving Compression for Latent Multi-Agent LLM Collaboration Authors: Yiping Li, Zhiyu An, Wan Du
-
DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference Authors: Yujie Zhang, Shivam Aggarwal, Tulika Mitra
-
SAEM: Stage-Aware Expert Management for Memory-Efficient MoE Inference in Chain-of-Thought Reasoning Authors: Yujie Zhang, Bin Gao, Tulika Mitra
-
Continuous Adversarial MeanFlow Transfer Authors: Yara Bahram, Zahra Dehghani, Mélodie Desbos, Eric Granger, Pablo Piantanida, Mohammadhadi Shateri
-
CD-LoRA: Consistency-Driven Low-Rank Adaptation for Multi-Task Fine-Tuning Authors: Qian Zha, Jinda Liu, Yuan Wu, Yi Chang
-
PowerSlider: Exploiting Phase Asymmetry for LLM Serving under Demand Response Authors: Yueying Li, Jiayang Chen, Yuanfan Chen, Leo Han, Haoran Qiu, Esha Choukse, Rodrigo Fonseca, Udit Gupta
-
Pandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimation Authors: Adam Fisch, Shubhendu Trivedi, Fantine Huot, William W. Cohen, Michael Kaisers, Mirella Lapata, Kate Larson, Jacob Eisenstein
-
Event-triggered Implicit Perturbation for Zeroth-Order Fine-Tuning of Spiking Transformers Authors: Tengteng Lei, Prabodh Katti, Rashi Dutt, Houssem Sifaou, Tan Peng, Osvaldo Simeone, Kai Xu, Bipin Rajendran
-
Two Stage Wireless Federated LoRA Fine-Tuning with Sparsified Orthogonal Updates Authors: Bumjun Kim, Wan Choi
-
DiD It in 87 Minutes: A Label-Free Softmax-to-Linear Adaptation of Vision Transformers for Object Detection Authors: Huaiyuan Qin, Gabriel James Goenawan, Zihang Lin, Muli Yang, Hongyuan Zhu
-
LLM4LLM: Bridging Kernel Benchmarks and Real Deployment via Closed-Loop Agentic Optimization Authors: Hui Zeng, Pengfei Yang, Yanxin Chen, Fusong Ju, Xinran Wei
-
HIERA: Workload-Aware Planning Across Implementation Spaces for GPU Kernel Optimization Authors: Jinghao Wang, Qiqi Gu, Chenpeng Wu, Jianguo Yao, Haibing Guan, Xijun Li
-
Target-Aware Calibration Data Selection for Preserving Uncertainty in Quantized Language Models Authors: Zhen Yang, Sizai Hou, Kaiwen Zheng, Yaofang Liu, Liang He, Yixuan Chen, Kangning Cui
MoE Training (1)
1. More Experts, Worse Dynamics: Inverse Scaling and Spectral Bias in Mixture-of-Experts State-Space Models
ArXiv ID: 2608.21840
Primary Topic: MoE Training
Also Matches: Architecture and Training Dynamics
Authors: Chandresh Pandey
Abstract: Mixture-of-Experts (MoE) architectures are commonly motivated as a way to increase expressivity by decomposing complex systems into simpler local dynamics. This intuition has recently been extended to spectral state-space models, where mixing stable operators is assumed to enable adaptation to heterogeneous or regime-switching time series. We critically evaluate this assumption in a controlled synthetic setting designed to isolate dynamical rather than representational challenges. We study a next-step prediction task on sequences composed of three regimes: chaotic dynamics generated by the Mackey-Glass system, a stable oscillatory regime, and a noise-dominated autoregressive regime. Across extensive ablations including capacity scaling, oracle routing, frozen-expert variants, and comparisons to output-level MoE baselines, operator-level mixture models consistently fail to outperform a single-expert baseline. Increasing the number of experts leads to inverse scaling, routing collapses or fails to induce meaningful specialization, and even perfect regime supervision does not prevent degradation in global performance. Furthermore, we show that apparent improvements in mean squared error on chaotic trajectories can be misleading. Phase-space analysis reveals that lower error often arises from temporal smoothing that destroys the geometry of the underlying attractor rather than from faithful modeling of the dynamics. These results identify a likely limitation of operator interpolation under the studied parameterization and training protocol, and underscore the need for geometry-aware evaluation when assessing regime-switching dynamical systems.
Comment: Analyzes routing collapse and inverse expert scaling during operator-level MoE training.
Topic Match: Routing and frozen-expert ablations investigate MoE training failures and operator-mixing dynamics, with conclusions limited to controlled synthetic settings.
Relevance: 8 Novelty: 7
Large-Scale Training Systems and Efficiency (5)
1. Unveiling the Depth-Performance Dilemma in Split-Federated Fine-tuning of LLMs
ArXiv ID: 2608.22188
Primary Topic: Large-Scale Training Systems and Efficiency
Also Matches: Architecture and Training Dynamics
Authors: Hariharan Ramesh, Someshwaran Murugaiyan, Jyotikrishna Dass
Abstract: Split Federated Fine-tuning (SFF) is a promising paradigm for scaling Large Language Models (LLMs) by partitioning model depth between resource-constrained clients and a centralized server. While system incentives for throughput and privacy favor deep partitions, the impact of such configurations on model utility remains poorly understood. In this work, we identify and characterize the Depth-Performance Dilemma: the regime that maximizes system efficiency is precisely where fine-tuning quality collapses. Through a comprehensive audit across four model scales (GPT-2 to Llama-3-8B) and diverse benchmarks, we demonstrate that deeper partitions provide monotonic gains in throughput and privacy at the cost of catastrophic performance plateaus. We evaluate a suite of state-of-the-art federated adapter aggregation methods including AVG, STACK, SVD, and FREEZE, revealing that while these techniques are effective in standard Federated Learning, they fail to mitigate the artifacts unique to split architectures. Finally, we provide a mechanistic diagnosis for this failure, tracing the collapse to the near-isometric topology of Transformers, which allows aggregation noise to propagate without attenuation until it triggers Attention Collapse in the server partition. Our findings challenge the prevailing assumption that partition depth is a utility-neutral tuning knob and provide a structural foundation for stable distributed LLM fine-tuning.
Comment: Characterizes split-federated partition depth as a systems-quality tradeoff and traces collapse to attention dynamics.
Topic Match: Distributed model partitioning is the primary systems contribution, supported by a mechanistic account of training failure.
Relevance: 8 Novelty: 7
2. FeLoG: Scalable and Efficient Distributed Graph Embedding with Feedback Loop Mechanism
ArXiv ID: 2606.22180
Primary Topic: Large-Scale Training Systems and Efficiency
Authors: Peng Fang, Arijit Khan, Ziqiang Wu, Zhenli Li, Yibo Zhou, Fang Wang, Dan Feng
Abstract: Graph embedding maps graph nodes into low-dimensional vectors to support applications such as recommendation, fraud detection, and graph-based retrieval-augmented generation (GraphRAG). As graphs scale to billions of edges, scalable and efficient graph embedding has become increasingly important. Existing frameworks commonly adopt a sampling-training paradigm, in which mini-batches are constructed by sampling nodes and their neighbors. However, sampling is typically decoupled from evolving embedding quality, causing redundant exploration of well-trained regions while under-sampling undertrained nodes. At the system level, such decoupling further leads to excessive communication, serialized execution, and low resource utilization in distributed environments. We present FeLoG, a feedback loop-driven system for scalable distributed graph embedding. (1) FeLoG introduces feedback-coupled sampling and training, dynamically prioritizing undertrained nodes according to real-time embedding-quality feedback, thereby reducing redundant computation and accelerating convergence. (2) It employs activity-aware communication that compresses frequently occurring node sequences to reduce intra-machine PCIe traffic and selectively synchronizes frequently updated embeddings to reduce inter-machine communication. (3) It adopts a round-interleaved pipeline that overlaps next-round sampling with current-round training to improve CPU-GPU utilization. Experiments against six state-of-the-art baselines on large-scale graphs show that FeLoG achieves an average speedup of 27.9x, reduces communication cost by more than 53.1%, and sustains over 80% CPU-GPU utilization.
Comment: Activity-aware communication and round-interleaved execution reduce distributed embedding-training overhead.
Topic Match: The core contribution is distributed training communication and scheduling, with applicability demonstrated on graph embeddings.
Relevance: 7 Novelty: 7
3. Two-level domain-decomposition AdaGrad method for scalable training of graph neural networks
ArXiv ID: 2608.22575
Primary Topic: Large-Scale Training Systems and Efficiency
Authors: Laurynas Varnas, Julien Herrmann, Alexander Heinlein, Serge Gratton, Alena KopaniÄáková
Abstract: Graph neural networks (GNNs) have emerged as a powerful framework for learning from graph-structured data. However, their efficient training remains challenging, particularly in distributed computing environments. This challenge arises from the use of message passing, which couples all graph nodes, leading to expensive optimization steps, high memory requirements, and substantial communication overhead. To alleviate these limitations, we propose a novel domain-decomposition (DD) variant of AG2m, an AdaGrad method enhanced with second-order curvature information and momentum, denoted by DD-AG2m. The proposed DD-AG2m alternates between AG2m optimization on the original (global) graph and AG2m optimization on the partitioned graphs. To incorporate global information at reduced cost, we further introduce a two-level variant (2DD-AG2m) that performs global optimization steps on a coarse graph obtained by randomly subsampling nodes within each subdomain. Numerical experiments spanning graph classification, node-level regression, and spatiotemporal forecasting tasks demonstrate that the proposed DD methods reduce the computational cost required to achieve the same predictive performance by a factor of 4-8. Moreover, for the fixed computational cost, they improve the predictive performance of GNNs by up to 22% compared with the baseline AG2m.
Comment: Two-level domain-decomposition optimization reduces training cost through local updates and coarse-graph global steps.
Topic Match: The core contribution is a multilevel training optimizer, with its applicability demonstrated on graph neural networks.
Relevance: 7 Novelty: 7
4. Thermo-FL: Thermal-Aware Robust Federated Fine-Tuning of Large Language Models for Edge AI
ArXiv ID: 2608.21172
Primary Topic: Large-Scale Training Systems and Efficiency
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Shiva Shrestha, Kazi Shaharair Sharif, Zongxing Xie, Jiajing Huang, Anhao Xiang, Honghui Xu
Abstract: Federated fine-tuning enables large language models to adapt on edge devices without centralizing private data, but practical deployments must address hardware instability and adversarial update corruption together. Thermally constrained clients may throttle, slow local training, or delay synchronous aggregation, while Byzantine clients and communication-layer adversaries can corrupt the updates used to form the global model. To address these challenges, we present Thermo-FL, a thermal-aware federated LoRA fine-tuning framework that uses device temperature as an active control signal for local adapter training and sparse update transmission. On the client side, Thermo-FL adjusts the active LoRA-layer fraction and transmitted update density as devices heat or cool, reducing workload under thermal stress. On the server side, Thermo-FL introduces TERRA, a robust aggregation pipeline for dynamically sparse LoRA updates that combines norm filtering, mask-aware directional validation, adaptive active-coordinate clipping, and mask-aware aggregation. We evaluate Thermo-FL using both a large-scale emulator and a Jetson-based physical testbed. In the emulator, Thermo-FL improves robustness under adversarial sparse aggregation and achieves the strongest BoolQ accuracy across clean and attack settings while remaining competitive on GSM8K. In the physical prototype, Thermo-FL stabilizes device temperature, reduces compressed upload size through bitmap sparse encoding, and preserves GSM8K utility under sign-flip/scale and MITM perturbations. These results show that secure edge LLM adaptation should jointly consider hardware behavior, workload regulation, sparse communication, and aggregation robustness.
Comment: Adapts LoRA workload and sparse communication online using device temperature, with mask-aware robust aggregation.
Topic Match: The distributed coordination, sparse communication, and aggregation algorithm make training systems the primary fit.
Relevance: 6 Novelty: 6
5. Model-Consistent Byzantine-Resilient Decentralized Federated Learning for Collaborative Missions
ArXiv ID: 2608.22552
Primary Topic: Large-Scale Training Systems and Efficiency
Authors: Yue Li, Sudip Bhujel, Cameron Lira, Ning Wang, Yang Xiao
Abstract: Decentralized federated learning (DFL) is a promising paradigm for autonomous nodes to collaboratively train AI models without relying on a central server. However, existing DFL solutions do not guarantee global model consistency, a critical requirement for collaborative mission-critical scenarios where model divergence undermines decision uniformity and safety. This lack of consistency also amplifies vulnerability to Byzantine adversaries, who exploit the decentralized network topology and weak synchrony to perform equivocation and model poisoning attacks against individual victims. This paper introduces DFL-C, a novel Byzantine-resilient DFL architecture that enables decentralized nodes to perform collaborative training with global model consistency. At its core, DFL-C integrates an asynchronous common subset (ACS) consensus protocol into the DFL workflow to ensure all nodes aggregate a uniform set of model updates to establish global model consistency, despite individual Byzantine equivocation. DFL-C further implements a dual-domain trust scoring mechanism to provide resilience against data-domain Byzantine manipulations including model poisoning attacks. This mechanism complements the consensus protocol, significantly reducing the latter's runtime. Our experimental results demonstrate that DFL-C maintains model accuracy while achieving global model consistency under Byzantine behaviors with moderate consensus overhead. Notably, when compared with the state-of-the-art DFL solution BALANCE (Fang et al.) that does not provide model consistency, DFL-C achieves better model accuracy against untargeted model poisoning attacks and comparable resilience against backdoor attacks, with the advantage widened under non-IID scenarios.
Comment: Uses asynchronous consensus to make decentralized learners aggregate the same updates despite Byzantine equivocation.
Topic Match: Consensus-backed aggregation is a distributed-training contribution, although the focus on federated fault tolerance makes its connection to large-model pretraining peripheral.
Relevance: 6 Novelty: 6
Architecture and Training Dynamics (24)
1. Improving Few-Step Language Flows with Untied Self-Conditioning
ArXiv ID: 2608.22244
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Bocheng Li, Linli Xu
Abstract: Flow-matching language models refine all token positions in parallel and can trade sampling steps for latency, yet generation quality still degrades sharply with few sampling steps. We trace a source of this degradation to a train--inference mismatch in previous-prediction self-conditioning: during training, the self-conditioning input is computed from the current noisy state with no intervening solver step; during sampling, the solver folds the previous prediction into the latent before that same prediction reappears as the explicit self-conditioning input. This coupling, absent during training, creates redundancy that grows with step width. We show that the mismatch degrades both the self-conditioning input and the solver update, and derive a correction for each from the model's own structure. From the frozen projection weights we identify directions along which the self-conditioning input is redundant with the latent and dampen them; from the solver's integration structure we derive that a step-average prediction is needed and approximate it from prediction history, with scale set by offline trajectory statistics. The resulting sampler, Untied Self-Conditioning, requires no retraining and uses one evaluation per step. At 8 sampling steps on LangFlow, it reduces OpenWebText generative perplexity from $531$ to~$62$ ($8.6\times$); under an adapted Arena-Hard-Auto~v2 protocol, its outputs are preferred in $96\%$ of pairwise comparisons. On ELF-B it reduces generative perplexity from $71$ to~$43$. Improvements hold from 8 to 256 sampling steps.
Comment: Derives corrections for a train-inference mismatch caused by coupling self-conditioning predictions with solver updates.
Topic Match: Mechanistic analysis of self-conditioning and solver coupling drives the method, with improved generation quality per model evaluation providing a second efficiency match.
Relevance: 9 Novelty: 8
2. Variational Structure at the Edge of Stability
ArXiv ID: 2608.21660
Primary Topic: Architecture and Training Dynamics
Authors: Eric Regis
Abstract: When discrete-time optimizers operate at the edge of stability, they exhibit near-two-periodic behavior. These oscillatory dynamics are reminiscent of conservative systems, such as the dynamics generated by symplectic integrators. However, a precise formulation of the connection between discrete-time optimizers at the edge of stability and discrete mechanics remains underexplored. Recently, Litman introduced the "edge coupling": a functional on consecutive gradient descent iterates whose critical points encode the fixed points and two-point orbits of the gradient descent dynamics. Here we extend the edge coupling to heavy-ball and Nesterov momentum. We show that its critical points characterize the fixed points and two-point orbits, with its Hessian characterizing their stability. We also show that the edge coupling can be identified with the symmetric Verlet action, formalizing the connection between the edge of stability and discrete mechanics.
Comment: Extends edge coupling to momentum optimizers, using its Hessian to characterize fixed-point and period-two stability.
Topic Match: Directly analyzes optimizer dynamics at the edge of stability and connects their oscillations to discrete variational mechanics.
Relevance: 9 Novelty: 7
3. Dynamic Multi-Byte Prediction With Hierarchical Language Models
ArXiv ID: 2608.15454
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Abraham Toluwase Owodunni, Chibuzor Okocha, Christan Grant, Tomasz Limisiewicz, Sachin Kumar
Abstract: Byte-level hierarchical language models (LMs) have recently emerged as a robust alternative to their popular counterparts that use subword tokenization. However, generating one byte at a time remains a bottleneck for inference speed. To address this, we introduce multi-byte prediction (MBP), which generates multiple bytes in parallel, speeding up inference with minimal performance impact and no additional parameters. MBP builds on the popular multi-token prediction (MTP) paradigm with two crucial innovations. First, we introduce a variable-length prediction window that aligns with the latent tokens, or segments, of a hierarchical LM. Second, we implement a novel attention-masking scheme that enables parallel byte prediction without violating causality. We show that multi-byte prediction strikes a Pareto-optimal trade-off across multiple generative tasks, instruction following, question answering, summarization, and machine translation, achieving the best trade-off between performance and inference throughput.
Comment: Latent-aligned prediction windows and causal attention masks enable parallel multi-byte generation without additional parameters.
Topic Match: Changes the hierarchical LM's prediction and attention mechanisms, producing direct inference-throughput benefits.
Relevance: 9 Novelty: 7
4. Closing the Curvature Gap: Full Transformer Hessians
ArXiv ID: 2510.16927
Primary Topic: Architecture and Training Dynamics
Authors: Egor Petrov, Nikita Kiselev, Vladislav Meshkov, Andrey Grabovoy
Abstract: The optimization landscape of Transformer models remains poorly understood despite their widespread adoption. While recent studies have derived curvature properties for isolated self-attention mechanisms, a comprehensive theoretical characterization of the full Transformer block, accounting for the interactions between Layer Normalization, Feed-Forward Networks (FFNs), and residual connections, is missing. In this work, we close this gap by deriving the exact, closed-form Hessian for the complete Transformer block under arbitrary twice-differentiable loss functions. We utilize rigorous matrix calculus to handle the non-linearities of LayerNorm and row-wise activations, establishing explicit spectral norm bounds for the resulting Hessian blocks. Our analysis reveals how different architectural components contribute distinct curvature mechanisms, identifying the specific curvature contributions of particular sub-layers. Furthermore, empirical validation against automatic differentiation confirms the exactness of the derived formulas up to numerical precision and shows substantial computational speedups for the closed-form Jacobian evaluations.
Comment: Exact full-block Transformer Hessians expose the curvature contributions of LayerNorm, FFNs, and residual connections.
Topic Match: Directly analyzes how interacting Transformer components shape optimization curvature, including explicit spectral norm bounds.
Relevance: 9 Novelty: 7
5. TANGO: Token-Aggregated Nonlinear Gating Operators for Natural and Formal Language Modeling
ArXiv ID: 2608.22117
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Joshua Nunley
Abstract: A standard Transformer block separates cross-token interaction in self-attention from a nonlinear feed-forward network applied independently at each position. We introduce the TANGO model (Token-Aggregated Nonlinear Gating Operators), which replaces these two sublayers with one cross-token gated residual update. Each source token produces a SwiGLU gate vector. Query-key similarities determine a weighted average of source gates for each destination, and the resulting gate rescales projected destination features. TANGO assigns a separate weight to every causally visible source and is quadratic in sequence length. The WANGO model (Windowed Aggregation of Nonlinear Gating Operators) retains the same unnormalized scores within a recent window and uses positive feature-map prefix statistics for older sources, giving linear sequence-length complexity for fixed window and feature dimensions. We compare TANGO and WANGO with Recurrent and Untied Transformer++, full-attention GAU, and FLASH. All models have approximately 44.3M nonembedding parameters and are trained in three matched runs. TANGO, WANGO, and Recurrent Transformer++ apply one shared block four times; the other architectures use four independent blocks. TANGO obtains the lowest mean validation negative log-likelihood on FineWeb-Edu, Lean, and DeepMind Mathematics, although it has the largest analytical forward-pass operation count. WANGO obtains the lowest mean FineWeb-Edu NLL among the architectures with computation linear in sequence length and outperforms Recurrent Transformer++ at nearly the same analytical forward-pass multiply-accumulate count.
Comment: Replaces separate attention and FFN sublayers with a cross-token gated residual operator.
Topic Match: The central contribution is a new sequence-modeling block; WANGO additionally provides linear sequence-length computation, with evidence at 44.3M nonembedding parameters.
Relevance: 9 Novelty: 7
6. Provable Edge-of-Stability for Adam on a One-Dimensional Quadratic
ArXiv ID: 2608.20638
Primary Topic: Architecture and Training Dynamics
Authors: Yiman Fong, Heng Yang
Abstract: The edge-of-stability (EoS) phenomenon of Adam has been widely observed, while its underlying dynamical mechanism is not yet fully understood. We study uncorrected Adam on a one-dimensional quadratic, a clean setting where constant curvature isolates the optimizer-induced dynamics behind the EoS. We characterize the resulting dynamics across the parameter space. In broad regimes, we prove that Adam exhibits a restoring tendency toward its frozen stability threshold $2(1+β_1)/[η(1-β_1)]$. We also identify settings in which this edge-seeking mechanism breaks down, including strictly subcritical periodic orbits and specially tuned trajectories that converge to the optimum while remaining uniformly supercritical. These results give a concrete dynamical explanation for Adam's EoS in a setting free of evolving loss geometry, while also exposing its limitations.
Comment: Proves Adam's restoring dynamics toward its stability threshold and characterizes exceptions.
Topic Match: Directly explains an optimizer-induced training-stability mechanism, although the analysis is restricted to a one-dimensional quadratic.
Relevance: 8 Novelty: 8
7. From Attention Masks to Inert Zero-Vector Tokens: OAttention and O-Closure for Token Dynamics
ArXiv ID: 2608.21174
Primary Topic: Architecture and Training Dynamics
Authors: Heyang Gong
Abstract: Attention masks are relation-level controls: they specify which query--source pairs may interact. They do not provide a representation-carried token state that is non-participating at the attention boundary. We assign each token hidden carrier (h_i) an active-presence coefficient (p_i=\lVert h_i\rVert^2/(Ï+\lVert h_i\rVert^2)). The same coefficient has two roles: it gates information emitted by token (i), and it determines the mass with which token (i) enters computations shared with other tokens. OAttention is the support-coupled attention realization of this rule. It gates the receiver output by (p_i) and weights source (j) by (p_j) in both the attention numerator and partition, while retaining the standard score, visibility relation, exponential competition, and value aggregation. This makes the zero-vector token a zero element and yields exact null-receiver, null-source insertion, self-attention insertion, and empty-support properties. The same token-level presence gives local O-components (OFFN, ONorm, and OInject), presence-weighted OStandardize, the O-Closure law (M(H\oplus0)=M(H)\oplus0), and an OTransformer by residual and compositional closure. The canonical operator is checked by contract tests and a GPU evaluation. In a zero-fine-tuning retrofit of a cloned pretrained TabPFN v3 regressor, calibrated hidden-carrier OAttention and Full-O variants change mean RMSE by $+0.088\%$ and $+0.177\%$, respectively, over 18 matched dataset--seed cases. A two-block ablation shows that OAttention alone does not preserve a NULL state through ordinary host components, whereas the OTransformer path does. These are scoped tests of exactness, active-path compatibility, and compositional necessity; they do not establish universal no-loss, arbitrary-host closure, learned attraction to the origin, or a general semantics for missing values.
Comment: Introduces presence-gated attention and residual components that make zero-vector tokens exactly inert under composition.
Topic Match: Its core contribution is a new attention-and-residual computational primitive with exact compositional behavior.
Relevance: 8 Novelty: 8
8. Minimax Optimality of Score-Entropy Discrete Diffusion
ArXiv ID: 2608.20635
Primary Topic: Architecture and Training Dynamics
Authors: Cholyeon Cho, Yuchen Wu
Abstract: Discrete diffusion models have demonstrated strong performance across a range of datasets, including natural language data and graph-structured data. Among many variants, score-entropy discrete diffusion (SEDD) has achieved particularly strong empirical results. In SEDD, new samples are generated by iteratively evaluating a sequence of concrete score functions, which are learned by minimizing a score-entropy loss. While much of the prior theoretical literature on discrete diffusion has focused on the sampling efficiency of SEDD under the assumption of small score estimation error, recent work has begun to investigate the finite-sample properties of score estimation itself. In this work, we take a different route by investigating the fundamental statistical limits of concrete score estimation. We focus on uniform and masking discrete diffusions, two of the most widely adopted discrete diffusion models. We establish a minimax lower bound under the score-entropy loss, and propose an MLE-based thresholding estimator that matches this lower bound up to constant and polylogarithmic factors that depend on neighboring density ratios. We further show that, for any target distribution, this density ratio is naturally controlled under both uniform and masking discrete diffusion models, yielding nearly matching minimax lower and upper bounds for the aggregated score estimation error. Our results imply that, with appropriate initialization and discretization, SEDD can achieve nearly optimal minimax sample complexity, as measured by the KL divergence between the target and generated distributions.
Comment: Establishes nearly matching minimax bounds for concrete-score estimation in masking and uniform discrete diffusion.
Topic Match: The paper provides foundational analysis of a generative-model training objective and its sample-complexity dynamics.
Relevance: 7 Novelty: 8
9. Finite-Nudge Equilibrium Propagation in Thermal Ensembles
ArXiv ID: 2511.22024
Primary Topic: Architecture and Training Dynamics
Authors: Elon Litman
Abstract: We liberate Equilibrium Propagation (EP) from the limit of infinitesimal perturbations by establishing a finite-nudge foundation for local credit assignment. By modeling network states as Gibbs-Boltzmann distributions rather than deterministic points, we prove that the gradient of the difference in Helmholtz free energy between a nudged and free phase is exactly the difference in expected local energy derivatives. This validates the classic Contrastive Hebbian Learning update as an exact gradient estimator for arbitrary finite nudging, requiring neither infinitesimal approximations nor convexity. In the zero-temperature limit, we prove that the same identity reduces to the deterministic contrastive rule around any local energy basin without assuming a unique global minimum, and a subsequent small-nudge limit recovers traditional EP. Finally, we derive an equivalent representation of the same gradient as an integral of the loss--energy covariance over nudging strength, which generalizes infinitesimal EP to strong error signals that its small-nudge approximation cannot support. Numerical experiments corroborate that finite nudging provides a practical signal-to-noise advantage over infinitesimal methods during training.
Comment: Proves an exact finite-nudge free-energy gradient identity for local equilibrium-propagation updates.
Topic Match: The core contribution is a mechanistic theory of credit assignment and gradient estimation, although large-model training applicability remains unestablished.
Relevance: 7 Novelty: 8
10. Learning with Boolean threshold functions
ArXiv ID: 2602.17493
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Veit Elser, Manish Krishan Lal
Abstract: We develop a method for training neural networks on Boolean data in which the values at all nodes are strictly $\pm 1$, and the resulting models are typically equivalent to networks whose nonzero weights are also $\pm 1$. The method replaces loss minimization with a nonconvex constraint formulation. Each node implements a Boolean threshold function (BTF), and training is expressed through a divide-and-concur decomposition into two complementary constraints: one enforces local BTF consistency between inputs, weights, and output; the other imposes architectural concurrence, equating neuron outputs with downstream inputs and enforcing weight equality across training-data instantiations of the network. The reflect-reflect-relax (RRR) projection algorithm is used to reconcile these constraints. Each BTF constraint includes a lower bound on the margin. When this bound is sufficiently large, the learned representations are provably sparse and equivalent to networks composed of simple logical gates with $\pm 1$ weights. Across a range of tasks -- including multiplier-circuit discovery, binary autoencoding, logic-network inference, and cellular automata learning -- the method achieves exact solutions or strong generalization in regimes where standard gradient-based methods struggle. These results demonstrate that projection-based constraint satisfaction provides a viable and conceptually distinct foundation for learning in discrete neural systems, with implications for interpretability and efficient inference.
Comment: Trains strictly binary networks through projection-based constraint satisfaction, with margin conditions yielding provably sparse logical-gate representations.
Topic Match: The discrete-network training formulation is primary; binary computation and provable sparsity provide a secondary efficiency connection, with large-model scaling unestablished.
Relevance: 7 Novelty: 8
11. Control Allocation in Neural Network Optimization: Joint Affine Control of Weight and Bias Updates
ArXiv ID: 2608.02991
Primary Topic: Architecture and Training Dynamics
Authors: Zhang Gongyue, Sheng Yixuan, Wang Zhiyong, Liu Donghan, Ren Weihong, Liu Honghai
Abstract: Optimization algorithms determine not only the magnitude of a neural-network update but also how that update is distributed across parameter channels. We study whether this distribution can be treated as a controllable quantity independently of global training progress. We define operational update allocation through normalized channel energies and analyze two scalar controls: a coordinate-preconditioning exponent and an affine spectral exponent that scales the bias column of an augmented weight--bias matrix. At a frozen state, a common nonzero step-size multiplier leaves normalized allocation unchanged; the coordinate exponent yields affine pairwise log-odds with an explicit inverse; and the affine exponent induces a rank-one positive-semidefinite Gram perturbation and a logistic raw-participation law. We further separate raw affine participation, spectral gain, and the decoded physical bias update, and show that finite polynomial spectral iterations preserve singular subspaces. Same-state replay verifies the exact control laws. On a five-seed controlled benchmark, intermediate controls improve held-out and worst-group metrics, whereas excessive affine control causes underfitting. A four-task single-seed transfer study provides descriptive corroboration. These results establish instantaneous allocation control and a bounded empirical operating regime, but do not imply a task-independent generalization ordering.
Comment: Treats weight-versus-bias update allocation as an independently controllable optimization variable with exact control laws.
Topic Match: The paper's main contribution is a mechanistic analysis and control method for neural-network update dynamics.
Relevance: 7 Novelty: 7
12. MentorPulse: Refreshing Cross-Model Latent Guidance for Long-Form Generation
ArXiv ID: 2608.20927
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Ziwu Liu, Guozhong Li, Chen Qiu, Weiyang Kong, Panos Kalnis
Abstract: Cross-model latent guidance lets a frozen large mentor encode an input once and a frozen small student generate from the resulting signal. Existing methods keep this signal fixed, assuming it stays useful as the output grows; we show this fails in long-form generation. On multi-turn instruction following, static guidance pushes a 4B student's constraint satisfaction 2.5 points below its no-guidance baseline; a training-free refresh every 16 tokens changes only the memory content and restores a 2.0-point gain over that baseline. We propose MentorPulse to keep guidance fresh at practical cost: it compresses mentor states into a capped slot memory, incrementally processes newly generated tokens, and updates the memory that the student reads through gated cross-attention without resetting the student's KV cache. Windowed Refresh Training exposes the bridge to prefix-conditioned memory. Across thirteen datasets, MentorPulse closes 52.2% of the mentor-student gap on macro average, outperforming C2C, T2T, and equal-budget LoRA, with the largest gains on long outputs. It performs best on all eleven mentor-student pairs from three model families, with margins that narrow as the capability gap grows, and a lightweight read-pattern check predicts the gain before deployment. Measured costs identify refresh intervals that dominate text guidance on long outputs.
Comment: Refreshes compressed mentor-state guidance during decoding through gated cross-attention without resetting the student KV cache.
Topic Match: Dynamic cross-model computation through refreshed latent memory is the core architectural contribution, with efficiency as a secondary benefit.
Relevance: 7 Novelty: 7
13. Free-Probability Kernels for Zero-Rollout Hyperparameter Selection in Reservoir Computing
ArXiv ID: 2608.20998
Primary Topic: Architecture and Training Dynamics
Authors: Sara Malacarne, Andrea Ceni, Claudio Gallicchio
Abstract: Reservoir computing (RC) couples a fixed recurrent dynamical system with a trained lightweight readout, but this efficiency is partly lost during hyperparameter selection: the recurrent gain, input scale, and leakage rate determine the reservoir's stability and temporal processing regime and are usually tuned through many rollouts. We introduce a deterministic, pilot-informed selector for leaky linear reservoirs followed by coordinate-wise nonlinear features. Free probability yields cross-lag propagation coefficients that summarize how the reservoir mixes past inputs. In the large-width limit, these coefficients define a deterministic temporal kernel that approximates the finite-reservoir feature geometry. Kernel ridge regression on a short labelled pilot sequence therefore ranks candidate operating regimes without instantiating or rolling out a reservoir, and the selected configuration transfers across widths. Across ten synthetic temporal benchmarks, zero-rollout selection obtains a mean deployment score of $0.772$, compared with $0.774$ for exhaustive simulation-based search, while avoiding $156\,600$ selection rollouts. With a small rollout budget, the proposed ranking provides the strongest mean performance at every tested budget and reaches the exhaustive reference using $4.8\%$ of its rollout cost. On four public electricity-transformer-temperature (ETT) forecasting datasets, five retained candidates recover the exhaustive operating point on three datasets. On multivariate cellular-traffic forecasting, 15 rollouts per cell reach the 462-rollout exhaustive reference and outperform random search and Bayesian optimization at low budgets. These results position free-probability kernels as deterministic surrogates for selecting reservoir operating regimes when validation rollouts are scarce.
Comment: Analytic temporal kernels predict recurrent propagation and select reservoir operating regimes without candidate rollouts.
Topic Match: Mechanistic analysis of recurrent computation drives hyperparameter selection and transfer across widths, with evidence limited to fixed linear reservoirs with nonlinear features.
Relevance: 7 Novelty: 7
14. Read, Write, Relax: Why Neural PDE Surrogates Need Both Global and Local Processing
ArXiv ID: 2608.21677
Primary Topic: Architecture and Training Dynamics
Authors: Anuj Kumar, Heiko Zimmermann, Josiah Bjorgaard, Jacan Chaplais, Nikolaos Bouklas, Matteo Salvador, Alexander Lavin
Abstract: Recent mesh-based simulation advances have, in no small part, relied on neural surrogates of two distinct families: global models that route information through a small set of latent tokens, and local models that perform message passing across mesh edges. Consistent with both classes is the inability to perform beyond low-dimensional problems and small-scale or oversimplified meshes, the simulation regimes where industrial problems reside. Our work shows this explicitly and presents a unified formulation. In global approaches, latent-token attention acts as a spatial low-pass filter, while local message passing lacks the global reach necessary to propagate information across large mesh spaces. Viewed through the error, the two operators are the halves of a multigrid cycle: one corrects errors at the lower end of the spectrum, the other at the higher end, and neither can do the other's job. We introduce Read-Write-Relax (RWR), which interleaves latent attention with message-passing relaxation under a unified formulation. The interleaved processor lowers error across the entire spectrum, making RWR the most accurate model in nearly every comparison across our industrial and public benchmarks. It is also markedly data-efficient in the scarce-data regimes, accurate on the engineering quantities of interest, and scales full-field predictions to challenging, large-scale problems.
Comment: Explains global attention and local message passing as complementary spectral corrections and interleaves them into a multigrid-like processor.
Topic Match: The core contribution is a mechanistically motivated global/local architecture, although its analysis and validation focus on mesh-based PDE surrogates.
Relevance: 7 Novelty: 7
15. Divisive Normalization Shapes Low-Rank Slow Manifolds for Continuous Working Memory
ArXiv ID: 2608.01947
Primary Topic: Architecture and Training Dynamics
Authors: Zhaotian Gu, Jie Su, Weiwei Wang, Chang Liu, Tianyi Qian, Dahui Wang
Abstract: The ability to robustly maintain and update continuous variables is a hallmark of working memory. While classical continuous attractor networks suffer from severe fine-tuning fragility, standard artificial recurrent neural networks (RNNs) like GRUs and LSTMs typically fail to stably learn continuous manifolds, instead shattering the state space into discretized point attractors. To bridge this gap, we draw inspiration from divisive normalization, a canonical neural computation widely observed across cortical circuits, and propose the Recurrent Divisive Normalization Network (RDNN), a minimal and algebraically isolated model of dynamic division. Through dynamical systems analysis on canonical working memory tasks, we demonstrate that this biophysical constraint allows the network to converge to robust, high-fidelity slow manifolds. Furthermore, we analyze the gradient dynamics of divisive normalization during Backpropagation Through Time (BPTT), showing that it introduces an activity-dependent local gradient scaling. This scaling dampens parameter updates in highly active regimes, which empirically aligns with a significant self-compression of the network's effective rank, confining the recurrent dynamics to a tight, low-dimensional subspace while avoiding the optimization pathologies associated with explicit low-rank factorization. Finally, ablations demonstrate that while subtractive inhibition can maintain static memories, divisive normalization is mathematically essential to prevent manifold shattering under time-varying inputs. Our findings identify divisive normalization not merely as a biological artifact, but as a critical computational mechanism for learning high-fidelity continuous representations.
Comment: Analyzes how recurrent divisive normalization rescales BPTT gradients and supports stable, low-rank continuous dynamics.
Topic Match: The normalization mechanism and its gradient dynamics are substantive architectural contributions, with evidence limited to canonical working-memory tasks.
Relevance: 7 Novelty: 7
16. Continuous-Time Quantum Walks based Graph Neural Network
ArXiv ID: 2608.20738
Primary Topic: Architecture and Training Dynamics
Authors: Yuliang Zhan, Zefeng Gao, Jian Li, Yang Liu, Hao sun
Abstract: Graph Neural Networks (GNNs) are widely used on graph-structured data, but most suffer from two key weaknesses. First, message passing behaves as a low-pass filter under the homophily assumption, leading to poor performance on heterophilic graphs. Second, stacking layers drives node features toward constants, causing over-smoothing. Existing methods usually address these issues separately, while the few joint solutions rely largely on empirical heuristics, and many over-smoothing remedies sacrifice model expressiveness. We propose \textbf{CTQW-GNN}, a GNN based on Continuous-Time Quantum Walks (CTQW), to address both issues with theoretical justification. Its design exploits two properties of the CTQW propagator $e^{-\mathrm{i}Ht}$. First, it is unitary and has eigenvalues on the unit circle, so no frequency component is damped, counteracting the low-pass bias. Second, unitarity preserves feature norms and prevents the Dirichlet energy from decaying exponentially with depth, thereby mitigating over-smoothing. CTQW-GNN combines three complementary aggregation modules. \textit{CTQW-based Aggregation} evolves node features through the unitary propagator, preserving mid- and high-frequency signals for heterophilic graphs while preventing Dirichlet-energy collapse. \textit{CTQW-Attention Aggregation} constructs a multi-hop neighbor graph from CTQW amplitudes and applies attention over it, enabling access to distant homophilic nodes missed by single-hop aggregation. \textit{LF Aggregation} uses a standard low-pass GAT branch to retain strong performance on homophilic graphs, where pure CTQW aggregation can be suboptimal. We further provide a spectral-gap analysis explaining energy preservation and a Lieb--Robinson-type bound that gives a principled rule for selecting the walk time $t$.
Comment: Unitary graph propagation preserves spectral components and limits depth-induced over-smoothing.
Topic Match: Introduces an aggregation architecture with spectral guarantees on signal preservation across depth, within the narrower setting of GNNs.
Relevance: 7 Novelty: 7
17. Where Cognition Lives: Dissecting Emergent from Computed Function in a Minimal Complete Cognitive Architecture
ArXiv ID: 2608.22347
Primary Topic: Architecture and Training Dynamics
Authors: Francisco M. Arrabal-Campos, Francisco G. Montoya, Alfredo Alcayde, Ignacio Fernández
Abstract: A cognitive architecture is more than the module that reasons: it must also decide how long to think and what deserves the effort. We built a minimal but complete system - a recurrent reasoner with adaptive halting, a homeostatic control field, and a value module - and asked of each part: does this function emerge from gradient descent, or must it be computed? Competence emerges. Stopping appears to emerge too, and to be worth more than everything decidable in advance, but that appearance is instrumentation: payoff at matched mean compute climbs from 0.467 (uniform) through 0.546 (difficulty) to 0.698 (ex-ante value), and the further climb to 0.921 (posterior self-observation) does not survive audit. PonderNet-style halting returns a halting-weighted mixture of hidden states while forced-depth baselines return one, and the language head is trained on the mixture alone; equalizing the readout annihilates the apparent advantage of native execution (residual +0.000 [0.000, 0.000]). Value does not emerge: trained couplings capture zero of a payoff an explicit allocator captures completely (+0.151, routing correlation +0.79), so the second-order decisions that pay must be computed, at least where value is orthogonal to content, as here by construction. On a frozen LLM actuator the same instruments show self-consistency voting to be a measured bound (+0.0236 [+0.0150, +0.0326]) and inter-sample agreement nearly worthless as a stopping signal, its mass concentrating on wrong answers. Every null we assert carries a mechanism and a positive control, and the protocol is part of the contribution. Executing our own falsifiable prediction, value under commitment pays +0.1312 [+0.1124, +0.1502] in a cliff-cost family, some seven times the smooth-family estimate - not because the cliff shifts information ex ante, but because it multiplies the attainable range fivefold (5.1x [3.4, 8.2]).
Comment: Shows that an apparent adaptive-halting advantage disappears when recurrent readout mechanisms are matched.
Topic Match: Directly analyzes adaptive recurrent computation and halting mechanisms, although the main evidence comes from a deliberately minimal architecture.
Relevance: 7 Novelty: 7
18. Loss-Parameterized Fisher Width Along Learning Trajectories
ArXiv ID: 2608.21561
Primary Topic: Architecture and Training Dynamics
Authors: Vu Khac Ky
Abstract: Fisher width measures the Gaussian width of a probe set after deformation by the local Fisher geometry. We study its evolution along learning trajectories and ask when training loss can serve as an effective coordinate for this quantity. We first derive an exact trace--shape factorization and a deterministic stability bound for fixed compact probes. In a population Gaussian-teacher logistic model, the teacher-aligned state is extremal on every loss level below $\log 2$: it has minimal parameter norm and maximizes both Fisher trace and Euclidean-ball Fisher width. We then show that population gradient flow asymptotically selects this branch, with explicit rates for the aligned and orthogonal coordinates. This yields, for $d\geq2$, [ \frac{w_F(B_2^d;θ(t))} {\sqrt{L(θ(t))}} \longrightarrow \frac{\sqrt6}Ï\mathbb E[Ï_{d-1}]. ] Controlled full-Fisher experiments support the matched-loss branch and the population predictions. In a nonlinear MLP with a diagonal model-Fisher approximation, GD and SGD remain close at matched loss, whereas Adam follows a substantially displaced branch; the fixed probes tested retain highly similar temporal shapes. These results support a branchwise, rather than universal, loss parametrization of Fisher width.
Comment: Characterizes optimizer-dependent Fisher geometry along learning trajectories at matched training loss.
Topic Match: The central result concerns optimizer-dependent training geometry, with evidence concentrated on logistic regression and small MLPs.
Relevance: 7 Novelty: 7
19. Dataset Complexity Shapes Finite-Distance Loss Geometry in Neural Networks
ArXiv ID: 2608.22361
Primary Topic: Architecture and Training Dynamics
Authors: Jaeyong Bae, Hawoong Jeong
Abstract: Finite datasets can share the same size and low-order statistics while differing strongly in structural complexity. We connect this dataset complexity to loss-landscape geometry by pairing local label mixing across neighborhood scales with local entropy around trained neural-network solutions. Adapted from the Franz--Parisi construction in spin-glass theory, local entropy measures the effective volume of low-loss, solution-like parameter configurations at each distance from a reference. We estimate it in finite networks using adaptive sequential Monte Carlo. In a controlled synthetic sweep, greater dataset complexity produces a larger decrease in local entropy near the reference. Farther away, its radial derivative becomes weak and nearly common across conditions. Dataset complexity therefore changes where the effective solution volume contracts, rather than making it decrease uniformly faster. Experiments on real image data show the same qualitative trend, with label randomization further amplifying the effect. These results show that dataset structure shapes how low-loss neighborhoods are organized across finite distances from trained solutions.
Comment: Connects dataset complexity to distance-dependent contraction of low-loss parameter neighborhoods.
Topic Match: Studies optimization-landscape geometry around trained solutions, with evidence from finite networks rather than large-model training runs.
Relevance: 7 Novelty: 6
20. Training, learning and inference: unified dynamics of neural systems
ArXiv ID: 2608.20965
Primary Topic: Architecture and Training Dynamics
Authors: Mian Wang
Abstract: We define an atomic generation fact f=(u,tau,omega,z;rho), recording the origin, realized transformation, concrete occurrence, generated result and relation role. Compiled into a Generation-Fact Graph (GFG), these facts provide an AI-native, compilable scientific fact substrate preserving generation histories. We establish a GFG-based recursive scientific process in which analysis, intervention, replay and validation form facts for later cycles. Using nanoGPT, we establish unified training-learning dynamics. Training is the evolution of a parameter-optimizer system with state and memory: each actual training action enters the receiving state and produces a finite-amplitude nonlinear functional response conditioned by that state and target-specific update geometry. Learning is the persistent reorganization of distributed functional support by these responses; capability formation, maintenance, decline or recovery becomes observable when target-specific states are evaluated against their readout boundaries. Three primary coordinates - target-boundary state, target-specific update geometry and parameter-Adam receiving state - yield a second-order predictor operating before post-update outputs are read. On held-out runs, it achieved 91.43% accuracy and 91.49% macro-averaged recall across four transitions. We further establish inference as a frozen projection of training-learning dynamics. Component gating and rollback show causal recruitment and non-additive combination of query-conditioned support formed during training, deriving organizational conditions realized by Attention. Controlled feedback indicates possible double-edged reinforcement effects. ResNet/CIFAR-100 and diffusion/CIFAR-10 experiments confirm receiving-state-conditioned responses, persistent support reorganization and frozen inference projection beyond nanoGPT.
Comment: A second-order predictor links training-update responses to parameter geometry and Adam state.
Topic Match: Its optimizer-state-conditioned account of update effects addresses training dynamics, although evidence centers on nanoGPT and CIFAR experiments.
Relevance: 7 Novelty: 6
21. SPARCL: Spectral Partitioned Analytic Continual Learning
ArXiv ID: 2608.21307
Primary Topic: Architecture and Training Dynamics
Authors: James Hartley, Zeropy Surio, Daniel Whitmore, Hannah Clarke, Thomas Reed
Abstract: Analytic continual learning has emerged as a strong exemplar-free alternative to gradient-based class-incremental learning because it replaces iterative optimization with closed-form ridge updates. Yet the usual forgetting narrative, centered on stochastic gradient overwriting, does not explain why analytic methods still drift on old classes despite exact recursive solvers. We identify the culprit as spectral interference: the joint ridge classifier for all tasks shares the inverse autocorrelation operator $(R+λI)^{-1}$, so incoming task samples that load onto old dominant eigendirections dilute the spectrum and perturb old-class logits even when old labels are never revisited. Based on this view, we propose SPARCL, a spectral partitioned analytic continual learner that decomposes the running autocorrelation into a high-energy core and a residual complement, freezes old-class classifier components in the core subspace, and updates only the residual block through recursive least squares with an optional residual random-projection expansion. This yields a simple closed-form update with a provable invariance guarantee for the core contribution of old logits. Across CIFAR-100, CUB-200, ImageNet-R, and ImageNet-A under a frozen ViT-B/16 protocol, SPARCL closes most of the gap from classical analytic learners to strong representation matchers, while remaining complementary to sparse feature-decorrelation approaches such as Fly-CL.
Comment: Spectral partitioning prevents recursive ridge updates from altering protected old-class logit components.
Topic Match: The spectral-interference analysis and constrained update rule address training dynamics, with evidence limited to frozen-feature continual classifiers.
Relevance: 6 Novelty: 7
22. Learning in PINNs: Phase transition, diffusion equilibrium, and generalization
ArXiv ID: 2403.18494
Primary Topic: Architecture and Training Dynamics
Authors: Sokratis J. Anagnostopoulos, Juan Diego Toscano, Nikolaos Stergiopulos, George Em Karniadakis
Abstract: We investigate the learning dynamics of fully-connected neural networks through the lens of the neural gradient signal-to-noise ratio (SNR), examining the behavior of first-order optimizers in non-convex objectives. Interpreting the drift/diffusion phases as proposed in the information bottleneck theory, we identify a third phase termed "diffusion equilibrium" (DE), a stable training phase characterized by highly-ordered neural gradients across the sample space. This phase is marked by an abrupt transition, where sample-wise gradients align (SNR increases), and stable optimizer convergence. Moreover, we find that when homogeneous residuals are also met across the sample space during the DE phase, this leads to better generalization, as the optimization steps are equally sensitive to each sample. Based on this observation, we propose a sample-wise re-weighting scheme, which considerably improves the residual homogeneity and generalization in quadratic loss functions, by targeting the problematic samples with large residuals and vanishing gradients. Finally, we explore the information compression phenomenon, pinpointing a significant saturation-induced compression of activations at the DE phase transition, driven by the sample-wise gradient directional alignment. Interestingly, it is during the saturation of activations that the model converges, with deeper layers experiencing negligible information loss. Supported by experimental examples on physics-informed neural networks (PINNs), which highlight the critical role of gradient agreement due to their inherent PDE-based interdependence of samples, our findings suggest that when both sample-wise gradients and residuals are ordered, this leads to faster convergence and better generalization. Identifying phase transitions could improve deep learning optimization strategies, enhancing physics-informed methods and machine learning performance.
Comment: Relates sample-gradient alignment and residual homogeneity to a stable convergence phase.
Topic Match: Mechanistic convergence analysis directly connects to training dynamics, although the evidence centers on PINNs and quadratic-loss training rather than large models.
Relevance: 6 Novelty: 7
23. Bounded Precision-Geometry Scaling for Robust Multi-Task Learning under Loss Scale Mismatch
ArXiv ID: 2608.21653
Primary Topic: Architecture and Training Dynamics
Authors: Krishna Subedi
Abstract: Multi-task learning often combines losses that span several orders of magnitude, causing homoscedastic uncertainty weighting to degrade severely. We propose Bounded Precision-Geometry Scaling (BPGS), a method that maps each task's log-variance through a bounded sigmoid parameterisation anchored to detached batch loss statistics, and decouples network optimisation from uncertainty optimisation. Its normalised task weights are provably invariant to uniform rescaling under non-degenerate loss scales. We evaluate BPGS on synthetic stress tests and three real-world benchmarks: NYUv2 dense prediction, Yeast multi-label classification, and RF1 multi-target regression. Under pure loss rescaling from $\times 1$ to $\times 1000$, its macro score changes from 0.777 to 0.778, whereas Kendall weighting drops from 0.780 to 0.637; $\ell_1$-normalising Kendall's weights does not close the gap. On NYUv2, BPGS records the lowest depth absolute relative error (0.223), depth RMSE (0.790), and total loss (1.891) among all compared methods, including Nash-MTL. Sensitivity studies on batch size and calibration show small variation across the tested ranges, and runtime overhead relative to Kendall is under 1%. BPGS posts the highest Yeast micro-F1 (0.616) and is competitive on RF1, though PCGrad leads RMSE and MAE there. These findings establish BPGS as a scale-robust alternative to homoscedastic uncertainty weighting, notably effective when loss-scale disparities dominate multi-task optimisation.
Comment: Bounded uncertainty weights stabilize multitask optimization under large loss-scale mismatches.
Topic Match: Loss-scale robustness connects to optimization dynamics, but the contribution and evidence concern conventional multitask learning with limited large-model training relevance.
Relevance: 6 Novelty: 6
24. Socialized Division and Collaboration: Rethinking Class-Incremental Learning under Optimization Conflicts
ArXiv ID: 2608.21044
Primary Topic: Architecture and Training Dynamics
Authors: Xinjie Yao, Zhihe Fan, Yunqi Zhu, Jiaqi Zhou, Dengyu Zhao, Zhoupeng Guo, Yan Fan, Guosong Jiang, Pengfei Zhu
Abstract: Class-incremental learning is commonly instantiated as a single-model paradigm, where a unified model sequentially adapts to an unbounded stream of sessions. While effective under mild distributional shifts, this formulation becomes strained when successive sessions induce incompatible optimization directions, leading to destructive interference and catastrophic forgetting. We argue that such forgetting reflects a structural limitation of enforcing heterogeneous learning dynamics within a single parameter space. Motivated by social solidarity theory, we propose Socialized Division and Collaboration (SDC) as a reformulation of continual learning that decomposes session learning across specialized models in response to optimization conflicts, while enabling coordinated collaboration. To support this formulation with a principled allocation mechanism, we introduce an energy-based session-model compatibility criterion grounded in Helmholtz free energy, which guides adaptive session allocation and model evolution under conflicting objectives. This framework integrates session assignment, model evolution, and collaborative inference into a unified pipeline, offering an alternative to monolithic continual learning formulations and highlighting a broader design principle for learning under persistent optimization conflicts.
Comment: Uses an energy-based compatibility criterion to allocate learning sessions across evolving specialized models.
Topic Match: Adaptive model allocation is a modular learning mechanism, but its class-incremental scope and unspecified large-model behavior limit relevance.
Relevance: 6 Novelty: 6
Efficiency, Compression, and Large-Scale Training (26)
1. Beyond Sparse Weights: When Is Attention Compressible?
ArXiv ID: 2608.21541
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Chiwun Yang, Xiaoyu Li
Abstract: KV-cache compression is often justified by attention maps with a few large weights. This is incomplete: large weights may not contain most of the mass, omitted values can cancel, and preserving the attention output may not preserve the task. We separate these questions. Global score gaps -- not threshold counts -- determine how many tokens are needed to retain a target mass. For a realized row, the weighted sum of omitted values is the exact missing statistic. A controlled retrieval--aggregation model explains when truncation helps and when it hurts. These results motivate CertKV, a training-free compressor that reserves one tail-summary slot per head and allocates the rest by value dispersion. Under matched budgets, CertKV is top-two in seven of nine LongBench-v2 settings, remains in the leading compressed tier on 128K RULER, and realizes a ten-fold cache budget in a packed Llama prototype. Compressibility depends on the mass, values, future queries, and task -- not on a sparse-looking map alone.
Comment: KV-cache compression preserves omitted-value information through per-head tail summaries and dispersion-based allocation.
Topic Match: Attention mass and value geometry directly motivate a cache-compression mechanism that preserves tail contributions under a fixed memory budget.
Relevance: 9 Novelty: 8
2. A JoLT for the KV cache: Near-lossless KV cache compression via joint Lagrangian allocation of Tucker ranks and a rotated residual for llms
ArXiv ID: 2607.12550
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Rahul Krishnan, Volker Schulz
Abstract: The key-value (KV) cache has become the dominant memory cost of transformer inference: it grows with batch size, context length, and depth, and at long context it, rather than the model weights, sets the throughput ceiling. Existing reductions fall into two families. Low-rank methods factor two-dimensional slices of the cache, either per-head matrices or cross-layer feature blocks, and quantization methods lower the bit-width of every entry. Neither exploits the fact that the cache at a layer is naturally a third-order tensor whose three axes, the heads, the tokens, and the features, carry very different amounts of redundancy. We take this tensor view directly. Our method, JoLT (Joint Lagrangian Tucker), applies a partial Tucker decomposition that compresses only the token and feature axes while leaving the head and layer axes intact, then restores the energy that truncation discards with a rotated low-bit residual: a random orthogonal rotation followed by low-bit quantization. A single Lagrangian dual allocates the Tucker ranks and the residual bit-widths together, per layer group and separately for keys and values, under one byte budget. The result is a near-lossless 2-3x compression. Perplexity stays near-lossless on both a grouped-query-attention model (Mistral-7B-v0.3) and a multi-head-attention model (LLaMA-2-13B), and GSM8K accuracy and needle-in-a-haystack retrieval hold at the uncompressed baseline at 2x on both architectures and through 3x on the GQA model. At 2x, JoLT reconstructs the cache to relative Frobenius error 0.009 (K) and 0.006 (V) on both architectures. A randomized-SVD variant, FlashJoLT, delivers a 5-13x compression-time speedup at 1024-token context and matched quality.
Comment: Jointly allocates Tucker ranks and quantized residual bits under a KV-cache memory budget.
Topic Match: Introduces a concrete KV-cache compression mechanism that directly reduces transformer inference memory.
Relevance: 9 Novelty: 7
3. Jacobian-guided Noise Injection for Quantization Robustness in Large Language Models
ArXiv ID: 2608.20988
Primary Topic: Efficiency, Compression, and Large-Scale Training
Also Matches: Architecture and Training Dynamics
Authors: Deepanshu Pandey, Arnav Chavan, Nahush Lele, Sankalp Dayal, Deepak Gupta
Abstract: Quantization of Large Language Models (LLMs) is often hindered by the sensitivity of the self-attention mechanism to discretization errors. We identify the softmax operator as a bottleneck for quantization stability due to its sensitivity to outliers and state-dependent Jacobian. We theoretically establish that suppressing the norm of this Jacobian helps in bounding quantization-induced performance degradation. Based on this, we propose Jacobian-Guided Noise Injection, a training strategy that injects zero-mean Gaussian noise into pre-attention logits, with variance derived directly from the Jacobian Frobenius norm. Unlike prior approaches that rely on heuristic or penalise jacobian directly, our method provides a way to identify the optimal noise variance based on the local attention sensitivity. We evaluate the method on SOTA LLM architectures, where it demonstrates improved robustness over popular PTQ methods. Empirical analysis reveals that the proposed method gives up to +37% relative gains on Top-1 accuracy on ImageNet-1K for SigLIP and improves relative perplexity by upto 40% on WikiText for language models in low bit quantisation settings, proving the efficacy of the approach.
Comment: Trains attention for low-bit robustness using logit-noise variance derived from the softmax Jacobian norm.
Topic Match: Quantization robustness is central, while the attention-Jacobian analysis provides a secondary training-dynamics contribution.
Relevance: 9 Novelty: 7
4. NeuroPrefetcher: Storage-Aware Sparse LLM Inference via Delta Prefetching
ArXiv ID: 2608.22643
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Nobel Dhar, Md Romyull Islam, Xuechen Zhang, Gongjin Sun, Sahidul Islam, Bobin Deng, Kun Suo
Abstract: Deploying large language models on edge devices is increasingly limited by a widening gap between model size and available memory. Existing approaches such as quantization, smaller models, and offloading can raise the effective memory limit, but they still assume that the model can be compressed or partitioned to fit within some budget. We target the harder model-exceeds-memory setting, in which the model remains larger than resident memory throughout execution and storage becomes an active source of weights on the critical path. We observe that MLP activity during autoregressive decoding has strong temporal locality: approximately 82-85% of active neurons persist from one token to the next. This means that most sparse weights needed for the current token are already resident, and only the newly needed rows must be fetched from storage. We present NeuroPrefetcher, a storage-backed LLM inference system that exploits this property through predictive delta prefetching. After layer 0, a single GPU-resident predictor, occupying 2.86% of base model parameters, predicts sparse activity for all downstream MLP layers in one forward pass. The runtime compares these predictions against resident GPU buffers and issues application-scheduled NVMe reads only for incoming delta rows, replacing reactive operating-system demand paging with explicit, model-aware weight movement. On real unified-memory edge hardware, NeuroPrefetcher achieves 7.9-12.0x speedup over llama.cpp across constrained memory budgets.
Comment: Predictive delta prefetching exploits temporal MLP sparsity to fetch only newly required weight rows.
Topic Match: The core mechanism reduces storage traffic for sparse LLM inference when model weights exceed resident memory.
Relevance: 9 Novelty: 7
5. HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models
ArXiv ID: 2509.23928
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Zhinan Xie, Peisong Wang, Shuang Qiu, Jian Cheng
Abstract: Speculative decoding has proven effective for accelerating inference in Large Language Models (LLMs), yet its extension to Vision-Language Models (VLMs) remains limited by the computational burden and semantic inconsistency introduced by visual tokens. Recent studies reveal that visual tokens in large VLMs are highly redundant, and most of them can be removed without compromising generation quality. Motivated by this observation, we propose HiViS (Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models), a framework that utilizes the target VLM as a semantic fusion model, allowing the drafter to obtain visual information without explicitly processing visual tokens, ensuring that the drafter's prefill sequence length matches that of the textual tokens. Furthermore, HiViS employs a time-step-aware aligned training scheme that allows the drafter to autonomously propagate and refine instructive visual-textual semantics during independent drafting, guided by step-dependent bias-correction residuals. Extensive experiments across representative VLMs and benchmarks demonstrate that HiViS achieves significant improvements in average acceptance length and speedup ratio.
Comment: Bypassing visual tokens in the speculative drafter reduces prefill work while target-derived semantics support draft acceptance.
Topic Match: Directly accelerates large VLM generation through visual-token bypass and a trained semantic-alignment mechanism for speculative decoding.
Relevance: 9 Novelty: 7
6. SSDi8: Accurate and Efficient 8-bit Quantization for State Space Duality
ArXiv ID: 2608.21952
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Hyunwoo Kim, Byoungchan Ko, Minseok Kang, Minwoo Kim, Dongjin Lee, Jaehoon Lee, Sungroh Yoon, Dahuin Jung
Abstract: Recent advances in sequence modeling have highlighted Mamba as a state space architecture offering efficient long-range dependency modeling and providing a viable alternative to Transformers. Building upon this, Mamba-2 introduces the Structured State Space Duality (SSD), which integrates recurrent and attention modes to achieve efficiency and scalability. However, this architectural expansion substantially increases memory and latency overhead, underscoring the need for efficient compression strategies tailored to SSD. In this work, we present SSDi8, the first post-training quantization framework specifically designed for SSD to maintain a persistent INT8 path. SSDi8 introduces a reformulation that decouples element-wise multiplications from matrix multiplications, enabling reuse of quantized activations across modules. Moreover, SSDi8 adaptively quantizes channel-varying activations at cost-effective points, further reducing latency. On the accuracy side, SSDi8 explicitly leverages the intrinsic dimensional decomposition of SSD, exploiting distinct outlier distributions across axes, and incorporates an error correction term based on per-channel error statistics. Comprehensive experiments demonstrate that SSDi8 achieves accuracy comparable to FP16 while delivering up to 1.4x speedup in W4A8 and W8A8 settings. We further validate its robustness in resource-constrained environments by deploying it on the Orin NX device.
Comment: An SSD operator reformulation enables quantized activation reuse and a persistent INT8 computation path.
Topic Match: Operator reformulation, axis-aware quantization, and error correction directly reduce Mamba-2 execution cost while preserving accuracy.
Relevance: 9 Novelty: 7
7. COEC: Calibrated Orthogonal-Equivalence Compensation for Structured Pruning of Large Language Models
ArXiv ID: 2608.21142
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Peiqi Yu, Nam Ling, Wei Wang, Wei Jiang
Abstract: Structured pruning reduces the size and inference cost of large language models (LLMs) by removing weight columns, but the resulting output error can degrade accuracy. Existing training-free compensation methods use an additive bias or a single orthogonal rotation on the output side of the retained weight. These corrections leave its input singular frame unchanged and therefore limit how the retained weight can adapt after column removal. We propose COEC (Calibrated Orthogonal-Equivalence Compensation), a training-free compensation framework that applies alternating left and right orthogonal rotations to the retained weight. The right rotation is optimized on a reduced Stiefel manifold, while singular values are rescaled using generalized cross-validation to select the regularization strength for each layer. COEC further tempers the calibration Gram matrix to reduce the dominance of high-energy activation directions and introduces an alignment penalty that preserves the geometric relation between adjacent attention projections.All components use second-order statistics from a small calibration set and require neither backpropagation through the LLM nor retraining of the model parameters. COEC is independent of the column pruning criterion and can be applied to multiple structured pruning methods. Experiments on the Llama-3, Llama-3.1, and Qwen2.5 model families across multiple structured sparsity levels show that COEC improves perplexity on every model and zero-shot accuracy in most settings over existing compensation methods, with larger gains at higher sparsity. These results show that post-pruning compensation can recover part of the performance lost to column removal.
Comment: Training-free structured LLM pruning compensation through two-sided orthogonal rotations.
Topic Match: Directly improves structured compression by recovering pruning accuracy without retraining.
Relevance: 9 Novelty: 7
8. Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs
ArXiv ID: 2608.21134
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Luka Ribar, Jeevan Bhoot, Douglas Orr
Abstract: Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. We present a framework for quantizing VLMs for efficient inference on resource-constrained hardware. Our approach combines a quantization pipeline that uses the model itself to generate training data and does not require access to the training setup, with a novel 2.7-bit-per-parameter format supporting efficient execution on Arm CPUs. We validate our approach by compressing the Llama 3.2 11B Vision Instruct model to 3.7 GB with 8-bit activations, preserving strong performance on a set of standard visual question answering tasks.
Comment: Introduces an Arm-executable 2.7-bit weight format with 8-bit activations for large vision-language models.
Topic Match: A new low-bit representation and quantization pipeline directly target large-model memory requirements and efficient CPU execution.
Relevance: 9 Novelty: 7
9. Width-Independent Compressibility of Deep Neural Networks
ArXiv ID: 2608.21752
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Hong-Yi Wang, Mingze Wang, Liu Ziyin
Abstract: It has long been known that well-trained neural networks can be compressed very strongly without affecting their performance, an important phenomenon that remains poorly understood. We prove a uniform compressibility theorem for deep multilayer perceptrons with analytic activations. For a deep, wide fixed teacher network, there exists a narrow (same depth) network that approximately represents the same function as the original. The reachable compressed width is strikingly independent of the original width, but is $O((\log(1/\varepsilon))^{d_{in}})$, where $\varepsilon$ is the error budget and $d_{in}$ is the effective input dimension. Our construction involves a novel derivative-matching technique which is aware of the low-dimensional input, and a layer-wise reweighting that preserves the input-output mapping.
Comment: Proves same-depth MLP compression to a width independent of teacher width through derivative matching and layer reweighting.
Topic Match: Neural-network compressibility is the central contribution, although the guarantee requires analytic activations and depends strongly on effective input dimension.
Relevance: 8 Novelty: 8
10. Tensor Seeks Layout: Formalizing Layout Selection for ML Compilers
ArXiv ID: 2608.21555
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Clemens Eisenhofer, Yuwen Jia, Daniel Kroening, Sergey Pupyrev
Abstract: Modern machine learning compilers select tensor memory layouts to minimize execution cost under hardware constraints. Layout selection is global: an operator may be fastest under one layout while its consumers prefer another, and aligning these preferences requires explicit layout conversions that can hurt model performance. Despite its practical importance, layout selection lacks a formal basis, so current compilers rely on ad-hoc heuristics. This paper presents the first formal study of layout selection in machine learning compilers. We formulate the problem as combinatorial optimization over dataflow graphs, minimizing the sum of operator execution costs and the per-tensor cost of these conversions. Our theoretical analysis shows that optimal layout selection is computationally hard, even for programs containing only matrix multiplications over two-dimensional tensors. We design an optimal polynomial-time algorithm for dataflow graphs of bounded treewidth. For general instances, we give a weighted MaxSAT encoding that an off-the-shelf solver can optimize. The formulation unifies several existing layout optimization strategies, including XLA's layout assignment, partition dimension selection in systolic array compilers, and layout planning in mobile GPU optimizers. We implement the formalization in a production compiler for an AI accelerator and measure the execution time of the compiled models under greedy heuristics, the compiler's rule-based strategy, and an optimal solver. Simple heuristics degrade execution time by up to $5\times$ on some workloads. Where the compiler's cost model is accurate, the solver matches or beats the rule-based strategy. On workloads with complex data movement it falls behind, and since the solver minimizes the stated objective exactly, that gap isolates cost-model error from search quality, showing where compiler effort actually pays off.
Comment: Global tensor-layout optimization accounts jointly for operator execution and layout-conversion costs.
Topic Match: Formal compiler algorithms optimize tensor memory layout and data movement, making model-execution efficiency the strongest fit.
Relevance: 8 Novelty: 8
11. Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs
ArXiv ID: 2608.20953
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Bakbergen Ryskulov, Iker GarcÃa-Ferrero, David Montero, David Jansen, Ali Hashemi, Jezabel R. Garcia, Antonio Tiene, Román Orús
Abstract: Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a fraction of their parameters and quantized to 4 bits. Together these steps degrade reasoning, mathematics, coding, and long-context behavior enough to require a recovery, or healing, stage before deployment. The default recipe, quantization-aware training (QAT), re-fits the compressed, quantized model to hard labels; in our pipeline it converged slowly and collapsed past its peak. We adopted Quantization-Aware Healing (QAH) instead. Because a structurally compressed model is never independently trained at full precision, its bfloat16 checkpoint is a distillation-recovered approximation of the original; QAH distills the 4-bit student directly from the original, uncompressed model. On a GPT-OSS 120B to 60B to MXFP4 pipeline, the QAH student matches or beats its bfloat16 source on 7 of 9 benchmarks at roughly 4 times less weight memory and half the teacher's parameter count, and is released open-weight as Hypernova-60B. Against a matched QAT baseline it reaches a comparable peak about 7 times faster and stays stable under continued training, without hand-tuned early stopping. We also report deployment lessons, including a large, reproducible quality gap between distributed-training backends. Our aim is a recipe deployable without a multi-week hyper-parameter search.
Comment: Distills compressed 4-bit LLMs directly from the original teacher to improve recovery-training speed and stability.
Topic Match: Recovery training for jointly compressed and quantized large models directly addresses compression quality and training cost.
Relevance: 9 Novelty: 6
12. Scaling Laws for Task-Specific LLM Distillation
ArXiv ID: 2606.24747
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Lavinia Ghita, Dhruv Desai, Ioana Boier
Abstract: Large Language Models (LLMs) achieve strong performance across a growing range of domains, yet their scale poses deployment challenges in applications where latency and cost constraints are critical. This paper derives empirical scaling laws for domain-specific LLM compression, quantifying how in-domain and general knowledge performance scale with dataset size, compression ratio, supervision format, and iterative pruning schedule. Using quantitative finance as our application domain, we compare logit-based and LoRA-based distillation under iterative structural pruning, introducing a blended chain-of-thought supervision loss that stabilizes KL-divergence distillation over reasoning traces. In-domain task quality degrades predictably under compression while general-knowledge benchmarks collapse well before the same point; supervision format is the key driver of this tradeoff, with chain-of-thought supervision actively recovering general knowledge that pruning erases. We release the headline dataset FinHeadlineMix, scaling law results, and practical recommendations to provide a reusable framework for domain-specific compression decisions.
Comment: Derives empirical scaling laws connecting dataset size, compression ratio, supervision format, and iterative structural pruning.
Topic Match: The core contribution quantifies compression behavior and tradeoffs under distillation and structural pruning.
Relevance: 8 Novelty: 7
13. GraniKV: Asymmetric Granularity KV-Cache Paging for Multi-Agent Systems with Long Shared Prefix
ArXiv ID: 2608.15584
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Jinhyun Jeon, Sungjoo Yoo
Abstract: Production paged-serving engines apply uniform paging granularity to the KV cache, even though the two regions of a multi-agent workload have opposite storage requirements: a long shared prefix demands contiguity, while the per-request suffix demands fine-grained allocation. We present \textbf{GraniKV}, a KV-cache layer that allocates the shared prefix in a contiguous HOT pool and the suffix in a token-level COLD pool, combined with a per-step dispatcher which selects the appropriate backend among dual backends for each regime (compute-, memory-, or communication-bound). To the best of our knowledge, GraniKV is the first system to apply asymmetric paging granularity to the KV cache of a production paged-serving engine. At $L_p{=}16$\,K shared tokens GraniKV reaches $\mathbf{2.16\times}$, $\mathbf{1.98\times}$, and $\mathbf{1.57\times}$ output-token throughput over the production baseline on Llama-3.1-8B/TP=1, Qwen-2.5-14B/TP=2, and Qwen-2.5-32B/TP=4. The gain decomposes: cascade attention integration contributes the majority at saturation; the asymmetric storage layer adds $1.05$--$1.15\times$ end-to-end while being what makes the batched-GEMM prefix backend possible at all. Under heterogeneous multi-agent serving with \emph{distinct} prompts of different lengths, the attribution inverts: GraniKV sustains $\mathbf{1.95\times}$ while batch-global cascade collapses to parity --- the storage layer alone carries the win in the regime that motivates the paper.
Comment: Introduces asymmetric KV-cache paging and regime-aware dispatch that materially increase long-prefix serving throughput.
Topic Match: The central innovation is a new KV-cache storage and execution design that substantially changes inference cost.
Relevance: 8 Novelty: 7
14. When Less Latent Leads to Better Relay: Information-Preserving Compression for Latent Multi-Agent LLM Collaboration
ArXiv ID: 2604.13349
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Yiping Li, Zhiyu An, Wan Du
Abstract: Multi-agent LLM systems are moving beyond discrete-token messages toward richer relays that preserve internal state. Recent work such as LatentMAS transmits full key-value (KV) caches between agents but pays a high memory and communication cost. We adapt KV-cache eviction to this setting and introduce \textbf{Orthogonal BackFill (OBF)}, which injects a low-rank residual from the discarded KV states back into the retained ones, orthogonal to what is already kept. With only $9.9\%$-$20.2\%$ of the prompt KV retained, compressed relay cuts bandwidth by $4.7\times$ and GPU memory by $8\%$ at under $5\%$ wall-clock overhead, and stays close to full relay in accuracy across nine benchmarks, ahead of it on several. OBF matches or improves over headwise eviction on all nine, and its gain is proportional to the accuracy gap eviction opens against full relay ($r{=}0.78$ across three model scales), so it gives back part of what eviction takes. Code is available at https://github.com/markli404/When-Less-Latent-Leads-to-Better-Relay.
Comment: Orthogonal low-rank backfilling preserves discarded KV information while reducing cache-relay bandwidth and memory.
Topic Match: The core contribution is a substantive KV-cache compression mechanism, making it relevant despite its multi-agent relay setting.
Relevance: 8 Novelty: 7
15. DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference
ArXiv ID: 2501.10375
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Yujie Zhang, Shivam Aggarwal, Tulika Mitra
Abstract: Mixture-of-Experts (MoE) models, though highly effective for various machine learning tasks, face significant deployment challenges on memory-constrained devices. While GPUs offer fast inference, their limited memory compared to CPUs means not all experts can be stored on the GPU simultaneously, necessitating frequent, costly data transfers from CPU memory, often negating GPU speed advantages. To address this, we present DAOP, an on-device MoE inference engine to optimize parallel GPU-CPU execution. DAOP dynamically allocates experts between CPU and GPU based on per-sequence activation patterns, and selectively pre-calculates predicted experts on CPUs to minimize transfer latency. This approach enables efficient resource utilization across various expert cache ratios while maintaining model accuracy through a novel graceful degradation mechanism. Comprehensive evaluations across various datasets show that DAOP outperforms traditional expert caching and prefetching methods by up to 8.20x and offloading techniques by 1.35x while maintaining accuracy.
Comment: Activation-aware expert placement and predictive CPU expert computation reduce MoE inference offloading latency.
Topic Match: The contribution is a new heterogeneous expert-execution scheme that reduces inference cost; it does not modify MoE training.
Relevance: 8 Novelty: 7
16. SAEM: Stage-Aware Expert Management for Memory-Efficient MoE Inference in Chain-of-Thought Reasoning
ArXiv ID: 2608.21614
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Yujie Zhang, Bin Gao, Tulika Mitra
Abstract: Chain-of-thought (CoT) prompting improves LLM reasoning by decomposing complex problems into intermediate steps, but its sequential nature increases decoding latency and memory usage. Mixture-of-Experts (MoE) models scale capacity through sparse expert activation, yet their full expert weights often exceed GPU memory and require costly GPU-CPU transfers. Existing runtimes treat all tokens uniformly, overlooking a key structural property of CoT traces: consecutive reasoning stages exhibit coherent and predictable expert activation patterns. Ignoring this stage-level regularity leads to inefficient caching and unnecessary data movement. We propose SAEM, a stage-aware MoE inference runtime that detects reasoning stage boundaries and exploits stage-level activation coherence to guide expert placement. SAEM combines stage-aware caching, expert-aligned token repacking, and in-situ CPU execution to reduce data transfer and kernel fragmentation. On mathematical and scientific reasoning workloads, SAEM achieves an average 1.33x throughput improvement over the strongest state-of-the-art caching and offloading baselines under constrained GPU memory, rising to 1.54x when calibration data matches the workload, demonstrating the effectiveness of stage-aware, locality-driven MoE inference for CoT reasoning.
Comment: Stage-aware expert caching exploits coherent activation patterns to reduce MoE weight transfers.
Topic Match: The core contribution is a memory-efficiency mechanism for large-model inference through stage-dependent expert placement and execution.
Relevance: 8 Novelty: 7
17. Continuous Adversarial MeanFlow Transfer
ArXiv ID: 2608.19540
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Yara Bahram, Zahra Dehghani, Mélodie Desbos, Eric Granger, Pablo Piantanida, Mohammadhadi Shateri
Abstract: Training fast generators on new domains with limited data remains challenging for two reasons. First, adapting a pretrained diffusion or flow model to a new domain leaves its costly multi-step sampling unaddressed, and existing acceleration methods are tied to the source parameterization--$ε$, $x$, $v$, or $u$--leaving heterogeneous pretrained models with no common acceleration target. Second, while adversarial refinement is proven effective for few-step quality, it is formulated only for instantaneous-velocity flows, not for the finite-interval average velocities that MeanFlow (MF) models predict. We address both problems. We propose MeanFlow-Transfer, which maps heterogeneous source outputs into a shared velocity representation, uses it to initialize an MF generator from the source weights, and optimizes an MF objective on the target domain. This unifies adaptation and acceleration in a single training loop across a broad range of pretrained models. We then introduce Continuous Adversarial MeanFlow, a post-training stage that extends continuous adversarial flow models from instantaneous velocities to MF's finite-interval average velocities. CAMF contrasts changes in a learned potential between real and predicted interval endpoints, recovering fine detail that MF regression averages away, and reduces to the instantaneous criterion in the vanishing-interval limit. Adapting four ImageNet-based source models--DiT ($ε$), SiT ($v$), JiT ($x$), iMF ($u$)--to five target domains, MF-T with CAMF matches or exceeds the fine-tuned teacher in FID and FDD at up to $125\times$ fewer Neural Function Evaluations (NFEs), while CAMF improves MF-T's few-step FID by $29\%$ on average.
Comment: Finite-interval adversarial MeanFlow training converts heterogeneous pretrained diffusion models into few-step generators.
Topic Match: The core contribution is a general sampling-acceleration mechanism across source parameterizations, reporting up to 125x fewer neural function evaluations.
Relevance: 8 Novelty: 7
18. CD-LoRA: Consistency-Driven Low-Rank Adaptation for Multi-Task Fine-Tuning
ArXiv ID: 2608.21909
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Qian Zha, Jinda Liu, Yuan Wu, Yi Chang
Abstract: While Multi-Task Learning (MTL) is essential for adapting Large Language Models (LLMs) to diverse domains, prevailing LoRA-based methods rely on complex routing mechanisms that partition task-specific knowledge. In this work, we reveal that such routing-based designs are prone to a training-inference discrepancy, where stochastic routing decisions under distribution shifts compromise inference stability. Driven by a second-order Taylor analysis that exposes the instability induced by routing variance, we challenge the training-inference discrepancy and propose Consistency-Driven Low-Rank Adaptation (CD-LoRA). By eliminating routers entirely, CD-LoRA employs a consistency-driven alignment mechanism to enforce representation congruence across tasks in a shared low-rank space. This paradigm fosters robust, task-agnostic features without explicit partitioning overhead. Extensive experiments show that CD-LoRA consistently outperforms state-of-the-art multi-adapter baselines, offering a simpler, router-free, and more stable solution for multi-task PEFT. The code is available at the anonymous link https://github.com/zhaqian21/CD-LoRA.
Comment: Router-free multi-task LoRA aligns tasks within a shared low-rank parameter space.
Topic Match: The core contribution redesigns parameter-efficient adaptation, removing adapter-routing overhead and addressing routing-induced inconsistency.
Relevance: 8 Novelty: 6
19. PowerSlider: Exploiting Phase Asymmetry for LLM Serving under Demand Response
ArXiv ID: 2608.21719
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Yueying Li, Jiayang Chen, Yuanfan Chen, Leo Han, Haoran Qiu, Esha Choukse, Rodrigo Fonseca, Udit Gupta
Abstract: AI inference clusters are increasingly constrained by instantaneous power, not just energy: grid operators condition new capacity on demand response, imposing time-varying power caps. Existing LLM serving systems optimize a static energy objective or shed fixed priority tiers under load; either way, goodput collapses when the power envelope moves. An LLM pipeline is not a uniform load: compute-bound prefill loses throughput almost linearly with GPU frequency, memory-bound answer decode sustains it down to $0.57\times$ nominal, and reasoning's thinking phase couples KV-cache capacity to scheduling -- so a cap should be steered to where each watt costs the least performance. PowerSlider does so with a new Flex SLO contract that turns bounded user slack into an optimization constraint, prefill--think--answer disaggregation exposing per-stage frequency and KV control, and a Karush--Kuhn--Tucker (KKT) online solver re-solving within 7.7 ms of every cap change, backed by a consolidated fail-safe that power-gates drained instances when DVFS bottoms out on static power. On SGLang with production traces, \sys{} sustains 78.3\% online goodput at a 30\% cap reduction versus 47.6\% for the best of five baselines ($1.64\times$), holds latency-critical tails within $1.3\times$ of nominal (baselines: $2.3$--$6\times$, up to $12\times$), and delivers 92\% mean goodput through a replayed CAISO grid-emergency day bottoming at $0.41\times$ (54\% at the trough; every baseline below 7\%).
Comment: Exploits phase-specific frequency sensitivity and KV-cache constraints to allocate limited power across LLM inference stages.
Topic Match: Its online power-allocation algorithm introduces a substantive LLM execution-efficiency mechanism, although the scope is serving.
Relevance: 7 Novelty: 7
20. Pandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimation
ArXiv ID: 2608.20316
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Adam Fisch, Shubhendu Trivedi, Fantine Huot, William W. Cohen, Michael Kaisers, Mirella Lapata, Kate Larson, Jacob Eisenstein
Abstract: Heterogeneous AI systems composed of multiple models, architectures, harnesses, or inference-time settings can improve quality and efficiency by routing queries to the specialist who can answer most effectively at the lowest cost. Routing requires estimating each specialist's expected return, but this value estimation has a cost. Cheap estimators (e.g., embedding-based predictors) are fast but noisy, while accurate estimators (e.g., fine-tuned models with access to retrieval results or partial reasoning traces) are expensive. We formalize this tradeoff as an instance of Pandora's Box, the classical problem of optimal search with costly inspection. Under a Gaussian signal model, the resulting policies have closed-form value-of-information expressions that determine, for each specialist and input, whether refining the value estimate is worth its cost. We call the centralized policy Pandora's Router. We extend this to a decentralized setting, Pandora's Bidder, where specialists independently decide whether to invest in self-assessment before accepting an offered price to claim a query. Experiments across three domains---a standard multi-LLM benchmark, retrieval-augmented specialists, and LLMs with variable inference-time reasoning---show that Pandora's Router matches the routing quality of exhaustive estimation, while querying the expensive estimator far less often. In the decentralized setting, value-of-information reasoning improves allocative efficiency when competing estimates are accurate; when competing estimates are noisy, however, it can increase the strategic specialist's utility at the expense of others.
Comment: Uses value-of-information policies to avoid unnecessary expensive quality estimates when routing LLM queries.
Topic Match: The substantive efficiency mechanism allocates inference and estimation budgets across whole-model specialists, placing it within inference-cost optimization.
Relevance: 7 Novelty: 7
21. Event-triggered Implicit Perturbation for Zeroth-Order Fine-Tuning of Spiking Transformers
ArXiv ID: 2608.21223
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Tengteng Lei, Prabodh Katti, Rashi Dutt, Houssem Sifaou, Tan Peng, Osvaldo Simeone, Kai Xu, Bipin Rajendran
Abstract: Zeroth-order (ZO) optimization estimates gradients using only forward-pass evaluations, making it suitable for fine-tuning non-differentiable, event-driven spiking neural networks (SNNs). However, its deployment on in-memory computing (IMC) accelerators is constrained by the repeated read-modify-write (RMW) operations arising from explicit weight perturbation and the prohibitive hardware footprint of random number generators (RNGs) for statistically independent per-weight perturbations. To address these challenges, we propose an implicit-perturbation ZO (IPZO) architecture in which perturbation sums computed by an event-triggered perturbation generation unit (PGU) are combined with the weighted sums produced by the IMC array, eliminating perturbation-induced RMW operations while preserving weight-stationary execution of IMC. By exploiting spike sparsity, the PGU generates and accumulates perturbation contributions only for spike-activated weight rows, reducing the required row dimension of the RNG array. An address-driven XOR recombination scheme (PGU-XOR) is further introduced to mitigate the spatial correlations caused by direct RNG reuse (PGU-Reuse). The results show that (1) PGU-XOR matches software RNGs in accuracy on Spikingformer/CIFAR-10 (76.41% vs. 76.53%) and perplexity (PPL) on SpikeGPT/WikiText-2 (54.20 vs. 53.23), whereas PGU-Reuse degrades accuracy by 9.56 percentage points and increases PPL by 11.8; (2) implemented in a TSMC 16-nm CMOS technology, PGU-XOR incurs 40.3%-46.0% area and 15.2%-48.9% energy overhead per matrix-vector multiplication relative to PGU-Reuse, yet its faster convergence reduces the total perturbation energy to 0.51x that of PGU-Reuse at iso-accuracy; (3) IPZO reduces the perturbation energy to 0.46x-0.83x that of conventional explicit weight perturbation for a batch size of B=64 and T=4 time steps, with the advantage growing as BT decreases.
Comment: Implicit, spike-triggered perturbations eliminate perturbation-induced weight rewrites during zeroth-order training.
Topic Match: The core contribution is a hardware mechanism reducing training energy and perturbation-generation overhead; its demonstrated scope is specialized spiking models.
Relevance: 7 Novelty: 7
22. Two Stage Wireless Federated LoRA Fine-Tuning with Sparsified Orthogonal Updates
ArXiv ID: 2505.00333
Primary Topic: Efficiency, Compression, and Large-Scale Training
Also Matches: Large-Scale Training Systems and Efficiency
Authors: Bumjun Kim, Wan Choi
Abstract: Federated fine-tuning with low-rank adaptation (LoRA) communicates only two low-rank matrices instead of the full model, but existing methods typically fix the LoRA rank in advance as a manually tuned hyperparameter. In wireless networks, however, the rank determines both adaptation capacity and uplink payload, while the deliverable payload varies with the fading channel. To address this coupling, we formulate wireless federated LoRA fine-tuning as a two-timescale design that separates the \emph{offline-optimized rank}, i.e., the rank of the shared LoRA structure selected before training from statistical channel information, from the per-iteration sparsification and bandwidth decisions adapted to instantaneous CSI. For per-iteration adaptation, we propose sparsified orthogonal fine-tuning (\textbf{SOFT}), which promotes near-orthogonality among rank components so that the product of the corresponding column and row norms approximates each component's singular value. This SVD-free score guides component-wise payload allocation and within-component entry selection without forming the full matrix product. We further derive a convergence bound linking the two timescales and develop a two-stage federated algorithm (\textbf{TSFA}) that selects the offline-optimized rank offline and jointly optimizes sparsification and bandwidth online via Lyapunov optimization.
Comment: SVD-free importance scores for near-orthogonal LoRA components enable communication-aware sparse updates.
Topic Match: LoRA-update sparsification is the core efficiency mechanism, coupled to a federated training algorithm that jointly manages rank, payload, and bandwidth.
Relevance: 7 Novelty: 6
23. DiD It in 87 Minutes: A Label-Free Softmax-to-Linear Adaptation of Vision Transformers for Object Detection
ArXiv ID: 2608.22368
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Huaiyuan Qin, Gabriel James Goenawan, Zihang Lin, Muli Yang, Hongyuan Zhu
Abstract: While linear attention is a compelling mechanism for high-resolution object detection due to its reduced cost for global token mixing, converting the Softmax-attention ViT backbone of a trained detector into a linear-attention one is not a trivial drop-in replacement. Directly swapping the attention operator leads to severe performance degradation, and generic label-free distillation, though effective for classification, often fails on detection tasks. We argue that the central challenge is \textit{detector-interface preservation}: the converted backbone must reproduce the exact feature tensors expected by the fixed downstream detector, rather than merely imitating internal Softmax hidden states. To address this, we introduce Detector-Interface Distillation (DiD), a label-free conversion method that exclusively trains the linear-attention backbone by aligning detector-facing interface tensors with those of a frozen Softmax teacher. On DOTA-v1.5, DiD substantially outperforms established baselines and matches supervised, fully trained linear models. Adaptation completes in roughly 87 minutes on 4 GPUs, and the linearized backbone cuts inference latency by ~62% and peak memory by ~49%. We hope our findings offer the community a simple, label-free route to reusing trained Softmax detectors as efficient linear ones, and encourage interface-aware objectives in future architecture-conversion work.
Comment: Interface-preserving distillation converts pretrained softmax attention into cheaper linear attention.
Topic Match: Efficient architecture conversion is the core contribution, with substantial reported latency and memory reductions; the objective and validation remain detection-specific.
Relevance: 7 Novelty: 6
24. LLM4LLM: Bridging Kernel Benchmarks and Real Deployment via Closed-Loop Agentic Optimization
ArXiv ID: 2608.21836
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Hui Zeng, Pengfei Yang, Yanxin Chen, Fusong Ju, Xinran Wei
Abstract: Large language models have become increasingly capable agents for low-level code and kernel optimization, but isolated kernel benchmarks provide only a proxy for the deployment behavior that matters in language-model inference. We identify a benchmark-to-deployment gap: candidate kernels that appear correct and fast in standalone harnesses can exhibit different performance, safety, or phase behavior after integration into a real inference workload. We introduce LLM4LLM, a deployment-aware closed-loop optimization framework that starts from a target inference script, extracts phase-aware optimization tasks, searches with an experience-guided episodic agent, and accepts patches through in-model validation. Across ten language-model inference workloads on A100 and H100 GPUs, LLM4LLM improves end-to-end latency for every evaluated model, achieving 3.91$\times$/6.98$\times$ geometric-mean speedups on A100/H100; as supporting kernel-level evidence, it also attains up to 2.745$\times$ GeoMean speedup on KernelBench Level 2.
Comment: Optimizes LLM kernels through phase-aware task extraction and end-to-end validation of candidate patches inside the model.
Topic Match: The deployment-aware kernel-optimization loop directly targets large-model execution cost, although its demonstrated benefits concern inference.
Relevance: 7 Novelty: 6
25. HIERA: Workload-Aware Planning Across Implementation Spaces for GPU Kernel Optimization
ArXiv ID: 2608.21157
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Jinghao Wang, Qiqi Gu, Chenpeng Wu, Jianguo Yao, Haibing Guan, Xijun Li
Abstract: High-performance GPU kernels underpin modern deep learning and scientific computing. As workloads become increasingly diverse and GPU hardware evolves rapidly, developing efficient methods for automated GPU kernel generation and optimization has become increasingly important. Existing LLM-based methods typically optimize within a fixed implementation space, limiting either optimization flexibility or search efficiency. We propose \textsc{HIERA}, a hierarchical search-space planning framework for GPU kernel optimization. \textsc{HIERA} constructs contract-augmented task specifications, selects an appropriate implementation space across PyTorch operators, CUDA libraries, and custom CUDA kernels, and uses profiling feedback and expert knowledge to guide structured iterative refinement. Experiments on KernelBench across multiple various workload levels and base LLMs show that \textsc{HIERA} delivers stronger overall implementation validity, sample efficiency, and optimization performance than existing training-free methods, while remaining competitive with the training-based CUDA-L1 without additional model training. A case study on a specialized stencil operator from scientific computing further achieves a (1.53\times) speedup over cuDNN, demonstrating the potentiality of the general framework beyond standard machine-learning workloads.
Comment: Hierarchical search across operators, CUDA libraries, and custom kernels improves GPU implementation selection.
Topic Match: The core method improves GPU computation efficiency through implementation-space planning; end-to-end large-model training savings are not established in the abstract.
Relevance: 7 Novelty: 6
26. Target-Aware Calibration Data Selection for Preserving Uncertainty in Quantized Language Models
ArXiv ID: 2608.21019
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Zhen Yang, Sizai Hou, Kaiwen Zheng, Yaofang Liu, Liang He, Yixuan Chen, Kangning Cui
Abstract: Quantization is widely used to deploy large language models, but its effect on uncertainty behavior, such as confidence, margins, and abstention, is rarely treated as a primary objective. We frame calibration-data selection for quantization as a target-dependent uncertainty-preservation problem. Different deployments emphasize different regions of the input distribution, yet prior work mainly optimizes accuracy-oriented compression metrics or adjusts scores after quantization. We formalize this goal with distributional and boundary preservation risks, and provide a simple mixture-mismatch argument explaining why no single calibration recipe should be expected to fit all targets. We introduce Doubt-Preserving Quantization (DPQ), a lightweight pre-quantization recipe family that uses full-precision predictions to construct target-aligned calibration mixtures of high-doubt examples and generic anchors. Across 8 language models, 9 NLP benchmarks, and 22 comparison methods, the leading fixed recipe changes with the preservation target: DPQ-r75 leads on SQuAD2 answerability-boundary preservation, while milder or single-signal variants, including DPQ-r50, confidence-only, and entropy-only, better preserve broad multiple-choice QA behavior. These results show that calibration data should be selected for the specific full-precision score behavior a deployment needs to preserve, rather than treated as a fixed quantization detail.
Comment: Target-aware calibration mixtures preserve full-precision uncertainty behavior under quantization.
Topic Match: The contribution refines quantization calibration-data selection to preserve uncertainty at existing compression settings.
Relevance: 7 Novelty: 6
Paper Selection Prompt
System Prompt
You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.
User Prompt
Instructions
Respond in JSONL. Output exactly one JSON object per paper, one per line:
{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}
Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only:
daily_hot,new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return[]. -daily_hotmeans the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. -new_frontiermeans the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.Scoring Criteria
Relevance and Novelty are independent axes. Score both from 1 to 10.
Relevance Scoring
- 9-10: directly centered on the target foundational topics; highest when the core contribution is clearly within them.
- 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
- 5-6: touches the target topics, but the main contribution is elsewhere.
- 3-4: largely outside the target topics, often application-focused or domain-specific.
- 1-2: unrelated.
Important: Broad frontier relevance, major launch status, or daily buzz is not enough for a high Relevance score here. Those cases belong in the hotspot digest unless the paper strongly matches the specialized paper topics.
Novelty Scoring
- 9-10: new paradigm, theory, or major methodological breakthrough.
- 7-8: substantial methodological advance or strong new insight.
- 5-6: meaningful but incremental extension or refinement.
- 3-4: minor, narrow, or mostly engineering or domain-specific improvement.
- 1-2: little originality; mainly standard application of existing methods.
Topic Registry
Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - moe_training: MoE Training - Mixture-of-Experts training: routing, load balancing, expert granularity and shared experts, training stability, MoE parallelism and communication, and MoE kernels. - training_systems: Large-Scale Training Systems and Efficiency - Large-scale training systems and efficiency: distributed training algorithms, optimizers, and communication. - architecture_training: Architecture and Training Dynamics - Core architectural or computational mechanisms, dynamic computation, and training-stability dynamics. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, memory or cache efficiency, and large-scale training systems that materially change cost or behavior.
Papers
[PAPER LIST HERE]
Relevant Topics
This is a training-side feed. The centre of gravity is Mixture-of-Experts training and the large-scale training systems it sits on; architecture and efficiency work is kept when it bears on how a large model is trained or what it costs to train.
Keep a paper when its CORE CONTRIBUTION falls in one of the four topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.
MoE Training (primary) - Keep: expert routing and gating mechanisms; load balancing, with or without auxiliary losses; expert granularity, fine-grained and shared experts; MoE training stability (loss spikes, router collapse, z-loss and related fixes); expert and hybrid parallelism; all-to-all dispatch/combine and the communication schedules MoE needs; MoE kernels and grouped-GEMM style implementations; expert capacity, token dropping, and upcycling dense checkpoints into MoE. - Filter: papers that merely train on top of a MoE model without a new routing, balancing, stability, parallelism, or kernel contribution; "mixture of experts" in the classical ensemble or recommender sense.
Large-Scale Training Systems - Keep: distributed training algorithms (data, tensor, pipeline, sequence, and expert parallelism, FSDP and ZeRO-style sharding); optimizers and preconditioners for large-scale pretraining; communication algorithms, overlap, and collective design; throughput and MFU work that materially changes what a training run costs or whether it converges; scaling-law work that informs how such runs are configured. - Filter: routine infrastructure tuning with no new algorithmic or systems idea.
Architecture and Training Dynamics - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, dynamic or modular computation); optimisation and training-dynamics analysis that explains why large models train the way they do. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight.
Efficiency and Compression - Keep: quantization, sparsity, pruning, low-rank adaptation, KV-cache and memory-efficient designs, and other work that materially changes the cost of running or training a large model, especially when the mechanism is new rather than a tuned variant. - Filter: straightforward application of standard efficiency methods, and deployment or serving engineering with no new idea.
Leave these to the hotspot digest unless the core contribution clearly falls in one of the four topics above: - major model or product releases - agent, tooling, and RAG launches - benchmark, leaderboard, or evaluation-only papers - representation-learning theory, interpretability, and feature-formation analysis - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning) - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains