Previous Day 2026-05-22
Monthly Overview 2026-05
Next Day 2026-05-26

This is a remedial run for missed papers from 05/22/2026 to 05/24/2026.

Results generated on 09/11/2026.

Personalized Daily ArXiv Papers 2026-05-25

Model Metric Usage Papers
Prompt Completion Total Total arXiv Scanned Relevant
local agent CLI Tokens not reported not reported not reported 883 883 33
Cost not reported not reported not reported

Token counts are not reported for this run. 9 of 14 model calls succeeded, 6,099s of model wall clock.

Topic Coverage:

TopicPapers
Large-Scale Training Systems and Efficiency5
Architecture and Training Dynamics13
Efficiency, Compression, and Large-Scale Training15

Table of contents by topic:

Large-Scale Training Systems and Efficiency (5)

  1. LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws Authors: Xu Ouyang, Deyi Liu, Yuhang Cai, Jing Liu, Yuan Yang, Chen Zheng, Thomas Hartvigsen, Yiyuan Ma

  2. Move on Muon : A Hamiltonian probability gradient flow perspective of Muon optimizer Authors: Aratrika Mustafi, Soumya Mukherjee, Bharath K. Sriperumbudur

  3. Strong Teacher Not Needed? On Distillation in LLM Pretraining Authors: Taiming Lu, Zhuang Liu

  4. Scaling Laws for Dynamic Mini-Batch SGD in Sketched Linear Regression Authors: Ziyan Chen, Zhongzhu Zhou, Ding-Xuan Zhou

  5. Efficient DP-SGD for LLMs with Randomized Clipping Authors: Enayat Ullah, Sai Aparna Aketi, Devansh Gupta, Huanyu Zhang, Meisam Razaviyayn

Architecture and Training Dynamics (13)

  1. Momentum Streams for Optimizer-Inspired Transformers Authors: Jingchu Gai, Nai-Chieh Huang, Jiayun Wu

  2. Non-normal spectral signatures of instability in neural network training dynamics Authors: Souvik Ghosh

  3. Neural Bayesian Sequential Routing Authors: Yongchao Huang

  4. Muon in Vision Transformers: Optimizer-Recipe Interactions and Gradient Spectra Authors: Ben S. Southworth, Shuai Jiang, Daniel McBride, Eric C. Cyr, Stephen Thomas

  5. Multi-Gate Residuals Authors: Zhizhan Zheng, Feiyun Zhang, Shuchun Liu, Tian Xia, Xi Liu, Dasheng Hu, Hongquan Zhou

  6. Improved Belief-Attention in Vision Task Authors: Guoqiang Zhang

  7. Spectral Asymptotics of Neural Network Loss Landscapes: An Exact Decomposition of the Curvature Exponent Authors: Anherutowa Calvo

  8. S$^3$GNN: Efficient Global Mixing and Local Message Passing for Long-Range Graph Learning Authors: Dai Shi, Luke Thompson, Linhan Luo, Lequan Lin, Andi Han, Junbin Gao, José Miguel Hernández Lobato

  9. Training-Free Looped Transformers Authors: Lizhang Chen, Jonathan Li, Chen Liang, Ni Lao, Qiang Liu

  10. Balancing Multimodal Learning through Label Space Reshaping Authors: Xiaoyu Ma, Weijie Zhang, Yuanhao Gao, Han Miao, Yongjian Deng, Hao Chen

  11. Feature Learning in Wide Neural Networks under $μ$P: Identifiability and Sparse-Dictionary Decomposition of the Mean-Field Limit Authors: Akmal Xodarev

  12. Grow-Prune-Freeze Networks: Adaptive & Continual Learning Technique for Olfactory Navigation Authors: Kordel K. France, Ovidiu Daescu

  13. ChainzRule: Sample-Efficient, Robust Deep Learning Across Tabular, NLP, and Vision Tasks Authors: Rowan Martnishn

Efficiency, Compression, and Large-Scale Training (15)

  1. Signs Beat Floats: Low-Rank Double-Binary Adaptation for On-Device Fine-Tuning Authors: Yoshihiko Fujisawa, Yuma Ichikawa, Yudai Fujimoto, Akira Sakai, Katsuki Fujisawa

  2. Adaptive Mass-Segmented KV Compression for Long-Context Reasoning Authors: Junzhe Yang, Xiaoyu Shen

  3. Pruning Deep Neural Networks via the Marchenko--Pastur Distribution Authors: Leonid Berlyand, Theo Bourdais, Houman Owhadi, Yitzchak Shmalo

  4. Approaching I/O-optimality for Approximate Attention Authors: Pál András Papp, Aleksandros Sobczyk, Anastasios Zouzias

  5. MX-SAFE: Versatile Inference- and Training-Proof Microscaling Format with On-the-Fly Exponent and Mantissa Bit Allocation Authors: Dahoon Park, Jahyun Koo, Sangwoo Hwang, Jaeha Kung

  6. Inference Time Context Sparsity: Illusion or Opportunity? Authors: Sahil Joshi, Prithvi Dixit, Agniva Chowdhury, Anshumali Shrivastava, Joseph E. Gonzalez, Ion Stoica, Kumar Krishna Agrawal, Aditya Desai

  7. Scale When Needed: Adaptive Neuron-level Mixed Precision Quantization Aware Training Authors: Ayush K. Varshney, Konstantinos Vandikas, Šarūnas Girdzijauskas, Adam Orucu, Aneta Vulgarakis Feljan

  8. Polynomial Context-Truncation Sensitivity in Autoregressive Language Models: Sequential Wyner-Ziv Bounds for KV Cache Compression Authors: Munsik Kim

  9. Beyond the Frontier: Stochastic Backtracking for Efficient Test-Time Scaling Authors: Dao Tran, Duc Anh Le, Ngoc Luu, Quan Pham, Tung Pham, Hung Bui

  10. Context Features Are Cheap: Rank-Aware Decomposition for Efficient Feature Interaction in Recommender Systems Authors: Yevgeny Tkach

  11. One-Forcing: Towards Stable One-Step Autoregressive Video Generation Authors: Jiaqi Feng, Justin Cui, Yuanhao Ban, Cho-Jui Hsieh

  12. ReLoRA: Knowledge-Reusing Adaptation for Fast Rollout of Evolving LLM Services Authors: Yang Xu, Zihuai Xu, Hongli Xu, Yunming Liao, Zhiwei Yao, Xitong Fu

  13. PALoRA: Projection-Adaptive LoRA for Preserving Reasoning in Large Language Models Authors: Mustafa Hayri Bilgin, Mariam Barry, Albert Bifet, Azzedine Idir Ait Said, Soumya Banerjee

  14. EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture Authors: Bowen Duan, Cong Guo, Chiyue Wei, Haoxuan Shan, Yuzhe Fu, Xinhua Chen, Yifan Xu, Ziyue Zhang, Changchun Zhou, Hai Li, Yiran Chen

  15. SonicSampler: Unified Tile-Aware Kernels for LLM Sampling and Speculative Verification Authors: Pragaash Ponnusamy, Shivam Sahni, Jue Wang, Tri Dao


Large-Scale Training Systems and Efficiency (5)

1. LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws

ArXiv ID: 2605.23901

Primary Topic: Large-Scale Training Systems and Efficiency

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Xu Ouyang, Deyi Liu, Yuhang Cai, Jing Liu, Yuan Yang, Chen Zheng, Thomas Hartvigsen, Yiyuan Ma

Abstract: Existing scaling laws for Large Language Models (LLMs), predominantly monotonic power laws, fail to explain emerging non-monotonic phenomena such as catastrophic overtraining and quantization-induced degradation, where performance deteriorates despite increased compute. We propose the Shannon Scaling Law, a unified theoretical framework that models LLM training as information transmission over a noisy channel, grounded in the Shannon-Hartley theorem. By mapping model parameters to channel bandwidth and training tokens to signal power, our formulation explicitly captures the interaction between learning signal and intrinsic noise. This perspective reveals a fundamental Shannon capacity for LLMs: scaling model size or data without preserving a sufficient signal-to-noise ratio (SNR) inevitably amplifies noise, inducing a transition from monotonic improvement to U-shaped performance degradation. We validate our theory through experiments on Pythia and OLMo2 under perturbations, including Gaussian noise, quantization and supervised fine-tuning on math, QA and code tasks. The Shannon Scaling Law consistently outperforms classical scaling laws and recent perturbation-aware laws, achieving strong $R^2$ scores and accurately capturing loss basins missed by prior approaches. It also extrapolates: fitted on $\leq$6.9B Pythia models with $\leq$180B tokens, it predicts the unseen 12B model up to 307B tokens at pooled $R^2{=}0.847$, while monotonic baselines collapse.

Comment: Models pretraining as transmission over a noisy channel, predicting a U-shaped capacity limit that explains catastrophic overtraining and quantization degradation where monotonic power laws fail.

Topic Match: Scaling-law formulation that directly informs how parameter and token budgets should be configured.

Relevance: 8 Novelty: 7


2. Move on Muon : A Hamiltonian probability gradient flow perspective of Muon optimizer

ArXiv ID: 2605.23871

Primary Topic: Large-Scale Training Systems and Efficiency

Also Matches: Architecture and Training Dynamics

Authors: Aratrika Mustafi, Soumya Mukherjee, Bharath K. Sriperumbudur

Abstract: We develop a gradient flow on the space of probability measures defined on matrix-valued parameters induced by regularized Muon, an analytically smoothed version of the idealized Muon optimizer. The key observation is that the regularized orthogonalization map is the gradient of a smooth Fenchel-dual smoothing of the nuclear norm. This identifies the (regularized) Muon update as a mirror/prox step in the update variable, with momentum acting as the dual coordinate. We use this structure to lift Muon from a single matrix parameter to finite-particle probability objectives of the form $J(ρ)=R\left(\int F d ρ\right)$, a setting motivated by mean-field descriptions of neural-network training, and derive the inertial continuous-time limit. Using this structure, we derive the finite-particle continuous-time limit under the inertial scaling of step size and momentum, and then pass to a phase-space mean-field equation over probability laws on parameter-momentum pairs. The resulting flow can be shown to be a damped Hamiltonian probability dynamics whose kinetic energy is induced by the regularized Muon mirror potential. We prove an exact Hamiltonian dissipation identity, showing that the Hamiltonian energy decreases monotonically. While the target objective itself need not be monotone along the inertial Muon dynamics, under additional gradient-dominance, bounded-momentum, and curvature/alignment assumptions, we obtain continuous and discrete-time exponential convergence rates for the objective gap. We also study the well-posedness of the mean-field limit equation and establish propagation of chaos guarantees for the interacting particle system. Finally, we extend the formulation to Hilbert-valued feature maps on product matrix spaces, yielding a blockwise Muon probability flow applicable to smooth transformer mixture-of-experts models.

Comment: Identifies the regularized orthogonalization map as the gradient of a smoothed nuclear norm, recasting Muon as a mirror step and proving an exact Hamiltonian dissipation identity with convergence rates.

Topic Match: Theory for an optimizer now used in large-scale pretraining, which is the optimizer strand of the training-systems topic.

Relevance: 8 Novelty: 7


3. Strong Teacher Not Needed? On Distillation in LLM Pretraining

ArXiv ID: 2605.23857

Primary Topic: Large-Scale Training Systems and Efficiency

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Taiming Lu, Zhuang Liu

Abstract: Knowledge distillation generally assumes a strong-to-weak relationship where stronger teachers yield better students. In this work, we examine this assumption about distillation in large language model pretraining. By varying architecture sizes and training token budgets, we create strong-to-weak, same-level, and weak-to-strong teacher-student relationships, and study distillation's effectiveness under each. We find that the teacher need not be strong: with proper mixing of the language modeling and knowledge distillation losses, even small and undertrained teachers improve larger students. At the same time, a stronger teacher is not always better: pushing the teacher further, through more parameters or more training tokens, can saturate or even reverse the distillation gains. We further observe that distillation improves generalization (out-of-distribution and downstream performance) more readily than in-domain fitting. Together, these results challenge the common belief that distillation pretraining always requires a strong teacher.

Comment: Sweeps teacher size and token budget in distillation pretraining to show weak or undertrained teachers still help larger students, and that stronger teachers can saturate or reverse gains.

Topic Match: A pretraining-recipe finding that changes how distillation budgets should be allocated in large runs.

Relevance: 7 Novelty: 7


4. Scaling Laws for Dynamic Mini-Batch SGD in Sketched Linear Regression

ArXiv ID: 2605.24316

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Ziyan Chen, Zhongzhu Zhou, Ding-Xuan Zhou

Abstract: Mini-batching is central to large-scale optimization, yet its role in statistical scaling laws remains limited. We study one-pass and multi-pass batch SGD for sketched linear regression under power-law spectral and source conditions. Our analysis reveals a two-horizon phenomenon induced by warmup--stable--decay schedules: deterministic learning is governed by the full optimization trajectory, while stochastic error retains only a shorter terminal memory. For dynamic batch schedules, the individual batch sizes enter through influence-weighted summaries that measure how strongly each update affects the final risk. Consequently, batching leaves the approximation and optimization-bias laws unchanged at a fixed update horizon, but controls the one-pass variance and the multi-pass fluctuation around full-batch gradient descent. We obtain matching one-pass variance bounds and nearly matching multi-pass fluctuation bounds, recover static-batch and full-batch behavior as special cases, and derive an oracle square-root rule for allocating a fixed iteration budget. These results identify WSD horizon separation and final-risk influence as the mechanisms governing dynamic mini-batch scaling.

Comment: Shows warmup-stable-decay schedules induce a two-horizon separation where stochastic error keeps only terminal memory, and derives a square-root rule for allocating batch sizes under a fixed update budget.

Topic Match: Batch-size and schedule scaling laws speak directly to how large pretraining runs are configured, albeit in a sketched linear regression model.

Relevance: 7 Novelty: 6


5. Efficient DP-SGD for LLMs with Randomized Clipping

ArXiv ID: 2605.24879

Primary Topic: Large-Scale Training Systems and Efficiency

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Enayat Ullah, Sai Aparna Aketi, Devansh Gupta, Huanyu Zhang, Meisam Razaviyayn

Abstract: Large language models (LLMs) are trained on vast datasets that may contain sensitive information. Differential privacy (DP), the de facto standard for formal privacy guarantees, provides a principled framework for training LLMs with provable privacy protection. However, state-of-the-art DP training implementations rely on fast gradient clipping techniques with memory overhead $O(B \min{T^2, d^2})$, where $B$ is the batch size, $T$ is the sequence length, and $d$ is the model width. This becomes prohibitive as both model size and context length grow. We propose DP-SGD-RC, a novel variant of DP-SGD with randomized clipping that reduces memory and compute complexity. DP-SGD-RC leverages stochastic trace estimation methods, specifically Hutchinson's estimator[Hutchinson, 1989] and its improved variant, Hutch++[Meyer et al., 2021], to reduce the memory footprint of per-sample gradient norm estimation. We provide a tight privacy analysis showing that DP-SGD-RC achieves noise multipliers competitive with deterministic clipping. Experiments fine-tuning Llama~3.2-1B on long-context benchmarks spanning classification, question answering, and summarization tasks demonstrate that DP-SGD-RC matches baseline utility while significantly reducing memory and compute requirements.

Comment: Replaces exact per-sample gradient-norm computation with Hutchinson/Hutch++ trace estimation, removing the O(B min{T^2,d^2}) clipping memory overhead that makes DP-SGD prohibitive at long context.

Topic Match: It is an optimizer variant whose contribution is the memory and compute complexity of the training step, with a matching privacy analysis.

Relevance: 6 Novelty: 6


Architecture and Training Dynamics (13)

1. Momentum Streams for Optimizer-Inspired Transformers

ArXiv ID: 2605.24425

Primary Topic: Architecture and Training Dynamics

Also Matches: Large-Scale Training Systems and Efficiency

Authors: Jingchu Gai, Nai-Chieh Huang, Jiayun Wu

Abstract: The residual update of a pre-norm Transformer layer admits an interpretation as one step of a first-order optimizer acting on a surrogate token energy, wherein the attention and MLP sublayers function as gradient oracles. Based on this observation, we build a family of optimizer-inspired Transformers (triple-momentum, Adam/AdamW, Muon, SOAP) and compare them under matched compute. In our main pretraining experiment, the triple-momentum TMMFormer achieves the lowest validation loss, outperforming the vanilla Transformer and prior architectural variants. A controlled ablation and supporting theory show that momentum, not preconditioning, is the main source of the gain. We further show that TMMFormer and other momentum-based designs reach flatter minima than the vanilla Transformer, which leads to less forgetting and better generalization.

Comment: Reinterprets the pre-norm residual update as a first-order optimizer step and builds transformer blocks from triple-momentum, Muon, and SOAP, with ablations isolating momentum rather than preconditioning as the source of gains.

Topic Match: A residual-design family derived from optimizer structure, compared under matched pretraining compute.

Relevance: 8 Novelty: 7


2. Non-normal spectral signatures of instability in neural network training dynamics

ArXiv ID: 2605.23476

Primary Topic: Architecture and Training Dynamics

Also Matches: Large-Scale Training Systems and Efficiency

Authors: Souvik Ghosh

Abstract: Training instabilities in deep networks - loss spikes, oscillatory convergence, and gradient pathologies - are empirically prevalent but lack a rigorous operator-theoretic explanation. We show that the linearized update operators for practically used optimizers are generically non-normal: for Adam, non-normality is controlled by the commutator [H, M] between the Hessian and the diagonal adaptive preconditioner, while for SGD with momentum it arises from the augmented state-space structure of the update map. Applying non-normal stability theory to these operators, we derive a conservative pseudospectral precursor bound in which κ(V) serves as an early-warning indicator of transient amplification even when the spectral radius remains below one, and we establish that exceptional points of the update operator appear as the κ(V) -> \infty limiting case of this framework. Numerical experiments on two-layer networks confirm that the spectral radius ρ(J) provides no separation between stable and unstable training phases while κ(V) separates them by approximately one order of magnitude, complementing the classical sharpness criterion with a continuous severity measure of non-normal amplification. These results establish non-Hermitian operator theory as a useful and underexplored framework for neural network optimization stability, offering a diagnostic language and proof-of-concept benchmark for understanding adaptive optimization stability.

Comment: Shows Adam's linearized update operator is non-normal through the Hessian-preconditioner commutator, and that eigenvector conditioning separates stable from unstable phases where spectral radius does not.

Topic Match: An operator-theoretic account of loss spikes and gradient pathologies, i.e. the training-stability strand, with an early-warning indicator.

Relevance: 8 Novelty: 7


3. Neural Bayesian Sequential Routing

ArXiv ID: 2605.26147

Primary Topic: Architecture and Training Dynamics

Authors: Yongchao Huang

Abstract: Human decision-making is sequential and uncertainty-aware, yet standard neural networks often rely on static, dense forward computation with limited visibility into evidence acquisition, uncertainty evolution, or when computation should stop. We introduce \textbf{Neural Bayesian Sequential Routing (NBSR)}, a framework that models neural inference as active evidence accumulation over a hierarchical Directed Acyclic Graph (DAG). Within a Dirichlet--Categorical conjugate framework, neural experts query a persistent global knowledge oracle to extract positive evidence vectors, which act as pseudo-counts and update a Dirichlet belief state by exact conjugate addition. Coupled with a Gumbel-Softmax Straight-Through estimator, this update enables hard, path-dependent routing while preserving surrogate gradients for end-to-end training. The resulting Dirichlet precision and entropy provide mechanisms for uncertainty quantification, entropy-based early exiting, OOD abstention, and cost-aware evidence acquisition. We prove that, under strictly positive evidence extraction, total Dirichlet precision increases monotonically along any valid trajectory and marginal predictive variance is bounded, formalizing sequential ``hypothesis sharpening''; under idealized capacity and optimization assumptions, the terminal Dirichlet expectation recovers the Bayes-optimal conditional distribution. Empirical evaluations across visual categorization, structured medical diagnosis, language modeling, partially observable control, and cost-aware Bayesian experimental design show that NBSR achieves competitive predictive performance while providing transparent routing traces, path-dependent evidence attribution, uncertainty-aware decision control, and resource-rational inference. Overall, NBSR offers a mathematically grounded framework for interpretable, modular, and resource-rational agentic AI.

Comment: Surrogate-gradient training enables hard, path-dependent routing through a modular neural graph.

Topic Match: The central contribution is a trainable conditional-computation architecture with belief-dependent routing and adaptive stopping.

Relevance: 8 Novelty: 7


4. Muon in Vision Transformers: Optimizer-Recipe Interactions and Gradient Spectra

ArXiv ID: 2605.24770

Primary Topic: Architecture and Training Dynamics

Authors: Ben S. Southworth, Shuai Jiang, Daniel McBride, Eric C. Cyr, Stephen Thomas

Abstract: Muon is a recently developed matrix-aware optimizer that has shown strong results in transformer training, but its behavior in vision transformers (ViTs) is not yet well understood. We study Muon for ViT training, largely on ImageNet-100 and Pl@ntNet-300K, comparing against AdamW under standard vision recipes involving mixup, cutmix, smoothing, and random augmentation and erasing. Muon consistently outperforms AdamW, with especially large gains on long-tailed Pl@ntNet macro top-1. These gains are also recipe-dependent, where Muon benefits much more than AdamW from advanced and significant data augmentation techniques. To understand this interaction, we analyze the singular-value structure of matrix gradients throughout the ViT. Within Muon training runs, removing heavy data augmentation induces a late-training spectral concentration and mode collapse in gradient matrices, primarily in deep MLP-down blocks. Under a fixed "full" augmentation recipe, the clearest Muon-AdamW contrast appears instead in QKV gradients, where AdamW gradient energy remains concentrated in a much narrower basis while Muon spreads energy across substantially more singular modes. Muon in ViTs is therefore best understood as an optimizer-recipe interaction. Under a fixed recipe, Muon differs from AdamW most clearly in attention projections, where its gradients consist of a broader spectral basis. Within Muon, a full training recipe is important for preventing late spectral concentration and mode collapse in deep feedforward blocks. We further demonstrate efficacy in training ViTs on image segmentation and masked autoencoder models, where Muon outperforms AdamW in all settings considered.

Comment: Relates Muon's augmentation sensitivity to gradient spectral concentration and late-training mode collapse.

Topic Match: Gradient-spectrum analysis explains interactions between optimizer choice and training recipes in transformers.

Relevance: 8 Novelty: 7


5. Multi-Gate Residuals

ArXiv ID: 2605.23259

Primary Topic: Architecture and Training Dynamics

Also Matches: Large-Scale Training Systems and Efficiency

Authors: Zhizhan Zheng, Feiyun Zhang, Shuchun Liu, Tian Xia, Xi Liu, Dasheng Hu, Hongquan Zhou

Abstract: While Attention Residuals has shown some effectiveness in addressing the widespread issue of unbounded activation growth across deep residual layers, it inevitably incurs significant communication overhead. To circumvent this bottleneck, we propose Multi-Gate Residuals (MGR), which stabilizes activation scales without additional communication burden. It utilizes a straightforward scoring and gating mechanism to maintain multi-stream context, coupled with Attention Pooling to extract hidden states from the stream states. Empirical experiments demonstrate that MGR is practical for large-scale training and deployment, offering tangible performance improvements over existing architectures.

Comment: Controls unbounded activation growth across deep residual stacks with a scoring-and-gating multi-stream residual that avoids the communication cost of attention residuals.

Topic Match: Residual and normalization design aimed at activation-scale stability under large-scale training, with communication overhead as the explicit constraint.

Relevance: 8 Novelty: 6


6. Improved Belief-Attention in Vision Task

ArXiv ID: 2606.00077

Primary Topic: Architecture and Training Dynamics

Authors: Guoqiang Zhang

Abstract: Recently, Belief-Attention \cite{Guoqiang25BeliefAttention} has been proposed by first performing an orthogonal projection of the softmax-based weighted summation of $V$ vectors with respect to the original $V$ vectors and then taking the perpendicular component as the residual signal in Transformer for performance improvement. In this paper, we first conduct an ablation study showing the projected component also carries information about the token correlation, which should not be ignored. We then propose to extend Belief-Attention by making use of both the perpendicular and projected components. In particular, the projected component goes through certain activation function and then a linear mapping before merging with the considered token. Conceptually speaking, the neural block for the projected component can be viewed as a two-layer feedforward network (FFN) within the new attention block. It is also noted that standard attention captures the token correlation via the inner-product matrix $QK^T$. We propose to introduce an additional inner-product matrix $ZZ^T$ to $QK^T$ to capture richer token correlation. We refer to the new module as Belief2-Attention. It can be easily shown that Belief2-Attention is more expressive than standard Attention. We then verify the effectiveness of Belief2-Attention for vision tasks of image classification and segmentation.

Comment: Extends attention by separately processing projected and perpendicular value components before residual integration.

Topic Match: The core contribution changes attention computation and residual signals, with validation currently focused on vision transformers.

Relevance: 8 Novelty: 6


7. Spectral Asymptotics of Neural Network Loss Landscapes: An Exact Decomposition of the Curvature Exponent

ArXiv ID: 2606.02596

Primary Topic: Architecture and Training Dynamics

Also Matches: Large-Scale Training Systems and Efficiency

Authors: Anherutowa Calvo

Abstract: The curvature exponent $α$ in $h_k \propto σ_k^α$ -- governing how Hessian eigenvalues scale with gradient singular values -- varies systematically across layer types ($α\approx 2$ for convolutions, $\approx 1$ for transformer attention, $< 1$ for MLP up-projections). Why? We prove the Spectral Alignment Decomposition: $α= 2 + d\logΦ_k / d\logσ_k$, where $Φ_k$ measures alignment between Kronecker factor eigenbases and gradient singular directions. This reduces "why does $α$ vary?" to a geometric question we answer for LayerNorm, residual connections, and softmax heads. The decomposition implies a spectral transfer identity $s = αγ$ linking curvature exponent, effective gradient rank-decay $γ$, and Hessian decay exponent $s$. The identity is algebraic; its empirical content is that $α$ and $γ$, fit on independent data (HVPs vs. SVD), recover $s$ to ~2% median error across 93 layers, five architectures, and three datasets -- with no free parameters. A zeta-function bound on participation ratio shows curvature concentrates onto effectively one direction per layer. As a proof of concept, we derive the architecture-adaptive preconditioner $T(σ;α)$ and show that Spectral Newton -- implementing $T$ in the gradient singular basis -- outperforms AdamW on vision benchmarks where $α\approx 2$.

Comment: Proves the curvature exponent decomposes into an alignment term explaining why it differs across convolution, attention, and MLP layers, and derives an architecture-adaptive preconditioner from it.

Topic Match: Loss-landscape curvature per layer type is training-dynamics analysis, and it yields a concrete preconditioner.

Relevance: 7 Novelty: 7


8. S$^3$GNN: Efficient Global Mixing and Local Message Passing for Long-Range Graph Learning

ArXiv ID: 2605.23467

Primary Topic: Architecture and Training Dynamics

Authors: Dai Shi, Luke Thompson, Linhan Luo, Lequan Lin, Andi Han, Junbin Gao, José Miguel Hernández Lobato

Abstract: Message-passing neural networks (MPNNs) often suffer from an information bottleneck when capturing long-range dependencies, leading to the oversquashing (OSQ) phenomenon. Alongside spatial connectivity enrichment (e.g., rewiring), recent studies have shown that spectral filtering can yield strong long-range learning outcomes, as spectral operators enable global information mixing that alleviates OSQ. These approaches achieve this either by stabilizing the Jacobian energies in deep propagation or by guaranteeing OSQ mitigation under strong theoretical assumptions. We revisit these conclusions and show that the associated Jacobian sensitivity lower bound is generally difficult to achieve in practice. We then propose S$^3$GNN, which mitigates OSQ without such restrictive assumptions by lightweightly reintroducing omitted components with substantially lower computational complexity, while standard stability constraints on feature transformations remain effective under our new dynamics. Extensive experiments across diverse domains (e.g., long-range benchmarks, KGQA, and mesh-based fluid dynamics) demonstrate that S$^3$GNN achieves up to an order-of-magnitude error reduction with up to 50\% fewer parameters. Our code can be found in https://github.com/EEthanShi/S3-GNN.git.

Comment: Lightweight spectral global mixing mitigates graph oversquashing under less restrictive propagation assumptions.

Topic Match: A new graph propagation mechanism and analysis of its stability make architecture the substantive match, with narrower relevance to graph models.

Relevance: 7 Novelty: 7


9. Training-Free Looped Transformers

ArXiv ID: 2605.23872

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Lizhang Chen, Jonathan Li, Chen Liang, Ni Lao, Qiang Liu

Abstract: We introduce training-free looped transformers, in which a lightweight inference-time wrapper loops a contiguous mid-stack block of layers of a frozen checkpoint without additional fine-tuning, continued training, or architectural changes. Unlike prior looped transformer methods that train with the looped structure end-to-end, we retrofit recurrence onto pretrained models at test time. We show that naive block reapplication usually degrades performance, highlighting the importance of the loop application strategy. Motivated by viewing a pre-norm transformer block as a forward Euler step on an ODE, we instead treat looping as a refinement of the same approximation, replacing one large update with smaller damped sub-steps. Across seven dense, sparse MoE, and MLA+MoE model families, our method improves Qwen3-4B-Instruct by +2.64 pp on MMLU-Pro, Qwen3-30B-A3B-Instruct by +1.14 pp on CommonsenseQA, and Moonlight-16B-A3B-Instruct by +1.20 pp on OpenBookQA.

Comment: Retrofits recurrence onto a frozen mid-stack block by treating pre-norm layers as forward-Euler steps and replacing one update with damped sub-steps, a dynamic-depth mechanism evaluated on dense, MoE and MLA+MoE families.

Topic Match: The core contribution is an architectural/dynamic-computation mechanism with an ODE-based justification for why looping helps.

Relevance: 7 Novelty: 7


10. Balancing Multimodal Learning through Label Space Reshaping

ArXiv ID: 2605.28869

Primary Topic: Architecture and Training Dynamics

Authors: Xiaoyu Ma, Weijie Zhang, Yuanhao Gao, Han Miao, Yongjian Deng, Hao Chen

Abstract: Multimodal learning often suffers from modality imbalance, where modalities that converge faster dominate optimization while others remain undertrained. Existing approaches typically mitigate this issue by strengthening the weak modality or adjusting optimization gradients. However, such strategies mainly compensate for optimization rate discrepancies, often at the expense of the strong modality's optimization capacity, without analyzing how these discrepancies arise at the modality level. Based on theoretical insights and empirical observations, we argue that the discrepancy of learning pace arises from differences in the mapping difficulty between modality-specific feature space and the shared label space. To address this issue, we propose Balanced Multimodal Label Reshaping (BMLR), the first method that promotes multimodal balance from the label-side design. BMLR reshapes the cross-modal label space to equalize mapping difficulty across modalities, thereby facilitating modality interaction and injecting richer inter-class information into each modality. Extensive experiments across multiple architectures demonstrate that BMLR consistently improves multimodal performance and exhibits strong compatibility with diverse model designs. The source code will be released soon.

Comment: Label-space reshaping balances modality-specific convergence by equalizing feature-to-label mapping difficulty.

Topic Match: The learning-pace analysis touches training dynamics, but the method focuses on supervised multimodal label geometry.

Relevance: 6 Novelty: 7


11. Feature Learning in Wide Neural Networks under $μ$P: Identifiability and Sparse-Dictionary Decomposition of the Mean-Field Limit

ArXiv ID: 2605.24710

Primary Topic: Architecture and Training Dynamics

Also Matches: Large-Scale Training Systems and Efficiency

Authors: Akmal Xodarev

Abstract: We establish four structural results for feature learning in wide two-layer neural networks under the Maximal Update Parametrization ($μ$P). First, we prove global existence and uniqueness of the mean-field limit of noisy gradient descent under $μ$P, identifying the maximal admissible weight $w^$ on the moment sequence of the initialization as the reciprocal parameter-moment-growth boundary, and hence the largest weighted moment class propagated by the flow. The finite-particle approximation has uniform-in-time squared-Wasserstein rate $O(N^{-1})$. Second, we characterize identifiability of the mean-field limit: two admissible parameter measures induce the same network function in $L^2$ exactly when their active components agree modulo the finite-rank realization symmetry of the architecture. The orbit depth $D^_{\mathrm{orb}}$ is separated from the moment-variety depth $D^_{\mathrm{var}}$. Third, under the Barron-Hermite target condition the active support of the long-time limit measure admits a sparse-dictionary decomposition: it is supported on at most $S^$ atoms modulo finite-rank realization symmetry, with $S^$ bounded by an explicit coefficient-threshold number. Fourth, we derive the total feature-learning-error decomposition into statistical, optimization, propagation-of-chaos, and sparse-residual components, with a target-dependent Hermite/Barron tail replacing any initialization-only residual. The four results are tied together by an architectural identity: the triple $(w^, D^_{\mathrm{orb}}, S^)$ -- the maximal admissible weight, the orbit identifiability depth, and the sparse-dictionary depth at which the target is realizable -- is the natural learning cell of the architecture-data pair $(σ, ρ)$. The proofs are self-contained except for standard results from $μ$P and mean-field Langevin theory.

Comment: Establishes existence, uniqueness, and identifiability of the mean-field limit of noisy gradient descent under maximal update parametrization, with an O(1/N) finite-particle rate.

Topic Match: Feature-learning theory under muP is training-dynamics analysis, though restricted to wide two-layer networks.

Relevance: 6 Novelty: 6


12. Grow-Prune-Freeze Networks: Adaptive & Continual Learning Technique for Olfactory Navigation

ArXiv ID: 2605.25170

Primary Topic: Architecture and Training Dynamics

Authors: Kordel K. France, Ovidiu Daescu

Abstract: Training data for olfaction is scattered through disparate, non-standardized datasets that limit the ability to build representative world models. Olfactory navigation is a highly dynamic and non-stationary task that benefits from real-time continual learning. We introduce an adaptive framework called Grow-Prune-Freeze (GPF) networks that enable an agent to continually learn through growing, pruning, and freezing early layers of its policy in response to world complexity. Grounding GPFs in non-linear random matrix theory, we show that the work of Pennington & Worth (2017) can be extended from single hidden layers to n-layer continual-learning models, and that eigenvalue composition of network weights is preserved as successive layers are added. We show that GPFs based on Expected SARSA achieve a 94% success rate on turbulent plume navigation - a partially observable, non-stationary task representative of the "big world" challenges that motivate adaptive learning in robotics - and provide supporting methodology for applying GPFs in other world models. Further experiments amount evidence that GPFs may generalize well to other machine learning tasks such as reinforcement learning in Atari, image classification, and autoregressive language models. We open source all code and data to encourage improvements on and more research in olfactory robotics.

Comment: Adaptive growth, pruning, and freezing change network capacity during continual training.

Topic Match: Adaptive network structure fits architectural training mechanisms, although the main evidence concerns olfactory-policy learning and large-model applicability remains preliminary.

Relevance: 6 Novelty: 6


13. ChainzRule: Sample-Efficient, Robust Deep Learning Across Tabular, NLP, and Vision Tasks

ArXiv ID: 2605.24340

Primary Topic: Architecture and Training Dynamics

Authors: Rowan Martnishn

Abstract: Production deep learning systems across enterprise domains operate under constraints that academic benchmarks routinely obscure: labeled data is expensive, inference budgets are tight, and models that cannot explain their behavior are difficult to trust and maintain. We present ChainzRule (CR), a neural architecture replacing typical activations with learnable polynomial layers governed by Differential Regularization (DREG), a layer-wise Jacobian penalty computed analytically during the forward pass at standard inference cost. The core claim is that bounding intermediate derivatives forces the network toward low-frequency, structurally stable representations, simultaneously reducing dependence on labeled data volume, improving robustness to distribution shift, and providing a measurable, gradient-based handle on model behavior. Evaluated across five domains, CR achieves $85.71\% \pm 2.01\%$ on Pima Diabetes (statistically superior to SVM and XGBoost), $46.20\% \pm 0.37\%$ on SST-5 sentiment classification with a frozen encoder (superior to RNTN using approximately 5\% of its training data), $55.79\%$ on SST-5 with a fine-tuned BERT backbone (versus BERT-base linear head at $54.9\%$), $70.17\%$ on Yelp Full ordinal regression with 3.2M parameters versus a 10-model average of $66.35\%$, and $+2.32\%$ mean corruption accuracy on CIFAR-10-C. All results with reported $p$-values fall below the $α= 0.05$ threshold after Bonferroni correction. CR maintains a gradient tail ratio $τ$ (p99/mean) of $1.01$--$1.02$ against $1.07$--$1.09$ for all typical activation function baselines across every data fraction, a structural invariant we propose as the mechanistic driver of sample efficiency and a deployment-time proxy for model reliability.

Comment: Learnable polynomial layers use analytic Jacobian penalties to constrain intermediate derivatives.

Topic Match: The activation design and derivative regularization are architectural training mechanisms, with evidence focused on supervised sample efficiency rather than large-scale pretraining.

Relevance: 6 Novelty: 6


Efficiency, Compression, and Large-Scale Training (15)

1. Signs Beat Floats: Low-Rank Double-Binary Adaptation for On-Device Fine-Tuning

ArXiv ID: 2605.24058

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Yoshihiko Fujisawa, Yuma Ichikawa, Yudai Fujimoto, Akira Sakai, Katsuki Fujisawa

Abstract: On-device adaptation of large language models commonly keeps a quantized base model frozen while training and deploying a small, task-specific LoRA adapter. In the unmerged adapter-mode setting, however, the adapter is more than a compact storage module; it introduces an additional dense floating-point branch, maintains a trainable state for local updates, and acts as a unit of communication and hot-swapping.We introduce LoRDBA, a LoRA-compatible adapter that replaces both low-rank factors with binary sign carriers while representing magnitudes through lightweight, channel-wise scales, converting the dense adapter branch into two sign-accumulation matrix multiplications interleaved with channel-wise scaling. A finite-sample analysis shows that reconstruction quality is governed by the residual-to-magnitude ratio of the original LoRA factors. In adapter-mode experiments, LoRDBA outperforms low-bit baselines at matched model sizes while matching fp16 LoRA quality in selected regimes. The unmerged adapter incurs at most 8% prefill latency overhead at matched rank r=16 despite an over 10x reduction in adapter footprint, with moderate training memory overhead of approximately 1.6x that of fp16 LoRA.

Comment: Binarizes both low-rank adapter factors while retaining channel-wise scales, reducing adapter storage and changing its matrix arithmetic.

Topic Match: Adapter compression and execution efficiency are central; reported training memory is approximately 1.6 times that of fp16 LoRA.

Relevance: 9 Novelty: 7


2. Adaptive Mass-Segmented KV Compression for Long-Context Reasoning

ArXiv ID: 2605.23200

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Junzhe Yang, Xiaoyu Shen

Abstract: The linear growth of the Key-Value (KV) cache is a critical bottleneck in long-form LLM inference. Existing KV compression methods mitigate this by evicting tokens based on importance scores. However, we show that their reliance on global Top-k selection triggers Region Wipe-out: the severe eviction of contiguous reasoning blocks that derails logical coherence. To address this, we propose Adaptive Mass-Segmented (AMS) KV Compression, a framework that shifts the paradigm from token-level competition to region-aware quota allocation. AMS adaptively partitions the KV cache based on the spatial distribution of attention mass, ensuring structurally vital reasoning segments receive guaranteed memory quotas. To ensure stability during iterative decoding, an EMA-based smoothing mechanism is incorporated to prevent jitter in segment boundaries. Crucially, AMS is a universal plug-and-play layer that is orthogonal to existing scorers. It can be seamlessly integrated into representative methods such as TOVA, Expected Attention, KeyDiff, R-KV and TriAttention. AMS is also system-compatible with modern paged-KV serving frameworks such as vLLM, supporting efficient gather-and-compact KV execution without introducing additional steady-state attention overhead. Extensive experiments across a diverse suite of tasks, including mathematical reasoning (MATH500, AIME, GSM8K), code completion, open-domain QA, and sparse retrieval, demonstrate that AMS consistently mitigates structural fragmentation and boosts model performance.

Comment: Allocates KV-cache quotas by attention-mass regions to prevent eviction of entire contiguous reasoning blocks.

Topic Match: The core mechanism changes cache eviction and memory allocation during long-context decoding.

Relevance: 9 Novelty: 7


3. Pruning Deep Neural Networks via the Marchenko--Pastur Distribution

ArXiv ID: 2606.02608

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Leonid Berlyand, Theo Bourdais, Houman Owhadi, Yitzchak Shmalo

Abstract: We study a Marchenko--Pastur (MP) random-matrix approach to pruning deep neural networks with very small post-pruning fine-tuning budgets. The main practical contribution is accuracy retention under short calibration and fine-tuning schedules, rather than a long post-pruning reoptimization pipeline. The theory gives deterministic data-path certificates: if the removed component $R$ has small propagated logit effect $L_s | R ψ_1(s) |\infty$, pruning decreases an elastic-net objective and preserves samples whose dense margin exceeds twice the perturbation. The zero-budget case gives perfect pruning; a prune--restore extension models weight restoration inside a fixed sparse-execution pattern; and an additive $L_2$-regularized model shows admissible random-like components vanish at the training limit, with persistent spikes stabilizing as the MP bulk collapses. Under iid-Gaussian sufficient conditions, the fitted MP edge $σ+$ gives a high-probability layerwise budget signal. On ImageNet-1k, after only three distillation epochs, ViT-B/16 $2{:}4{+}$ToMe reaches $83.41\%$ top-1 ($-1.70$ pp from dense) at $59.81\%$ sparse-execution MAC reduction, with $1.388\times$ best-observed A40 native-$2{:}4$ backend speedup for the same checkpoint and ToMe graph; a separate no-ToMe A100 endpoint gives $2.705\times$. At structured sparsity, ViT-B/16 $6{:}12$ reaches $83.74\%$, ViT-L/16 $8{:}16$ dense+permutation reaches $85.33\%$ ($-0.51$ pp), and ConvNeXtV2-Base $12{:}16$ reaches $86.35\%$ ($-0.37$ pp). For CNNs, ResNet50 $8{:}16$ dense+permutation reaches $75.87\%$ ($-0.26$ pp), and ResNet152d CAST-conv+permutation reaches $81.33\%$ ($-1.53$ pp) at ${\sim}50\%$ MAC accounting with a $1.62\times$ A40 im2col$+2{:}4$ sparse-GEMM audit.

Comment: Uses Marchenko-Pastur spectral structure and perturbation certificates to guide pruning under short fine-tuning budgets.

Topic Match: The central contribution is a principled pruning mechanism targeting sparse execution and inexpensive accuracy recovery.

Relevance: 9 Novelty: 7


4. Approaching I/O-optimality for Approximate Attention

ArXiv ID: 2605.23751

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Pál András Papp, Aleksandros Sobczyk, Anastasios Zouzias

Abstract: We revisit the I/O complexity of attention in large language models. Given query-key-value matrices $Q,K,V\in\mathbb{R}^{n\times d}$, and a machine with fast memory size $M$, the goal is to compute the "attention matrix" $A=\text{softmax}(Q K ^{\top}/\sqrt{d}) V$ with the minimal number of data transfers between fast and slow memory. Existing methods in the literature, most notably FlashAttention and its variants, incur an I/O cost that depends quadratically on $n$, while a trivial lower bound only requires $Ω(nd)$ I/O's to read the inputs and write the output. In this work, we present a technique for computing attention where the I/O cost only depends almost-linearly on $n$ in most parameter regimes. This is achieved by developing I/O-efficient algorithms inspired by the recent approximate attention framework of Alman and Song. We also prove corresponding lower bounds in each parameter regime to show that our algorithms are indeed close to I/O-optimal.

Comment: Attention algorithms whose I/O between fast and slow memory is almost linear in sequence length in most parameter regimes, with matching lower bounds, breaking the quadratic I/O of FlashAttention-style tiling.

Topic Match: Memory-hierarchy cost of the attention kernel is exactly the efficiency topic, and the result is a new algorithm plus optimality proof rather than a tuned variant.

Relevance: 8 Novelty: 8


5. MX-SAFE: Versatile Inference- and Training-Proof Microscaling Format with On-the-Fly Exponent and Mantissa Bit Allocation

ArXiv ID: 2605.24391

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Large-Scale Training Systems and Efficiency

Authors: Dahoon Park, Jahyun Koo, Sangwoo Hwang, Jaeha Kung

Abstract: As the demand for deep learning grows, cost reduction through quantization has become essential for both training and inference. In 2022, the Open Compute Project (OCP) consortium standardized narrow precision formats for deep learning, called the microscaling (MX) format. The MX format is a hardware-friendly dynamic quantization scheme that effectively reduces the data size by sharing an 8-bit exponent across multiple operands. The MX format can be categorized into two types with their own strengths: (i) MXINT which focuses on a high precision consisting only of mantissa bits and (ii) MXFP which focuses on a wider dynamic range by allowing local exponent bits. In this work, we present a versatile MXFP format, called MX-SAFE (MXSF in short), that adaptively uses two modes, i.e., a wider mantissa mode (FP8 E2M5) and a subnormal FP mode (FP5 E3M2), to support both training and direct-cast inference. Furthermore, we propose a tile-based block design to increase hardware efficiency by reducing the burden of re-quantization process during the training with the MXSF format. Owing to the use of the proposed MXSF format, 0.05%/11.1% and 3.55%/3.57% improvements in accuracy, on average, for inference/full-training compared to MXFP8 E2M5 and MXFP8 E4M3 are observed, respectively. Moreover, we present a training-inference accelerator that supports the MXSF format and it achieves similar accuracy to the BF16 baseline while using 24.9% less total energy consumption.

Comment: A microscaling format that switches on the fly between a wide-mantissa and a subnormal mode so one numeric format serves both full training and direct-cast inference, with tile-based blocks cutting re-quantization cost.

Topic Match: Low-precision numerics that materially change training cost, not just inference quantization.

Relevance: 8 Novelty: 7


6. Inference Time Context Sparsity: Illusion or Opportunity?

ArXiv ID: 2605.24168

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Sahil Joshi, Prithvi Dixit, Agniva Chowdhury, Anshumali Shrivastava, Joseph E. Gonzalez, Ion Stoica, Kumar Krishna Agrawal, Aditya Desai

Abstract: Sparsity has long been a central theme in LLM efficiency, but its role in context processing remains unresolved. As LLM workloads shift toward longer contexts and agentic interactions, the compute and memory bottlenecks of attention become increasingly critical, raising the question of whether these constraints are fundamental. Our position is that these constraints are artificial and unnecessary, and that the future of LLM inference lies in extreme but principled sparsity along the context dimension. This position is supported by several strands of empirical and theoretical evidence. First, we find the insistence on dense attention unreasonable, since in a long context a query effectively projects O(N) attention information into a hidden space of dimension d << N, making the process inherently lossy. Second, we perform an extensive study of sparsity in LLMs spanning 20 models across five model families, varying context lengths, and different sparsity levels. We empirically demonstrate a strong trend: current LLMs, despite not being trained for context sparsity, are remarkably robust to inference-time decode sparsity across tasks of varying complexity, including retrieval, multi-hop QA, mathematical reasoning, and agentic coding. Importantly, we also show that current hardware is already sufficient to realize substantial gains from this sparsity. For example, our sparse decode kernels accelerate large-context processing by up to 10x over FlashInfer at 50x sparsity levels on hardware such as the H100. Overall, these results position extreme context sparsity not as a heuristic, but as a principled foundation for LLM inference, training, and architecture design: one that is both feasible and beneficial, and a compelling direction for future systems.

Comment: Argues dense attention is inherently lossy since a query projects O(N) context into dimension d<<N, backs it with a 20-model sparsity study, and shows sparse decode kernels reaching 10x over FlashInfer at 50x sparsity on H100.

Topic Match: Context-dimension sparsity plus the accompanying kernels directly change memory and compute cost, and the position is aimed at architecture and training design, not just serving.

Relevance: 8 Novelty: 7


7. Scale When Needed: Adaptive Neuron-level Mixed Precision Quantization Aware Training

ArXiv ID: 2605.25054

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Ayush K. Varshney, Konstantinos Vandikas, Šarūnas Girdzijauskas, Adam Orucu, Aneta Vulgarakis Feljan

Abstract: Deploying deep neural networks on resource-constrained 6G edge devices demands aggressive compression with minimal accuracy loss. Quantization-Aware Training (QAT) has emerged as a leading compression approach; however, existing mixed-precision methods typically operate at coarse layer- or channel-level granularity. These methods often rely on heuristic or search-based bit-allocation strategies, which may overlook fine-grained variability at the neuron level. We propose Neuron-Level Mixed-Precision QAT (NMP-QAT), where each neuron independently learns its own discrete precision during training. Starting from low-bit precision, NMP-QAT expands bit-width only when training signals demand it, via differentiable surrogates and straight-through estimators, while preserving a fully discrete inference graph. This adaptability extends to both weights and activations, reducing memory movement. Evaluated on telecom and non-telecom datasets across MLP and tabular foundation model architectures, NMP-QAT achieves superior compression-accuracy trade-offs over mixed-precision QAT baselines, making it well-suited for Green AI deployments at the network edge.

Comment: Each neuron learns discrete weight and activation precision through adaptive bit-width expansion.

Topic Match: Trainable precision allocation directly fits quantization and compression, with validation concentrated on MLPs and tabular foundation models.

Relevance: 8 Novelty: 6


8. Polynomial Context-Truncation Sensitivity in Autoregressive Language Models: Sequential Wyner-Ziv Bounds for KV Cache Compression

ArXiv ID: 2605.25085

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Munsik Kim

Abstract: We study the rate-distortion limits of online KV cache compression in autoregressive language models, formulating it as sequential Wyner-Ziv source coding on the filtration induced by the model, with the next-step query as decoder side information. Empirically, across four models spanning two families and $0.5$-$3$B parameters, we find that the next-token distribution's sensitivity to context truncation decays \emph{polynomially} rather than \emph{geometrically}: a power law improves on an exponential fit by an order of magnitude in extrapolation, the fitted exponent is recovered independently from a sink-plus-recent KL measurement, and the decay is verified to be free of positional-encoding artifacts by a position-preserving ablation. Under a corresponding \emph{polynomial truncation-sensitivity} assumption, our main result characterizes the per-token memory requirement of \emph{suffix-only} cache policies: a sliding-window scheme attains distortion $\varepsilon$ with window $w = O(\varepsilon^{-1/α})$, and -- under an additional two-sided Bayes-risk condition -- a converse shows $w = Ω(\varepsilon^{-1/α})$ is necessary within this policy class, so the scaling is $Θ(\varepsilon^{-1/α})$ for suffix-only policies. Whether recurrent or propagating cache summaries can beat this scaling is left open. An explicit block-Markov scheme achieves the upper bound; its rate-of-convergence exponent matches the converse under additional forward-decay and regularity hypotheses (not implied by truncation sensitivity alone), and differs by a factor of two otherwise. Empirically, the polynomial law predicts the degradation curves of concrete cache policies: recency-based eviction (sliding, sink-plus-recent) suppresses distortion by roughly two orders of magnitude over random retention at equal budget, with a power-law decay in the budget.

Comment: Measures that sensitivity to context truncation decays polynomially rather than geometrically, then derives matching upper and lower bounds showing suffix-only cache policies need window size scaling as a power of the distortion target.

Topic Match: Rate-distortion limits for KV-cache compression give the memory-efficiency topic an information-theoretic floor rather than another heuristic.

Relevance: 7 Novelty: 7


9. Beyond the Frontier: Stochastic Backtracking for Efficient Test-Time Scaling

ArXiv ID: 2605.25143

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Dao Tran, Duc Anh Le, Ngoc Luu, Quan Pham, Tung Pham, Hung Bui

Abstract: Test-time scaling improves language model reasoning by spending additional compute to explore multiple solution trajectories. The key challenge is to maximize accuracy while minimizing the total number of generated tokens during reasoning. Recent PRM-guided methods score intermediate prefixes to steer this search, but most are frontier-only: they keep only the current active prefixes and irreversibly prune or resample away the rest using noisy PRM scores. This can cause premature commitment, diversity collapse, and the loss of prefixes that still admit correct continuations. We introduce stochastic backtracking over a persistent pool of historical prefixes, allowing test-time compute to revisit previously generated states instead of only expanding the current frontier. To make this efficient, we propose two complementary mechanisms. Subpool Selection strengthens greedy PRM-guided search by applying Top-N selection within random subpools, giving historical prefixes a chance to bypass over-scored frontier candidates. Power Backtrack Sequential Monte Carlo extends SMC-style resampling to the persistent pool using powered PRM scores and mixture-corrected weights. Across mathematical reasoning benchmarks and model scales, our methods consistently achieve higher accuracy per token count, and the same level of accuracy using only a fraction of the token count in comparison to strong PRM-guided baselines, demonstrating that persistent-pool stochastic backtracking provides a simple and effective way to improve the accuracy-token trade-off in test-time scaling.

Comment: Stochastic backtracking reuses historical prefixes to improve reasoning accuracy per generated token.

Topic Match: New search and resampling mechanisms reduce inference token cost, providing a narrower match through computational efficiency.

Relevance: 7 Novelty: 7


10. Context Features Are Cheap: Rank-Aware Decomposition for Efficient Feature Interaction in Recommender Systems

ArXiv ID: 2605.27450

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Yevgeny Tkach

Abstract: Modern industrial recommender systems use a deep ranking model to score N candidates against the same user and context features. Standard implementations broadcast context features early in the forward pass, redundantly computing context-only operations N times per request. We present a rank-aware decomposition applicable to the dominant interaction mechanisms in modern recommender architectures-Factorization Machine (FM) pairwise products, Deep Cross Network (DCNv2) cross layers, self-attention, and fully connected (FC) projection layers-built on a single algebraic principle: any linear or bilinear operation over a rank-partitioned input admits an exact block decomposition that moves context-only computation from once-per-candidate to once-per-request, identity-equivalent to the original model. Closed-form analysis and controlled ablation verify that savings scale quadratically with the number of context features. Applied to a production DLRM-style ranker without any architectural change, the decomposition increases per-pod throughput by 87.5% (a 47% reduction in peak pod count) at identical model predictions. The identity-equivalent decomposition applies only at the first layer of cross networks and self-attention, since each layer mixes ranks in its output. To extend savings across depth, we further introduce rDCN, an architectural variant of DCNv2 that maintains rank discipline across depth and matches DCNv2 accuracy within training noise at 67% fewer total FLOPs, and sketch an analogous architectural variant for self-attention.

Comment: Exact block decomposition reuses context-only feature computations across ranked candidates.

Topic Match: Eliminating redundant computation is the core contribution; rank-preserving cross layers add an architectural mechanism, although validation centers on recommender rankers.

Relevance: 7 Novelty: 6


11. One-Forcing: Towards Stable One-Step Autoregressive Video Generation

ArXiv ID: 2605.23458

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Jiaqi Feng, Justin Cui, Yuanhao Ban, Cho-Jui Hsieh

Abstract: Recent advances have substantially improved real-time interactive video generation in the autoregressive regime. However, most existing few-step autoregressive video generation methods, often distilled from a corresponding many-step teacher, default to a 4-step sampling configuration, which still incurs considerable latency during deployment and suffers from severe quality degradation when the number of sampling steps is further reduced, particularly in the one-step setting. Trajectory-style consistency distillation methods often produce videos with weak dynamics, while DMD-based approaches, such as Self-Forcing, tend to yield blurry frames. To address this challenge, we propose One-Forcing, a simple yet effective approach which augments the DMD objective with an auxiliary GAN loss for high-quality and efficient one-step video generation. Experiments on VBench show that One-Forcing achieves a total score of 83.76, establishing state-of-the-art performance among one-step causal video generation methods and remaining competitive with strong many-step approaches. We further demonstrate that one-step framewise autoregressive generation can be achieved stably with merely one-third of the training cost of the chunkwise model, a setting that prior methods have failed to achieve successfully.

Comment: Adversarially regularized distribution-matching distillation stabilizes one-step autoregressive video generation.

Topic Match: The training objective directly reduces sampling steps and enables a cheaper training configuration, with evidence limited to video generation.

Relevance: 7 Novelty: 6


12. ReLoRA: Knowledge-Reusing Adaptation for Fast Rollout of Evolving LLM Services

ArXiv ID: 2606.02606

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Yang Xu, Zihuai Xu, Hongli Xu, Yunming Liao, Zhiwei Yao, Xitong Fu

Abstract: Large Language Models (LLMs) are increasingly deployed as continuously evolving services, where frequent base-model updates may invalidate previously deployed task-specific Low-Rank Adaptation (LoRA) adapters. For service providers managing numerous downstream model services, retraining each LoRA adapter from scratch for every updated base model is computationally prohibitive and delays service rollout. Meanwhile, the simpler alternative, i.e., naively applying the original LoRA adapter to the updated base model, often leads to degraded service quality due to adapter-backbone incompatibility. To address this problem, we propose ReLoRA, a knowledge-reusing re-adaptation framework that efficiently restores service-ready LoRA adapters for evolving LLM services while preserving or improving task performance. Specifically, ReLoRA comprises two key optimization steps: 1) Adaptive LoRA initialization leverages Bayesian optimization to construct a compatibility-aware starting point by fusing information from both the previously deployed task adapter and the base model's evolution; 2) Fine-tuning with scheduled regularization first rapidly steers the adapter to a high-quality region via strong regularization, followed by relaxed regularization for task-specific refinement. This design enables rapid service-quality recovery with reduced re-adaptation overhead. Extensive experiments demonstrate that ReLoRA reduces time-to-readiness by up to 8.9$\times$ and improves accuracy by up to 4.6\% compared to baselines.

Comment: Reuses existing LoRA adapters through compatibility-aware initialization to reduce re-adaptation compute after base-model updates.

Topic Match: Cross-version adapter reuse directly reduces repeated fine-tuning cost, with a narrower focus on maintaining evolving model services.

Relevance: 7 Novelty: 6


13. PALoRA: Projection-Adaptive LoRA for Preserving Reasoning in Large Language Models

ArXiv ID: 2605.24549

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Mustafa Hayri Bilgin, Mariam Barry, Albert Bifet, Azzedine Idir Ait Said, Soumya Banerjee

Abstract: Efficiently updating Large Language Models (LLMs) with new or evolving factual knowledge remains a central challenge, as even parameter-efficient adaptation can erode previously acquired reasoning abilities. This tension reflects a plasticity-stability dilemma: models must incorporate new knowledge while preserving skill-critical representations. In this work, we study this trade-off through the spectral structure of multilayer perceptron weight matrices. We show, both theoretically and empirically, that information essential for reasoning is not localized only in dominant singular directions, but is instead distributed across the singular spectrum. Motivated by this observation, we introduce PALoRA, a two-stage framework for knowledge injection with reduced interference. PALoRA first trains a Singular Value Fine-Tuning (SVF) expert on a reasoning dataset and uses its learned singular scaling vector as a frozen geometric probe to identify components that are critical for the target skill. It then performs factual knowledge injection with Low-Rank Adaptation (LoRA) under a structural orthogonality constraint, ensuring that updates avoid the identified skill-relevant subspace. Across Llama 3.1 8B and Mistral 7B, and across mathematical, coding, and scientific reasoning benchmarks, PALoRA preserves on average 95% of the SVF expert's reasoning performance while maintaining competitive factual recall. It consistently improves skill retention over prior spectral Parameter-Efficient Fine-Tuning (PEFT) methods while adding less than 0.006% parameter overhead.

Comment: Orthogonally constrained low-rank updates avoid skill-critical singular subspaces identified by a trained probe.

Topic Match: The constrained adapter fits parameter-efficient adaptation, although its main benefit is post-training skill retention and additional compute savings are not established.

Relevance: 6 Novelty: 7


14. EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture

ArXiv ID: 2605.24144

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Bowen Duan, Cong Guo, Chiyue Wei, Haoxuan Shan, Yuzhe Fu, Xinhua Chen, Yifan Xu, Ziyue Zhang, Changchun Zhou, Hai Li, Yiran Chen

Abstract: Large Language Models (LLMs) have achieved impressive performance across diverse domains but remain inefficient during the autoregressive decoding phase. Unlike the prefill stage, which employs compute-bound GEMM operations, decoding executes a sequence of small GEMV-like computations that are memory-bound and underutilize modern accelerators. Weight-only vector quantization (VQ) has emerged as an effective compression technique that clusters model weights into a shared codebook and replaces the original weight matrix with low-precision indices, enabling 2-bit-level weight compression. While this approach substantially reduces model size and memory bandwidth, it still suffers from two critical inefficiencies: the low utilization of GEMV computation and frequent memory conflicts during codebook lookups. This paper presents EVA, an efficient vector-quantization-based architecture that addresses both computational and memory bottlenecks in LLM decoding. EVA builds on a simple yet effective insight that combines input-codebook computation with conflict-free memory access. Instead of reconstructing quantized weights from indices, EVA directly performs dot products between input vectors and the weight codebook, transforming LLM decoding from GEMV to GEMM computation. It then performs structured lookups from an intermediate output buffer, eliminating memory bank conflicts. We further design a hardware-software co-optimized architecture specialized for LLM decoding while remaining compatible with conventional prefill execution. Evaluations show that EVA achieves up to 11.17$\times$ speedup and 7.17$\times$ higher energy efficiency compared with the SOTA lookup-based architecture, while preserving arithmetic precision after vector quantization. Our code is available at https://github.com/dbw6/Eva.git.

Comment: Turns vector-quantized decoding from GEMV into GEMM by dot-producting inputs against the codebook directly, eliminating index reconstruction and bank conflicts, with a co-designed accelerator.

Topic Match: A new quantization-execution mechanism that changes the memory-bound cost of running a large model, though it targets inference hardware.

Relevance: 6 Novelty: 6


15. SonicSampler: Unified Tile-Aware Kernels for LLM Sampling and Speculative Verification

ArXiv ID: 2607.20475

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Pragaash Ponnusamy, Shivam Sahni, Jue Wang, Tri Dao

Abstract: Sampling in LLM inference comprises a combinatorial set of logit processing, token selection, and verification operations for speculative decoding. However, existing implementations either accelerate only subsets of this pipeline, rely on multiple kernel launches, or assume homogeneous sampling behavior across a batch, limiting support for dynamic serving workloads and preventing efficient CUDA Graph execution. We present $\textbf{SonicSampler}$, a unified suite of tile-aware Triton kernels that vertically fuses the complete sampling pipeline into a fixed, workload-aware execution model. Our kernels support dynamic per-request sampling behaviors, including grammar-constrained decoding, repetition, frequency and presence penalties, logit bias, temperature scaling, top-$k$ / top-$p$ / min-$p$ filtering, and speculative verification - within a single batched kernel while remaining fully CUDA Graph-compatible. Central to our approach is a novel hierarchical two-stage top-$k$ algorithm that achieves up to $\textbf{10x speedup}$ over competitive baselines and exploits the low-entropy structure of LLM outputs to enable efficient selection over large vocabularies. Across heterogeneous speculative decoding workloads, SonicSampler achieves up to $\textbf{16x speedup}$ over state-of-the-art baselines while preserving flexible batched execution.

Comment: Fuses logit processing, token selection, and speculative verification into one CUDA-Graph-compatible batched kernel, with a hierarchical two-stage top-k exploiting low-entropy output distributions.

Topic Match: A genuinely new kernel algorithm rather than tuning, though it targets the sampling path of inference.

Relevance: 6 Novelty: 6


Paper Selection Prompt

System Prompt

You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.

User Prompt

Instructions

Respond in JSONL. Output exactly one JSON object per paper, one per line:

{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}

Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only: daily_hot, new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return []. - daily_hot means the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. - new_frontier means the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.

Scoring Criteria

Relevance and Novelty are independent axes. Score both from 1 to 10.

Relevance Scoring

  • 9-10: directly centered on the target foundational topics; highest when the core contribution is clearly within them.
  • 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
  • 5-6: touches the target topics, but the main contribution is elsewhere.
  • 3-4: largely outside the target topics, often application-focused or domain-specific.
  • 1-2: unrelated.

Important: Broad frontier relevance, major launch status, or daily buzz is not enough for a high Relevance score here. Those cases belong in the hotspot digest unless the paper strongly matches the specialized paper topics.

Novelty Scoring

  • 9-10: new paradigm, theory, or major methodological breakthrough.
  • 7-8: substantial methodological advance or strong new insight.
  • 5-6: meaningful but incremental extension or refinement.
  • 3-4: minor, narrow, or mostly engineering or domain-specific improvement.
  • 1-2: little originality; mainly standard application of existing methods.

Topic Registry

Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - moe_training: MoE Training - Mixture-of-Experts training: routing, load balancing, expert granularity and shared experts, training stability, MoE parallelism and communication, and MoE kernels. - training_systems: Large-Scale Training Systems and Efficiency - Large-scale training systems and efficiency: distributed training algorithms, optimizers, and communication. - architecture_training: Architecture and Training Dynamics - Core architectural or computational mechanisms, dynamic computation, and training-stability dynamics. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, memory or cache efficiency, and large-scale training systems that materially change cost or behavior.

Papers

[PAPER LIST HERE]

Relevant Topics

This is a training-side feed. The centre of gravity is Mixture-of-Experts training and the large-scale training systems it sits on; architecture and efficiency work is kept when it bears on how a large model is trained or what it costs to train.

Keep a paper when its CORE CONTRIBUTION falls in one of the four topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.

  1. MoE Training (primary) - Keep: expert routing and gating mechanisms; load balancing, with or without auxiliary losses; expert granularity, fine-grained and shared experts; MoE training stability (loss spikes, router collapse, z-loss and related fixes); expert and hybrid parallelism; all-to-all dispatch/combine and the communication schedules MoE needs; MoE kernels and grouped-GEMM style implementations; expert capacity, token dropping, and upcycling dense checkpoints into MoE. - Filter: papers that merely train on top of a MoE model without a new routing, balancing, stability, parallelism, or kernel contribution; "mixture of experts" in the classical ensemble or recommender sense.

  2. Large-Scale Training Systems - Keep: distributed training algorithms (data, tensor, pipeline, sequence, and expert parallelism, FSDP and ZeRO-style sharding); optimizers and preconditioners for large-scale pretraining; communication algorithms, overlap, and collective design; throughput and MFU work that materially changes what a training run costs or whether it converges; scaling-law work that informs how such runs are configured. - Filter: routine infrastructure tuning with no new algorithmic or systems idea.

  3. Architecture and Training Dynamics - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, dynamic or modular computation); optimisation and training-dynamics analysis that explains why large models train the way they do. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight.

  4. Efficiency and Compression - Keep: quantization, sparsity, pruning, low-rank adaptation, KV-cache and memory-efficient designs, and other work that materially changes the cost of running or training a large model, especially when the mechanism is new rather than a tuned variant. - Filter: straightforward application of standard efficiency methods, and deployment or serving engineering with no new idea.

Leave these to the hotspot digest unless the core contribution clearly falls in one of the four topics above: - major model or product releases - agent, tooling, and RAG launches - benchmark, leaderboard, or evaluation-only papers - representation-learning theory, interpretability, and feature-formation analysis - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning) - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains