This is a remedial run for missed papers from 02/02/2026 to 02/02/2026.
Results generated on 09/11/2026.
Personalized Daily ArXiv Papers 2026-02-03
| Model | Metric | Usage | Papers | ||||
|---|---|---|---|---|---|---|---|
| Prompt | Completion | Total | Total arXiv | Scanned | Relevant | ||
local agent CLI |
Tokens | not reported | not reported | not reported | 692 | 692 | 40 |
| Cost | not reported | not reported | not reported | ||||
Token counts are not reported for this run. 7 of 11 model calls succeeded, 4,375s of model wall clock.
Topic Coverage:
| Topic | Papers |
|---|---|
| MoE Training | 3 |
| Large-Scale Training Systems and Efficiency | 7 |
| Architecture and Training Dynamics | 14 |
| Efficiency, Compression, and Large-Scale Training | 16 |
Table of contents by topic:
MoE Training (3)
-
Certain Head, Uncertain Tail: Expert-Sample for Test-Time Scaling in Fine-Grained MoE Authors: Yuanteng Chen, Peisong Wang, Nanxin Zeng, Yuantian Shao, Shuang Qiu, Gang Li, Jing Liu, Jian Cheng
-
SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning Authors: Zhen-Hao Xie, Jun-Tao Tang, Yu-Cheng Shi, Han-Jia Ye, De-Chuan Zhan, Da-Wei Zhou
-
Expert-Data Alignment Governs Generation Quality in Decentralized Diffusion Models Authors: Marcos Villagra, Bidhan Roy, Raihan Seraj, Zhiying Jiang
Large-Scale Training Systems and Efficiency (7)
-
DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers Authors: Ionut-Vlad Modoranu, Philip Zmushko, Erik Schultheis, Mher Safaryan, Dan Alistarh
-
SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning Authors: Qifan Yu, Xinyu Ma, Zhijian Zhuo, Minrui Wang, Deyi Liu, Shiyi Zhan, Yiyuan Ma, Liang Xiang, Xingyan Bin, Di He
-
The Effect of Mini-Batch Noise on the Implicit Bias of Adam Authors: Matias D. Cattaneo, Boris Shigida
-
Decentralized SGD with Controlled Disagreement Finds Flatter Minima Authors: Zesen Wang, Mikael Johansson
-
ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning Authors: Jingwei Song, Meng Chen, Jie Xiao, Qingnan Ren, Jiaqi Huang, Yangshen Deng, Chris Tong, Wanyi Chen, Suli Wang, Zhisheng Chen, Ziqian Bi, Shuo Lu, Yiqun Duan, Xu Wang, Rymon Yu, Lynn Ai, Eric Yang, Tianyu Shi
-
Stein-Rule Shrinkage for Stochastic Gradient Estimation in High Dimensions Authors: M. Arashi, M. Amintoosi
-
Grappa: Gradient-Only Communication for Scalable Graph Neural Network Training Authors: Chongyang Xu, Christoph Siebenbrunner, Laurent Bindschaedler
Architecture and Training Dynamics (14)
-
Softmax Linear Attention: Reclaiming Global Competition Authors: Mingwei Xu, Xuan Lin, Xinnan Guo, Wanqing Xu, Wanyun Cui
-
Poly-attention: a general scheme for higher-order self-attention Authors: Sayak Chakrabarti, Toniann Pitassi, Josh Alman
-
A Provable Expressiveness Hierarchy in Hybrid Linear-Full Attention Authors: Xiaowei Ye, Xiaoyu He, Chao Liao, Chen Wu, Pinyan Lu
-
An Empirical Study on Noisy Data and LLM Pretraining Loss Divergence Authors: Qizhen Zhang, Ankush Garg, Jakob Foerster, Niladri Chatterji, Kshitiz Malik, Mike Lewis
-
State Rank Dynamics in Linear Attention LLMs Authors: Ao Sun, Hongtao Zhang, Heng Zhou, Yixuan Ma, Yiran Qin, Tongrui Su, Yan Liu, Zhanyu Ma, Jun Xu, Jiuchong Gao, Jinghua Hao, Renqing He
-
CoMeT: Collaborative Memory Transformer for Efficient Long Context Modeling Authors: Runsong Zhao, Shilei Liu, Jiwei Tang, Langming Liu, Haibin Chen, Weidong Zhang, Yujin Yuan, Tong Xiao, Jingbo Zhu, Wenbo Su, Bo Zheng
-
HopFormer: Sparse Graph Transformers with Explicit Receptive Field Control Authors: Sanggeon Yun, Raheeb Hassan, Ryozo Masukawa, Sungheon Jeong, Mohsen Imani
-
Training-Free Self-Correction for Multimodal Masked Diffusion Models Authors: Yidong Ouyang, Panwen Hu, Zhengyan Wan, Zhe Wang, Liyan Xie, Dmitriy Bespalov, Ying Nian Wu, Guang Cheng, Hongyuan Zha, Qiang Sun
-
Breaking the Reversal Curse in Autoregressive Language Models via Identity Bridge Authors: Xutao Ma, Yixiao Huang, Hanlin Zhu, Somayeh Sojoudi
-
On the Spatiotemporal Dynamics of Generalization in Neural Networks Authors: Zichao Wei
-
A Theoretical Analysis of Why Masked Diffusion Models Mitigate the Reversal Curse Authors: Moongyu Jeon, Sangwoo Shin, BumJun Kim, Kyelim Lee, Albert No
-
Beyond Precision: Training-Inference Mismatch is an Optimization Problem and Simple LR Scheduling Fixes It Authors: Yaxiang Zhang, Yingru Li, Jiacai Liu, Jiawei Xu, Ziniu Li, Qian Liu, Haoyuan Li
-
Late-Stage Generalization Collapse in Grokking: Detecting anti-grokking with Weightwatcher Authors: Hari K Prakash, Charles H Martin
-
Geometric Analysis of Token Selection in Multi-Head Attention Authors: Timur Mudarisov, Mikhal Burtsev, Tatiana Petrova, Radu State
Efficiency, Compression, and Large-Scale Training (16)
-
Dissecting Outlier Dynamics in LLM NVFP4 Pretraining Authors: Peijie Dong, Ruibo Fan, Yuechen Tao, Di Mou, Wenhu Hu, Zhenheng Tang, Yinghao Yu, Jiamang Wang, Wenbo Su, Guodong Yang, Liping Zhang, Xiaowen Chu, Baochun Li, Bo Li
-
Revisiting Adaptive Rounding with Vectorized Reparameterization for LLM Quantization Authors: Yuli Zhou, Qingxuan Chen, Luca Benini, Guolei Sun, Yawei Li
-
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs Authors: Yoonjun Cho, Dongjae Jeon, Soeun Kim, Moongyu Jeon, Albert No
-
More Than a Quick Glance: Overcoming the Greedy Bias in KV-Cache Compression Authors: Aryan Sood, Tanvi Sharma, Vansh Agrawal
-
Every Bit Counts: A Theoretical Study of Precision-Expressivity Tradeoffs in Quantized Transformers Authors: Sayak Chakrabarti, Toniann Pitassi, Josh Alman
-
BAPS: A Fine-Grained Low-Precision Scheme for Softmax in Attention via Block-Aware Precision reScaling Authors: Zisheng Ye, Xiaoyu He, Maoyuan Song, Guoliang Qiu, Chao Liao, Chen Wu, Yonggang Sun, Zhichun Li, Xiaoru Xie, Yuanyong Luo, Hu Liu, Pinyan Lu, Heng Liao
-
Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention Authors: Dvir Samuel, Issar Tzachor, Matan Levy, Michael Green, Gal Chechik, Rami Ben-Ari
-
Two-Stage Grid Optimization for Group-wise Quantization of LLMs Authors: Junhan Kim, Gukryeol Lee, Seungwoo Son, Jeewook Kim, Yongkweon Jeon
-
IntraSlice: Towards High-Performance Structural Pruning with Block-Intra PCA for LLMs Authors: Meng Li, Peisong Wang, Yuantian Shao, Qinghao Hu, Hongjian Fang, Yifan Zhang, Zhihui Wei, Jian Cheng
-
Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM Compression Authors: Ali Abbasi, Chayne Thrash, Haoran Qin, Shansita Sharma, Sepehr Seifi, Soheil Kolouri
-
FlexRank: Nested Low-Rank Knowledge Decomposition for Adaptive Model Deployment Authors: Riccardo Zaccone, Stefanos Laskaridis, Marco Ciccone, Samuel Horváth
-
The Multiple Ticket Hypothesis: Random Sparse Subnetworks Suffice for RLVR Authors: Israel Adewuyi, Solomon Okibe, Vladmir Ivanov
-
TraceNAS: Zero-shot LLM Pruning via Gradient Trace Correlation Authors: Prajna G. Malettira, Manish Nagaraj, Arjun Roy, Shubham Negi, Kaushik Roy
-
You Need an Encoder for Native Position-Independent Caching Authors: Shiju Zhao, Junhao Hu, Jiaqi Zheng, Guihai Chen
-
Prism: Efficient Test-Time Scaling via Hierarchical Search and Self-Verification for Discrete Diffusion Language Models Authors: Jinbin Bai, Yixuan Li, Yuchen Zhu, Yi Xin, Qingyu Shi, Aosong Feng, Xiaohong Liu, Molei Tao, Jianru Xue, Xiangtai Li, Ming-Hsuan Yang
-
ReasonCACHE: Teaching LLMs To Reason Without Weight Updates Authors: Sharut Gupta, Phillip Isola, Stefanie Jegelka, David Lopez-Paz, Kartik Ahuja, Mark Ibrahim, Mohammad Pezeshki
MoE Training (3)
1. Certain Head, Uncertain Tail: Expert-Sample for Test-Time Scaling in Fine-Grained MoE
ArXiv ID: 2602.02443
Primary Topic: MoE Training
Authors: Yuanteng Chen, Peisong Wang, Nanxin Zeng, Yuantian Shao, Shuang Qiu, Gang Li, Jing Liu, Jian Cheng
Abstract: Test-time scaling improves LLM performance by generating multiple candidate solutions, yet token-level sampling requires temperature tuning that trades off diversity against stability. Fine-grained MoE, featuring hundreds of well-trained experts per layer and multi-expert activation per token, offers an unexplored alternative through its rich routing space. We empirically characterize fine-grained MoE routing and uncover an informative pattern: router scores exhibit a certain head of high-confidence experts followed by an uncertain tail of low-confidence candidates. While single-run greedy accuracy remains stable when fewer experts are activated, multi-sample pass@n degrades significantly-suggesting that the certain head governs core reasoning capability while the uncertain tail correlates with reasoning diversity. Motivated by these findings, we propose Expert-Sample, a training-free method that preserves high-confidence selections while injecting controlled stochasticity into the uncertain tail, enabling diverse generation without destabilizing outputs. Evaluated on multiple fine-grained MoE models across math, knowledge reasoning, and code tasks, Expert-Sample consistently improves pass@n and verification-based accuracy. On Qwen3-30B-A3B-Instruct evaluated on GPQA-Diamond with 32 parallel samples, pass@32 rises from 85.4% to 91.9%, and accuracy improves from 59.1% to 62.6% with Best-of-N verification.
Comment: Characterises fine-grained MoE router scores as a confident head plus an uncertain tail, then injects stochasticity only into the tail so routing itself becomes the diversity knob instead of token temperature.
Topic Match: The contribution is an empirical characterisation of fine-grained expert routing plus a routing-space sampling mechanism, which is MoE routing work even though it is applied at test time.
Relevance: 8 Novelty: 7
2. SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning
ArXiv ID: 2602.01990
Primary Topic: MoE Training
Authors: Zhen-Hao Xie, Jun-Tao Tang, Yu-Cheng Shi, Han-Jia Ye, De-Chuan Zhan, Da-Wei Zhou
Abstract: Multimodal Large Language Models (MLLMs) achieve strong performance through instruction tuning, but real-world deployment requires them to continually expand their capabilities, making Multimodal Continual Instruction Tuning (MCIT) essential. Recent methods leverage sparse expert routing to promote task specialization, but we find that the expert routing process suffers from drift as the data distribution evolves. For example, a grounding query that previously activated localization experts may instead be routed to irrelevant experts after learning OCR tasks. Meanwhile, the grounding-related experts can be overwritten by new tasks and lose their original functionality. Such failure reflects two problems: router drift, where expert selection becomes inconsistent over time, and expert drift, where shared experts are overwritten across tasks. Therefore, we propose StAbilized Mixture-of-Experts (SAME) for MCIT. To address router drift, SAME stabilizes expert selection by decomposing routing dynamics into orthogonal subspaces and updating only task-relevant directions. To mitigate expert drift, we regulate expert updates via curvature-aware scaling using historical input covariance in a rehearsal-free manner. SAME also introduces adaptive expert activation to freeze selected experts during training, reducing redundant computation and cross-task interference. We also introduce a new benchmark to evaluate MCIT with long task sequence, and extensive experiments demonstrate SAME's SOTA performance. Code is available at https://github.com/LAMDA-CL/Prism.
Comment: Diagnoses router drift and expert overwriting, then stabilizes routing via orthogonal-subspace updates and curvature-scaled expert updates with adaptive expert freezing.
Topic Match: The contribution is a routing-stability and expert-update mechanism for sparse MoE, even though the setting is multimodal continual instruction tuning.
Relevance: 7 Novelty: 6
3. Expert-Data Alignment Governs Generation Quality in Decentralized Diffusion Models
ArXiv ID: 2602.02685
Primary Topic: MoE Training
Authors: Marcos Villagra, Bidhan Roy, Raihan Seraj, Zhiying Jiang
Abstract: Decentralized Diffusion Models (DDMs) route denoising through experts trained independently on disjoint data clusters, which can strongly disagree in their predictions. What governs the quality of generations in such systems? We present the first ever systematic investigation of this question. A priori, the expectation is that minimizing denoising trajectory sensitivity -- minimizing how perturbations amplify during sampling -- should govern generation quality. We demonstrate this hypothesis is incorrect: a stability-quality dissociation. Full ensemble routing, which combines all expert predictions at each step, achieves the most stable sampling dynamics and best numerical convergence while producing the worst generation quality (FID 47.9 vs. 22.6 for sparse Top-2 routing). Instead, we identify expert-data alignment as the governing principle: generation quality depends on routing inputs to experts whose training distribution covers the current denoising state. Across two distinct DDM systems, we validate expert-data alignment using (i) data-cluster distance analysis, confirming sparse routing selects experts with data clusters closest to the current denoising state, and (ii) per-expert analysis, showing selected experts produce more accurate predictions than non-selected ones, and (iii) expert disagreement analysis, showing quality degrades when experts disagree. For DDM deployment, our findings establish that routing should prioritize expert-data alignment over numerical stability metrics.
Comment: Shows routing stability and generation quality dissociate: full ensemble routing is the most numerically stable yet worst quality (FID 47.9 vs 22.6 for Top-2), and expert-data alignment is what actually governs quality.
Topic Match: It is a routing-principle study of sparse versus dense expert combination, including per-expert accuracy and disagreement analysis, though in decentralized diffusion models rather than LLMs.
Relevance: 6 Novelty: 6
Large-Scale Training Systems and Efficiency (7)
1. DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers
ArXiv ID: 2602.02016
Primary Topic: Large-Scale Training Systems and Efficiency
Authors: Ionut-Vlad Modoranu, Philip Zmushko, Erik Schultheis, Mher Safaryan, Dan Alistarh
Abstract: Shampoo is one of the leading approximate second-order optimizers: a variant of it has won the MLCommons AlgoPerf competition, and it has been shown to produce models with lower activation outliers that are easier to compress. Yet, applying Shampoo currently comes at the cost of significant computational slowdown, due to its expensive internal operations. In this paper, we take a significant step to address this shortcoming by proposing \method (for \textbf{D}istributed \textbf{A}ccelerated \textbf{SH}ampoo), a faster implementation of Distributed Shampoo based on two main new techniques: First, we show that preconditioner blocks can be stacked into 3D tensors to significantly improve GPU utilization; second, we introduce the Newton-DB iteration and the Chebyshev polynomial approximations as novel and faster approaches for computing the inverse matrix roots required by Shampoo. Along with these algorithmic contributions, we provide a first in-depth analysis of how matrix scaling critically affects Shampoo convergence. On the practical side, our GPU-aware implementation achieves up to $5.6\times$ faster optimizer steps compared to the well-optimized Distributed Shampoo, while Newton-DB attains the lowest validation perplexity per iteration among all tested methods. Our code is available at https://github.com/IST-DASLab/DASH.
Comment: Makes Distributed Shampoo practical: preconditioner blocks batched into 3D tensors for GPU utilisation, plus Newton-DB and Chebyshev inverse-root solvers and a matrix-scaling convergence analysis, for 5.6x faster optimizer steps.
Topic Match: A second-order preconditioner for large-scale pretraining with both algorithmic and GPU-implementation contributions, exactly the optimizer/systems target.
Relevance: 9 Novelty: 7
2. SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning
ArXiv ID: 2602.02472
Primary Topic: Large-Scale Training Systems and Efficiency
Also Matches: MoE Training, Architecture and Training Dynamics
Authors: Qifan Yu, Xinyu Ma, Zhijian Zhuo, Minrui Wang, Deyi Liu, Shiyi Zhan, Yiyuan Ma, Liang Xiang, Xingyan Bin, Di He
Abstract: Progressive Learning (PL) reduces pre-training computational overhead by gradually increasing model scale. While prior work has extensively explored depth expansion, width expansion remains significantly understudied, with the few existing methods limited to the early stages of training. However, expanding width during the mid-stage is essential for maximizing computational savings, yet it remains a formidable challenge due to severe training instabilities. Empirically, we show that naive initialization at this stage disrupts activation statistics, triggering loss spikes, while copy-based initialization introduces gradient symmetry that hinders feature diversity. To address these issues, we propose SPARKLING (balancing {S}ignal {P}reservation {A}nd symmet{R}y brea{K}ing for width-progressive {L}earn{ING}), a novel framework for mid-stage width expansion. Our method achieves signal preservation via RMS-scale consistency, stabilizing activation statistics during expansion. Symmetry breaking is ensured through asymmetric optimizer state reset and asymmetric learning rate re-warmup. Extensive experiments on dense and Mixture-of-Experts (MoE) models demonstrate that, across multiple width axes and optimizer families, SPARKLING consistently outperforms training from scratch and reduces training cost by up to 35% under $2\times$ width expansion.
Comment: Mid-training width expansion made stable via RMS-scale signal preservation plus asymmetric optimizer-state reset and LR re-warmup, cutting pretraining cost up to 35% on dense and MoE models.
Topic Match: It materially changes what a pretraining run costs and diagnoses the loss-spike and gradient-symmetry failure modes that block width growth.
Relevance: 9 Novelty: 7
3. The Effect of Mini-Batch Noise on the Implicit Bias of Adam
ArXiv ID: 2602.01642
Primary Topic: Large-Scale Training Systems and Efficiency
Also Matches: Architecture and Training Dynamics
Authors: Matias D. Cattaneo, Boris Shigida
Abstract: With limited high-quality data and growing compute, multi-epoch training is gaining back its importance across sub-areas of deep learning. Adam(W), versions of which are go-to optimizers for many tasks such as next token prediction, has two momentum hyperparameters $(β_1, β_2)$ controlling memory and one very important hyperparameter, batch size, controlling (in particular) the amount mini-batch noise. We introduce a theoretical framework to understand how mini-batch noise influences the implicit bias of memory in Adam (depending on $β_1$, $β_2$) towards sharper or flatter regions of the loss landscape, which is commonly observed to correlate with the generalization gap in multi-epoch training. We find that in the case of large batch sizes, higher $β_2$ increases the magnitude of anti-regularization by memory (hurting generalization), but as the batch size becomes smaller, the dependence of (anti-)regulariation on $β_2$ is reversed. A similar monotonicity shift (in the opposite direction) happens in $β_1$. In particular, the commonly "default" pair $(β_1, β_2) = (0.9, 0.999)$ is a good choice if batches are small; for larger batches, in many settings moving $β_1$ closer to $β_2$ is much better in terms of validation accuracy in multi-epoch training. Moreover, our theoretical derivations connect the scale of the batch size at which the shift happens to the scale of the critical batch size. We illustrate this effect in experiments with small-scale data in the about-to-overfit regime.
Comment: Theory for how mini-batch noise flips the direction in which Adam's beta1/beta2 bias training toward sharp or flat regions, tying the crossover batch size to the critical batch size.
Topic Match: It is optimizer analysis that directly informs how large runs set momentum and batch size, the configuration-facing side of training systems.
Relevance: 8 Novelty: 7
4. Decentralized SGD with Controlled Disagreement Finds Flatter Minima
ArXiv ID: 2602.02899
Primary Topic: Large-Scale Training Systems and Efficiency
Also Matches: Architecture and Training Dynamics
Authors: Zesen Wang, Mikael Johansson
Abstract: Decentralized training is often regarded as inferior to centralized training because the consensus errors between workers are thought to undermine convergence and generalization. This work challenges this view by introducing decentralized SGD with Adaptive Consensus (DSGD-AC), which uses a time-dependent scaling mechanism to maintain consensus errors throughout the training. We show that adaptive consensus changes the stationary variance of disagreement modes by balancing two effects: it preserves consensus-error magnitude through weaker graph damping while still allowing curvature-dependent damping to shape the disagreement directions. This balance can produce a stronger Hessian-weighted loss-envelope penalty around the deployed model, even when normalized Hessian alignment is weaker than in standard DSGD. Empirical results on image classification show that DSGD-AC reaches flatter solutions and higher test accuracy than standard DSGD and even centralized SGD. Together, these results support consensus errors as a useful implicit regularizer and open a new perspective on the design of decentralized learning algorithms.
Comment: Deliberately sustains consensus error with time-dependent scaling so weaker graph damping keeps disagreement magnitude while curvature-dependent damping shapes its directions, yielding flatter minima than centralized SGD.
Topic Match: A decentralized training algorithm with an analysis of how consensus error acts as implicit regularisation, directly in the distributed-training topic.
Relevance: 8 Novelty: 6
5. ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning
ArXiv ID: 2602.02192
Primary Topic: Large-Scale Training Systems and Efficiency
Authors: Jingwei Song, Meng Chen, Jie Xiao, Qingnan Ren, Jiaqi Huang, Yangshen Deng, Chris Tong, Wanyi Chen, Suli Wang, Zhisheng Chen, Ziqian Bi, Shuo Lu, Yiqun Duan, Xu Wang, Rymon Yu, Lynn Ai, Eric Yang, Tianyu Shi
Abstract: Reinforcement learning (RL) is a critical stage in post-training large language models (LLMs), involving repeated interaction between rollout generation, reward evaluation, and centralized learning. Distributing rollout execution offers opportunities to leverage more cost-efficient inference resources, but introduces challenges in wide-area coordination and policy dissemination. We present ECHO-2, a distributed RL framework for post-training with remote inference workers and non-negligible dissemination latency. ECHO-2 combines centralized learning with distributed rollouts and treats bounded policy staleness as a user-controlled parameter, enabling rollout generation, dissemination, and training to overlap. We introduce an overlap-based capacity model that relates training time, dissemination latency, and rollout throughput, yielding a practical provisioning rule for sustaining learner utilization. To mitigate dissemination bottlenecks and lower cost, ECHO-2 employs peer-assisted pipelined broadcast and cost-aware activation of heterogeneous workers. Experiments on GRPO post-training of LLMs ranging from 4B to 32B parameters under real wide-area bandwidth regimes show that ECHO-2 significantly improves cost efficiency while preserving RL reward comparable to strong baselines.
Comment: Treats bounded policy staleness as a tunable parameter and derives an overlap-based capacity model relating dissemination latency, training time and rollout throughput, with peer-assisted pipelined broadcast for wide-area workers.
Topic Match: A distributed-training system with a provisioning model and a communication/broadcast design; the RL setting does not change that the contribution is systems-level.
Relevance: 7 Novelty: 6
6. Stein-Rule Shrinkage for Stochastic Gradient Estimation in High Dimensions
ArXiv ID: 2602.01777
Primary Topic: Large-Scale Training Systems and Efficiency
Authors: M. Arashi, M. Amintoosi
Abstract: Stochastic gradient methods are central to large-scale learning, but they treat mini-batch gradients as unbiased estimators, which classical decision theory shows are inadmissible in high dimensions. We formulate gradient computation as a high-dimensional estimation problem and introduce a framework based on Stein-rule shrinkage. We construct a gradient estimator that adaptively contracts noisy mini-batch gradients toward a stable estimator derived from historical momentum. The shrinkage intensity is determined in a data-driven manner using an online estimate of gradient noise variance, leveraging statistics from adaptive optimizers. Under a Gaussian noise model, we show our estimator uniformly dominates the standard stochastic gradient under squared error loss and is minimax-optimal. We incorporate this into the Adam optimizer, yielding SR-Adam, a practical algorithm with negligible computational cost. Empirical evaluations on CIFAR10 and CIFAR100 across multiple levels of input noise show consistent improvements over Adam in the large-batch regime. Ablation studies indicate that gains arise primarily from selectively applying shrinkage to high-dimensional convolutional layers, while indiscriminate shrinkage across all parameters degrades performance. These results illustrate that classical shrinkage principles provide a principled approach to improving stochastic gradient estimation in deep learning.
Comment: Noise-adaptive Stein shrinkage contracts minibatch gradients toward momentum within Adam.
Topic Match: The core contribution is a gradient estimator and optimizer update, although validation is limited to CNNs rather than large-model pretraining.
Relevance: 7 Novelty: 6
7. Grappa: Gradient-Only Communication for Scalable Graph Neural Network Training
ArXiv ID: 2602.01872
Primary Topic: Large-Scale Training Systems and Efficiency
Authors: Chongyang Xu, Christoph Siebenbrunner, Laurent Bindschaedler
Abstract: Cross-partition edges dominate the cost of distributed GNN training: fetching remote features and activations per iteration overwhelms the network as graphs deepen and partition counts grow. Grappa is a distributed GNN training framework that enforces gradient-only communication: during each iteration, partitions train in isolation and exchange only gradients for the global update. To recover accuracy lost to isolation, Grappa (i) periodically repartitions to expose new neighborhoods and (ii) applies a lightweight coverage-corrected gradient aggregation inspired by importance sampling. We present an asymptotically unbiased estimator for gradient correction, which we use to develop a minimum-distance batch-level variant that is compatible with common deep-learning packages. We also introduce a shrinkage version that improves stability in practice. Empirical results on real and synthetic graphs show that Grappa trains GNNs 4x faster on average (up to 13x) than state-of-the-art systems, achieves better accuracy especially for deeper models, and sustains training at the trillion-edge scale on commodity hardware. Grappa is model-agnostic, supports full-graph and mini-batch training, and does not rely on high-bandwidth interconnects or caching.
Comment: Eliminates per-iteration remote feature and activation fetches entirely, training partitions in isolation and exchanging only gradients, with periodic repartitioning plus an asymptotically unbiased coverage-corrected aggregation estimator.
Topic Match: A distributed-training communication algorithm with a gradient-correction estimator; the GNN setting narrows it but the idea is systems-level.
Relevance: 6 Novelty: 6
Architecture and Training Dynamics (14)
1. Softmax Linear Attention: Reclaiming Global Competition
ArXiv ID: 2602.01744
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Mingwei Xu, Xuan Lin, Xinnan Guo, Wanqing Xu, Wanyun Cui
Abstract: While linear attention reduces the quadratic complexity of standard Transformers to linear time, it often lags behind in expressivity due to the removal of softmax normalization. This omission eliminates \emph{global competition}, a critical mechanism that enables models to sharply focus on relevant information amidst long-context noise. In this work, we propose \textbf{Softmax Linear Attention (SLA)}, a framework designed to restore this competitive selection without sacrificing efficiency. By lifting the softmax operation from the token level to the head level, SLA leverages attention heads as coarse semantic slots, applying a competitive gating mechanism to dynamically select the most relevant subspaces. This reintroduces the ``winner-take-all'' dynamics essential for precise retrieval and robust long-context understanding. Distinct from prior methods that focus on refining local kernel functions, SLA adopts a broader perspective by exploiting the higher-level multi-head aggregation structure. Extensive experiments demonstrate that SLA consistently enhances state-of-the-art linear baselines (RetNet, GLA, GDN) across language modeling and long-context benchmarks, particularly in challenging retrieval scenarios where it significantly boosts robustness against noise, validating its capability to restore precise focus while maintaining linear complexity.
Comment: Lifts softmax from token level to head level so linear attention regains winner-take-all global competition, improving RetNet/GLA/GDN on long-context retrieval.
Topic Match: A new attention mechanism for linear-complexity sequence modelling, exactly the architectural-mechanism topic.
Relevance: 9 Novelty: 7
2. Poly-attention: a general scheme for higher-order self-attention
ArXiv ID: 2602.02422
Primary Topic: Architecture and Training Dynamics
Authors: Sayak Chakrabarti, Toniann Pitassi, Josh Alman
Abstract: The self-attention mechanism, at the heart of the Transformer model, is able to effectively model pairwise interactions between tokens. However, numerous recent works have shown that it is unable to perform basic tasks involving detecting triples of correlated tokens, or compositional tasks where multiple input tokens need to be referenced to generate a result. Some higher-dimensional alternatives to self-attention have been proposed to address this, including higher-order attention and Strassen attention, which can perform some of these polyadic tasks in exchange for slower, superquadratic running times. In this work, we define a vast class of generalizations of self-attention, which we call poly-attention mechanisms. Our mechanisms can incorporate arbitrary higher-order (tensor) computations as well as arbitrary relationship structures between the input tokens, and they include the aforementioned alternatives as special cases. We then systematically study their computational complexity and representational strength, including giving new algorithms and matching complexity-theoretic lower bounds on the time complexity of computing the attention matrix exactly as well as approximately, and tightly determining which polyadic tasks they can each perform. Our results give interesting trade-offs between different desiderata for these mechanisms, including a tight relationship between how expressive a mechanism is, and how large the coefficients in the model may be so that the mechanism can be approximated in almost-linear time. Notably, we give a new attention mechanism which can be computed exactly in quadratic time, and which can perform function composition for any fixed number of functions. Prior mechanisms, even for just composing two functions, could only be computed in superquadratic time, and our new lower bounds show that faster algorithms for them are not possible.
Comment: Defines a general class of higher-order attention mechanisms with matching upper and lower bounds, including a new mechanism computable exactly in quadratic time that can do function composition.
Topic Match: Core contribution is an attention-mechanism family plus complexity/expressiveness characterisation, squarely an architectural-mechanism result.
Relevance: 8 Novelty: 8
3. A Provable Expressiveness Hierarchy in Hybrid Linear-Full Attention
ArXiv ID: 2602.01763
Primary Topic: Architecture and Training Dynamics
Authors: Xiaowei Ye, Xiaoyu He, Chao Liao, Chen Wu, Pinyan Lu
Abstract: Transformers serve as the foundation of most modern large language models. To mitigate the quadratic complexity of standard full attention, various efficient attention mechanisms, such as linear and hybrid attention, have been developed. A fundamental gap remains: their expressive power relative to full attention lacks a rigorous theoretical characterization. In this work, we theoretically characterize the performance differences among these attention mechanisms. Our theory applies to all linear attention variants that can be formulated as a recurrence, including Mamba, DeltaNet, etc. Specifically, we establish an expressiveness hierarchy: for the sequential function composition-a multi-step reasoning task that must occur within a model's forward pass, an ($L+1$)-layer full attention network is sufficient, whereas any hybrid network interleaving $L-1$ layers of full attention with a substantially larger number ($2^{3L^2}$) of linear attention layers cannot solve it. This result demonstrates a clear separation in expressive power between the two types of attention. Our work provides the first provable separation between hybrid attention and standard full attention, offering a theoretical perspective for understanding the fundamental capabilities and limitations of different attention mechanisms.
Comment: First provable separation between hybrid linear-full attention and full attention: L-1 full layers plus exponentially many linear layers cannot do sequential function composition that L+1 full layers can.
Topic Match: An expressiveness hierarchy over attention variants formulated as recurrences (Mamba, DeltaNet) is core architectural-mechanism analysis.
Relevance: 8 Novelty: 8
4. An Empirical Study on Noisy Data and LLM Pretraining Loss Divergence
ArXiv ID: 2602.02400
Primary Topic: Architecture and Training Dynamics
Also Matches: Large-Scale Training Systems and Efficiency
Authors: Qizhen Zhang, Ankush Garg, Jakob Foerster, Niladri Chatterji, Kshitiz Malik, Mike Lewis
Abstract: Large-scale pretraining datasets drive the success of large language models (LLMs). However, these web-scale corpora inevitably contain large amounts of noisy data due to unregulated web content or randomness inherent in data. Although LLM pretrainers often speculate that such noise contributes to instabilities in large-scale LLM pretraining and, in the worst cases, loss divergence, this phenomenon remains poorly understood.In this work, we present a systematic empirical study of whether noisy data causes LLM pretraining divergences and how it does so. By injecting controlled synthetic uniformly random noise into otherwise clean datasets, we analyze training dynamics across model sizes ranging from 480M to 5.2B parameters. We show that noisy data indeed induces training loss divergence, and that the probability of divergence depends strongly on the noise type, amount of noise, and model scale. We further find that noise-induced divergences exhibit activation patterns distinct from those caused by high learning rates, and we provide diagnostics that differentiate these two failure modes. Together, these results provide a large-scale, controlled characterization of how noisy data affects loss divergence in LLM pretraining.
Comment: Controlled noise injection across 480M-5.2B models shows how noise type, dose, and scale drive pretraining loss divergence, with activation diagnostics separating it from LR-induced divergence.
Topic Match: A controlled study of large-model training instability and its failure signatures, which is the training-dynamics topic rather than any downstream task.
Relevance: 8 Novelty: 6
5. State Rank Dynamics in Linear Attention LLMs
ArXiv ID: 2602.02195
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Ao Sun, Hongtao Zhang, Heng Zhou, Yixuan Ma, Yiran Qin, Tongrui Su, Yan Liu, Zhanyu Ma, Jun Xu, Jiuchong Gao, Jinghua Hao, Renqing He
Abstract: Linear Attention Large Language Models (LLMs) offer a compelling recurrent formulation that compresses context into a fixed-size state matrix, enabling constant-time inference. However, the internal dynamics of this compressed state remain largely opaque. In this work, we present a comprehensive study on the runtime state dynamics of state-of-the-art Linear Attention models. We uncover a fundamental phenomenon termed State Rank Stratification, characterized by a distinct spectral bifurcation among linear attention heads: while one group maintains an effective rank oscillating near zero, the other exhibits rapid growth that converges to an upper bound. Extensive experiments across diverse inference contexts reveal that these dynamics remain strikingly consistent, indicating that the identity of a head,whether low-rank or high-rank,is an intrinsic structural property acquired during pre-training, rather than a transient state dependent on the input data. Furthermore, our diagnostic probes reveal a surprising functional divergence: low-rank heads are indispensable for model reasoning, whereas high-rank heads exhibit significant redundancy. Leveraging this insight, we propose Joint Rank-Norm Pruning, a zero-shot strategy that achieves a 38.9\% reduction in KV-cache overhead while largely maintaining model accuracy.
Comment: Finds a spectral bifurcation of linear-attention heads into persistent low- and high-rank groups fixed at pretraining, where low-rank heads carry reasoning and high-rank ones are redundant, enabling 38.9% state-cache pruning.
Topic Match: The substance is a mechanistic property of the recurrent state in linear attention, with pruning as a downstream consequence.
Relevance: 7 Novelty: 6
6. CoMeT: Collaborative Memory Transformer for Efficient Long Context Modeling
ArXiv ID: 2602.01766
Primary Topic: Architecture and Training Dynamics
Also Matches: Large-Scale Training Systems and Efficiency, Efficiency, Compression, and Large-Scale Training
Authors: Runsong Zhao, Shilei Liu, Jiwei Tang, Langming Liu, Haibin Chen, Weidong Zhang, Yujin Yuan, Tong Xiao, Jingbo Zhu, Wenbo Su, Bo Zheng
Abstract: The quadratic complexity and indefinitely growing key-value (KV) cache of standard Transformers pose a major barrier to long-context processing. To overcome this, we introduce the Collaborative Memory Transformer (CoMeT), a novel architecture that enables LLMs to handle arbitrarily long sequences with constant memory usage and linear time complexity. Designed as an efficient, plug-in module, CoMeT can be integrated into pre-trained models with only minimal fine-tuning. It operates on sequential data chunks, using a dual-memory system to manage context: a temporary memory on a FIFO queue for recent events, and a global memory with a gated update rule for long-range dependencies. These memories then act as a dynamic soft prompt for the next chunk. To enable efficient fine-tuning on extremely long contexts, we introduce a novel layer-level pipeline parallelism strategy. The effectiveness of our approach is remarkable: a model equipped with CoMeT and fine-tuned on 32k contexts can accurately retrieve a passkey from any position within a 1M token sequence. On the SCROLLS benchmark, CoMeT surpasses other efficient methods and achieves performance comparable to a full-attention baseline on summarization tasks. Its practical effectiveness is further validated on real-world agent and user behavior QA tasks. The code is available at: https://github.com/LivingFutureLab/Comet
Comment: Chunked dual-memory module (FIFO temporary memory plus gated global memory acting as a soft prompt) giving constant memory and linear time, trained long-context via a layer-level pipeline parallelism strategy.
Topic Match: The core is a recurrent-memory sequence-modelling mechanism that removes KV growth, with the pipeline-parallel recipe as supporting systems work.
Relevance: 7 Novelty: 6
7. HopFormer: Sparse Graph Transformers with Explicit Receptive Field Control
ArXiv ID: 2602.02268
Primary Topic: Architecture and Training Dynamics
Authors: Sanggeon Yun, Raheeb Hassan, Ryozo Masukawa, Sungheon Jeong, Mohsen Imani
Abstract: Graph Transformers typically rely on explicit positional or structural encodings and dense global attention to incorporate graph topology. In this work, we show that neither is essential. We introduce HopFormer, a graph Transformer that injects structure exclusively through head-specific n-hop masked sparse attention, without the use of positional encodings or architectural modifications. This design provides explicit and interpretable control over receptive fields while enabling genuinely sparse attention whose computational cost scales linearly with mask sparsity. Through extensive experiments on both node-level and graph-level benchmarks, we demonstrate that our approach achieves competitive or superior performance across diverse graph structures. Our results further reveal that dense global attention is often unnecessary: on graphs with strong small-world properties, localized attention yields more stable and consistently high performance, while on graphs with weaker small-world effects, global attention offers diminishing returns. Together, these findings challenge prevailing assumptions in graph Transformer design and highlight sparsity-controlled attention as a principled and efficient alternative.
Comment: Head-specific n-hop sparse attention controls receptive fields without explicit positional encodings.
Topic Match: The core contribution is a structural attention mechanism with mechanistic findings about locality, although its evidence is confined to graph Transformers.
Relevance: 7 Novelty: 6
8. Training-Free Self-Correction for Multimodal Masked Diffusion Models
ArXiv ID: 2602.02927
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Yidong Ouyang, Panwen Hu, Zhengyan Wan, Zhe Wang, Liyan Xie, Dmitriy Bespalov, Ying Nian Wu, Guang Cheng, Hongyuan Zha, Qiang Sun
Abstract: Masked diffusion models have emerged as a powerful framework for text and multimodal generation. However, their sampling procedure updates multiple tokens simultaneously and treats generated tokens as immutable, which may lead to error accumulation when early mistakes cannot be revised. In this work, we revisit existing self-correction methods and identify limitations stemming from additional training requirements or reliance on misaligned likelihood estimates. We propose a training-free self-correction framework that exploits the inductive biases of pre-trained masked diffusion models. Without modifying model parameters or introducing auxiliary evaluators, our method significantly improves generation quality on text-to-image generation and multimodal understanding tasks with reduced sampling steps. Moreover, the proposed framework generalizes across different masked diffusion architectures, highlighting its robustness and practical applicability. Code can be found in https://github.com/huge123/FreeCorrection.
Comment: Training-free self-correction improves masked-diffusion generation while reducing sampling steps.
Topic Match: The contribution changes the masked-diffusion sampling procedure across backbones, with an accompanying reduction in generation cost.
Relevance: 7 Novelty: 6
9. Breaking the Reversal Curse in Autoregressive Language Models via Identity Bridge
ArXiv ID: 2602.02470
Primary Topic: Architecture and Training Dynamics
Authors: Xutao Ma, Yixiao Huang, Hanlin Zhu, Somayeh Sojoudi
Abstract: Autoregressive large language models (LLMs) have achieved remarkable success in many complex tasks, yet they can still fail in very simple logical reasoning such as the "reversal curse" -- when trained on forward knowledge data of the form "$A \rightarrow B$" (e.g., Alice's husband is Bob), the model is unable to deduce the reversal knowledge "$B \leftarrow A$" (e.g., Bob's wife is Alice) during test. Extensive prior research suggests that this failure is an inherent, fundamental limit of autoregressive causal LLMs, indicating that these models tend to memorize factual-level knowledge rather than capture higher-level rules. In this paper, we challenge this view by showing that this seemingly fundamental limit can be mitigated by slightly tweaking the training data with a simple regularization data recipe called the Identity Bridge of the form "$A \to A$" (e.g., The name of Alice is Alice). Theoretically, we prove that under this recipe, even a one-layer transformer can break the reversal curse by analyzing the implicit bias of gradient descent. Empirically, we show that a 1B pretrained language model finetuned with the proposed data recipe achieves a 50% success rate on reversal tasks, in stark contrast to a near-zero success rate when trained solely on forward-knowledge data. Our work provides a novel theoretical foundation for the reversal curse and offers a principled, low-cost path to encouraging LLMs to learn higher-level rules from data.
Comment: Shows an 'A to A' identity-bridge data recipe breaks the reversal curse, with a proof via the implicit bias of gradient descent in a one-layer transformer.
Topic Match: The mechanism is an implicit-bias/training-dynamics argument about what gradient descent stores, not a task-specific fix.
Relevance: 6 Novelty: 7
10. On the Spatiotemporal Dynamics of Generalization in Neural Networks
ArXiv ID: 2602.01651
Primary Topic: Architecture and Training Dynamics
Authors: Zichao Wei
Abstract: Why do neural networks fail to generalize addition from 16-digit to 32-digit numbers, while a child who learns the rule can apply it to arbitrarily long sequences? We argue that this failure is not an engineering problem but a violation of physical postulates. Drawing inspiration from physics, we identify three constraints that any generalizing system must satisfy: (1) Locality -- information propagates at finite speed; (2) Symmetry -- the laws of computation are invariant across space and time; (3) Stability -- the system converges to discrete attractors that resist noise accumulation. From these postulates, we derive -- rather than design -- the Spatiotemporal Evolution with Attractor Dynamics (SEAD) architecture: a neural cellular automaton where local convolutional rules are iterated until convergence. Experiments on three tasks validate our theory: (1) Parity -- demonstrating perfect length generalization via light-cone propagation; (2) Addition -- achieving scale-invariant inference from L=16 to L=1 million with 100% accuracy, exhibiting input-adaptive computation; (3) Rule 110 -- learning a Turing-complete cellular automaton without trajectory divergence. Our results suggest that the gap between statistical learning and logical reasoning can be bridged -- not by scaling parameters, but by respecting the physics of computation.
Comment: Derives a neural cellular-automaton architecture from locality, symmetry and attractor-stability postulates, achieving exact length generalisation on addition from L=16 to L=1M.
Topic Match: The contribution is an architectural mechanism with input-adaptive iterated computation, derived rather than tuned.
Relevance: 6 Novelty: 7
11. A Theoretical Analysis of Why Masked Diffusion Models Mitigate the Reversal Curse
ArXiv ID: 2602.02133
Primary Topic: Architecture and Training Dynamics
Authors: Moongyu Jeon, Sangwoo Shin, BumJun Kim, Kyelim Lee, Albert No
Abstract: Autoregressive language models (ARMs) suffer from the reversal curse: after learning ''$A$ is $B$,'' they often fail on the reverse query ''$B$ is $A$.'' Masked diffusion language models (MDMs) exhibit this failure in a much weaker form, but the underlying reason has remained unclear. A common explanation attributes this mitigation to their any-order masked training objective. However, observing ''$[\mathbf{M}]$ is $B$'' during training teaches recovery of $A$ from $B$ in one positional configuration, and does not by itself explain why the learned evidence should transfer to the reverse prompt ''$B$ is $[\mathbf{M}]$.'' We provide a theoretical analysis showing that this transfer arises from a parameter-level coupling between forward and reverse positional conditionals: shared Transformer parameters store token-pair evidence, while relative positional encodings route attention through queries and keys without changing the value-side evidence being retrieved. In a one-layer MDM, we prove that forward masked training strengthens evidence that is reusable in reverse queries, induces correlated forward--reverse attention routes, and yields a positively aligned shared-storage gradient component that decreases the reverse loss to first order. Controlled one-layer experiments and large-scale LLaDA/Dream experiments verify these signatures and show that they translate into improved reverse prediction.
Comment: Attributes masked diffusion models' weaker reversal curse to parameter-level coupling: shared parameters store token-pair evidence on the value side while relative positional encodings reroute queries and keys, giving a positively aligned reverse-loss gradient.
Topic Match: A mechanistic training-dynamics account of how a training objective and positional design interact, verified from one-layer models up to LLaDA/Dream.
Relevance: 6 Novelty: 7
12. Beyond Precision: Training-Inference Mismatch is an Optimization Problem and Simple LR Scheduling Fixes It
ArXiv ID: 2602.01826
Primary Topic: Architecture and Training Dynamics
Also Matches: Large-Scale Training Systems and Efficiency
Authors: Yaxiang Zhang, Yingru Li, Jiacai Liu, Jiawei Xu, Ziniu Li, Qian Liu, Haoyuan Li
Abstract: Reinforcement Learning (RL) for training Large Language Models is notoriously unstable. While recent studies attribute this to "training inference mismatch stemming" from inconsistent hybrid engines, standard remedies, such as Importance Sampling, might fail during extended training runs. In this work, we analyze this instability through the lens of optimization, demonstrating that gradient noise and training-inference mismatch escalate in tandem as training progresses. Meanwhile, we find that the mismatch can be effectively suppressed by shrinking the update size. Taken together, we deduce that the mismatch is not merely a static numerical discrepancy, but a dynamic failure coupled with the model's optimization. Based on this insight, we propose a simple yet effective solution: a specialized Learning Rate (LR) scheduler. Instead of pre-defined decay schedule in traditional LR scheduler, our method dynamically triggers LR decay based on response length, which we identify as a reliable early-warning signal for impending instability. Empirical evidence suggests that by reducing the learning rate as gradient noise rises, we can consistently stabilize RL training and keep the training-inference mismatch at a safe level.
Comment: Reframes training-inference mismatch as an optimisation failure coupled to rising gradient noise, and stabilises it with a response-length-triggered learning-rate decay rather than importance sampling.
Topic Match: The core claim is a training-dynamics diagnosis (update size and gradient noise drive the mismatch) with an LR-schedule remedy, so training dynamics fits best.
Relevance: 6 Novelty: 6
13. Late-Stage Generalization Collapse in Grokking: Detecting anti-grokking with Weightwatcher
ArXiv ID: 2602.02859
Primary Topic: Architecture and Training Dynamics
Authors: Hari K Prakash, Charles H Martin
Abstract: \emph{Memorization} in neural networks lacks a precise operational definition and is often inferred from the grokking regime, where training accuracy saturates while test accuracy remains very low. We identify a previously unreported third phase of grokking in this training regime: \emph{anti-grokking}, a late-stage collapse of generalization. We revisit two canonical grokking setups: a 3-layer MLP trained on a subset of MNIST and a transformer trained on modular addition, but extended training far beyond standard. In both cases, after models transition from pre-grokking to successful generalization, test accuracy collapses back to chance while training accuracy remains perfect, indicating a distinct post-generalization failure mode. To diagnose anti-grokking, we use the open-source \texttt{WeightWatcher} tool based on HTSR/SETOL theory. The primary signal is the emergence of \emph{Correlation Traps}: anomalously large eigenvalues beyond the Marchenko--Pastur bulk in the empirical spectral density of shuffled weight matrices, which are predicted to impair generalization. As a secondary signal, anti-grokking corresponds to the average HTSR layer quality metric $α$ deviating from $2.0$. Neither metric requires access to the test or training data. We compare these signals to alternative grokking diagnostic, including $\ell_2$ norms, Activation Sparsity, Absolute Weight Entropy, and Local Circuit Complexity. These track pre-grokking and grokking but fail to identify anti-grokking. Finally, we show that Correlation Traps can induce catastrophic forgetting and/or prototype memorization, and observe similar pathologies in large-scale LLMs, like OSS GPT 20/120B.
Comment: Identifies a late-stage 'anti-grokking' collapse where test accuracy returns to chance, detected by correlation traps in the shuffled-weight spectral density without touching train or test data.
Topic Match: A previously unreported training-dynamics phase plus a data-free weight-spectrum diagnostic for it.
Relevance: 6 Novelty: 6
14. Geometric Analysis of Token Selection in Multi-Head Attention
ArXiv ID: 2602.01893
Primary Topic: Architecture and Training Dynamics
Authors: Timur Mudarisov, Mikhal Burtsev, Tatiana Petrova, Radu State
Abstract: We present a geometric framework for analysing multi-head attention in large language models (LLMs). Without altering the mechanism, we view standard attention through a top-N selection lens and study its behaviour directly in value-state space. We define geometric metrics - Precision, Recall, and F-score - to quantify separability between selected and non-selected tokens, and derive non-asymptotic bounds with explicit dependence on dimension and margin under empirically motivated assumptions (stable value norms with a compressed sink token, exponential similarity decay, and piecewise attention weight profiles). The theory predicts a small-N operating regime of strongest non-trivial separability and clarifies how sequence length and sink similarity shape the metrics. Empirically, across LLaMA-2-7B, Gemma-7B, and Mistral-7B, measurements closely track the theoretical envelopes: top-N selection sharpens separability, sink similarity correlates with Recall. We also found that in LLaMA-2-7B heads specialize into three regimes - Retriever, Mixer, Reset - with distinct geometric signatures. Overall, attention behaves as a structured geometric classifier with measurable criteria for token selection, offering head level interpretability and informing geometry-aware sparsification and design of attention in LLMs.
Comment: Geometric theory of top-N token selection in attention value-space, with non-asymptotic separability bounds that inform attention sparsification design.
Topic Match: Analyses the attention mechanism itself and derives head-level structure, though framed largely as interpretability rather than a new training mechanism.
Relevance: 6 Novelty: 6
Efficiency, Compression, and Large-Scale Training (16)
1. Dissecting Outlier Dynamics in LLM NVFP4 Pretraining
ArXiv ID: 2602.02047
Primary Topic: Efficiency, Compression, and Large-Scale Training
Also Matches: Large-Scale Training Systems and Efficiency, Architecture and Training Dynamics
Authors: Peijie Dong, Ruibo Fan, Yuechen Tao, Di Mou, Wenhu Hu, Zhenheng Tang, Yinghao Yu, Jiamang Wang, Wenbo Su, Guodong Yang, Liping Zhang, Xiaowen Chu, Baochun Li, Bo Li
Abstract: Training large language models using 4-bit arithmetic enhances throughput and memory efficiency. Yet, the limited dynamic range of FP4 increases sensitivity to outliers. While NVFP4 mitigates quantization error via hierarchical microscaling, a persistent loss gap remains compared to BF16. This study conducts a longitudinal analysis of outlier dynamics across architecture during NVFP4 pretraining, focusing on where they localize, why they occur, and how they evolve temporally. We find that, compared with Softmax Attention (SA), Linear Attention (LA) reduces per-tensor heavy tails but still exhibits persistent block-level spikes under block quantization. Our analysis attributes outliers to specific architectural components: Softmax in SA, gating in LA, and SwiGLU in FFN, with "post-QK" operations exhibiting higher sensitivity to quantization. Notably, outliers evolve from transient spikes early in training to a small set of persistent hot channels (i.e., channels with persistently large magnitudes) in later stages. Based on these findings, we introduce Hot-Channel Patch (HCP), an online compensation mechanism that identifies hot channels and reinjects residuals using hardware-efficient kernels. We then develop CHON, an NVFP4 training recipe integrating HCP with post-QK operation protection. On GLA-1.3B model trained for 60B tokens, CHON reduces the loss gap to BF16 from 0.94% to 0.58% while maintaining downstream accuracy.
Comment: Longitudinal study of where NVFP4 pretraining outliers localise (Softmax in softmax attention, gating in linear attention, SwiGLU in FFN, with post-QK most sensitive) and how transient spikes become persistent hot channels, fixed by online residual reinjection.
Topic Match: Low-precision pretraining is compression work that directly changes training cost, and the analysis plus hardware-efficient compensation kernel targets the BF16 loss gap in an actual training recipe.
Relevance: 9 Novelty: 7
2. Revisiting Adaptive Rounding with Vectorized Reparameterization for LLM Quantization
ArXiv ID: 2602.02151
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Yuli Zhou, Qingxuan Chen, Luca Benini, Guolei Sun, Yawei Li
Abstract: Adaptive Rounding has emerged as an alternative to round-to-nearest (RTN) for post-training quantization by enabling cross-element error cancellation. Yet, dense and element-wise rounding matrices are prohibitively expensive for billion-parameter large language models (LLMs). We revisit adaptive rounding from an efficiency perspective and propose VQRound, a parameter-efficient optimization framework that reparameterizes the rounding matrix into a compact codebook. Unlike low-rank alternatives, VQRound minimizes the element-wise worst-case error under $L_\infty$ norm, which is critical for handling heavy-tailed weight distributions in LLMs. Beyond reparameterization, we identify rounding initialization as a decisive factor and develop a lightweight end-to-end finetuning pipeline that optimizes codebooks across all layers using only 128 samples. Extensive experiments on OPT, LLaMA, LLaMA2, and Qwen3 models demonstrate that VQRound achieves better convergence than traditional adaptive rounding at the same number of steps while using as little as 0.2% of the trainable parameters. Our results show that adaptive rounding can be made both scalable and fast-fitting. The code is available at https://github.com/zhoustan/VQRound.
Comment: Compact codebooks replace element-wise adaptive-rounding parameters to make LLM quantization cheaper to optimize.
Topic Match: The core contribution is a quantization parameterization that reduces optimization memory while controlling rounding error.
Relevance: 9 Novelty: 7
3. Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
ArXiv ID: 2602.02001
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Yoonjun Cho, Dongjae Jeon, Soeun Kim, Moongyu Jeon, Albert No
Abstract: Quantization Error Reconstruction (QER) reduces accuracy loss in Post-Training Quantization (PTQ) by approximating weights as $\mathbf{W} \approx \mathbf{Q} + \mathbf{L}\mathbf{R}$, using a rank-$r$ correction to reconstruct quantization error. Prior methods devote the full rank budget to error reconstruction, which is suboptimal when $\mathbf{W}$ has intrinsic low-rank structure and quantization corrupts dominant directions. We propose Structured Residual Reconstruction (SRR), a rank-allocation framework that preserves the top-$k$ singular subspace of the activation-scaled weight before quantization, quantizes only the residual, and uses the remaining rank $r-k$ for error reconstruction. We derive a theory-guided criterion for selecting $k$ by balancing quantization-exposed energy and unrecoverable error under rank constraints. We further show that resulting $\mathbf{Q} + \mathbf{L}\mathbf{R}$ parameterization naturally supports Quantized Parameter-Efficient Fine-Tuning (QPEFT), and stabilizes fine-tuning via gradient scaling along preserved directions. Experiments demonstrate consistent perplexity reductions across diverse models and quantization settings in PTQ, along with a 5.9 percentage-point average gain on GLUE under 2-bit QPEFT. The project page is available at https://ai-isl.github.io/srr.
Comment: Allocates a fixed low-rank budget between preserving dominant weight directions and reconstructing quantization error.
Topic Match: Rank allocation directly improves low-bit weight reconstruction and supports more stable quantized parameter-efficient fine-tuning.
Relevance: 9 Novelty: 7
4. More Than a Quick Glance: Overcoming the Greedy Bias in KV-Cache Compression
ArXiv ID: 2602.02199
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Aryan Sood, Tanvi Sharma, Vansh Agrawal
Abstract: While Large Language Models (LLMs) can theoretically support extensive context windows, their actual deployment is constrained by the linear growth of Key-Value (KV) cache memory. Prevailing compression strategies mitigate this through various pruning mechanisms, yet trade-off semantic recall for memory efficiency. In this work, we present LASER-KV (Layer Accumulated Selection with Exact-LSH Recall), a framework designed to test the limits of KV compression under a strict accumulative budgeting policy. We deviate from the standard fixed summary size approach by implementing a block-wise accumulation strategy governed by a protection divisor (n). This allows us to isolate the effects of compression from sliding window artifacts. Our experiments on the Babilong benchmark reveal performance degradation in previous compression methods by 15-30% on various long context tasks. LASER-KV maintains stable performance, achieving superior accuracies by a margin of upto 10% at 128k. These findings challenge the prevailing assumption that attention scores alone are a sufficient proxy for token utility.
Comment: Block-wise accumulation changes KV-cache retention under explicit compression budgets.
Topic Match: KV-cache selection and memory budgeting are the central mechanisms, directly addressing long-context compression and recall.
Relevance: 9 Novelty: 6
5. Every Bit Counts: A Theoretical Study of Precision-Expressivity Tradeoffs in Quantized Transformers
ArXiv ID: 2602.02707
Primary Topic: Efficiency, Compression, and Large-Scale Training
Also Matches: Architecture and Training Dynamics
Authors: Sayak Chakrabarti, Toniann Pitassi, Josh Alman
Abstract: Quantization reduces the numerical precision of Transformer computations and is widely used to accelerate inference, yet its effect on expressivity remains poorly characterized. We demonstrate a fine-grained theoretical tradeoff between expressivity and precision: For every p we exhibit a function Γ, inspired by the equality function, and prove that a one-layer softmax Transformer can compute Γ, with p bits of precision, but not with p-1 bits of precision. This result concretely explains the widely observed phenomenon of empirical loss of expressivity when quantization is used. Practically, it suggests that tasks requiring equality-like comparisons (exact match, membership, etc.) are especially sensitive to quantization. Dropping even one bit can cross a threshold where the model cannot represent the needed comparison reliably. Thus, it paves the way for developing heuristics that will help practitioners choose how much quantization is possible: the precision should be chosen as a function of the length of equality to be checked for the specific task. Our proofs combine explicit finite-precision Transformer constructions with communication-complexity lower bounds, yielding a tight "one-bit" threshold.
Comment: Proves a tight one-bit expressivity threshold for one-layer softmax Transformers via communication-complexity lower bounds, giving a principled rule for choosing quantization precision.
Topic Match: A theoretical account of what quantization costs in representational power, directly grounding the compression topic.
Relevance: 8 Novelty: 7
6. BAPS: A Fine-Grained Low-Precision Scheme for Softmax in Attention via Block-Aware Precision reScaling
ArXiv ID: 2602.02071
Primary Topic: Efficiency, Compression, and Large-Scale Training
Also Matches: Architecture and Training Dynamics
Authors: Zisheng Ye, Xiaoyu He, Maoyuan Song, Guoliang Qiu, Chao Liao, Chen Wu, Yonggang Sun, Zhichun Li, Xiaoru Xie, Yuanyong Luo, Hu Liu, Pinyan Lu, Heng Liao
Abstract: As the performance gains from accelerating quantized matrix multiplication plateau, the softmax operation becomes the critical bottleneck in Transformer inference. This bottleneck stems from two hardware limitations: (1) limited data bandwidth between matrix and vector compute cores, and (2) the significant area cost of high-precision (FP32/16) exponentiation units (EXP2). To address these issues, we introduce a novel low-precision workflow that employs a specific 8-bit floating-point format (HiF8) and block-aware precision rescaling for softmax. Crucially, our algorithmic innovations make low-precision softmax feasible without the significant model accuracy loss that hampers direct low-precision approaches. Specifically, our design (i) halves the required data movement bandwidth by enabling matrix multiplication outputs constrained to 8-bit, and (ii) substantially reduces the EXP2 unit area by computing exponentiations in low (8-bit) precision. Extensive evaluation on language models and multi-modal models confirms the validity of our method. By alleviating the vector computation bottleneck, our work paves the way for doubling end-to-end inference throughput without increasing chip area, and offers a concrete co-design path for future low-precision hardware and software.
Comment: Block-aware precision rescaling with an 8-bit float format makes softmax low-precision without accuracy loss, halving matrix-to-vector bandwidth and shrinking EXP2 unit area.
Topic Match: A new quantization scheme for the attention softmax itself plus hardware co-design, attacking the bottleneck left after GEMM quantization.
Relevance: 8 Novelty: 7
7. Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention
ArXiv ID: 2602.01801
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Dvir Samuel, Issar Tzachor, Matan Levy, Michael Green, Gal Chechik, Rami Ben-Ari
Abstract: Autoregressive video diffusion models enable streaming generation, opening the door to long-form synthesis, video world models, and interactive neural game engines. However, their core attention layers become a major bottleneck at inference time: as generation progresses, the KV cache grows, causing both increasing latency and escalating GPU memory, which in turn restricts usable temporal context and harms long-range consistency. In this work, we study redundancy in autoregressive video diffusion and identify three persistent sources: near-duplicate cached keys across frames, slowly evolving (largely semantic) queries/keys that make many attention computations redundant, and cross-attention over long prompts where only a small subset of tokens matters per frame. Building on these observations, we propose a unified, training-free attention framework (FAST-AR) for FAST-AutoRegressive diffusion, consisting of three components: TempCache compresses the KV cache via temporal correspondence to bound cache growth; AnnCA accelerates cross-attention by selecting frame-relevant prompt tokens using fast approximate nearest neighbor (ANN) matching; and AnnSA sparsifies self-attention by restricting each query to semantically matched keys, also using a lightweight ANN. Together, these modules reduce attention, compute, and memory and are compatible with existing autoregressive diffusion backbones and world models. Experiments demonstrate up to x5 - x10 end-to-end speedups while preserving near-identical visual quality and, crucially, maintaining stable throughput and nearly constant peak GPU memory usage over long rollouts, where prior methods progressively slow down and suffer from increasing memory usage.
Comment: Temporal-correspondence KV-cache compression bounds memory growth during autoregressive video generation.
Topic Match: Cache compression and attention sparsification are the core contributions, directly reducing model execution costs despite the video-specific setting.
Relevance: 8 Novelty: 7
8. Two-Stage Grid Optimization for Group-wise Quantization of LLMs
ArXiv ID: 2602.02126
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Junhan Kim, Gukryeol Lee, Seungwoo Son, Jeewook Kim, Yongkweon Jeon
Abstract: Group-wise quantization is an effective strategy for mitigating accuracy degradation in low-bit quantization of large language models (LLMs). Among existing methods, GPTQ has been widely adopted due to its efficiency; however, it neglects input statistics and inter-group correlations when determining group scales, leading to a mismatch with its goal of minimizing layer-wise reconstruction loss. In this work, we propose a two-stage optimization framework for group scales that explicitly minimizes the layer-wise reconstruction loss. In the first stage, performed prior to GPTQ, we initialize each group scale to minimize the group-wise reconstruction loss, thereby incorporating input statistics. In the second stage, we freeze the integer weights obtained via GPTQ and refine the group scales to minimize the layer-wise reconstruction loss. To this end, we employ the coordinate descent algorithm and derive a closed-form update rule, which enables efficient refinement without costly numerical optimization. Notably, our derivation incorporates the quantization errors from preceding layers to prevent error accumulation. Experimental results demonstrate that our method consistently enhances group-wise quantization, achieving higher accuracy with negligible overhead.
Comment: Two-stage group-scale optimization for GPTQ-style quantization with a closed-form coordinate-descent update that accounts for propagated layer errors.
Topic Match: Core contribution is a new quantization scale-fitting objective and solver, directly in the compression topic.
Relevance: 8 Novelty: 6
9. IntraSlice: Towards High-Performance Structural Pruning with Block-Intra PCA for LLMs
ArXiv ID: 2602.01975
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Meng Li, Peisong Wang, Yuantian Shao, Qinghao Hu, Hongjian Fang, Yifan Zhang, Zhihui Wei, Jian Cheng
Abstract: Large Language Models (LLMs) achieve strong performance across diverse tasks but face deployment challenges due to their massive size. Structured pruning offers acceleration benefits but leads to significant performance degradation. Recent PCA-based pruning methods have alleviated this issue by retaining key activation components, but are only applied between modules in order to fuse the transformation matrix, which introduces extra parameters and severely disrupts activation distributions due to residual connections. To address these issues, we propose IntraSlice, a framework that applies block-wise module-intra PCA compression pruning. By leveraging the structural characteristics of Transformer modules, we design an approximate PCA method whose transformation matrices can be fully fused into the model without additional parameters. We also introduce a PCA-based global pruning ratio estimator that further considers the distribution of compressed activations, building on conventional module importance. We validate our method on Llama2, Llama3, and Phi series across various language benchmarks. Experimental results demonstrate that our approach achieves superior compression performance compared to recent baselines at the same compression ratio or inference speed.
Comment: Module-intra PCA pruning whose approximate transformation matrices fuse into the model with zero extra parameters, avoiding the residual-path activation disruption of inter-module PCA pruning.
Topic Match: A new structural-pruning mechanism plus a global rank-ratio estimator, which is compression-mechanism work.
Relevance: 7 Novelty: 6
10. Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM Compression
ArXiv ID: 2602.02848
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Ali Abbasi, Chayne Thrash, Haoran Qin, Shansita Sharma, Sepehr Seifi, Soheil Kolouri
Abstract: Advances in large language models have driven strong performance across many tasks, but their memory and compute costs still hinder deployment. SVD-based compression reduces storage and can speed up inference via low-rank factors, yet performance depends on how rank is allocated under a global compression ratio. Prior methods often use homogeneous ranks for similarly sized matrices, despite large differences in loss sensitivity, or rely on expensive iterative pre-truncation optimization to determine per matrix ranks. We propose \textbf{Zero Sum SVD} (\textbf{ZS-SVD}), a post-training method that performs \emph{global} singular component selection using activation whitening and first-order calibration loss estimates in whitened coordinates. \textbf{ZS-SVD} prunes components across the whole model with a \textbf{zero sum} rule that keeps the cumulative predicted loss change near zero, automatically yielding heterogeneous ranks without solving a rank allocation optimization. Motivated by evidence that gradients near pretrained solutions exhibit low rank structure, we also introduce an optional lightweight correction that applies a \textbf{single} projected gradient update after truncation, followed by re-truncation. Extensive experiments across multiple LLM architectures show consistent gains across diverse benchmarks and compression ratios. Code is available at https://github.com/mint-vu/Zero-Sum-SVD
Comment: Global singular-component selection under a zero-sum rule on first-order calibration loss in whitened coordinates, yielding heterogeneous per-matrix ranks without a rank-allocation search.
Topic Match: A new low-rank compression mechanism (rank allocation by loss sensitivity), directly in the compression topic.
Relevance: 7 Novelty: 6
11. FlexRank: Nested Low-Rank Knowledge Decomposition for Adaptive Model Deployment
ArXiv ID: 2602.02680
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Riccardo Zaccone, Stefanos Laskaridis, Marco Ciccone, Samuel Horváth
Abstract: The growing scale of deep neural networks, encompassing large language models (LLMs) and vision transformers (ViTs), has made training from scratch prohibitively expensive and deployment increasingly costly. These models are often used as computational monoliths with fixed cost, hindering adaptive deployment across different cost budgets. We argue that nested components, ordered by importance, can be extracted from pretrained models and selectively activated within the available computational budget. To this end, our proposed FlexRank method leverages low-rank weight decomposition with nested, importance-based consolidation to extract submodels of increasing capabilities. Our approach enables a "train-once, deploy-everywhere" paradigm offering a graceful trade-off between cost and performance without training from scratch for each budget - advancing practical deployment of large models.
Comment: Nested importance-ordered low-rank decomposition yields submodels at many compute budgets from one pretrained checkpoint without per-budget retraining.
Topic Match: Low-rank compression with a nested-consolidation mechanism that changes the cost of deploying a trained model.
Relevance: 7 Novelty: 6
12. The Multiple Ticket Hypothesis: Random Sparse Subnetworks Suffice for RLVR
ArXiv ID: 2602.01599
Primary Topic: Efficiency, Compression, and Large-Scale Training
Also Matches: Architecture and Training Dynamics
Authors: Israel Adewuyi, Solomon Okibe, Vladmir Ivanov
Abstract: The Lottery Ticket Hypothesis demonstrated that sparse subnetworks can match full-model performance, suggesting parameter redundancy. Meanwhile, in Reinforcement Learning with Verifiable Rewards (RLVR), recent work has shown that updates concentrate on a sparse subset of parameters, which further lends evidence to this underlying redundancy. We study the simplest possible way to exploit this redundancy: training only a randomly selected subset of parameters at extreme sparsities. Empirically, we find that training just 1\% of parameters matches or exceeds full-parameter RLVR finetuning across 3 models and 2 task domains. Moreover, different random masks show minimal overlap ($\leq 0.005$ Jaccard similarity) and yet all succeed, suggesting pretrained models contain many viable sparse subnetworks rather than one privileged set. We term this the Multiple Ticket Hypothesis. We explain this phenomenon through the implicit per-step KL constraint in RLVR, which restricts updates to a low-dimensional subspace, enabling arbitrary sparse masks to succeed.
Comment: Training a random 1% of parameters matches full RLVR fine-tuning, with different masks at near-zero Jaccard overlap all succeeding, explained via the implicit per-step KL constraint restricting updates to a low-dimensional subspace.
Topic Match: An extreme-sparsity training result with a mechanistic explanation, so sparsity/efficiency is the closest topic despite the RLVR setting.
Relevance: 6 Novelty: 6
13. TraceNAS: Zero-shot LLM Pruning via Gradient Trace Correlation
ArXiv ID: 2602.02891
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Prajna G. Malettira, Manish Nagaraj, Arjun Roy, Shubham Negi, Kaushik Roy
Abstract: Structured pruning is essential for efficient deployment of Large Language Models (LLMs). The varying sensitivity of LLM sub-blocks to pruning necessitates the identification of optimal non-uniformly pruned models. Existing methods evaluate the importance of layers, attention heads, or weight channels in isolation. Such localized focus ignores the complex global structural dependencies that exist across the model. Training-aware structured pruning addresses global dependencies, but its computational cost can be just as expensive as post-pruning training. To alleviate the computational burden of training-aware pruning and capture global structural dependencies, we propose TraceNAS, a training-free Neural Architecture Search (NAS) framework that jointly explores structured pruning of LLM depth and width. TraceNAS identifies pruned models that maintain a high degree of loss landscape alignment with the pretrained model using a scale-invariant zero-shot proxy, effectively selecting models that exhibit maximal performance potential during post-pruning training. TraceNAS is highly efficient, enabling high-fidelity discovery of pruned models on a single GPU in 8.5 hours, yielding a 10$\times$ reduction in GPU-hours compared to training-aware methods. Evaluations on the Llama and Qwen families demonstrate that TraceNAS is competitive with training-aware baselines across commonsense and reasoning benchmarks.
Comment: Training-free NAS over joint depth and width pruning using a scale-invariant loss-landscape-alignment proxy, cutting search to 8.5 GPU-hours versus training-aware methods.
Topic Match: Structured pruning with a new zero-shot selection proxy that accounts for global structural dependencies.
Relevance: 6 Novelty: 6
14. You Need an Encoder for Native Position-Independent Caching
ArXiv ID: 2602.01519
Primary Topic: Efficiency, Compression, and Large-Scale Training
Also Matches: Architecture and Training Dynamics
Authors: Shiju Zhao, Junhao Hu, Jiaqi Zheng, Guihai Chen
Abstract: The Key-Value (KV) cache of Large Language Models (LLMs) is prefix-based, making it highly inefficient for processing contexts retrieved in arbitrary order. Position-Independent Caching (PIC) has been proposed to enable KV reuse without positional constraints; however, existing approaches often incur substantial accuracy degradation, limiting their practical adoption. To address this issue, we propose native PIC by reintroducing the encoder to prevalent decoder-only LLMs and explicitly training it to support PIC. We further develop COMB, a PIC-aware caching system that integrates seamlessly with existing inference frameworks. Experimental results show that COMB reduces Time-to-First-Token (TTFT) by 51-94% and increases throughput by 3$\times$ with comparable accuracy. Furthermore, the quality improvement when using DeepSeek-V2-Lite-Chat demonstrates the applicability of COMB to other types of decoder-only LLMs. Our code is available at https://github.com/shijuzhao/Comb.
Comment: Reintroduces an encoder into decoder-only LLMs and trains it explicitly for position-independent KV reuse, cutting TTFT 51-94% without the accuracy loss of post-hoc PIC.
Topic Match: A KV-cache design change backed by an architectural addition, which is the memory/cache-efficiency target with a new mechanism.
Relevance: 6 Novelty: 6
15. Prism: Efficient Test-Time Scaling via Hierarchical Search and Self-Verification for Discrete Diffusion Language Models
ArXiv ID: 2602.01842
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Jinbin Bai, Yixuan Li, Yuchen Zhu, Yi Xin, Qingyu Shi, Aosong Feng, Xiaohong Liu, Molei Tao, Jianru Xue, Xiangtai Li, Ming-Hsuan Yang
Abstract: Inference-time compute has re-emerged as a practical way to improve LLM reasoning. Most test-time scaling (TTS) algorithms rely on autoregressive decoding, which is ill-suited to discrete diffusion language models (dLLMs) due to their parallel decoding over the entire sequence. As a result, developing effective and efficient TTS methods to unlock dLLMs' full generative potential remains an underexplored challenge. To address this, we propose Prism (Pruning, Remasking, and Integrated Self-verification Method), an efficient TTS framework for dLLMs that (i) performs Hierarchical Trajectory Search (HTS) which dynamically prunes and reallocates compute in an early-to-mid denoising window, (ii) introduces Local branching with partial remasking to explore diverse implementations while preserving high-confidence tokens, and (iii) replaces external verifiers with Self-Verified Feedback (SVF) obtained via self-evaluation prompts on intermediate completions. Across four mathematical reasoning and code generation benchmarks on three dLLMs, including LLaDA 8B Instruct, Dream 7B Instruct, and LLaDA 2.0-mini, our Prism achieves a favorable performance-efficiency trade-off, matching best-of-N performance with substantially fewer function evaluations (NFE). The code is released at https://github.com/viiika/Prism.
Comment: Hierarchical trajectory pruning reallocates denoising evaluations to improve diffusion-language-model reasoning efficiency.
Topic Match: The compute savings concern reasoning-time search and self-verification, making this peripheral to foundational model execution and training efficiency.
Relevance: 6 Novelty: 6
16. ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
ArXiv ID: 2602.02366
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Sharut Gupta, Phillip Isola, Stefanie Jegelka, David Lopez-Paz, Kartik Ahuja, Mark Ibrahim, Mohammad Pezeshki
Abstract: Can Large language models (LLMs) learn to reason without any weight update and only through in-context learning (ICL)? ICL is strikingly sample-efficient, often learning from only a handful of demonstrations, but complex reasoning tasks typically demand many training examples to learn from. However, naively scaling ICL by adding more demonstrations breaks down at this scale: attention costs grow quadratically, performance saturates or degrades with longer contexts, and the approach remains a shallow form of learning. Due to these limitations, practitioners predominantly rely on in-weight learning (IWL) to induce reasoning. In this work, we show that by using Prefix Tuning, LLMs can learn to reason without overloading the context window and without any weight updates. We introduce $\textbf{ReasonCACHE}$, an instantiation of this mechanism that distills demonstrations into a fixed key-value cache. Empirically, across challenging reasoning benchmarks, including GPQA-Diamond, ReasonCACHE outperforms standard ICL and matches or surpasses IWL approaches. Further, it achieves this all while being more efficient across three key axes: data, inference cost, and trainable parameters. We also theoretically prove that ReasonCACHE can be strictly more expressive than low-rank weight update since the latter ties expressivity to input rank, whereas ReasonCACHE bypasses this constraint by directly injecting key-values into the attention mechanism. Together, our findings identify ReasonCACHE as a middle path between in-context and in-weight learning, providing a scalable algorithm for learning reasoning skills beyond the context window without modifying parameters. Our project page: https://reasoncache.github.io/
Comment: Fixed-size learned KV prefixes reduce context and trainable-parameter costs when learning reasoning skills.
Topic Match: Parameter-efficient adaptation is relevant, but the abstract primarily applies established prefix tuning to reasoning, supplemented by an expressivity comparison.
Relevance: 6 Novelty: 6
Paper Selection Prompt
System Prompt
You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.
User Prompt
Instructions
Respond in JSONL. Output exactly one JSON object per paper, one per line:
{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}
Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only:
daily_hot,new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return[]. -daily_hotmeans the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. -new_frontiermeans the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.Scoring Criteria
Relevance and Novelty are independent axes. Score both from 1 to 10.
Relevance Scoring
- 9-10: directly centered on the target foundational topics; highest when the core contribution is clearly within them.
- 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
- 5-6: touches the target topics, but the main contribution is elsewhere.
- 3-4: largely outside the target topics, often application-focused or domain-specific.
- 1-2: unrelated.
Important: Broad frontier relevance, major launch status, or daily buzz is not enough for a high Relevance score here. Those cases belong in the hotspot digest unless the paper strongly matches the specialized paper topics.
Novelty Scoring
- 9-10: new paradigm, theory, or major methodological breakthrough.
- 7-8: substantial methodological advance or strong new insight.
- 5-6: meaningful but incremental extension or refinement.
- 3-4: minor, narrow, or mostly engineering or domain-specific improvement.
- 1-2: little originality; mainly standard application of existing methods.
Topic Registry
Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - moe_training: MoE Training - Mixture-of-Experts training: routing, load balancing, expert granularity and shared experts, training stability, MoE parallelism and communication, and MoE kernels. - training_systems: Large-Scale Training Systems and Efficiency - Large-scale training systems and efficiency: distributed training algorithms, optimizers, and communication. - architecture_training: Architecture and Training Dynamics - Core architectural or computational mechanisms, dynamic computation, and training-stability dynamics. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, memory or cache efficiency, and large-scale training systems that materially change cost or behavior.
Papers
[PAPER LIST HERE]
Relevant Topics
This is a training-side feed. The centre of gravity is Mixture-of-Experts training and the large-scale training systems it sits on; architecture and efficiency work is kept when it bears on how a large model is trained or what it costs to train.
Keep a paper when its CORE CONTRIBUTION falls in one of the four topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.
MoE Training (primary) - Keep: expert routing and gating mechanisms; load balancing, with or without auxiliary losses; expert granularity, fine-grained and shared experts; MoE training stability (loss spikes, router collapse, z-loss and related fixes); expert and hybrid parallelism; all-to-all dispatch/combine and the communication schedules MoE needs; MoE kernels and grouped-GEMM style implementations; expert capacity, token dropping, and upcycling dense checkpoints into MoE. - Filter: papers that merely train on top of a MoE model without a new routing, balancing, stability, parallelism, or kernel contribution; "mixture of experts" in the classical ensemble or recommender sense.
Large-Scale Training Systems - Keep: distributed training algorithms (data, tensor, pipeline, sequence, and expert parallelism, FSDP and ZeRO-style sharding); optimizers and preconditioners for large-scale pretraining; communication algorithms, overlap, and collective design; throughput and MFU work that materially changes what a training run costs or whether it converges; scaling-law work that informs how such runs are configured. - Filter: routine infrastructure tuning with no new algorithmic or systems idea.
Architecture and Training Dynamics - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, dynamic or modular computation); optimisation and training-dynamics analysis that explains why large models train the way they do. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight.
Efficiency and Compression - Keep: quantization, sparsity, pruning, low-rank adaptation, KV-cache and memory-efficient designs, and other work that materially changes the cost of running or training a large model, especially when the mechanism is new rather than a tuned variant. - Filter: straightforward application of standard efficiency methods, and deployment or serving engineering with no new idea.
Leave these to the hotspot digest unless the core contribution clearly falls in one of the four topics above: - major model or product releases - agent, tooling, and RAG launches - benchmark, leaderboard, or evaluation-only papers - representation-learning theory, interpretability, and feature-formation analysis - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning) - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains