Previous Day 2026-05-12
Monthly Overview 2026-05
Next Day 2026-05-14

This is a remedial run for missed papers from 05/12/2026 to 05/12/2026.

Results generated on 09/11/2026.

Personalized Daily ArXiv Papers 2026-05-13

Model Metric Usage Papers
Prompt Completion Total Total arXiv Scanned Relevant
local agent CLI Tokens not reported not reported not reported 695 695 26
Cost not reported not reported not reported

Token counts are not reported for this run. 6 of 7 model calls succeeded, 1,942s of model wall clock.

Topic Coverage:

TopicPapers
MoE Training1
Large-Scale Training Systems and Efficiency3
Architecture and Training Dynamics13
Efficiency, Compression, and Large-Scale Training9

Table of contents by topic:

MoE Training (1)

  1. ROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory Systems Authors: Wenyong Zhou, Yuannuo Feng, Yizhe Chen, Taiqiang Wu, Wendong Xu, Wenbo Qi, Zhengwu Liu, Wang Kang, Ngai Wong

Large-Scale Training Systems and Efficiency (3)

  1. Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters Authors: Alexander Yukhimchuk, Mladen Kolar, Martin Takáč, Sayantan Choudhury

  2. Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives Authors: Konstantinos Oikonomidis, Jan Quan, Kimon Antonakopoulos, Antonio Silveti-Falls, Volkan Cevher, Panagiotis Patrinos

  3. Scaling Laws for Mixture Pretraining Under Data Constraints Authors: Anastasiia Sedova, Skyler Seto, Natalie Schluter, Pierre Ablin

Architecture and Training Dynamics (13)

  1. The Routing and Filtering Structure of Attention Authors: Shafayeth Jamil, Rehan Kapadia

  2. Scaling Laws and Tradeoffs in Recurrent Networks of Expressive Neurons Authors: Aaron Spieler, Georg Martius, Anna Levina

  3. Center-Manifold Reduction of Learning at Bifurcations: Interference and Rich Learning in Recurrent Neural Networks Authors: James Hazelden, Eric Shea-Brown

  4. Elastic Attention Cores for Scalable Vision Transformers Authors: Alan Z. Song, Yinjie Chen, Mu Nan, Rui Zhang, Jiahang Cao, Weijian Mai, Muquan Yu, Hossein Adeli, Deva Ramanan, Michael J. Tarr, Andrew F. Luo

  5. A Unified Framework for Critical Scaling of Inverse Temperature in Self-Attention Authors: Tomohiro Hayase, Ryo Karakida

  6. Training-Inference Consistent Segmented Execution for Long-Context LLMs Authors: Xianpeng Shang, Jiang Li, Zehua Duo, Qianyi Cai, Xiangdong Su

  7. Detecting overfitting in Neural Networks during long-horizon grokking using Random Matrix Theory Authors: Hari K. Prakash, Charles H Martin

  8. Emergence of Frontier Superposition: Möbius attractor and Cascade Supervision Authors: Hongyu Gu, Jingwen Fu

  9. Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization Authors: Zhehang Du, Hangfeng He, Weijie Su

  10. Early Data Exposure Improves Robustness to Subsequent Fine-Tuning Authors: Lawrence Feng, Gaurav R. Ghosal, Jacob Mitchell Springer, Ziqian Zhong, Aditi Raghunathan

  11. TCP-SSM: Efficient Vision State Space Models with Token-Conditioned Poles Authors: Sara Shoouri, Morteza Tavakoli Taba, Hun-Seok Kim

  12. Improving the Performance and Learning Stability of Parallelizable RNNs Designed for Ultra-Low Power Applications Authors: Julien Brandoit, Arthur Fyon, Damien Ernst, Guillaume Drion

  13. Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems Reconstruction Authors: Florian Hess, Florian Götz, Daniel Durstewitz

Efficiency, Compression, and Large-Scale Training (9)

  1. Search Your Block Floating Point Scales! Authors: Tanmaey Gupta, Hayden Prairie, Xiaoxia Wu, Reyna Abhyankar, Qingyang Wu, Austin Silveria, Pragaash Ponnusamy, Jue Wang, Ben Athiwaratkun, Leon Song, Tri Dao, Daniel Y. Fu, Chris De Sa

  2. FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression Authors: Namyoon Lee, Yongjune Kim

  3. OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models Authors: Yuchen Deng, Zidang Cai, Hai-Tao Zheng, Jie Wang, Feidiao Yang, Yuxing Han

  4. SOAR: Scale Optimization for Accurate Reconstruction in NVFP4 Quantization Authors: Chengzhu Bao, Xianglong Yan, Zhiteng Li, Guangshuo Qin, Guanghua Yu, Yulun Zhang

  5. LOFT: Low-Rank Orthogonal Fine-Tuning via Task-Aware Support Selection Authors: Lanxin Zhao, Bamdev Mishra, Pratik Jawanpuria, Lequan Lin, Dai Shi, Junbin Gao, Andi Han

  6. Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusion Authors: Chien Van Nguyen, Chaitra Hegde, Van Cuong Pham, Ryan A. Rossi, Franck Dernoncourt, Thien Huu Nguyen

  7. D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting Authors: Tianyu Wu, Yu Yao, Zhenting Qi, Han Zheng, Zhuohan Wang, Haoran Ma, Lawrence Liao, Himabindu Lakkaraju, Ju Li, Yilun Du

  8. Not How Many, But Which: Parameter Placement in Low-Rank Adaptation Authors: Arijit Sehanobish, Charles Lovering

  9. Multi-Token Residual Prediction Authors: Yufeng Xu, Zishuo Bao, Qian Wang, Zeshen Zhang, Haoqi Zhang, Bowen Peng, Ang Li, Rahul Chalamala, Yucheng Lu


MoE Training (1)

1. ROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory Systems

ArXiv ID: 2605.11800

Primary Topic: MoE Training

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Wenyong Zhou, Yuannuo Feng, Yizhe Chen, Taiqiang Wu, Wendong Xu, Wenbo Qi, Zhengwu Liu, Wang Kang, Ngai Wong

Abstract: Large language models (LLMs) with mixture-of-experts (MoE) architectures achieve remarkable scalability by sparsely activating a subset of experts per token, yet their frequent expert switching creates memory bandwidth bottlenecks that compute-in-memory (CIM) architectures are well-suited to mitigate. However, analog CIM systems suffer from inherent hardware imperfections that perturb stored weights, and its negative impact on MoE-based LLMs in noisy CIM environments remains unexplored. In this work, we present the first systematic investigation of MoE-based LLMs under noise model calibrated with real chip measurements, revealing that hardware noise critically disrupts expert load balance and renders clean-trained routing decisions consistently suboptimal. Based on these findings, we propose ROMER, a post-training calibration framework that (1) replaces underactivated experts with high-frequency ones to restore load balance, and (2) recalibrates router logits via percentile-based normalization to stabilize routing under noise. Extensive experiments across multiple benchmarks demonstrate that ROMER achieves up to 58.6\%, 58.8\%, and 59.8\% reduction in perplexity under real-chip noise conditions for DeepSeek-MoE, Qwen-MoE, and OLMoE, respectively, establishing its effectiveness and generalizability across diverse MoE architectures.

Comment: Post-training router calibration plus expert replacement to restore MoE expert load balance when hardware noise perturbs weights; targets routing and balancing directly, though at inference time on analog CIM rather than during training.

Topic Match: The contribution is a load-balance restoration and router-logit recalibration mechanism for MoE models, which is squarely a routing/balancing method.

Relevance: 7 Novelty: 6


Large-Scale Training Systems and Efficiency (3)

1. Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters

ArXiv ID: 2605.11838

Primary Topic: Large-Scale Training Systems and Efficiency

Also Matches: Architecture and Training Dynamics

Authors: Alexander Yukhimchuk, Mladen Kolar, Martin Takáč, Sayantan Choudhury

Abstract: Gradient clipping is a standard safeguard for training neural networks under noisy, heavy-tailed stochastic gradients; yet, most clipping rules treat all parameters as vectors and ignore the matrix structure of modern architectures. We show empirically that data outliers often amplify only a small number of leading singular values in layer-wise gradient matrices, while the rest of the spectrum remains largely unchanged. Motivated by this phenomenon, we propose spectral clipping, which stabilizes training by clamping singular values that exceed a threshold while preserving the singular directions. This framework generalizes classical gradient norm clipping and can be easily integrated into existing optimizers. We provide a convergence analysis for non-convex optimization with spectrally clipped SGD, yielding the optimal $\mathcal{O}\left(K^{\frac{2 - 2α}{3α- 2}}\right)$ rate for heavy-tailed noise. To minimize hyperparameter tuning, we introduce layer-wise adaptive thresholds based on moving averages or sliding-window quantiles of the top singular values. Finally, we develop efficient implementations that clip only the top $r$ singular values via randomized truncated SVD, avoiding full decompositions for large layers. We demonstrate competitive performance across synthetic heavy-tailed settings and neural network training tasks.

Comment: Spectral gradient clipping that clamps leading singular values of layer gradient matrices, with heavy-tailed convergence rates and truncated-SVD implementations.

Topic Match: A new optimizer-side stabilisation rule for matrix-valued parameters, i.e. large-scale training algorithmics.

Relevance: 8 Novelty: 7


2. Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives

ArXiv ID: 2605.11850

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Konstantinos Oikonomidis, Jan Quan, Kimon Antonakopoulos, Antonio Silveti-Falls, Volkan Cevher, Panagiotis Patrinos

Abstract: In this work, we develop proximal preconditioned gradient methods with a focus on spectral gradient methods providing a proximal extension to the Muon and Scion optimizers. We introduce a family of stochastic algorithms that can handle a wide variety of convex and nonconvex constraints and study its convergence under heavy-tailed noise, through a novel analysis tailored to the geometry of the proposed methods. We further propose a variance-reduced version, which achieves faster convergence under standard noise assumptions. Finally, we show that the polynomial iterations used in Muon are more accurately captured by a nonlinear preconditioner than by the ideal matrix sign, leading to a convergence analysis that more faithfully reflects practical implementations.

Comment: Proximal preconditioned gradient methods extending Muon/Scion to constraints, with heavy-tailed convergence analysis and a nonlinear-preconditioner model of Muon's polynomial iteration.

Topic Match: Directly about spectral optimizers used for large-scale pretraining and how their practical implementations actually converge.

Relevance: 8 Novelty: 7


3. Scaling Laws for Mixture Pretraining Under Data Constraints

ArXiv ID: 2605.12715

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Anastasiia Sedova, Skyler Seto, Natalie Schluter, Pierre Ablin

Abstract: As language models scale, the amount of data they require grows -- yet many target data sources, such as low-resource languages or specialized domains, are inherently limited in size. A common strategy is to mix this scarce but valuable target data with abundant generic data, which presents a fundamental trade-off: too little target data in the mixture underexposes the model to the target domain, while too much target data repeats the same examples excessively, yielding diminishing returns and eventual overfitting. We study this trade-off across more than 2,000 language-model training runs spanning multiple model and target dataset sizes, as well as several data types, including multilingual, domain-specific, and quality-filtered mixtures. Across all settings, we find that repetition is a central driver of target-domain performance, and that mixture training tolerates much higher repetition than single-source training: scarce target corpora can be reused 15-20 times, with the optimal number of repetitions depending on the target data size, compute budget, and model scale. Next, we introduce a repetition-aware mixture scaling law that accounts for the decreasing value of repeated target tokens and the regularizing role of generic data. Optimizing the scaling law provides a principled way to compute effective mixture configurations, yielding practical mixture recommendations for pretraining under data constraints.

Comment: Repetition-aware mixture scaling law fit over 2000+ runs, quantifying that scarce target corpora tolerate 15-20 repeats when diluted with generic data and giving optimal mixture ratios per compute budget.

Topic Match: This is scaling-law work that directly determines how a pretraining run should be configured under data constraints.

Relevance: 8 Novelty: 7


Architecture and Training Dynamics (13)

1. The Routing and Filtering Structure of Attention

ArXiv ID: 2605.18826

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Shafayeth Jamil, Rehan Kapadia

Abstract: The attention interaction matrix $QK^{\top}$ contains two entangled computations: a skew-symmetric component that redistributes information between positions (routing) and a symmetric component that scales mutual relevance (filtering). We decompose 1776 heads across five pretrained transformers and find routing operating at low rank, well below the routing capacity allocated by the weight kernel. We introduce $S$-$D$ attention as a diagnostic parameterization that disentangles routing from filtering by construction with guaranteed stability ($\mathrm{Re}(λ) \le 0$) and trains stably without layer normalization. When disentangled and unnormalized, routing self-organizes into a spectral cascade, effective rank $2$ at the first layer, expanding with depth across six scales from 7M to 355M parameters. The cascade predicts where attention can be simplified: linearizing the first seven layers of 125M $S$-$D$ attention costs ${<}5\%$ perplexity, whereas standard attention collapses under the same intervention. The linearizable region widens with depth. Replacing the first four layers with ELU+1 linear attention reaches within $1.4\%$ of baseline at full head dimension. Cascade-allocated architectures trade attention parameters for perplexity ($47\%-65\%$ fewer attention parameters at $+3.9\%$ to $+8.4\%$ PPL). The routing-filtering decomposition makes the spectral budget legible; the cascade makes it actionable.

Comment: Splits QK^T into skew-symmetric routing and symmetric filtering, finds routing occupies far lower rank than the weight kernel allocates, and turns the resulting spectral cascade into concrete layer-linearization and parameter-reallocation budgets.

Topic Match: It is a mechanistic decomposition of attention that both explains and predicts where the architecture can be simplified, measured across six model scales.

Relevance: 8 Novelty: 8


2. Scaling Laws and Tradeoffs in Recurrent Networks of Expressive Neurons

ArXiv ID: 2605.12049

Primary Topic: Architecture and Training Dynamics

Authors: Aaron Spieler, Georg Martius, Anna Levina

Abstract: Cortical neurons are complex, multi-timescale processors wired into recurrent circuits, shaped by long evolutionary pressure under stringent biological constraints. Mainstream machine learning, by contrast, predominantly builds models from extremely simple units, a default inherited from early neural-network theory. We treat this as a normative architectural question. How should one split a fixed parameter budget $P$ between the number of units $N$, per-unit effective complexity $k_e$, and per-unit connectivity $k_c$? What controls the optimal allocation? This calls for a model in which per-unit complexity can be tuned independently of width and connectivity. Accordingly, we introduce the ELM Network, whose recurrent layer is built from Expressive Leaky Memory (ELM) neurons, chosen to mirror functional components of cortical neurons. The architecture allows for individually adjusting $N$, $k_e$, and $k_c$ and trains stably across orders of magnitude in scale. We evaluate the model on two qualitatively different sequence benchmarks: the neuromorphic SHD-Adding task and Enwik8 character-level language modeling. Performance improves monotonically along each of the three axes individually. Under a fixed budget, a clear non-trivial optimum emerges in their tradeoff, and larger budgets favor both more and more complex neurons. A closed-form information-theoretic model captures these tradeoffs and attributes the diminishing returns at two ends to: per-neuron signal-to-noise saturation and across-neuron redundancy. A hyperparameter sweep spanning three orders of magnitude in trainable parameters traces a near-Pareto-frontier scaling law consistent with the framework. This suggests that the simple-unit default in ML is not obviously optimal once this tradeoff surface is probed, and offers a normative lens on cortex's reliance on complex spatio-temporal integrators.

Comment: Derives parameter-budget tradeoffs between recurrent neuron complexity, network width, and connectivity.

Topic Match: The recurrent architecture and explanatory scaling model directly examine how computational capacity should be allocated, with evidence from smaller sequence-modeling settings.

Relevance: 8 Novelty: 8


3. Center-Manifold Reduction of Learning at Bifurcations: Interference and Rich Learning in Recurrent Neural Networks

ArXiv ID: 2605.12763

Primary Topic: Architecture and Training Dynamics

Authors: James Hazelden, Eric Shea-Brown

Abstract: Rich learning in recurrent neural networks often proceeds through sudden transitions in latent dynamics, but there is little theory predicting how gradient descent behaves during these events. We study the local learning geometry near codimension-one bifurcations through the global empirical Neural Tangent Kernel (GeNTK). Under local center-manifold conditions, and when bifurcation-related sensitivity dominates bounded residual terms, we show that the global parameter-to-state Jacobian (D_θh) is approximated by a low-rank normal-form operator. The induced GeNTK and Fisher information matrix therefore become strongly amplified and anisotropic, concentrating toward a rank-one channel for the four scalar codimension-one bifurcations and a rank-two real channel for a Neimark--Sacker bifurcation. Controlled high-dimensional RNN experiments validate this operator reduction. In learned RNNs, the same low-rank concentration coincides with abrupt loss changes and subtask interference, while a local projection predicts the sign of these effects near isolated events. Finally, in an input-driven 15-task LeakyRNN, GeNTK amplification aligns with continuation-detected changes in the MemoryPro dynamics. These results suggest a tractable operator-level description of learning near dynamical transitions, together with scalable diagnostics for amplified low-dimensional learning geometry.

Comment: Reduces learning geometry near RNN bifurcations to amplified low-rank modes that explain abrupt loss changes and task interference.

Topic Match: The analysis directly connects recurrent dynamical transitions to gradient-descent geometry and optimization behavior, with validation in controlled RNN settings.

Relevance: 8 Novelty: 8


4. Elastic Attention Cores for Scalable Vision Transformers

ArXiv ID: 2605.12491

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Alan Z. Song, Yinjie Chen, Mu Nan, Rui Zhang, Jiahang Cao, Weijian Mai, Muquan Yu, Hossein Adeli, Deva Ramanan, Michael J. Tarr, Andrew F. Luo

Abstract: Vision Transformers (ViTs) achieve strong data-driven scaling by leveraging all-to-all self-attention. However, this flexibility incurs a computational cost that scales quadratically with image resolution, limiting ViTs in high-resolution domains. Underlying this approach is the assumption that pairwise token interactions are necessary for learning rich visual-semantic representations. In this work, we challenge this assumption, demonstrating that effective visual representations can be learned without any direct patch-to-patch interaction. We propose VECA (Visual Elastic Core Attention), a vision transformer architecture that uses efficient linear-time core-periphery structured attention enabled by a small set of learned cores. In VECA, these cores act as a communication interface: patch tokens exchange information exclusively through the core tokens, which are initialized from scratch and propagated across layers. Because the $N$ image patches only directly interact with a resolution invariant set of $C$ learned "core" embeddings, this yields linear complexity $O(N)$ for predetermined $C$, which bypasses quadratic scaling. Compared to prior cross-attention architectures, VECA maintains and iteratively updates the full set of $N$ input tokens, avoiding a small $C$-way bottleneck. Combined with nested training along the core axis, our model can elastically trade off compute and accuracy during inference. Across classification and dense tasks, VECA achieves performance competitive with the latest vision foundation models while reducing computational cost. Our results establish elastic core-periphery attention as a scalable alternative building block for Vision Transformers.

Comment: Core-periphery attention where patches interact only through a resolution-invariant set of learned core tokens, giving linear-time attention with elastic compute.

Topic Match: A new attention mechanism with nested elastic training, i.e. an architectural mechanism rather than a task application.

Relevance: 7 Novelty: 7


5. A Unified Framework for Critical Scaling of Inverse Temperature in Self-Attention

ArXiv ID: 2605.12697

Primary Topic: Architecture and Training Dynamics

Authors: Tomohiro Hayase, Ryo Karakida

Abstract: Length-dependent logit rescaling is widely used to stabilize long-context self-attention, but existing analyses and methods suggest conflicting inverse-temperature laws for the context length $n$, ranging from $(\log n)^{1/2}$ to $\log n$ and $(\log n)^2$. We provide a general theory showing that the desirable scale is determined by the gap-counting function $N_n$ of each attention row. Counting how many competitors lie within each gap from the maximum, we define an upper-tail accumulation scale and prove that it gives the critical inverse-temperature scale for softmax concentration: below this scale, the top competitors remain unseparated, whereas above it, the attention entropy collapses. This framework unifies prior scaling laws as different $N_n$ and yields a direct diagnostic for attention-score families, from idealized theoretical models to more practical transformers.

Comment: Derives the critical inverse-temperature scale from each attention row's gap-counting function, unifying the conflicting sqrt(log n), log n and (log n)^2 laws and giving a threshold below which top competitors stay unseparated and above which entropy collapses.

Topic Match: It is a mechanistic theory of softmax attention concentration that directly governs how logit scaling must be set for long-context stability.

Relevance: 7 Novelty: 7


6. Training-Inference Consistent Segmented Execution for Long-Context LLMs

ArXiv ID: 2605.11744

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Xianpeng Shang, Jiang Li, Zehua Duo, Qianyi Cai, Xiangdong Su

Abstract: Transformer-based large language models face severe scalability challenges in long-context generation due to the computational and memory costs of full-context attention. Under practical computation and memory constraints, many inference-efficient long-context methods improve efficiency by adopting bounded-context or segment-level execution only during inference, while continuing to train models under full-context attention, resulting in a mismatch between training and inference execution and state-transition semantics. Based on this insight, we propose a training-inference consistent segment-level generation framework, in which training and inference follow the same segment-level forward execution semantics. During training, consistency with inference is enforced by restricting gradient propagation to KV states carried over from the immediately preceding segment, while permitting head-specific access to past KV states during the forward pass without involving them in gradient propagation. Across long-context benchmarks, our approach achieves performance comparable to full-context attention, while achieving competitive latency-memory trade-offs against strong inference-efficient baselines, and substantially improving scalability at very long context lengths (e.g., approximately 6x lower peak prefill memory at 128K compared to full-context attention with FlashAttention).

Comment: Trains under the same segment-level forward semantics used at inference, restricting gradient flow to the previous segment's KV while allowing head-specific gradient-free access to older states.

Topic Match: An attention-and-state-transition training mechanism that removes the train/inference mismatch, with large prefill-memory consequences.

Relevance: 7 Novelty: 6


7. Detecting overfitting in Neural Networks during long-horizon grokking using Random Matrix Theory

ArXiv ID: 2605.12394

Primary Topic: Architecture and Training Dynamics

Authors: Hari K. Prakash, Charles H Martin

Abstract: Training Neural Networks (NNs) without overfitting is difficult; detecting that overfitting is difficult as well. We present a novel Random Matrix Theory method that detects the onset of overfitting in deep learning models without access to train or test data. For each model layer, we randomize each weight matrix element-wise, $\mathbf{W} \to \mathbf{W}^{\mathrm{rand}}$, fit the randomized empirical spectral distribution with a Marchenko-Pastur distribution, and identify large outliers that violate self-averaging. We call these outliers Correlation Traps. During the onset of overfitting, which we call the "anti-grokking" phase in long-horizon grokking, Correlation Traps form and grow in number and scale as test accuracy decreases while train accuracy remains high. Traps may be benign or may harm generalization; we provide an empirical approach to distinguish between them by passing random data through the trained model and evaluating the JS divergence of output logits. Our findings show that anti-grokking is an additional grokking phase with high train accuracy and decreasing test accuracy, structurally distinct from pre-grokking through its Correlation Traps. More broadly, we find that some foundation-scale LLMs exhibit the same Correlation Traps, indicating potentially harmful overfitting.

Comment: Random-matrix method detecting overfitting onset from weight spectra alone via Correlation Traps that violate self-averaging, identifying an anti-grokking phase.

Topic Match: Training-dynamics analysis of weight spectra that explains a distinct late-training phase, with data-free diagnostics.

Relevance: 6 Novelty: 7


8. Emergence of Frontier Superposition: Möbius attractor and Cascade Supervision

ArXiv ID: 2605.18820

Primary Topic: Architecture and Training Dynamics

Authors: Hongyu Gu, Jingwen Fu

Abstract: Superposition allows Transformers to reason in depth, carrying an entire reasoning frontier in parallel through a bounded-depth forward pass instead of unrolling serial chain-of-thought tokens. While Zhu et al. (2025) hand-crafted an equal-weight breadth-first frontier in a single residual stream for graph reachability, it remained open whether gradient descent could ever find this target amidst permutation-symmetric saddles. We close this gap on Reachability-by-Superposition over Erdős-Rényi graphs by isolating architectural and supervisional contributions. Architecturally, we identify a Möbius attractor: under $S_n$-symmetry in the tree regime, layerwise dynamics reduce to a 1D Möbius map whose zero set is a codimension-one manifold of global optima containing the equal-weight superposition state. On the supervision side, we identify Cascade Supervision: a loss class whose backward pass simultaneously delivers (A) selectivity bootstrap, (B) gradient persistence across depth, and (C) per-step discrimination (e.g., \mathcal{L}{sup} and \mathcal{L} in the graph fan-out and stall before the manifold is reached. Our thesis: Möbius attractor + Cascade Supervision = emergence of superposition reasoning. The parameter-free decay law predicts a final-step cosine of 0.35 vs. 0.71 (end-to-end vs. cascade) at depth D=3; experiments confirm 0.37 vs. 0.69, matching within 0.02 at every step.}). End-to-end supervision fails condition (B) and is provably insufficient: internal gradients at layer c decay as (np)^{-(D-c-2)/2

Comment: Shows end-to-end supervision provably cannot reach the superposition solution because layer-c gradients decay as a power of graph fan-out, and identifies the loss class whose backward pass keeps gradients alive across depth.

Topic Match: It is a training-dynamics result explaining when gradient descent reaches a target internal representation, with a quantitative decay law confirmed empirically.

Relevance: 6 Novelty: 7


9. Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization

ArXiv ID: 2605.12756

Primary Topic: Architecture and Training Dynamics

Authors: Zhehang Du, Hangfeng He, Weijie Su

Abstract: Large language models (LLMs) are pretrained by minimizing the cross-entropy loss for next-token prediction. In this paper, we study whether this optimization strategy can induce geometric structure in the learned model weights and context embeddings. We approach this problem by analyzing a constrained layer-peeled optimization program, which serves as a mathematically tractable surrogate for LLMs by treating the output projection matrix and last-layer context embeddings as optimization variables. Our analysis of this nonconvex optimization program demonstrates that symmetries in the target next-token distributions are transferred to the global minimizers of the layer-peeled model in a precise group-theoretic sense. Specifically, we prove that when the target tokens exhibit a cyclic-shift symmetry (such as the seven days of the week or the twelve months of the year), the optimal logit matrix is exactly circulant, and the Gram matrices of both the output projections and the context embeddings form circulant geometries as well. Next, for exchangeable target distributions invariant under the symmetric group and, more generally, under two-transitive group actions, we show that the global optimal output projection matrix forms a simplex equiangular tight frame, while the optimal logit matrix and context embeddings inherit the permutation symmetries present in the input data. A key technical step is to reduce the constrained nonconvex factorized problem to an explicit logit-level convex characterization for cyclic symmetry and to a symmetry-based lower bound for permutation symmetry, together with a sharp characterization of the optimal factorization. Finally, we empirically demonstrate that open-source LLMs naturally exhibit symmetries consistent with our theoretical predictions, despite being trained without any explicit regularization promoting such geometric structure.

Comment: Proves that cyclic and permutation symmetries in the target next-token distribution transfer to the global minimizers of a layer-peeled surrogate, forcing circulant logits and simplex ETF projections.

Topic Match: It explains what geometric structure cross-entropy pretraining induces in weights and embeddings, which is training-dynamics analysis of large models.

Relevance: 6 Novelty: 7


10. Early Data Exposure Improves Robustness to Subsequent Fine-Tuning

ArXiv ID: 2605.12705

Primary Topic: Architecture and Training Dynamics

Authors: Lawrence Feng, Gaurav R. Ghosal, Jacob Mitchell Springer, Ziqian Zhong, Aditi Raghunathan

Abstract: How can we train models whose post-trained capabilities survive subsequent fine-tuning? Rather than focusing on downstream interventions to mitigate forgetting of upstream capabilities, we study how upstream training choices - that is, the manner in which a capability is acquired - shape how robustly that capability is retained. We investigate this question in a controlled three-stage language-model pipeline: pretraining, post-training to acquire a target capability, and downstream fine-tuning on a new objective. Across 135M and 1B models, two post-training domains, and two downstream fine-tuning tasks, we find that immediate post-training performance does not reliably predict retention after subsequent fine-tuning: training recipes that look equivalent immediately after post-training can retain the target capability very differently after subsequent fine-tuning. In particular, early exposure - mixing post-training data into pretraining - consistently improves the frontier between retained upstream performance and downstream performance. In compute-matched experiments, where the target data must be allocated between pretraining and post-training, we find that the optimum lies at neither extreme. Together with our other empirical and theoretical findings, this supports the view that post-training drives immediate specialization while early exposure improves robustness to later forgetting. Replay and dropout, typically used to mitigate forgetting as it occurs during fine-tuning, provide complementary gains to early exposure when applied during post-training. Our findings suggest that robustness to subsequent fine-tuning should be treated as a first-class objective of upstream training, addressed preventatively through choices like early exposure rather than reactively during fine-tuning itself.

Comment: Compute-matched study showing that mixing post-training data into pretraining changes how capabilities survive later fine-tuning.

Topic Match: Training-dynamics analysis of how data allocation across pretraining stages shapes retention, informing how runs are configured.

Relevance: 6 Novelty: 6


11. TCP-SSM: Efficient Vision State Space Models with Token-Conditioned Poles

ArXiv ID: 2605.11563

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Sara Shoouri, Morteza Tavakoli Taba, Hun-Seok Kim

Abstract: State Space Models (SSMs) have emerged as a compelling alternative to attention models for long-range vision tasks, offering input-dependent recurrence with linear complexity. However, most efficient SSM variants reduce computation cost by modifying scan routes, resolutions, or traversal patterns, while largely leaving the recurrent dynamics implicit. Consequently, the model's state-dependent memory behavior is difficult to control, particularly in compact backbones where long scan paths can exceed the effective memory horizon. We propose Token-Conditioned Poles SSM (TCP-SSM), a structured selective SSM framework that improves efficiency while making recurrence dynamics explicit and interpretable through stable poles. TCP-SSM builds each scan operator with 1) real poles that model monotone or sign-alternating decay, and 2) complex-conjugate poles that capture damped oscillatory responses. Using bounded radius and angle modulation, TCP-SSM converts shared base poles into token-dependent poles, allowing each scan step to adapt its memory behavior to the current visual token while preserving pole stability. For practical scalability, we integrate grouped pole sharing with a lightweight low-rank input pathway, yielding an efficient scan operator that preserves linear-time scan complexity. Across image classification, semantic segmentation, and object detection, TCP-SSM reduces SSM computation complexity up to 44% in Vision Mamba-style models while maintaining or surpassing baseline accuracy.

Comment: Makes SSM recurrence explicit through stable real and complex-conjugate poles modulated per token, with grouped pole sharing and a low-rank input path preserving linear-time scan.

Topic Match: A new selective state-space recurrence parameterization, i.e. an architectural mechanism, though evaluated on vision backbones.

Relevance: 6 Novelty: 6


12. Improving the Performance and Learning Stability of Parallelizable RNNs Designed for Ultra-Low Power Applications

ArXiv ID: 2605.11855

Primary Topic: Architecture and Training Dynamics

Authors: Julien Brandoit, Arthur Fyon, Damien Ernst, Guillaume Drion

Abstract: Sequence learning is dominated by Transformers and parallelizable recurrent neural networks (RNNs) such as state-space models, yet learning long-term dependencies remains challenging, and state-of-the-art designs trade power consumption for performance. The Bistable Memory Recurrent Unit (BMRU) was introduced to enable hardware-software co-design of ultra-low power RNNs: quantized states with hysteresis provide persistent memory while mapping directly to analog primitives. However, BMRU performance lags behind parallelizable RNNs on complex sequential tasks. In this paper, we identify gradient blocking during state updates as a key limitation and propose a cumulative update formulation that restores gradient flow while preserving persistent memory, creating skip-connections through time. This leads to the Cumulative Memory Recurrent Unit (CMRU) and its relaxed variant, the $α$CMRU. Experiments show that the cumulative formulation dramatically improves convergence stability and reduces initialization sensitivity. The CMRU and $α$CMRU match or outperform Linear Recurrent Units (LRUs) and minimal Gated Recurrent Units (minGRUs) across diverse benchmarks at small model sizes, with particular advantages on tasks requiring discrete long-range retention, while the CMRU retains quantized states, persistent memory, and noise-resilient dynamics essential for analog implementation.

Comment: Identifies gradient blocking in bistable recurrent state updates and fixes it with a cumulative update that creates skip connections through time while keeping quantized persistent memory.

Topic Match: A recurrent-architecture change diagnosed from gradient flow, improving convergence stability and initialization sensitivity.

Relevance: 6 Novelty: 6


13. Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems Reconstruction

ArXiv ID: 2605.12683

Primary Topic: Architecture and Training Dynamics

Also Matches: Large-Scale Training Systems and Efficiency

Authors: Florian Hess, Florian Götz, Daniel Durstewitz

Abstract: Reconstructing nonlinear dynamical systems (DS) from data (DSR) is a fundamental challenge in science and engineering, but it inherently relies on sequential models. Recent breakthroughs for sequential models have produced algorithms that parallelize computation along sequence length $T$, achieving logarithmic time complexity, $\mathcal{O}(\log T)$. Since sequence lengths have been practically limited due to the linear runtime complexity $\mathcal{O}(T)$ of classical backpropagation through time, this opens new avenues for DSR. This paper studies two prominent classes of parallel-in-time algorithms for this task, both of which leverage parallel associative scans as their core computational primitive. The first class comprises models with linear yet non-autonomous dynamics and a nonlinear readout, such as modern State Space Models (SSMs), while the second consists of general nonlinear models which can be parallelized using the DEER framework. We find that the linear training-time recurrence of the first class of models imposes limitations that often hinder learning of accurate nonlinear dynamics. To address this, we augment DEER with Generalized Teacher Forcing (GTF), a novel variant within the more general nonlinear framework that ensures stable and effective learning of nonlinear dynamics across arbitrary sequence lengths. Using GTF-DEER, we investigate the benefits of training on extremely long sequences ($T>10^4$) for DSR. Our results show that access to such long trajectories significantly improves DSR if the data features long time scales. This work establishes GTF-DEER as a robust tool for data-driven discovery and underscores the largely untapped potential of long-sequence learning in modeling complex DS.

Comment: Shows the linear training-time recurrence of SSM-style models blocks learning of genuinely nonlinear dynamics, and augments the DEER parallel scan with generalized teacher forcing so nonlinear models train stably at sequence lengths beyond 10^4.

Topic Match: The core contribution is about what recurrent and state-space sequence models can learn when training is parallelized along sequence length.

Relevance: 6 Novelty: 6


Efficiency, Compression, and Large-Scale Training (9)

1. Search Your Block Floating Point Scales!

ArXiv ID: 2605.12464

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Tanmaey Gupta, Hayden Prairie, Xiaoxia Wu, Reyna Abhyankar, Qingyang Wu, Austin Silveria, Pragaash Ponnusamy, Jue Wang, Ben Athiwaratkun, Leon Song, Tri Dao, Daniel Y. Fu, Chris De Sa

Abstract: Quantization has emerged as a standard technique for accelerating inference for generative models by enabling faster low-precision computations and reduced memory transfers. Recently, GPU accelerators have added first-class support for microscaling Block Floating Point (BFP) formats. Standard BFP algorithms use a fixed scale based on the maximum magnitude of the block. We observe that this scale choice can be suboptimal with respect to quantization errors. In this work, we propose ScaleSearch, an alternative strategy for selecting these scale factors: using a fine-grained search leveraging the mantissa bits in microscaling formats to minimize the quantization error for the given distribution. ScaleSearch can be integrated with existing quantization methods such as Post Training Quantization and low-precision attention, and is shown to improve their performance. Additionally, we introduce ScaleSearchAttention, an accelerated NVFP4-based attention algorithm, which uses ScaleSearch and adapted prior techniques to ensure near-0 performance loss for causal language modeling. Experiments show that ScaleSearch reduces quantization error by 27% for NVFP4 and improves language model PTQ by up to 15 points for MATH500 (Qwen3-8B), while ScaleSearchAttention improves Wikitext-2 PPL by upto 0.77 points for Llama 3.1 70B. The proposed methods closely match baseline performance while providing quantization accuracy improvements.

Comment: ScaleSearch selects microscaling block-float scale factors by fine-grained mantissa search instead of block-max, plus an NVFP4 attention kernel.

Topic Match: New scale-selection mechanism for hardware-native low-precision formats, central to quantization efficiency.

Relevance: 8 Novelty: 7


2. FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression

ArXiv ID: 2605.11478

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Namyoon Lee, Yongjune Kim

Abstract: Long-context inference is increasingly a memory-traffic problem. The culprit is the key--value (KV) cache: it grows with context length, batch size, layers, and heads, and it is read at every decoding step. Rotation-based scalar codecs meet this systems constraint by storing a norm, applying a shared random rotation, and quantizing one coordinate at a time. They are universal and random-access, but they discard the geometry created by the normalization step. After a Haar rotation, a block of $k$ consecutive coordinates is not a product source; it is a spherical-Beta source on the unit ball. We introduce \textsc{FibQuant}, a universal fixed-rate vector quantizer that keeps the same normalize--rotate--store interface while replacing scalar tables by a shared radial--angular codebook matched to this canonical source. The codebook combines Beta-quantile radii, Fibonacci\,/\,Roberts--Kronecker quasi-uniform directions, and multi-restart Lloyd--Max refinement. We prove that the resulting vector code strictly improves on its scalar product specialization at matched rate, with a high-rate gain that separates into a cell-shaping factor and a density-matching factor. The same construction gives a dense rate axis, including fractional-bit and sub-one-bit operating points, without calibration or variable-length addresses. On GPT-2 small KV caches, \textsc{FibQuant} traces a memory--fidelity frontier from $5\times$ compression at $0.99$ attention cosine similarity to $34\times$ at $0.95$. End-to-end on TinyLlama-1.1B, it is within $0.10$ perplexity of fp16 at $4\times$ compression and has $3.6\times$ lower perplexity than scalar \textsc{TurboQuant} at $b = 2$ ($8\times$ compression), where scalar random-access quantization begins to fail.

Comment: Replaces scalar rotation-based KV codecs with a radial-angular vector codebook matched to the spherical-Beta source induced by normalize-then-rotate, keeping random access at fixed rate.

Topic Match: KV-cache compression with a new quantizer construction and a proved gain over its scalar specialization.

Relevance: 8 Novelty: 7


3. OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models

ArXiv ID: 2605.12056

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Yuchen Deng, Zidang Cai, Hai-Tao Zheng, Jie Wang, Feidiao Yang, Yuxing Han

Abstract: Omnimodal large language models (Omni-LLMs) show strong capability in audio-video understanding, but their practical deployment remains limited by high inference cost of long video streams and dense audio sequences. Despite recent progress, existing compression methods for Omni-LLMs typically rely on fixed or native compression units, which can disrupt cross-modal correspondence and the complementary information required for audio-video reasoning, making it difficult to improve inference efficiency while stably preserving performance. To address this, we propose OmniRefine, a training-free two-stage framework for efficient audio-visual token compression in Omni-LLMs. First, Correspondence-Preserving Chunk Refinement refines native chunk boundaries into cross-modally aligned compression units through frame-audio similarity and dynamic programming. Second, Modality-Aware Cooperative Compression jointly compresses video and audio tokens within each refined unit to reduce redundancy while preserving critical evidence. Extensive experiments show that OmniRefine achieves a better efficiency-performance trade-off than strong baselines and maintains stable performance under lower compression ratios. On WorldSense, it still reaches 46.7% accuracy at a 44% token retention ratio, nearly matching the full-token baseline. The code and interface will be released to facilitate further research.

Comment: Refines cross-modal chunk boundaries to support joint audio-video token compression while preserving correspondence.

Topic Match: The core contribution is a correspondence-aware token-compression mechanism that reduces large multimodal model input length and inference cost.

Relevance: 8 Novelty: 7


4. SOAR: Scale Optimization for Accurate Reconstruction in NVFP4 Quantization

ArXiv ID: 2605.12245

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Chengzhu Bao, Xianglong Yan, Zhiteng Li, Guangshuo Qin, Guanghua Yu, Yulun Zhang

Abstract: NVFP4 has recently emerged as an efficient 4-bit microscaling format for large language models (LLMs), offering superior numerical fidelity with native hardware support. However, existing methods often yield suboptimal performance due to inflexible scale selection and the coupled treatment of quantization and dequantization scales. To address these issues, we propose Scale Optimization for Accurate Reconstruction (SOAR), a novel post-training quantization framework that improves the accuracy of NVFP4 quantization. At its core, SOAR features Closed-form Joint Scale Optimization (CJSO), which jointly optimizes global and block-wise scales via analytical solutions derived from reconstruction error minimization. Furthermore, it incorporates Decoupled Scale Search (DSS). DSS decouples the high-precision quantization scale from its constrained dequantization counterpart, and performs discrete search to mitigate precision loss from scale quantization. Extensive experiments across multiple LLMs show that our method consistently outperforms existing NVFP4 quantization baselines, achieving superior accuracy under the same memory footprint with no additional hardware overhead. The code and models will be available at https://github.com/steven-bao1/SOAR.

Comment: Closed-form joint optimisation of global and block-wise NVFP4 scales plus decoupled quantize/dequantize scale search.

Topic Match: Core contribution is a new 4-bit microscaling quantization mechanism, squarely in compression.

Relevance: 8 Novelty: 6


5. LOFT: Low-Rank Orthogonal Fine-Tuning via Task-Aware Support Selection

ArXiv ID: 2605.11872

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Lanxin Zhao, Bamdev Mishra, Pratik Jawanpuria, Lequan Lin, Dai Shi, Junbin Gao, Andi Han

Abstract: Orthogonal parameter-efficient fine-tuning (PEFT) adapts pretrained weights through structure-preserving multiplicative transformations, but existing methods often conflate two distinct design choices: the subspace in which adaptation occurs and the transformation applied within that subspace. This paper introduces LOFT, a low-rank orthogonal fine-tuning framework that explicitly separates these two components. By viewing orthogonal adaptation as a multiplicative subspace rotation, LOFT provides a unified formulation that recovers representative orthogonal PEFT methods, including coordinate-, butterfly-, Householder-, and principal-subspace-based variants. More importantly, this perspective exposes support selection as a central design axis rather than a byproduct of a particular parameterization. We develop a first-order analysis showing that useful adaptation supports should be informed by the downstream training signal, motivating practical task-aware support selection strategies. Across language understanding, visual transfer, mathematical reasoning, and multilingual out-of-distribution adaptation, LOFT recovers principal-subspace orthogonal adaptation while gradient-informed supports improve the efficiency-performance trade-off under matched parameter, memory, and compute budgets. These results suggest that principled support selection is an important direction for improving orthogonal PEFT.

Comment: Unified low-rank orthogonal PEFT that separates adaptation subspace from the transformation and makes gradient-informed support selection the design axis.

Topic Match: Low-rank adaptation mechanism under matched parameter/memory budgets, the compression-and-efficiency topic.

Relevance: 7 Novelty: 6


6. Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusion

ArXiv ID: 2605.12825

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Chien Van Nguyen, Chaitra Hegde, Van Cuong Pham, Ryan A. Rossi, Franck Dernoncourt, Thien Huu Nguyen

Abstract: We introduce Orthrus, a simple and efficient dual-architecture framework that unifies the exact generation fidelity of autoregressive Large Language Models (LLMs) with the high-speed parallel token generation of diffusion models. The sequential nature of standard autoregressive decoding represents a fundamental bottleneck for high-throughput inference. While diffusion language models attempt to break this barrier via parallel generation, they suffer from significant performance degradation, high training costs, and a lack of rigorous convergence guarantees. Orthrus resolves this dichotomy natively. Designed to seamlessly integrate into existing Transformers, the framework augments a frozen LLM with a lightweight, trainable module to create a parallel diffusion view alongside the standard autoregressive view. In this unified system, both views attend to the exact same high-fidelity Key-Value (KV) cache; the autoregressive head executes context pre-filling to construct accurate KV representations, while the diffusion head executes parallel generation. By employing an exact consensus mechanism between the two views, Orthrus guarantees lossless inference, delivering up to a 7.8x speedup with only an O(1) memory cache overhead and minimal parameter additions.

Comment: Attaches a trainable diffusion head to a frozen LLM sharing the same KV cache, with an exact consensus rule giving lossless parallel decoding at O(1) cache overhead.

Topic Match: A memory-efficient dual-view decoding mechanism; the new idea is in cache sharing and consensus rather than in training.

Relevance: 6 Novelty: 7


7. D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting

ArXiv ID: 2605.18810

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Tianyu Wu, Yu Yao, Zhenting Qi, Han Zheng, Zhuohan Wang, Haoran Ma, Lawrence Liao, Himabindu Lakkaraju, Ju Li, Yilun Du

Abstract: Speculative decoding accelerates LLM inference by having a small drafter propose tokens that a larger target model verifies in parallel. Recent diffusion-based parallel drafters such as DFlash predict the full B-token block in one forward pass, enabling deeper drafters and longer accepted blocks. However, existing multi-token drafter objectives often use fixed position-dependent weighting schedules, such as head-dependent weights or block-position decays, which do not adapt as the positions limiting acceptance change during training. To address this, we derive per-position training weights from a differentiable surrogate of expected accepted draft length, matching the weight of each position to its log-probability gradient contribution. The resulting loss, D-PACE (Dynamic Position-Aware Cross-Entropy), shifts training signal toward positions that currently limit acceptance as the drafter improves. Across six benchmarks, two Qwen3-4B draft depths, two decoding temperatures, and two additional target models, D-PACE consistently improves both wall-clock speedup and average emitted length, with 2.3\% measured training-time overhead and no changes to the drafter architecture or inference procedure.

Comment: Per-position drafter loss weights derived from a differentiable surrogate of expected accepted length for parallel speculative decoding.

Topic Match: Changes inference cost via a new drafter training objective; efficiency rather than large-scale training.

Relevance: 6 Novelty: 6


8. Not How Many, But Which: Parameter Placement in Low-Rank Adaptation

ArXiv ID: 2605.12207

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Arijit Sehanobish, Charles Lovering

Abstract: We study the \textit{parameter placement problem}: given a fixed budget of $k$ trainable entries within the B matrix of a LoRA adapter (A frozen), does the choice of which $k$ matter? Under supervised fine-tuning, random and informed subsets achieve comparable performance. Under GRPO on base models, random placement fails to improve over the base model, while gradient-informed placement recovers standard LoRA accuracy. This regime dependence traces to gradient structure: SFT gradients are low-rank and directionally stable, so any subset accumulates coherent updates; GRPO gradients are high-rank and near-orthogonal across steps, so only elements with consistently signed gradients retain the learning signal. Our scoring procedure identifies these critical parameters in under 10 seconds at less than 0.5% of training cost. Selected parameters concentrate on residual-stream-writing projections (V, O, Down), stable across model families and scales (1.5B - 8B).

Comment: Shows which entries of a LoRA B matrix are trainable matters under GRPO but not SFT, tracing the difference to high-rank near-orthogonal gradients versus low-rank stable ones.

Topic Match: A low-rank adaptation result whose mechanism is a concrete gradient-structure insight about where the learning signal lives.

Relevance: 6 Novelty: 6


9. Multi-Token Residual Prediction

ArXiv ID: 2605.18817

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Yufeng Xu, Zishuo Bao, Qian Wang, Zeshen Zhang, Haoqi Zhang, Bowen Peng, Ang Li, Rahul Chalamala, Yucheng Lu

Abstract: Diffusion Language Models (DLMs) generate text by iteratively denoising masked token sequences, offering a tradeoff between parallelism and quality compared to autoregressive models. In current practice, the number of tokens decoded per step is controlled by a confidence threshold, and quality degrades monotonically as more tokens are denoised per step. We introduce Multi-token Residual Prediction (MRP), a lightweight module that enables dependency-aware multi-token denoising within a single backbone forward pass. MRP exploits a key property of the denoising process: the logit distributions at adjacent denoising steps are remarkably similar. Rather than running the backbone a second time to obtain the next-step logits, MRP predicts the residual between steps from the backbone's hidden states, effectively denoising more tokens per backbone forward at a fraction of the cost. We apply MRP across the two operating regimes of DLM decoding. In the high-quality-low-throughput static denoising regime, MRP serves as a drafter for speculative decoding: its proposals are verified against the backbone, yielding lossless acceleration of up to 1.4x in SGLang. In the low-quality-high-throughput dynamic denoising regime, MRP instead drives a remasking scheme that revokes over-eager reveals, recovering most of the accuracy lost to aggressive low-threshold decoding and improving accuracy by up to 22.6 points on code generation task HumanEval and 17.7 points on reasoning task GSM8K.

Comment: Exploits the similarity between adjacent denoising-step logits to predict the inter-step residual from hidden states, denoising more tokens per backbone pass in diffusion LMs.

Topic Match: A new lightweight module that materially changes decoding cost by avoiding a second backbone forward, which is a mechanism-level efficiency result.

Relevance: 6 Novelty: 6


Paper Selection Prompt

System Prompt

You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.

User Prompt

Instructions

Respond in JSONL. Output exactly one JSON object per paper, one per line:

{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}

Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only: daily_hot, new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return []. - daily_hot means the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. - new_frontier means the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.

Scoring Criteria

Relevance and Novelty are independent axes. Score both from 1 to 10.

Relevance Scoring

  • 9-10: directly centered on the target foundational topics; highest when the core contribution is clearly within them.
  • 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
  • 5-6: touches the target topics, but the main contribution is elsewhere.
  • 3-4: largely outside the target topics, often application-focused or domain-specific.
  • 1-2: unrelated.

Important: Broad frontier relevance, major launch status, or daily buzz is not enough for a high Relevance score here. Those cases belong in the hotspot digest unless the paper strongly matches the specialized paper topics.

Novelty Scoring

  • 9-10: new paradigm, theory, or major methodological breakthrough.
  • 7-8: substantial methodological advance or strong new insight.
  • 5-6: meaningful but incremental extension or refinement.
  • 3-4: minor, narrow, or mostly engineering or domain-specific improvement.
  • 1-2: little originality; mainly standard application of existing methods.

Topic Registry

Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - moe_training: MoE Training - Mixture-of-Experts training: routing, load balancing, expert granularity and shared experts, training stability, MoE parallelism and communication, and MoE kernels. - training_systems: Large-Scale Training Systems and Efficiency - Large-scale training systems and efficiency: distributed training algorithms, optimizers, and communication. - architecture_training: Architecture and Training Dynamics - Core architectural or computational mechanisms, dynamic computation, and training-stability dynamics. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, memory or cache efficiency, and large-scale training systems that materially change cost or behavior.

Papers

[PAPER LIST HERE]

Relevant Topics

This is a training-side feed. The centre of gravity is Mixture-of-Experts training and the large-scale training systems it sits on; architecture and efficiency work is kept when it bears on how a large model is trained or what it costs to train.

Keep a paper when its CORE CONTRIBUTION falls in one of the four topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.

  1. MoE Training (primary) - Keep: expert routing and gating mechanisms; load balancing, with or without auxiliary losses; expert granularity, fine-grained and shared experts; MoE training stability (loss spikes, router collapse, z-loss and related fixes); expert and hybrid parallelism; all-to-all dispatch/combine and the communication schedules MoE needs; MoE kernels and grouped-GEMM style implementations; expert capacity, token dropping, and upcycling dense checkpoints into MoE. - Filter: papers that merely train on top of a MoE model without a new routing, balancing, stability, parallelism, or kernel contribution; "mixture of experts" in the classical ensemble or recommender sense.

  2. Large-Scale Training Systems - Keep: distributed training algorithms (data, tensor, pipeline, sequence, and expert parallelism, FSDP and ZeRO-style sharding); optimizers and preconditioners for large-scale pretraining; communication algorithms, overlap, and collective design; throughput and MFU work that materially changes what a training run costs or whether it converges; scaling-law work that informs how such runs are configured. - Filter: routine infrastructure tuning with no new algorithmic or systems idea.

  3. Architecture and Training Dynamics - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, dynamic or modular computation); optimisation and training-dynamics analysis that explains why large models train the way they do. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight.

  4. Efficiency and Compression - Keep: quantization, sparsity, pruning, low-rank adaptation, KV-cache and memory-efficient designs, and other work that materially changes the cost of running or training a large model, especially when the mechanism is new rather than a tuned variant. - Filter: straightforward application of standard efficiency methods, and deployment or serving engineering with no new idea.

Leave these to the hotspot digest unless the core contribution clearly falls in one of the four topics above: - major model or product releases - agent, tooling, and RAG launches - benchmark, leaderboard, or evaluation-only papers - representation-learning theory, interpretability, and feature-formation analysis - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning) - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains