Previous Day 2026-09-02
Monthly Overview 2026-09
Next Day 2026-09-04

This is a remedial run for missed papers from 09/02/2026 to 09/02/2026.

Results generated on 09/14/2026.

Personalized Daily ArXiv Papers 2026-09-03

Model Metric Usage Papers
Prompt Completion Total Total arXiv Scanned Relevant
local agent CLI Tokens not reported not reported not reported 515 515 35
Cost not reported not reported not reported

Token counts are not reported for this run. 24 of 24 model calls succeeded, 3,193s of model wall clock.

Topic Coverage:

TopicPapers
MoE Training1
Large-Scale Training Systems and Efficiency4
Architecture and Training Dynamics7
Efficiency, Compression, and Large-Scale Training23

Table of contents by topic:

MoE Training (1)

  1. Evidence for Shared Routing Geometry and Dynamics in Sparse Mixture-of-Experts Authors: Kirill Labzin, Stepan Kulibaba, Artem Dzhalilov, Artem Gorokhov

Large-Scale Training Systems and Efficiency (4)

  1. Momentum in large-batch training: Polyak enlarges the critical batch size, Nesterov improves data efficiency Authors: Jia-Nan Wang, Zixun Huang, Kairui Li, Lei Wu

  2. BASP: Communication-Efficient Batch-Aware Sequence Parallelism for LLM Training Authors: Bigyan Ghimire, Jon C. Calhoun

  3. Gradient Prediction with Control Variates in the Cheap-Forward Regime Authors: Kamil Ciosek, Nicolò Felicioni, Juan Elenter, Ehsan Imani

  4. Nova: An End-to-End MLIR Compiler for Deep Learning Authors: Adwaid Suresh, Aparna A, Harshini V M, Jona Delcy C A, Killi Uma Maheswara Rao, Ram Charan Golla, Surendra Vendra

Architecture and Training Dynamics (7)

  1. Multi-Mask Diffusion Language Models for Few-Step Generation Authors: Sijin Chen, Yinuo Ren, Heyang Zhao, Ziheng Cheng, Quanquan Gu, Lexing Ying

  2. On the Expressive Power and Limitations of Multi-Layer SSMs Authors: Nikola Zubić, Qian Li, Yuyi Wang, Davide Scaramuzza

  3. MDM-Prime-v2: Binary Encoding and Index Shuffling Enable Scaling of Diffusion Language Models Authors: Chen-Hao Chao, Wei-Fang Sun, Junwei Quan, Chun-Yi Lee, Rahul G. Krishnan

  4. CHASE: Cache-Hole-Adapted Skip Exit for Looped State-Space Language Models Authors: Zhenxuan Yu, Takeshi Kojima, Yutaka Matsuo, Yusuke Iwasawa

  5. The Dynamics of Continuous Mixture Collapse in Language Models Authors: Ali Backour

  6. Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders Authors: Pierre-Antoine Lequeu, Camille Barboule, Benjamin Piwowarski

  7. Gradient Descent on Logistic Regression with Non-Separable Data and Large Step Sizes Authors: Si Yi Meng, Antonio Orvieto, Daniel Yiming Cao, Christopher De Sa

Efficiency, Compression, and Large-Scale Training (23)

  1. LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style Updates Authors: Dmitrii Andriianov, Andrey Veprikov, Aleksandr Beznosikov

  2. A Mathematical Theory of Reusable Neural Bases for Network Compression Authors: Binshuai Wang, Peng Wei

  3. HARP: Hadamard-Preconditioned Adaptive Rotation Processor for Extreme LLM Quantization Authors: Artur Zagitov, Gleb Molodtsov, Aleksandr Beznosikov

  4. HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models Authors: Renjie Xie, Juncheng Yang, Aoting Hu, Mingxi Zhang, Liyao Wu, Zheheng Hong, Wei Xu

  5. Enabling KV Caching of Shared Prefix for Diffusion Language Models Authors: Younghun Go, Jaehoon Han, Changyong Shin, Chuck Yoo, Gyeongsik Yang

  6. Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs Authors: Seungwoo Jung, Dohyeok Kwon, Seungmin Cha, Junseok Lee, Yeonho Yoo, Chuck Yoo, Gyeongsik Yang

  7. QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization Authors: Yipin Guo, Arun M George, Jie Fu, Tareq Mahmoud, Sixue Xing, Siddharth Joshi

  8. Train What You Deploy: Closing the MLP Reachability Gap in Low-Rank Clone Distillation Authors: Wenhui Chen, Zhifeng Li, Jie Zhou, Navan Preet Singh, Madalina Ciobanu, Chenghua Wang, Qingqing Mao, Ritankar Das

  9. XMerge: Cross-Axis Selection and Reconstructive Layer Merging for LLM Depth Compression Authors: Jundong Hu, Shekar Ramachandran

  10. TaRA: Training-Aware Low-Rank Adaptation Initialization Authors: Taehyeon Kim, Eunhyeok Park

  11. LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference Authors: Renyuan Liu, Yuyang Leng, Kaiyan Liu, Yuzhou Zhong, Shaohan Hu, Chun-Fu, Chen, Peijun Zhao, Heechul Yun, Shuochao Yao

  12. Routing Is Not Enough: Diagnosing Intra-Adapter Subspace Contention in MoE+LoRA Fine-Tuning Authors: Mehreen Hossain Chowdhury, Nowshin Mahjabin, Ahmed Shafin Ruhan, Md Azam Hossain, Abu Raihan Mostofa Kamal, Md Tahmid Rahman Laskar

  13. Scalable Kronecker-Fisher Approximation: Efficient Hessian Analysis for Billion-Parameter Language Models Compression Authors: Viacheslav Yusupov, Daria Cherniuk, Evgeny Frolov

  14. DLM-One: Diffusion Language Models for One-Step Sequence Generation Authors: Tianqi Chen, Shujian Zhang, Mingyuan Zhou

  15. Accelerating Unified Multimodal Models with Core-Expansion Routing and Unified Computation Scheduling Authors: Wengyi Zhan, Chenqian Yan, Songwei Liu, Mingbao Lin, Rongrong Ji

  16. Who Speaks for the Pruned? Visual Token Pruning as Coverage Optimization Authors: Qingchan Zhu, Weihang You, Hanqi Jiang, Changdi Yang, Tianming Liu, Geng Yuan

  17. Unfolding the Leech Lattice: Fused Multi-Shell Decoding and VRAM Layouts for 2-Bit LLM Weights Authors: Pier-Jean Malandrino

  18. Compute-in-Memory Attention: A Time-Domain Analog Softmax Circuit with RC-Tunable Temperature Authors: Ankur Singh, Ashish Gautam, Shruti R. Kulkarni, Guojing Cong

  19. Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens Authors: Yizhen Yao, Qinglin Zhu, Runcong Zhao, Xiangxiang Dai, Yanzheng Xiang, Yulan He, Lin Gui

  20. MeanField Surrogate Modeling for Scalable Runtime Scheduling of Concurrent Heterogeneous AI Inference on Shared GPUs Authors: Youssef Ennouri, Soonhoi Ha

  21. GeoSPRINT: Geometric Redundancy-Aware Step Pruning for Inference in Diffusion Trajectories Authors: Arpita Joshi

  22. Predict, Don't Iterate: Efficient Adaptive-Length Infilling for Diffusion Language Models Authors: Haobo Xu, Sirui Chen, Yuanchen Bei, Lingjie Chen, Yuchen Yan, Dongqi Fu, Jingrui He, Hanghang Tong

  23. Discriminative and Consistent Representation Distillation Authors: Nikos Giakoumoglou, Tania Stathaki


MoE Training (1)

1. Evidence for Shared Routing Geometry and Dynamics in Sparse Mixture-of-Experts

ArXiv ID: 2609.02404

Primary Topic: MoE Training

Also Matches: Architecture and Training Dynamics

Authors: Kirill Labzin, Stepan Kulibaba, Artem Dzhalilov, Artem Gorokhov

Abstract: Sparse mixture-of-experts (MoE) models use an independently parameterized router at each sparse layer to select experts for every token. Prior work has shown that routing decisions across depth can often be predicted from earlier routing signals, suggesting that routing is not fully independent across layers. However, the structure behind this predictability remains unclear. In this work, we provide evidence that routing-relevant states across layers share a common geometric structure that is obscured by layer-specific coordinate systems. We isolate the control subspace of each router and align these spaces into a shared canonical representation using generalized orthogonal Procrustes analysis. After alignment, a single linear transition reaches $R^2=0.39$--$0.71$ and retains 79--90\% of the predictive power of separately fitted layer-specific dynamics, indicating that much of routing-state evolution follows a reusable process across depth. We then ask whether this shared dynamics is specific to routing or simply reflects the smooth evolution of hidden representations. A matched-rank comparison shows that residual representations are often easier to predict across layers, while router-control states preserve the model's expert choices much more faithfully. This separates generic cross-layer predictability from routing-specific information. Finally, we test whether the predicted canonical states remain meaningful when used in place of native routing states. The transported states preserve local routing behavior, while learned state evolution reduces $Δ\mathrm{NLL}$ relative to simple persistence by 15.7\% on OLMoE and 6.2\% over a 10-router horizon on Phi.

Comment: Aligns router-control subspaces to identify reusable cross-layer expert-routing dynamics.

Topic Match: Expert-selection dynamics and routing-state substitution directly concern MoE routing; the shared geometric mechanism also fits architectural analysis.

Relevance: 9 Novelty: 7


Large-Scale Training Systems and Efficiency (4)

1. Momentum in large-batch training: Polyak enlarges the critical batch size, Nesterov improves data efficiency

ArXiv ID: 2609.02728

Primary Topic: Large-Scale Training Systems and Efficiency

Also Matches: Architecture and Training Dynamics

Authors: Jia-Nan Wang, Zixun Huang, Kairui Li, Lei Wu

Abstract: We study when and how momentum improves large-batch training in the one-pass regime, using power-law kernel regression as a tractable setting. We first characterize risk stability through the critical learning rate, defined as the largest learning rate for stable training, and obtain $η{\mathrm{SGD}}^{\mathrm{crit}}\eqsim 1$, $η\eqsim \min{1,B^β(1-ρ)}$, where $B$ is the batch size, $ρ$ is the momentum factor, and $β>1$ is the capacity exponent. Within this admissible region, we derive scaling laws for the full risk dynamics, capturing the progression from an early transient, through power-law decay, to a noise floor. We then minimize the final-step risk over the admissible learning rates and momentum factors under a fixed data budget, yielding a three-regime batch-size phase diagram that reveals how the role of momentum changes with batch size. Notably, Polyak enlarges the critical batch size, the largest batch size preserving the best small-batch data-scaling exponent, thereby enabling greater parallelism without sacrificing data efficiency. In contrast, Nesterov achieves better data efficiency in the large-batch regime because its look-ahead mechanism suppresses noise accumulation. Numerical experiments validate the predicted stability boundaries, risk dynamics, and batch-size phase diagram.}}^{\mathrm{crit}}\eqsim \min{1,B(1-ρ)}$, and $η_{\mathrm{Nesterov}}^{\mathrm{crit}

Comment: Kernel-regression scaling laws explain how Polyak and Nesterov momentum change critical batch size and data efficiency.

Topic Match: Batch-size, learning-rate, and momentum selection directly inform parallel training configuration; the supporting stability theory is established in power-law kernel regression.

Relevance: 9 Novelty: 8


2. BASP: Communication-Efficient Batch-Aware Sequence Parallelism for LLM Training

ArXiv ID: 2609.03151

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Bigyan Ghimire, Jon C. Calhoun

Abstract: Long-context reasoning for large language models (LLMs) is becoming increasingly important, but training over long sequences remains challenging due to massive memory and communication requirements. Sequence parallelism has emerged as an essential technique for addressing bottlenecks in long sequence LLM training. However, we observe that existing sequence parallelism methods are batch-agnostic and apply uniform sequence partitioning across all batch sizes, resulting in inefficient communication. In this paper, we introduce Batch- Aware Sequence Parallelism (BASP), a sequence parallelism approach that leverages batch structure to reduce communication overhead. BASP exploits batch structure by partitioning GPUs into disjoint sequence-parallel groups according to the micro- batch size. This design reduces the all-to-all communication group size, thereby localizing communication and improving training efficiency. Experimental results on an NVIDIA A100 cluster show that BASP improves end-to-end training time by up to 1.17 - 1.31x in Llama and Qwen models compared to standard sequence parallel baselines, while preserving identical model accuracy and memory usage.

Comment: Partitions sequence-parallel GPUs according to micro-batch structure to shrink all-to-all communication groups.

Topic Match: The core contribution is a distributed LLM training communication scheme, with reported end-to-end speedups at unchanged accuracy and memory usage.

Relevance: 10 Novelty: 6


3. Gradient Prediction with Control Variates in the Cheap-Forward Regime

ArXiv ID: 2511.05187

Primary Topic: Large-Scale Training Systems and Efficiency

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Kamil Ciosek, Nicolò Felicioni, Juan Elenter, Ehsan Imani

Abstract: We study whether otherwise-idle inference resources could reduce the scarce-GPU cost of training. Our analysis uses a simulated compute ledger in which fleet work is billed at a fraction of a scarce-GPU forward; all experiments run on a regular GPU. Our algorithm predicts gradients with a reduced-precision, inference-style reverse-mode program and combines many predictions with a few exact gradients through a control variate, so approximation error becomes variance rather than bias. On a 124M-parameter language model and selected short training windows, the method can lower simulated ledger cost relative to the tested baselines when fleet work is sufficiently cheap. Experiments spanning 10M-774M parameters show both transfers and failures. We do not test inference-only hardware, end-to-end distributed latency, or a full optimizer-by-batch-size baseline sweep.

Comment: Combines inexpensive reduced-precision gradient predictions with a few exact gradients through a control variate that converts approximation error into variance.

Topic Match: The gradient estimator directly targets training compute cost; reported savings depend on simulated fleet pricing and selected short training windows.

Relevance: 9 Novelty: 7


4. Nova: An End-to-End MLIR Compiler for Deep Learning

ArXiv ID: 2608.00029

Primary Topic: Large-Scale Training Systems and Efficiency

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Adwaid Suresh, Aparna A, Harshini V M, Jona Delcy C A, Killi Uma Maheswara Rao, Ram Charan Golla, Surendra Vendra

Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physical hardware. While high-level tensor frameworks provide flexible abstractions, their execution models inherently lack the whole-graph visibility required to maximize hardware utilization, often forcing a reliance on opaque, hand-written kernel libraries for complex operations like Attention. To bridge this gap, we present the next iteration of Nova, an automated end-to-end JIT compiler that achieves absolute control over hardware mapping by synthesizing fine-grained kernels directly from the computation's structure. In this work, we extend Nova's compilation pipeline to natively support full Transformer architectures. By capturing eager executions and unifying forward and backward passes into a single value-semantic dialect, Nova unlocks aggressive whole-graph optimizations. Rather than relying on rigid, pre-compiled library calls, Nova focuses on extensive cross-operator fusions, collapsing complex causal attention sub-graphs, element-wise operations, and memory-bound normalizations directly into single fused kernels to drastically reduce global memory roundtrips. In our evaluations training a full GPT-2 architecture on Ada 6000 GPUs, Nova demonstrates superior end-to-end throughput, averaging 441K tokens/second compared to 406K for our own eager execution and 405K for torch.compile. By drastically reducing memory-bound overheads through compiler-native fusion, Nova enables efficient full LLM compilation on modern hardware while strictly maintaining numerical parity.

Comment: Unifies forward and backward compilation and synthesizes fused kernels to reduce Transformer training memory traffic.

Topic Match: Compiler-driven execution and cross-operator fusion directly improve training throughput; reduced memory traffic also matches computational efficiency.

Relevance: 8 Novelty: 6


Architecture and Training Dynamics (7)

1. Multi-Mask Diffusion Language Models for Few-Step Generation

ArXiv ID: 2607.19686

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Sijin Chen, Yinuo Ren, Heyang Zhao, Ziheng Cheng, Quanquan Gu, Lexing Ying

Abstract: Masked diffusion models (MDMs) are a promising family of language generators, but achieving high-quality few-step generation remains challenging. In MDMs, all forward trajectories collapse to a single fully masked state, leaving no terminal entropy for consistency-style few-step generation. While recent few-step alternatives based on uniform-state diffusion avoid this degeneracy, it becomes harder to distinguish clean tokens from noise than MDMs, which usually harms modeling quality and training efficiency. In this work, we propose a multi-mask diffusion model (MultiMDM) that preserves the masking structure towards few-step generation. In the forward process, each clean token is first pushed towards a designated mask and then gradually mixes over the mask set. As a result, the backward process has a drafting capability by predicting a designated mask before refining to a clean token. We derive a closed-form ELBO training objective for MultiMDM that supports continual training from pretrained MDMs. In addition, we formulate a purely discrete-state consistency distillation scheme, with a shared-Gumbel coupling to reduce pathwise entropy. Experiments on pretraining and distillation show that MultiMDM provides an effective foundation for principled few-step generation.

Comment: Introduces a multi-mask diffusion process that preserves terminal entropy for few-step generation.

Topic Match: The core contribution changes the discrete diffusion process and derives its training objective; consistency distillation also reduces generation steps.

Relevance: 9 Novelty: 8


2. On the Expressive Power and Limitations of Multi-Layer SSMs

ArXiv ID: 2604.14501

Primary Topic: Architecture and Training Dynamics

Authors: Nikola Zubić, Qian Li, Yuyi Wang, Davide Scaramuzza

Abstract: We study how depth, finite precision, state dimension, and chain-of-thought (CoT) affect the expressive power of multi-layer state-space models (SSMs). For the explicit-table $K$-function-composition problem, a canonical benchmark for sequential information propagation, we prove that any $L$-layer SSM solving $(L+3)$-function composition must satisfy $d^2p=Ω(N/L^3)$, where $d$ is the state dimension and $p$ is the per-scalar precision. Conversely, $K$-function composition is solved exactly by a $(K+1)$-layer generalized SSM with $d=1$ and $p=Θ(\log N)$. This gives a worst-case depth hierarchy for this formal problem family. We then distinguish post-input reasoning, in which all thought tokens are generated after the input, from input-interleaved reasoning, in which thought tokens may be inserted while the input stream is being read. Post-input reasoning does not circumvent our communication-based lower-bound pipeline, whereas input-interleaved reasoning admits bidirectional simulations with general deterministic one-pass streaming algorithms at the granularity of persistent memory. Finally, width and precision are not interchangeable under exact step-preserving simulation in the base affine-state model, but become interchangeable through the streaming-memory characterization once input-interleaved reasoning is allowed.

Comment: Depth-state-precision bounds and streaming equivalences characterize the computational limits of multilayer SSMs.

Topic Match: The core results analyze a recurrent architecture and establish how depth, persistent state, and reasoning placement change its expressive capabilities.

Relevance: 9 Novelty: 8


3. MDM-Prime-v2: Binary Encoding and Index Shuffling Enable Scaling of Diffusion Language Models

ArXiv ID: 2603.16077

Primary Topic: Architecture and Training Dynamics

Authors: Chen-Hao Chao, Wei-Fang Sun, Junwei Quan, Chun-Yi Lee, Rahul G. Krishnan

Abstract: Masked diffusion models (MDM) exhibit superior generalization when learned using a Partial masking scheme (Prime). This approach converts tokens into sub-tokens and models the diffusion process at the sub-token level. We identify two limitations of the MDM-Prime framework. First, we find that the functional form of the subtokenizer significantly increases the cross-entropy loss in the objective when paired with commonly used Byte-Pair-Encoding (BPE) tokenizers. Second, we lack tools to guide the hyperparameter choice of the token granularity in the subtokenizer. To address these limitations, we analyze the optimal design of the subtokenizer that minimizes MDM-Prime training objective and develop MDM-Prime-v2, a masked diffusion language model which incorporates Binary Encoding and Index Shuffling. Our analysis characterizes how token granularity and sub-token entropy influence the training objective and downstream performance, providing principled criteria for subtokenizer design. When extending the model size to 1.1B parameters, MDM-Prime-v2 demonstrates superior average zero-shot accuracy across eight commonsense reasoning benchmarks, outperforming similar-sized baselines including GPT-Neo, OPT, Pythia, Bloom, SMDM, and TinyLLaMA.

Comment: Derives how subtoken granularity and entropy affect masked-diffusion training objectives, motivating binary encoding and index shuffling.

Topic Match: The core contribution connects subtokenizer design to the training objective and behavior of diffusion language models.

Relevance: 9 Novelty: 7


4. CHASE: Cache-Hole-Adapted Skip Exit for Looped State-Space Language Models

ArXiv ID: 2607.10110

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Zhenxuan Yu, Takeshi Kojima, Yutaka Matsuo, Yusuke Iwasawa

Abstract: Recent work on looped language models suggests that many reasoning problems benefit from greater computational depth rather than from additional independent parameters. Existing studies, however, focus almost exclusively on Transformer backbones, leaving open whether this principle also applies to state-space language models. We investigate Looped Mamba and Looped Hybrid Mamba-Transformer architectures, which repeatedly apply a shared Mamba (or hybrid) block to introduce explicit finite-depth recurrent computation. On two controlled reasoning tasks-Mano (modular-arithmetic manipulation) and p-hop induction-Looped Mamba consistently outperforms parameter-matched non-looped baselines and, in several settings, matches or exceeds non-looped models of equal effective depth. We then extend the study to language model pre-training under matched iso-parameter and iso-FLOPs protocols, which jointly disentangle the effects of parameter sharing and effective depth: looped models remain competitive on downstream benchmarks with substantially fewer distinct parameters, although deeper non-looped models retain an advantage in validation perplexity under strict iso-FLOPs comparisons. Finally, we adapt Ouro's two-stage exit gate to Looped Mamba for threshold-controlled selection among recurrent-step outputs. Executing such exits on a state-space backbone, however, leaves the recurrent state without its deeper updates, and validation perplexity then degrades severely. We therefore introduce a cache-hole adaptation that aligns continued training with skipped-state inference. At the scales studied, the adapted model keeps perplexity close to full computation and matches or exceeds full-compute exit-state selection on downstream benchmarks while executing roughly half of the recurrent steps, which translates into measured inference speedups once the prefill is compute-bound.

Comment: Cache-hole-aware continued training enables adaptive recurrent-step skipping in looped state-space models.

Topic Match: Shared recurrent depth and adaptive exits are core architectural mechanisms; adapting recurrent states to skipped computation also reduces execution cost.

Relevance: 9 Novelty: 7


5. The Dynamics of Continuous Mixture Collapse in Language Models

ArXiv ID: 2609.02049

Primary Topic: Architecture and Training Dynamics

Authors: Ali Backour

Abstract: LLMs latent-state reasoning methods replace discrete intermediate tokens with continuous states, such as weighted mixtures of token embeddings, to retain multiple possible reasoning directions rather than committing to one. Yet pretrained language models often fail to preserve these mixtures. We study why through a combination of theoretical analysis and controlled empirical investigations on a variety of models. We identify three independent, distinct sources of failure. First, transformer architectures already distort mixture geometry, and training substantially amplifies this effect. Moreover, the failure can occur even if the model transports mixtures perfectly linearly: the softmax readout and autoregressive feedback form a dynamical system that either amplifies small differences until one component of the mixture dominates or contracts different mixtures until they become indistinguishable. We verify this theoretical prediction empirically: the observed transition between contraction and amplification occurs near the theoretical threshold derived by our analysis, and pretrained-model rollouts lie predominantly on the amplifying side. Finally, we generalize to mixtures of many components and show that exact preservation generally requires context-dependent correction, whose required dimensionality can grow with the number of components.

Comment: Derives when softmax and autoregressive feedback amplify or erase continuous-state mixtures.

Topic Match: The core contribution explains architectural failure modes of continuous-state computation, including how training amplifies mixture distortion.

Relevance: 8 Novelty: 8


6. Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders

ArXiv ID: 2605.30022

Primary Topic: Architecture and Training Dynamics

Authors: Pierre-Antoine Lequeu, Camille Barboule, Benjamin Piwowarski

Abstract: Positional encoding (PE) underpins how permutation-invariant Transformers represent sequence order, yet how positional information is processed and stored remains poorly understood. Modern PE methods such as RoPE still struggle on tasks such as long-context understanding or retrieval \cite{chen-etal-2025-hope}. Hence, a better understanding of the internal positional mechanism could help design better PE. Building on evidence that positional and semantic signals occupy nearly orthogonal subspaces in trained Transformers, we modify an encoder Transformer to process three explicitly disentangled streams: semantic, absolute positional (AP) and relative positional (RP), and confine the masked-language-modeling (MLM) objective to the semantic stream. This decoupling enables a clean mechanistic study and yields three take-aways. (1) The isolated AP subspace spontaneously collapses into a low-frequency two-dimensional manifold that captures the structure of the document; (2) Attention heads specialize into structure and semantic-oriented groups, with RP exclusively supporting the latter; (3) Standard positional encodings do not robustly retain macroscopic structure: RoPE and RP only weakly encode it, and entangled AP loses it in the final layers under MLM pressure. The disentangled approach preserves positional encoding, which improves linguistic representation on 49 of the 65 linguistic phenomena of the Flash-Holmes probing benchmark.

Comment: Separates semantic, absolute-position, and relative-position streams to study how masked-language-model training preserves sequence structure.

Topic Match: Explicit stream separation changes the positional architecture and isolates its interaction with the training objective, providing a concrete architectural mechanism despite the narrow encoder setting.

Relevance: 8 Novelty: 7


7. Gradient Descent on Logistic Regression with Non-Separable Data and Large Step Sizes

ArXiv ID: 2406.05033

Primary Topic: Architecture and Training Dynamics

Authors: Si Yi Meng, Antonio Orvieto, Daniel Yiming Cao, Christopher De Sa

Abstract: We study gradient descent (GD) dynamics on logistic regression problems with large, constant step sizes. For linearly-separable data, it is known that GD converges to the minimizer with arbitrarily large step sizes, a property which no longer holds when the problem is not separable. In fact, the behaviour can be much more complex -- a sequence of period-doubling bifurcations begins at the critical step size $2/λ$, where $λ$ is the largest eigenvalue of the Hessian at the solution. Using a smaller-than-critical step size guarantees convergence if initialized nearby the solution: but does this suffice globally? In one dimension, we show that a step size less than $1/λ$ suffices for global convergence. However, for all step sizes between $1/λ$ and the critical step size $2/λ$, one can construct a dataset such that GD converges to a stable cycle. In higher dimensions, this is actually possible even for step sizes less than $1/λ$. Our results show that although local convergence is guaranteed for all step sizes less than the critical step size, global convergence is not, and GD may instead converge to a cycle depending on the initialization.

Comment: Proves that gradient-descent step sizes guaranteeing local convergence can still produce stable cycles from other initializations.

Topic Match: The distinction between local stability and global optimization dynamics directly fits training-dynamics analysis, although the results concern logistic regression.

Relevance: 7 Novelty: 8


Efficiency, Compression, and Large-Scale Training (23)

1. LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style Updates

ArXiv ID: 2609.02734

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Dmitrii Andriianov, Andrey Veprikov, Aleksandr Beznosikov

Abstract: Low-rank adaptation (LoRA) is the standard way to fine-tune large models, yet when its two factors are trained independently, the update ignores the geometry of the low-rank weight change it induces. We introduce LoRA-TSD, an optimizer that treats every LoRA step as a tangent vector of the fixed-rank matrix manifold and takes the spectral-norm steepest-descent step of Muon inside that tangent space, mapping the result back to the factors through a retraction native to the LoRA parametrization. The step avoids expensive operations on full weight matrices, and its retraction is up to $2.8\times$ cheaper than the truncated-SVD retraction used by prior manifold methods. We prove that the Frobenius-norm version of our surrogate recovers LoRA-Pro, and we identify the tangent-projected gradient, the Riemannian gradient of the manifold, as the stationarity measure natural to LoRA training and computable from the factor gradients alone. Under this measure we give the first global convergence guarantees for both LoRA-Pro and LoRA-TSD, with rates that drive the factor-gradient norms to zero. Across six commonsense and natural-language-inference benchmarks with Llama-3.2-1B, Llama-3.1-8B and Qwen3-32B, LoRA-TSD outperforms every competing LoRA optimizer and stays robust to the adapter rank. Code is available at https://github.com/brain-lab-research/LoRA-TSD.

Comment: Muon-style tangent-space updates optimize LoRA through efficient factor-space retractions without full-weight-matrix operations.

Topic Match: The central mechanism improves low-rank adaptation efficiency; its manifold geometry and convergence guarantees also address training dynamics.

Relevance: 9 Novelty: 8


2. A Mathematical Theory of Reusable Neural Bases for Network Compression

ArXiv ID: 2609.01550

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Binshuai Wang, Peng Wei

Abstract: As large AI models become increasingly prevalent across a wide range of applications, memory cost has become a critical bottleneck in both training and inference. To mitigate this issue, we introduce the Linear Reusable Neural Bases Architecture (LRNBA), a novel framework aimed at improving parameter efficiency and reducing memory cost. Inspired by recurrent neural network (RNN) designs, the core idea of our approach is to represent each network block as a linear combination of a shared set of neural bases, thereby enjoying highly network compression rate while maintaining stable training. The proposed architecture allows for the construction of significantly wider and deeper networks under the same parameter budget. Extensive experiments demonstrate that our model achieves comparable or even faster convergence and lower loss than classical architectures, while maintaining stable training dynamics.

Comment: Compresses network parameters by expressing each block as a linear combination of shared neural bases.

Topic Match: Parameter sharing and memory reduction are central; the reusable-basis architecture also provides a secondary architectural match.

Relevance: 9 Novelty: 7


3. HARP: Hadamard-Preconditioned Adaptive Rotation Processor for Extreme LLM Quantization

ArXiv ID: 2605.29843

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Artur Zagitov, Gleb Molodtsov, Aleksandr Beznosikov

Abstract: Post-training quantization (PTQ) is essential for deploying LLMs under memory and bandwidth constraints. However, extreme low-bit quantization remains highly sensitive to activation outliers and anisotropic weight curvature. Existing incoherence-based PTQ methods mitigate this issue with fixed randomized Hadamard transforms (RHTs), which improve quantization robustness but cannot adapt the rotated basis to the layer, calibration distribution, or quantizer. We introduce HARP (Hadamard-preconditioned Adaptive Rotation Processor), a learnable structured two-sided orthogonal processor that replaces fixed Hadamard mixing while preserving exact full-precision equivalence. HARP represents each rotation as a product of sparse butterfly-like block-orthogonal stages, supports non-power-of-two dimensions through Mixed-Radix schedules, and initializes to the RHT processor up to a fixed permutation. Fitted only on calibration data, HARP adapts the quantization basis to each layer and backend. Across 2--4-bit settings on Llama models from 1B to 70B, HARP consistently improves perplexity and yields its clearest zero-shot gains at 2 bits; a 2-bit Qwen3-8B experiment shows the same transfer beyond the Llama family. HARP also preserves deployment efficiency: on Llama 2 7B at 2 bits, it reaches 128 tok/s, retaining 90% of RHT throughput (142 tok/s) and running approximately $2.1\times$ faster than FP16 (61 tok/s).

Comment: Learns structured orthogonal rotations that adapt the quantization basis for 2–4-bit LLMs.

Topic Match: Adaptive low-bit quantization is the central contribution, with structured rotations improving compression accuracy while retaining efficient execution on large LLMs.

Relevance: 9 Novelty: 7


4. HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models

ArXiv ID: 2609.02029

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Renjie Xie, Juncheng Yang, Aoting Hu, Mingxi Zhang, Liyao Wu, Zheheng Hong, Wei Xu

Abstract: Long-context inference retains a growing key--value (KV) cache during decoding, which consumes substantial GPU memory and can reduce generation throughput. This bottleneck remains in hybrid language models because their residual global-attention layers can dominate context-dependent cache demand. We study how to allocate this state under an aggregate KV-residency budget. We introduce HeadWiseKV, a training-free framework that compresses the residual global KV caches of hybrid language models while preserving their native local, recurrent, and linear paths. It assigns each physical KV head a static, multilevel history window, making cache demand predictable before serving. We formulate this allocation as a restricted operational rate--distortion problem and propose SeqCalib as the core policy-generation algorithm in HeadWiseKV. SeqCalib processes layers in execution order and conditions each decision on the lower-layer policy used at deployment, thereby accounting for interactions across depth. A grouped-cache runtime materializes the selected policy as actual per-head KV residency rather than a mask over a full cache. We evaluate downstream quality across four hybrid long-context models and study physical residency and serving behavior on Qwen3.6-27B. HeadWiseKV retains near-Full-KV RULER and LoCoMo quality across the evaluated models. In the fixed-model systems study, it reduces sampled peak device memory by 8.59\% at a 112K context length and extends the largest verified successful context from 114K to 161K.

Comment: Allocates physical per-head KV-cache windows under a memory budget using calibration conditioned on preceding layers' compression policies.

Topic Match: Budgeted KV-cache compression is the core contribution, with an allocation algorithm and runtime that realize actual memory savings in hybrid LLMs.

Relevance: 9 Novelty: 7


5. Enabling KV Caching of Shared Prefix for Diffusion Language Models

ArXiv ID: 2606.07571

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Younghun Go, Jaehoon Han, Changyong Shin, Chuck Yoo, Gyeongsik Yang

Abstract: Key-value (KV) caching for shared prefixes is essential for high-throughput large language model (LLM) serving, but it faces critical challenges in emerging diffusion language models (DLMs). In DLMs, bidirectional attention means that updating any token dynamically alters the entire context and its corresponding KVs. Thus, existing caching techniques developed for LLMs, which assume that KVs remain invariant once computed, corrupt the shared prefix KVs. Our experiments show that applying these techniques to DLMs causes model accuracy to collapse to near zero. To unlock high-throughput DLM serving, we propose bidirectional prefix caching, BiCache, the first KV caching technique for shared prefixes in DLMs. BiCache is designed based on key observations from our comprehensive analysis: shared prefix KVs remain stable and reusable in shallow layers, while the depth of shallow layers depends on the fraction of shared prefix tokens in each request. Thus, BiCache dynamically identifies a safe layer depth for reusing shared prefix KVs and eliminates redundant computation. Evaluations demonstrate that BiCache significantly improves serving throughput by 36.3%-98.3% compared to existing techniques without accuracy collapse (only 0-1.8% difference).

Comment: Adaptively reuses shared-prefix KVs in diffusion-model layers where those states remain sufficiently stable.

Topic Match: The core contribution is a new depth-adaptive cache-reuse mechanism that reduces computation despite bidirectional attention.

Relevance: 9 Novelty: 7


6. Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs

ArXiv ID: 2609.00575

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Seungwoo Jung, Dohyeok Kwon, Seungmin Cha, Junseok Lee, Yeonho Yoo, Chuck Yoo, Gyeongsik Yang

Abstract: Mixture-of-experts (MoE) architectures scale large language models efficiently, but they demand massive GPU memory. To cope with such demand, models are commonly compressed to reduce their memory footprint. Residual sparsification is a representative compression technique that decomposes each projection matrix of an expert into a shared base matrix and per-expert residual matrix, and then compresses the residuals. Existing sparsification methods compress each residual matrix independently by minimizing its compression error, thereby minimizing the error of each projection matrix. However, our analysis shows that this objective is misaligned with preserving model accuracy after compression. In an expert, the final output is produced through computations coupled across multiple projections and hidden representations. Therefore, even small errors in individual matrices can propagate through hidden representations and projection interactions, leading to large expert output errors and accuracy degradation. To address this misalignment, we propose PARSER, a new residual sparsification method that shifts the compression objective from minimizing isolated matrix errors to preserving the expert output error. PARSER achieves this by introducing output importance, which measures the actual contribution to the expert output error. Our experiments show that, compared with existing methods, PARSER narrows the accuracy gap to the uncompressed model by 1.41$\times$ on Qwen and 1.44$\times$ on DeepSeek, while achieving the same peak memory reduction. Our code is available at https://github.com/OSSS-KU/PARSER.

Comment: Sparsifies MoE expert residuals using output importance that accounts for interactions across expert projections.

Topic Match: Expert-weight compression is the core contribution, with an output-aware objective that improves accuracy at the same memory reduction.

Relevance: 9 Novelty: 7


7. QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization

ArXiv ID: 2609.00224

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Yipin Guo, Arun M George, Jie Fu, Tareq Mahmoud, Sixue Xing, Siddharth Joshi

Abstract: Weight-only post-training quantization (PTQ) can alleviate the computational burden of serving large language models (LLMs) at scale. However, existing PTQ methods often fail to generalize across models and suffer severe accuracy loss below 2 bits. Many leverage unstructured sparsity to mitigate this loss, but at the cost of regularity and GPU-friendly execution. We present QTEA, a sub-2-bit PTQ framework that quantizes weights into ternary values and uses salient weights as residual error compensators. To maintain hardware efficiency, residuals are assigned to selected columns with semi-structured $1:4$ sparsity within the salient columns. We further add column-wise rescale refinement to GPTQ-style column-by-column quantization, alternately updating per-column scales and ternary assignments to reduce reconstruction error. We also identify order-dependent error propagation in GPTQ and introduce error decay to attenuate late-stage error accumulation. On Qwen3-14B, QTEA compresses all weights to an effective 1.7 bits per weight while improving average accuracy over the strongest ternary PTQ baseline by 16.7%. It also achieves 1.40$\times$ and 2.61$\times$ lower perplexity on WikiText and C4 respectively. This trend holds on Llama3-8B, where QTEA obtains a 6.6% accuracy gain and 1.34$\times$ / 1.95$\times$ lower perplexity on the same datasets. Finally, we develop a lookup-table based kernel that achieves 7.2$\times$ faster per-token generation over an FP16 baseline. Code is available at https://github.com/Intelligent-Microsystems-Lab/QTEA.

Comment: Enables sub-2-bit LLM quantization through ternary weights, structured residual compensation, and column-wise error control.

Topic Match: The core contribution combines extreme weight compression with a compatible lookup-table kernel to reduce LLM memory and generation cost.

Relevance: 9 Novelty: 7


8. Train What You Deploy: Closing the MLP Reachability Gap in Low-Rank Clone Distillation

ArXiv ID: 2609.02006

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Wenhui Chen, Zhifeng Li, Jie Zhou, Navan Preet Singh, Madalina Ciobanu, Chenghua Wang, Qingqing Mao, Ritankar Das

Abstract: A compressed student has two shapes that need not agree: the weight it deploys at inference and the weight family its training can reach. We show that a state-of-the-art weight-inheritance distiller, Low-Rank Clone (LRC), deploys a full-width student MLP but ties training to a teacher-induced slice, leaving 62.5-81.4% of each deployed matrix's independent linear degrees of freedom unreachable-paid for at inference, never trainable. Our principle is one line: train what you deploy. From the identical LRC warm start, we make the training object the entire deployed matrix, with no change in deployed shape, deployed parameter count, or inference FLOPs, via two mergeable realizations (Dense-LRC and CORE-LRC) that both collapse to one deployed weight. This recovers stranded capacity: taking the stronger realization per teacher, +2.36/+2.71/+10.45 Avg9 over matched-budget plain-LRC baselines across three teachers (Llama3.2-3B, Llama3.1-8B, Qwen2.5-3B), with the largest gain on the widest teacher (Qwen), where it reaches the original recipe's approx. 20B-token accuracy at 10B tokens (2x token efficiency); there the strictly same-lineage arm still recovers +6.39, the fully controlled figure. Controls strongly support attributing the gain to the enlarged reachable set, rather than to added parameters or the recipe. From approx. 10B distillation tokens plus a short SFT, a half-parameter 1.5B student matches its approx. 9T-token teacher's 9-task macro-average, within evaluation noise and with a residual MMLU deficit, and a 2.7B student beats Meta's own official compression of Llama3.1-8B at ~900x fewer compression tokens (a token count under unmatched recipes, not a compute claim). All results are from single-seed runs on the LRC backbone.

Comment: Removes teacher-imposed low-rank training restrictions so compressed students can optimize their full deployed MLP capacity.

Topic Match: Directly improves compressed-student training by expanding reachable weights without increasing deployed size; parameterization analysis explains the gains, although evidence is single-seed.

Relevance: 9 Novelty: 7


9. XMerge: Cross-Axis Selection and Reconstructive Layer Merging for LLM Depth Compression

ArXiv ID: 2609.02083

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Jundong Hu, Shekar Ramachandran

Abstract: Removing complete transformer layers preserves a standard serving architecture, but existing depth-compression methods can lose substantial quality, and the loss varies unpredictably across models. We introduce XMerge, a post-training method with two components. Cross-axis selection identifies a block with low relative-magnitude and angular hidden-state change, and local boundary reconstruction re-fits the adjacent surviving block to match the original two-block output. XMerge uses no task labels or end-to-end fine-tuning, and it introduces neither architectural changes nor additional inference-time parameters. Across seven Llama and Qwen backbones (0.5B-8B), five published baselines, and three layer-reduction levels, its advantage over baselines is largest at the most aggressive removal: at k=4 it ranks first on six of seven backbones on CORE (a 22-task aggregate) and, separately, on six of seven on MMLU (five of seven on both at once), while avoiding the large perplexity increases of several competing operators. In a task-level bootstrap, the 95% confidence intervals for the three largest CORE margins exclude zero; the remaining margins are consistent with ties. Across the 14 (model, regime) cells it is also the only evaluated operator that never collapses, ranking top-2 in both zero-shot and in-context regimes; on a first calibration probe (one backbone) it is the best-calibrated operator. Ablations show that local reconstruction provides most of the gain, while cross-axis fusion helps when the two selection axes disagree. The additional construction cost is recovered through per-token decode savings after roughly tens of thousands of requests.

Comment: Locally reconstructs surviving blocks after layer removal to preserve quality while reducing LLM depth and decoding cost.

Topic Match: The core contribution is structured LLM compression through layer selection and local reconstruction, requiring no end-to-end fine-tuning.

Relevance: 9 Novelty: 6


10. TaRA: Training-Aware Low-Rank Adaptation Initialization

ArXiv ID: 2609.02639

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Taehyeon Kim, Eunhyeok Park

Abstract: Low-Rank Adaptation (LoRA) has become a de facto standard for parameter-efficient fine-tuning (PEFT), yet its performance is highly sensitive to initialization due to the information bottleneck imposed by low-rank decomposition. Existing approaches attempt to construct high-quality LoRA initializations by exploiting principal components of pretrained weights, activations, or gradients. However, these methods do not directly account for the training dynamics of the full-rank model. In this paper, we propose Training-aware Low-Rank Adaptation Initialization (TaRA), a method that initializes LoRA such that the gradients induced by the low-rank factors closely approximate the gradient of the corresponding full-rank weight matrix. Derived from a mathematical formulation, TaRA improves gradient fidelity at the start of training while introducing negligible computational overhead. Across diverse and challenging fine-tuning tasks, TaRA consistently outperforms prior state-of-the-art methods, establishing a simple, robust, and scalable solution for effective LoRA initialization.

Comment: Initializes LoRA factors so their induced gradients better approximate full-rank weight gradients.

Topic Match: The core contribution is a general low-rank adaptation initialization method that improves gradient fidelity with negligible additional overhead.

Relevance: 9 Novelty: 6


11. LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference

ArXiv ID: 2609.03079

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Renyuan Liu, Yuyang Leng, Kaiyan Liu, Yuzhou Zhong, Shaohan Hu, Chun-Fu, Chen, Peijun Zhao, Heechul Yun, Shuochao Yao

Abstract: On-device LLM inference is attractive for privacy and responsiveness, but remains challenging on mobile and embedded devices because model weights far exceed available DRAM. Prior systems exploit activation sparsity and offload weights to SSD or flash storage, but face a fundamental systems trade-off: accurate sparse execution decisions require the latest context, whereas efficient computation-I/O overlap requires early prediction. As a result, existing designs either serialize execution or incur redundant weight fetches, extra computation, and large cache overheads. We present LeanStream, a streaming speculate-and-refine framework for efficient on-device LLM inference. LeanStream progressively refines computation, loading, and cache-retention priorities using partial GPU results, enabling fine-grained overlap between GPU execution and storage I/O. We implement LeanStream on both mobile and embedded platforms. Compared with prior on-device LLM inference systems, LeanStream reduces memory usage by 4.8$\times$ to 7.5$\times$ at the best throughput achieved by prior work, while further improving token generation throughput by 1.6$\times$ to 2.1$\times$.

Comment: Refines sparse weight-loading and cache priorities using partial GPU results to overlap execution with storage I/O.

Topic Match: The new weight-streaming mechanism materially reduces LLM inference memory and I/O stalls, although its focus is on-device execution.

Relevance: 8 Novelty: 7


12. Routing Is Not Enough: Diagnosing Intra-Adapter Subspace Contention in MoE+LoRA Fine-Tuning

ArXiv ID: 2609.03150

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Mehreen Hossain Chowdhury, Nowshin Mahjabin, Ahmed Shafin Ruhan, Md Azam Hossain, Abu Raihan Mostofa Kamal, Md Tahmid Rahman Laskar

Abstract: Multi-domain fine-tuning often combines MoE routing with LoRA, assuming that token-level routing separates domain-specific updates. We test this assumption in MoE+LoRA using Python code paired with biomedical text and mathematical reasoning. Although these domains show near-disjoint expert routing, adding biomedical data substantially increases code perplexity, indicating that routing separation alone may not prevent negative transfer. To localize the failure, we introduce Jaccard routing overlap and adapter-gradient cosine similarity, which measure expert sharing and update compatibility, respectively. These diagnostics indicate that interference arises mostly from nearly orthogonal domain gradients competing within the same low-rank adapter subspace. We address this issue with SpawnLoRA, which dynamically adds gated sub-adapters inside MoE experts when adapter-level contention is detected, while keeping the router fixed. We evaluate SpawnLoRA on Phi-tiny-MoE-instruct and OLMoE-1B-7B across multiple mixture settings and find that it effectively reduces negative transfer compared with standard and rank-adaptive LoRA. These results demonstrate that structural separation inside experts provides benefits beyond routing or rank expansion alone.

Comment: Dynamically spawns gated low-rank sub-adapters when domain gradients compete within an adapter subspace.

Topic Match: The methodological contribution restructures low-rank adaptation, with dynamic intra-expert modules providing a secondary architectural match; the MoE router remains fixed.

Relevance: 8 Novelty: 7


13. Scalable Kronecker-Fisher Approximation: Efficient Hessian Analysis for Billion-Parameter Language Models Compression

ArXiv ID: 2609.02451

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Viacheslav Yusupov, Daria Cherniuk, Evgeny Frolov

Abstract: In this paper, we propose a scalable Kronecker-based approximation that captures cross-layer interactions without storing the entire Fisher matrix, enabling practical Hessian analysis for billion-parameter networks where full computation is infeasible. Our approach reveals consistent vulnerability patterns: value projection layers exhibit the highest sensitivity and strongest cross-layer correlations across multiple model families, while other components exhibit architecture-specific behaviors. Through extensive experiments on quantization, sparsification, inter-layer corruption, and post-corruption fine-tuning, we demonstrate that our approximation strongly correlates with both performance degradation and recovery. Our framework provides a practical, theoretically grounded tool for identifying fragile components in large models, opening new avenues for guided compression and optimization strategies, such as mixed-precision allocation, layer-wise sparsity, and adaptive low-rank decomposition across layers and even individual weight groups.

Comment: A scalable cross-layer Fisher approximation estimates compression sensitivity without storing the full Fisher matrix.

Topic Match: The new approximation makes curvature-based compression analysis tractable for billion-parameter models; its demonstrated contribution is analysis rather than a complete compression algorithm.

Relevance: 8 Novelty: 7


14. DLM-One: Diffusion Language Models for One-Step Sequence Generation

ArXiv ID: 2506.00290

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Tianqi Chen, Shujian Zhang, Mingyuan Zhou

Abstract: This paper introduces DLM-One, a score-distillation-based framework for one-step sequence generation with continuous diffusion language models (DLMs). DLM-One eliminates iterative refinement by aligning the scores of a student model's outputs with the score function of a pretrained teacher DLM in the forward-diffused noisy space. We demonstrate that our framework is architecture-agnostic and robust across diverse continuous manifolds, including standard token embedding spaces and logit simplex spaces. Through experiments on multiple representative DLMs, we show that DLM-One achieves up to $\sim$2000$\times$ speedup in sampling steps and $\sim$500$\times$ in wall-clock time, while maintaining competitive performance on benchmark text generation tasks. We further analyze failure modes in language-domain diffusion distillation and propose an adversarially-regularized two-stage training scheme to prevent student degeneration. Our findings position one-step score distillation as a viable path for the efficient deployment of continuous diffusion models operating in continuous space for natural language processing.

Comment: Score distillation replaces iterative continuous-diffusion sampling with one-step sequence generation.

Topic Match: Eliminating iterative sampling directly reduces inference cost; analysis and prevention of student degeneration also address training stability.

Relevance: 8 Novelty: 7


15. Accelerating Unified Multimodal Models with Core-Expansion Routing and Unified Computation Scheduling

ArXiv ID: 2608.29291

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Wengyi Zhan, Chenqian Yan, Songwei Liu, Mingbao Lin, Rongrong Ji

Abstract: Unified multimodal models jointly support understanding and generation, but incur substantial redundant computation across tokens, layers, and generation timesteps. Through token-importance probing, we identify an asymmetric core-expansion structure: understanding exhibits a stable importance component, while generation largely shares this component but requires progress-dependent corrections. We therefore propose CE-Router, which uses a task-shared core scorer and progress-conditioned generation expansions, optimized through generation decomposition and cross-task core alignment. At inference, CE-Router compacts token computation and supplies a learned routing signal to Unified Computation Scheduling, which coordinates layer skipping, FFN pruning, diffusion-head cache reuse, and denoising-step early exit. Experiments on two representative UMM architectures demonstrate consistent quality--efficiency improvements across both tasks, retaining 98.03\% of dense understanding performance with a 1.93$\times$ end-to-end inference speedup.

Comment: Learned core-expansion routing coordinates sparse computation across tokens, layers, and generation timesteps.

Topic Match: The core contribution reduces unified-model inference computation through learned routing and coordinated skipping, pruning, and cache reuse, reporting 1.93x end-to-end speedup.

Relevance: 8 Novelty: 7


16. Who Speaks for the Pruned? Visual Token Pruning as Coverage Optimization

ArXiv ID: 2609.03158

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Qingchan Zhu, Weihang You, Hanqi Jiang, Changdi Yang, Tianming Liu, Geng Yuan

Abstract: Visual token pruning reduces the inference cost of vision-language models (VLMs), but most methods only ask which tokens to keep. This retained-token view can keep redundant high-scoring tokens while leaving discarded evidence without a close representative. We propose CoverPruner, a training-free pruner that asks the complementary demand-side question: after a token is removed, which surviving original token represents it for the target VLM? CoverPruner formulates pruning as Representational Coverage Maximization (RCM), covering the full projected visual-token set with query-weighted demand. It instantiates RCM with projector-space coverage and a lightweight first-layer attention probe. Across multiple VLM architectures and compression rates, CoverPruner achieves the best average accuracy among all compared methods, with the largest gains usually appearing under aggressive compression.

Comment: Formulates visual token pruning as query-weighted coverage so retained tokens represent discarded evidence.

Topic Match: The core is a token-compression mechanism for reducing VLM inference cost, validated across architectures and compression rates.

Relevance: 8 Novelty: 6


17. Unfolding the Leech Lattice: Fused Multi-Shell Decoding and VRAM Layouts for 2-Bit LLM Weights

ArXiv ID: 2609.02652

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Pier-Jean Malandrino

Abstract: Leech-lattice vector quantization holds the strongest reported 2-bit quality under its own evaluation protocol. Its kernel decodes one shell; we found no implementation of the multi-shell decoder the rate requires. This paper supplies one and measures its serving cost for decode-phase GEMV at batch 1. First, a serving path for the full 301-class codebook: an offline expansion into GPU layouts and a fused dequantize-plus-matvec kernel reading them without warp divergence, verified against f64. Second, the in-VRAM rate is a design axis distinct from the on-disk rate. Four bit-exact layouts timed in one process show binary bit planes beating one-hot masks on size and speed at constant bandwidth (4.80 bits per weight, 2.15x FP16). Below 4.3 bits a second, irregular stream enters; at 3.6 the decode stops being shifts and masks. Third, deployed four-bit (AWQ) and two-bit (QTIP) GEMV kernels run in the same process. The trellis kernel reads 2.40x fewer bytes than our served layout and runs 2.27x faster at near-equal fractions of their byte bounds: the time gap tracks the traffic gap, the price of unfolding a codebook too large for a lookup table. Fourth, the validity envelope: the trellis kernel outruns our no-weights control, so our launch geometry sets that floor, and on a second memory hierarchy every lattice arm falls below FP16. With the output head held identical across arms, the kernel-and-format path gains 1.11x, 1.29x and 1.41x end to end at 4B, 8B and 14B; with an int8 output head the served 4B reaches 87.0 tok/s in 2.60 GB. The quality cost, 1.38x perplexity and 14.7 MMLU points at 4B, shrinks across the three sizes measured.

Comment: Fuses multi-shell lattice weight decoding with matrix-vector multiplication and designs GPU layouts around the effective in-VRAM bit rate.

Topic Match: The decoder and layout contribution directly addresses compressed LLM inference costs, while exposing why nominal 2-bit storage does not imply 2-bit GPU residency.

Relevance: 8 Novelty: 6


18. Compute-in-Memory Attention: A Time-Domain Analog Softmax Circuit with RC-Tunable Temperature

ArXiv ID: 2609.04266

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Ankur Singh, Ashish Gautam, Shruti R. Kulkarni, Guojing Cong

Abstract: Softmax is a key operation in Transformer attention, but its exponentiation and normalization add significant overhead in compute-in-memory (CIM) accelerators, especially when analog attention scores must first be converted to the digital domain. This work presents a tunable-temperature analog softmax circuit in GlobalFoundries 22-nm fully depleted silicon-on-insulator (FDSOI) technology that operates directly on CIM-generated score voltages without intermediate analog-to-digital conversion. Each input score is converted into a time-domain event using a shared falling ramp. The corresponding comparator transition samples an RC-decaying reference to generate an exponential weight, which is then processed by an in-circuit normalization stage. In contrast to analog softmax circuits that rely on transistor weak-inversion behavior for exponentiation, the proposed architecture controls the softmax response through the ramp slope and RC time constant, enabling programmable effective temperature. The 128-element architecture is evaluated using transistor-level and post-layout extracted simulations, including multi-level input vectors, capacitance variation and mismatch, process and temperature variation, monte carlo analysis, and shared-interconnect parasitics. The complete 128-element implementation occupies 9453.42~$μ\mathrm{m}^{2}$ including the shared global ramp circuitry, while each replicated softmax element occupies 70.2~$μ\mathrm{m}^{2}$. The circuit achieves a 242.97-ns evaluation latency at 13.44~mW total power, corresponding to 25.5~pJ per output element. The simultaneous 128-element evaluation achieves an RMSE of 24.46~mV relative to the ideal softmax response. The extracted circuit characteristics are further incorporated into a MemTorch-based hardware-aware Transformer model, where the proposed softmax achieves a validation loss within 2.5\% of the ideal-softmax baseline.

Comment: Computes attention softmax directly from analog scores through tunable time-domain circuitry, eliminating intermediate analog-to-digital conversion.

Topic Match: The new attention execution primitive targets computational efficiency, although its evidence is limited to circuit simulations and hardware-aware modeling.

Relevance: 7 Novelty: 7


19. Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens

ArXiv ID: 2606.16847

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Yizhen Yao, Qinglin Zhu, Runcong Zhao, Xiangxiang Dai, Yanzheng Xiang, Yulan He, Lin Gui

Abstract: Diffusion Large Language Models (dLLMs) offer a promising avenue for parallel generation but face a trade-off between decoding speed and quality. While revocable decoding strategies attempt to mitigate errors by verifying and remasking tokens, they typically operate within a mixed-quality context. This leads to two critical failures: \textit{Error Propagation}, where new tokens absorb toxic information from erroneous context, and \textit{Local Error Reinforcement}, where errors mutually reinforce each other to evade detection. To alleviate these challenges, we propose ASRD (Anchor Supervised Revocable Decoding), a training-free framework that operates within the embedding space. ASRD explicitly decouples the decoding context into trusted \textit{Anchor Tokens}, which are identified via temporal consistency, and uncertain candidates. Leveraging a dynamic Anchor Tokens Cache, we introduce two complementary mechanisms: (1) Anchor-Guided Generation, which injects entropy-weighted anchor signals into masked positions to implicitly rectify attention toward the reliable global skeleton; and (2) Anchor-Perturbed Verification, which applies orthogonal perturbations to uncertain candidate tokens, destabilizing and remasking errors driven by fragile local consensus. Extensive experiments on math and coding benchmarks demonstrate that ASRD outperforms recent remasking baselines, achieving accuracy improvements of up to 6.4\% while accelerating inference throughput by up to 7.2$\times$.

Comment: Temporal-consistency anchors guide generation and perturbation-based remasking to accelerate diffusion decoding.

Topic Match: The new generation and verification mechanisms improve inference throughput, making efficiency the strongest fit despite the inference-only scope.

Relevance: 7 Novelty: 7


20. MeanField Surrogate Modeling for Scalable Runtime Scheduling of Concurrent Heterogeneous AI Inference on Shared GPUs

ArXiv ID: 2609.02109

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Youssef Ennouri, Soonhoi Ha

Abstract: Deploying heterogeneous AI models concurrently on a shared GPU introduces resource contention that complicates runtime scheduling. While surrogate models avoid costly online benchmarking, their profiling requirements typically grow combinatorially with the number of co-running models, limiting scalability. We propose a MeanField surrogate that predicts per-model performance from local configuration and aggregate GPU state rather than explicitly modeling all joint interactions. Experiments on concurrent LLM and vision workloads across $N \in {2,3,4,5,6}$ show high predictive accuracy ($R^2 \approx 0.96$) with an empirical sample budget that grows approximately linearly in $N$, in contrast to the combinatorial cost of fully joint profiling. Integrated into a genetic algorithm scheduler, the surrogate scales to an $N=5$ problem with 78,732 feasible joint configurations, remaining within 0.10% of the exhaustive search with zero SLA violations across eight dynamic workload scenarios, while complete online GA decisions take 26 ms median, about $5\times$ faster than exhaustive surrogate search.

Comment: Mean-field contention modeling replaces combinatorial co-run profiling with local configuration and aggregate GPU state.

Topic Match: A substantive scheduling-model idea reduces profiling and decision overhead for shared-GPU inference; the reported speedup concerns scheduler decisions, limiting its training-side relevance.

Relevance: 7 Novelty: 7


21. GeoSPRINT: Geometric Redundancy-Aware Step Pruning for Inference in Diffusion Trajectories

ArXiv ID: 2609.02160

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Arpita Joshi

Abstract: Diffusion models achieve high sample quality but remain expensive at inference time because sampling requires many sequential neural function evaluations (NFEs). Existing acceleration methods either use fixed step-skipping schedules, adapt step sizes based on local numerical error, or require additional training. We introduce GeoSPRINT (Geometric Step Pruning for Inference in Trajectories), a training-free framework for constructing non-uniform sampling schedules from the geometry of denoising trajectories. GeoSPRINT detects geometrically redundant steps using a hyperplanarity test in latent space, implemented efficiently via QR factorization, and converts the resulting redundancy profile into a sampling schedule that allocates more steps to high-curvature regions of the trajectory. In addition, we introduce the trajectory projection score $α_{\mathrm{traj}}$, a residual-variance metric that quantifies trajectory straightness and serves as a model-free diagnostic for rectified flow quality. Across CIFAR-10 ($32{\times}32$), LSUN Church ($256{\times}256$), and Stable Diffusion v1.5 ($512{\times}512$ latent), GeoSPRINT consistently improves over uniform DDIM (Denoising Diffusion Implicit Models) schedules at matched NFE budgets. On CIFAR-10, GeoSPRINT improves FID (Fréchet Inception Distance) by 0.7-1.1 over DDIM across 49-89 NFEs and surpasses DPM-Solver++ at NFE${\geq}30$ despite using a first-order DDIM solver. On LSUN Church, it reduces FID from 1.48 to 1.26 at 52 steps, and on Stable Diffusion v1.5 it achieves up to 1.93 FID improvement over DDIM. These results show that trajectory geometry provides a useful global signal for allocating inference steps and that schedule quality can substantially improve diffusion sampling efficiency without retraining.

Comment: Uses QR-based trajectory geometry to prune redundant diffusion steps and allocate nonuniform sampling budgets.

Topic Match: The core mechanism improves inference efficiency through geometry-guided sampling schedules, including on Stable Diffusion; its connection to training is peripheral.

Relevance: 7 Novelty: 6


22. Predict, Don't Iterate: Efficient Adaptive-Length Infilling for Diffusion Language Models

ArXiv ID: 2609.02108

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Haobo Xu, Sirui Chen, Yuanchen Bei, Lingjie Chen, Yuchen Yan, Dongqi Fu, Jingrui He, Hanghang Tong

Abstract: Diffusion language models (DLMs) have emerged as a promising alternative to the auto-regressive paradigm. With bidirectional attention and any-order generation, DLMs naturally fit infilling tasks, which require generating a middle span conditioned on both the prefix and the suffix. However, infilling is sensitive to the length of the span, while DLMs require the length to be fixed before generation. Although prior studies extend DLMs to dynamic lengths, they still suffer from two limitations. (i) Sensitivity to initial length. These methods require a preset length to initialize the search and are highly sensitive to this initial length, often yielding suboptimal results. (ii) Inference inefficiency. They either insert length-changing operations during generation or repeatedly search for an appropriate length using multi-step denoising confidence, both of which introduce substantial extra forward passes and computational cost. Therefore, we propose PILL (Probing-based InfiLling with preset-Length-free decoding), an efficient infilling method for DLMs that requires no preset initial length and adds far fewer extra forward passes than baselines, substantially reducing inference time. Experiments show that, across five DLMs spanning different families, architectures, and training recipes on eight infilling benchmarks, PILL improves over the strongest baseline by +4.8 average pass rate on code and +6.0 BLEU-2 on text, while running 1.82x faster than that baseline. The code is available at https://github.com/Hsu1023/PILL.

Comment: Uses probing-based length selection to reduce extra denoising forward passes in diffusion-language-model infilling.

Topic Match: The contribution directly reduces generation compute through adaptive-length decoding, with relevance narrowed by its infilling-specific scope.

Relevance: 7 Novelty: 6


23. Discriminative and Consistent Representation Distillation

ArXiv ID: 2407.11802

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Nikos Giakoumoglou, Tania Stathaki

Abstract: Knowledge Distillation (KD) transfers knowledge from a large teacher to a smaller student model. While contrastive objectives have proven effective for learning structured representations in self-supervised settings, their use in distillation is hindered by two practical shortcomings: the reliance on external memory banks for negative sampling, and fixed temperature hyperparameters that limit adaptability across training stages and teacher-student pairs. We therefore propose Discriminative and Consistent Representation Distillation (DCD), which combines contrastive instance discrimination with a consistency regularization term over the cross-model similarity matrix. The contrastive term aligns each student representation with its teacher counterpart, while the consistency term penalizes asymmetry between the row-normalized and column-normalized views of that matrix, constraining the off-diagonal structure that instance discrimination alone leaves free; we show that it vanishes precisely when this matrix is symmetric. We further introduce an efficient in-batch sampling that eliminates external memory banks, and learnable scale and bias parameters that adapt during training to control the sharpness and offset of the distillation signal. The method matches the training speed of standard KD while adding only 66K additional parameters. Through extensive experiments on CIFAR-100, ImageNet, and MS-COCO, together with cross-dataset transfer to STL-10 and Tiny ImageNet, we show that our approach achieves competitive performance in classification, object detection, and transfer, while substantially reducing memory consumption and training time compared to existing contrastive distillation methods.

Comment: Eliminates external memory banks from contrastive distillation through in-batch teacher-student comparisons.

Topic Match: Memory-efficient distillation supplies the compression connection; the contribution remains an incremental refinement evaluated on conventional vision models.

Relevance: 6 Novelty: 6


Paper Selection Prompt

System Prompt

You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.

User Prompt

Instructions

Respond in JSONL. Output exactly one JSON object per paper, one per line:

{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}

Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only: daily_hot, new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return []. - daily_hot means the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. - new_frontier means the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.

Scoring Criteria

Relevance and Novelty are independent axes. Score both from 1 to 10.

Relevance Scoring

  • 9-10: directly centered on the target foundational topics; highest when the core contribution is clearly within them.
  • 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
  • 5-6: touches the target topics, but the main contribution is elsewhere.
  • 3-4: largely outside the target topics, often application-focused or domain-specific.
  • 1-2: unrelated.

Important: Broad frontier relevance, major launch status, or daily buzz is not enough for a high Relevance score here. Those cases belong in the hotspot digest unless the paper strongly matches the specialized paper topics.

Novelty Scoring

  • 9-10: new paradigm, theory, or major methodological breakthrough.
  • 7-8: substantial methodological advance or strong new insight.
  • 5-6: meaningful but incremental extension or refinement.
  • 3-4: minor, narrow, or mostly engineering or domain-specific improvement.
  • 1-2: little originality; mainly standard application of existing methods.

Topic Registry

Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - moe_training: MoE Training - Mixture-of-Experts training: routing, load balancing, expert granularity and shared experts, training stability, MoE parallelism and communication, and MoE kernels. - training_systems: Large-Scale Training Systems and Efficiency - Large-scale training systems and efficiency: distributed training algorithms, optimizers, and communication. - architecture_training: Architecture and Training Dynamics - Core architectural or computational mechanisms, dynamic computation, and training-stability dynamics. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, memory or cache efficiency, and large-scale training systems that materially change cost or behavior.

Papers

[PAPER LIST HERE]

Relevant Topics

This is a training-side feed. The centre of gravity is Mixture-of-Experts training and the large-scale training systems it sits on; architecture and efficiency work is kept when it bears on how a large model is trained or what it costs to train.

Keep a paper when its CORE CONTRIBUTION falls in one of the four topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.

  1. MoE Training (primary) - Keep: expert routing and gating mechanisms; load balancing, with or without auxiliary losses; expert granularity, fine-grained and shared experts; MoE training stability (loss spikes, router collapse, z-loss and related fixes); expert and hybrid parallelism; all-to-all dispatch/combine and the communication schedules MoE needs; MoE kernels and grouped-GEMM style implementations; expert capacity, token dropping, and upcycling dense checkpoints into MoE. - Filter: papers that merely train on top of a MoE model without a new routing, balancing, stability, parallelism, or kernel contribution; "mixture of experts" in the classical ensemble or recommender sense.

  2. Large-Scale Training Systems - Keep: distributed training algorithms (data, tensor, pipeline, sequence, and expert parallelism, FSDP and ZeRO-style sharding); optimizers and preconditioners for large-scale pretraining; communication algorithms, overlap, and collective design; throughput and MFU work that materially changes what a training run costs or whether it converges; scaling-law work that informs how such runs are configured. - Filter: routine infrastructure tuning with no new algorithmic or systems idea.

  3. Architecture and Training Dynamics - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, dynamic or modular computation); optimisation and training-dynamics analysis that explains why large models train the way they do. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight.

  4. Efficiency and Compression - Keep: quantization, sparsity, pruning, low-rank adaptation, KV-cache and memory-efficient designs, and other work that materially changes the cost of running or training a large model, especially when the mechanism is new rather than a tuned variant. - Filter: straightforward application of standard efficiency methods, and deployment or serving engineering with no new idea.

Leave these to the hotspot digest unless the core contribution clearly falls in one of the four topics above: - major model or product releases - agent, tooling, and RAG launches - benchmark, leaderboard, or evaluation-only papers - representation-learning theory, interpretability, and feature-formation analysis - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning) - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains