Previous Day 2026-08-07
Monthly Overview 2026-08
Next Day 2026-08-11

This is a remedial run for missed papers from 08/07/2026 to 08/09/2026.

Results generated on 09/13/2026.

Personalized Daily ArXiv Papers 2026-08-10

Model Metric Usage Papers
Prompt Completion Total Total arXiv Scanned Relevant
local agent CLI Tokens not reported not reported not reported 1095 1095 93
Cost not reported not reported not reported

Token counts are not reported for this run. 12 of 26 model calls succeeded, 13,140s of model wall clock.

Topic Coverage:

TopicPapers
MoE Training4
Large-Scale Training Systems and Efficiency9
Architecture and Training Dynamics35
Efficiency, Compression, and Large-Scale Training45

Table of contents by topic:

MoE Training (4)

  1. Beyond Routing: Decoupling Expert Dispatch and Aggregation in Sparse Mixture-of-Experts Authors: Zongfei Li

  2. Fast LapSum: Exact Differentiable Top-k at Million Scale Authors: Łukasz Struski, Joanna Wojciechowicz, Jakub Antczak, Marcin Mazur, Kamil KsiÄ Å¼ek, Jacek Tabor

  3. Path-Lock Expert: Separating Reasoning Mode in Hybrid Thinking via Architecture-Level Separation Authors: Shouren Wang, Wang Yang, Chuang Ma, Debargha Ganguly, Vikash Singh, Chaoda Song, Xinpeng Li, Xianxuan Long, Vipin Chaudhary, Xiaotian Han

  4. EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inference Authors: Yize Wu, Ke Gao, Ling Li, Yanjun Wu

Large-Scale Training Systems and Efficiency (9)

  1. ZeroLock: Concurrent Memory-Efficient LLM Training via Modular Update Decoupling Authors: Wentao Dai, Xuanran Li, Yuxiang Zhang, Ming Tang, Chao Huang

  2. Periodic Asynchrony: An On-Policy Approach for Accelerating LLM Reinforcement Learning Authors: Jian Lu, Yi Luo

  3. Optimization as a Dynamical System: Generative Schedules from Latent ODEs Authors: Matt L. Sampson, Peter Melchior

  4. To Memorize or to Retrieve: Scaling the Interaction Between Pretraining and Retrieval Authors: Karan Singh, Michael Yu, Varun Gangal, Zhuofu Tao, Sachin Kumar, Emmy Liu, Steven Y. Feng

  5. Capacity Confounds and Coverage Guarantees in Adaptive Sub-model Federated Learning Authors: Alireza Moayedikia, Alicia Troncoso Lora

  6. SSTQ:Privacy-Preserving Vector Quantization via Subsampled Stochastic TurboQuant Authors: Adel Javanmard, David P. Woodruff, Vahab Mirrokni

  7. LGNNIC: Acceleration of Large-Scale GNN Training using SmartNICs Authors: Liad Gerstman, Aditya Dhakal, Dejan Milojicic, Avi Mendelson

  8. Automated Numerical Stability Analysis of Deep Learning Operators Authors: Xinye Chen

  9. Stream Learning: Partition-Fair Gossip Learning Without Tokens Authors: Fabien Mathieu, Alexandre Pham, Maria Gradinariu Potop-Butucaru, S{é}bastien Tixeuil

Architecture and Training Dynamics (35)

  1. TEMPER: Tensorized Efficient Manifold-constrained Parameterization for Expressive Residual Routing Authors: Yuxuan Gu, Wuyang Zhou, Huijun Xing, Danilo Mandic

  2. Skip a Layer or Loop It? Learning Program-of-Layers in LLMs Authors: Ziyue Li, Yang Li, Tianyi Zhou

  3. Modular TTT: Rethinking Test-Time Training as Composable Modules Authors: Bohao Tang, Zhen Qin, Yuqi Pan, Zheng Li, Pengfei Liu, Ya Zhang

  4. Faster Query-Key Learning Sharpens Attention in Self-Attention Models Authors: Rahul Vashisht, Harish G. Ramaswamy

  5. Barriers to Universal Reasoning With Transformers (And How to Overcome Them) Authors: Oliver Kraus, Yash Sarrof, Yuekun Yao, Alexander Koller, Michael Hahn

  6. Exact Network Surgery: Functional Invariance and Gradient Plasticity in Reactive Computational Graphs Authors: Abdallah Khemais

  7. Beyond Structural Symmetries: Linear Mode Connectivity via Neuron Identifiability Authors: Vincent Bürgin, Daniel Herbst, Ya-Wei Eileen Lin, Stefanie Jegelka

  8. Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers Authors: Ali Janati, Kaoutar El Maghraoui, Andrei Kanavalau, Anass Belfatmi

  9. Pay Attention to Attention Distribution: A New Local Lipschitz Bound for Transformers Authors: Nikolay Yudin, Sergei Kudriashov, Alexander Gaponov, Maxim Rakhuba

  10. Clustered Attractor Manifolds and Dynamical Condensation in Self-Attention Authors: Qucheng Gao, Zuyi Yang, Xiao Chen

  11. Shattered Compositionality: Counterintuitive Learning Dynamics of Transformers for Arithmetic Authors: Xingyu Zhao, Darsh Sharma, Rheeya Uppaal, Yiqiao Zhong

  12. Transformer Circuits Can Realize Clustering Algorithms Authors: Kenneth L. Clarkson, Lior Horesh, Takuya Ito, Charlotte Park, Parikshit Ram

  13. Rethinking Reasoning with MDLMs: Early Exits, Post-hoc Reasoning, and Beyond Authors: Zachary Horvitz, Raghav Singhal, Hao Zou, Carles Domingo-Enrich, Zhou Yu, Rajesh Ranganath, Kathleen McKeown

  14. Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation Authors: Kaichen Zhang, Wei Huang, Keming Wu, Bo Li, Xiaojuan Qi

  15. Second Order Drifting Models Authors: Drake Brown, Yuhao Huang, Shih-Hsin Wang, Bao Wang

  16. Two-Layer Linear Auto-Regressive Models Estimate Latent States Authors: Yahya Sattar, Sunmook Choi, Leo Maynard-Zhang, Yassir Jedra, Maryam Fazel, Sarah Dean

  17. Tokenisation over Bounded Alphabets is Hard Authors: Violeta Kastreva, Philip Whittington, Dennis Komm, Tiago Pimentel

  18. Symmetry Reveals Layerwise Dynamics: How Transformers Perform In-Context Classification Authors: Patrick Lutz, Themistoklis Haris, Arjun Chandra, Aditya Gangrade, Venkatesh Saligrama

  19. How Many Different Outputs Can a Transformer Generate? Authors: Maxime Meyer, Mario Michelessa, Caroline Chaux, Vincent Y. F. Tan

  20. On the Infinite Width and Depth Limits of Predictive Coding Networks Authors: Francesco Innocenti, El Mehdi Achour, Rafal Bogacz

  21. The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought Authors: Moritz Brösamle, Stephan Eckstein

  22. Is SwiGLU's Open Positive Tail Necessary? Evidence from Closed-Tail Gating with MemGLU Authors: Yuting Ge, Pengju Yang, Mingkai Nie

  23. Think Deep, Speak Once: Relit, A Recursive Latent Implicit Transformer Framework Authors: Abhishek Panwar, Maheep Singh, Saksham Bansal

  24. State Propagation Also Satisfies: A Complex-Valued State-Space Model for Deterministic State Tracking Authors: Xiaohe Li, Yang Lu

  25. Neuronal Attention Circuit (NAC) for Representation Learning Authors: Waleed Razzaq, Yun-Bo Zhao

  26. PRISM: Principled Reference Identification for Schrodinger Bridge Model Authors: Forouzan Fallah, Yezhou Yang

  27. Kolmogorov-Arnold Energy Models: Fast, Interpretable Generative Modeling Authors: Prithvi Raj

  28. No Unique Minimizer, No Problem: On the Consistency of Robust Neural Classifiers Authors: Subhabrata Majumdar, Anand Deo, Partha Pratim Saha, Abhik Ghosh

  29. Understanding Reasoning from Pretraining to Post-Training Authors: Jingyan Shen, Ang Li, Salman Rahman, Yifan Sun, Micah Goldblum, Matus Telgarsky, Pavel Izmailov

  30. Semantic Adapter Routing with Fine-Tuning Task Embeddings Authors: Enrico Cassano, Michał Brzozowski, Paolo Mandica, Zuzanna Dubanowska, Neo Christopher Chung

  31. Early Data Exposure Improves Robustness to Subsequent Fine-Tuning Authors: Lawrence Feng, Gaurav R. Ghosal, Jacob Mitchell Springer, Ziqian Zhong, Aditi Raghunathan

  32. Characteristic Learning for Provable One Step Generation Authors: Zhao Ding, Chenguang Duan, Yuling Jiao, Ruoxuan Li, Jerry Zhijian Yang, Pingwen Zhang

  33. Intelligence Foundation Model: A New Perspective to Approach Artificial General Intelligence Authors: Borui Cai, Yao Zhao

  34. Error Bound Analysis for the Regularized Loss of Deep Linear Neural Networks Authors: Po Chen, Rujun Jiang, Peng Wang

  35. LF${}^{2}$AR: Accounting for Layerwise Dynamics to Improve Multimodal Adaptation of Language Models Authors: Santiago Cuervo, Adel Moumen, Yanis Labrak, Sameer Khurana, Antoine Laurent, Mickael Rouvier, Phil Woodland, Ricard Marxer

Efficiency, Compression, and Large-Scale Training (45)

  1. LegoLM: Structured Weight Sharing for Large Language Models Authors: Joseph Bingham

  2. Every Cache Entry Earns Its Place: Global Allocation of Resolution and Coverage for KV Cache Compression Authors: Haolin Tian, Yuzhe Liu, Tonghan Wang

  3. Statistically-Lossless Quantization of Large Language Models Authors: Michael Helcig, Eldar Kurtic, Dan Alistarh

  4. Autonomy-of-Heads: Data-Free Sparse Attention from Frozen Query-Key Geometry Authors: Yehan Yang, Junyuan Shang, Yang Li, Guanqun Zhao, Shuohuan Wang, Dianhai Yu

  5. Shape Mutating Expert Compression:LorExperts and BTExperts Authors: Inesh Chakrabarti, Sourjya Roy, Bowen Bao, Thiago Crepaldi, Spandan Tiwari, Ashish Sirasao

  6. RotaryQuant: Fitting 120B MoE Models on Consumer Hardware via Fused Compressed-Space Attention Authors: Anthony. Lui, Mohamed. Elsaied, N. P. Savani

  7. Quantization Degradation in Large Language Models: A Signal-Noise Perspective Authors: Chenxi Zhou, Pengfei Cao, Jinyu Ye, Bohan Yu, Haida Yu, Jiang Li, Jun Zhao, Kang Liu

  8. Attn-QAT: 4-Bit Attention With Quantization-Aware Training Authors: Peiyuan Zhang, Matthew Noto, Wenxuan Tan, Chengquan Jiang, Will Lin, Wei Zhou, Hao Zhang

  9. SPECTRA: Pushing the KV Cache Beyond the 2-Bit Cliff via Spectral Transform Coding Authors: Jiamu Zhang, Liang Wu, Kelly Wan, Hanjie Chen, Liangjie Hong

  10. KronQ: LLM Quantization via Kronecker-Factored Hessian Authors: Donghyun Lee, Yuhang Li, Ruokai Yin, Priyadarshini Panda

  11. ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantization Authors: Yongge Ma, Guoan Wang, Feiyu Wang, Yaoming Li, Qian Zhang, Zihan Yan, Yinjun Han, Tong Yang

  12. CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit Weights Authors: Xuetian Gao

  13. Length-MAX Tokenizer for Language Models Authors: Dong Dong, Weijie Su

  14. LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment Authors: Linhan Xia, Rui Liu, Zhaofeng Zhang, Yihao Wang, Binrui Shen, Shengxin Zhu

  15. Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation Authors: Bruce Changlong Xu, Adarsh Kumarappan, Mu Zhou

  16. Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Models Authors: Ali Janati, Kaoutar El Maghraoui, Xinyi Luo, Wenyuan Shen, Owen Zou, Yankai Mao

  17. RippleKV: Cross-Layer KV Cache Allocation via Perturbation Propagation Authors: Dongjie Xu, Kai Qian, Julius, Weijie Shi, Yuxuan Sun, Minghua Tang, Fenglei Jin, Hanchi Dong, Jiajie Xu

  18. SimSD: Simple Speculative Decoding in Diffusion Language Models Authors: Junxia Cui, Haotian Ye, Runchu Tian, Hongcan Guo, Jinya Jiang, Haoru Li, Chaojie Ren, Yiming Huang, Kaijie Zhu, Zhongkai Yu, Kun Zhou, Jingbo Shang

  19. DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM Inference Authors: Asaad Althoubi

  20. Tied Trit-Planes: Constraining PTQTP to a Uniform Nine-Level Quantizer, with a Persistent Folded Format for Disk-Streamed Mixture-of-Experts Serving Authors: Matteo Grella

  21. Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference Authors: Sanjeev Rao Ganjihal

  22. IFCLoRA: Topology-Aware Rank Allocation for Parameter-Efficient Fine-Tuning Authors: Wei Zhang, Xinwu Liu, Yihang Cheng

  23. Unified Static-Dynamic Pruning for Efficient LLM Inference Authors: Jinhyeok Kim, Yejoon Lee, Jaeyoung Do

  24. Understanding Calibration and Truncation Error Propagation in Training-Free Low-Rank Compression for LLMs Authors: Mohanad Odema, Gabrielle De Micheli, Dayin Gou, Nilesh Malpeddi, Prathamesh Vaste, Jacob Song

  25. Dropping the Anchor: Statistical Context Summarization for Distributed Systems via Pulsar Attention Authors: Aryan Sood, Shantanu Acharya, Gaurav Kumar Nayak

  26. CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents Authors: Weizhong Huang, Jinchao Zhang, Xiawu Zheng

  27. CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG Authors: Gyuwan Kim, Cheoneum Park, Tao Yang

  28. Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models Authors: Minseok Kang, Hyunwoo Kim, Chanyoung Kim, Minwoo Kim, Jaekoo Lee, Dahuin Jung

  29. High-Layer Attention Pruning with Rescaling Authors: Songtao Liu, Peng Liu

  30. Diving into Kronecker Adapters: Component Design Matters Authors: Jiayu Bai, Danchen Yu, Zhenyu Liao, TianQi Hou, Feng Zhou, Robert C. Qiu, Zenan Ling

  31. OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference Authors: Yuzhe Gu, Xiyu Liang, Jiaojiao Zhao, Enmao Diao

  32. VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference Authors: Wenxu Jia, Dongjie Fu, Xize Cheng, Fangming Feng, Linjun Li, Wenshi Chen, Yingming Li, Zhou Zhao, Tao Jin

  33. Hybrid Policy Distillation for LLMs Authors: Wenhong Zhu, Ruobing Xie, Rui Wang, Pengfei Liu

  34. PTQ4SNN: Membrane-Aware Post-Training Quantization for Spiking Neural Networks Authors: Hui Xie, Tong Shi, Haotong Qin, Aishan Liu, Xiaode Liu, Jinyang Guo

  35. LibraSpec: Dynamic Diffusion-Based Speculative Decoding via Marginal-Gain-Driven Optimization Authors: Zexun Lin, Yuan Feng, Junlin Lv, Kevin S. Zhou, Xike Xie

  36. Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Authors: Kyeongyoon Lee, Hongyeob Kim, Youngeun Kim, Sungeun Hong

  37. Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills Authors: Agamdeep Singh, Srishti Gautam, Priyanshu Gupta, Nikita Mehrotra, Tanmay Bakshi, Sumit Gulwani

  38. RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs Authors: Qiyanhui Lu, Han Wu, Rongjian Xu, Tingzhang Luo, Cheng Fan, Xinghao Chen, Minjing Dong, Jufeng Yang, Jianyuan Guo

  39. MiCoPro: End-to-End Mixed Precision HW/SW Co-design with HW-aware Proxy Model Authors: Zijun Jiang, Yangdi Lyu

  40. LoCA: Spatially-Aware Low-Rank Convolutional Adaptation of Vision Foundation Models Authors: Sojung An, Junha Lee, Sujeong You, Nam Ik Cho, Donghyun Kim

  41. Domain-Aware Pruning: Sparsity and Domain Generalization via Regularized Probabilistic Masking Authors: Parham Sazdar, Mostafa Tavassolipour, Reshad Hosseini

  42. LoRSA: Toward Generalizable Parameter-Efficient Fine-Tuning for Biomedical Downstream Tasks Authors: Saed Moradi, Benyamin Ghojogh, M. Hadi Sepanj, Yimin Yang, Ashirbani Saha

  43. Addressable Memory for Video World Models Authors: Xindi Wu, Sven Elflein, James Lucas, Olga Russakovsky, Laura Leal-Taixé, Despoina Paschalidou, Jonathan Lorraine, Aljoša Ošep

  44. HEART: Exploiting Head Heterogeneity in Sparse Attention for Video Diffusion Authors: Xuzhe Zheng, Yuexiao Ma, Jing Xu, Xiawu Zheng, Rongrong Ji, Fei Chao

  45. CHIME: A Case for Efficient Long-Context Attention-FC Disaggregated Inference with DIMM-PIM Authors: Qingyuan Liu, Liyan Chen, Haocheng Wang, Yanning Yang, Dong Du, Zhigang Mao, Naifeng Jing, Yubin Xia, Haibo Chen


MoE Training (4)

1. Beyond Routing: Decoupling Expert Dispatch and Aggregation in Sparse Mixture-of-Experts

ArXiv ID: 2608.08853

Primary Topic: MoE Training

Also Matches: Architecture and Training Dynamics

Authors: Zongfei Li

Abstract: Sparse Mixture-of-Experts (MoE) routers commonly use the same scores both to select experts and to weight their already-computed outputs. We study whether these two roles, dispatch and aggregation, should be coupled. On pretrained OLMoE-1B-7B, we keep selected Top-8 expert IDs, expert computation, and total selected router mass fixed and change only within-set aggregation. A structured oracle improves full-horizon cross-entropy by 0.0160 +/- 0.0039 across three seeds; the router's top-scored expert is the counterfactual-best vertex only 17.2% of the time, with router-utility Spearman 0.030. We therefore train Fixed-Dispatch Adaptive Aggregation (FDAA), a 301K-parameter post-compute head optimized directly with the language-modeling objective while freezing the backbone, router, and experts. On OLMoE, FDAA improves fresh WikiText-103 test by Delta CE = -0.1523 +/- 0.0031 across three seeds, and mixed-domain training gives robust gains on WikiText-103, C4, and held-out Penn Treebank under frozen confirmatory evaluation. We also replicate the fixed-dispatch audit on DeepSeek-V2-Lite, which uses Top-6 routed experts plus shared experts. Best-vertex headroom remains significant on WikiText and C4, while router Top1 identifies the best selected expert in only 12.5% and 16.7% of audited examples. In a one-seed mixed-domain replication, FDAA improves locked WikiText and PTB, while C4 is statistically neutral. These results support a cross-architecture distinction between expert selection and expert commitment.

Comment: Decouples expert dispatch from trainable output aggregation while holding selected experts fixed.

Topic Match: Directly revises how sparse MoE outputs are weighted and trains the new aggregation mechanism with the language-modeling objective, supporting a broader architectural distinction between selection and combination.

Relevance: 10 Novelty: 8


2. Fast LapSum: Exact Differentiable Top-k at Million Scale

ArXiv ID: 2608.06912

Primary Topic: MoE Training

Also Matches: Architecture and Training Dynamics, Efficiency, Compression, and Large-Scale Training

Authors: Łukasz Struski, Joanna Wojciechowicz, Jakub Antczak, Marcin Mazur, Kamil KsiÄ Å¼ek, Jacek Tabor

Abstract: The top-$k$ operation is a fundamental building block of modern sparse computation, enabling token routing, expert activation, memory selection, and attention pruning. Yet standard hard top-$k$ blocks gradients, while existing continuous (soft) relaxations remain too costly for large-scale models. We introduce Fast LapSum, an exact-budget soft top-$k$ primitive whose GPU solver runs in linear time after sorting. Unlike prior linear-time methods such as DFTopK, which relax the normalization constraint, Fast LapSum is, to our knowledge, the first method to preserve an exact selection mass of $k$ while remaining fully differentiable end-to-end. Our solver combines a linear-time threshold computation with an analytical vector--Jacobian product, and for extreme scales employs probabilistic bracketing to sort only the uncertain middle band of kernel-noised scores. The resulting overhead is almost negligible: the solver processes $10^6$, $10^7$, and $10^8$ scores in $0.41$, $1.15$, and $5.23$\,ms, respectively. This makes exact soft top-$k$ practical for sparse routing, retrieval, and large-scale optimization. We demonstrate Fast LapSum on two demanding applications operating over millions of coordinates inside the training loop: generating megapixel sparse adversarial examples with an exact soft budget of ${\sim}0.02\%$ of an image's pixels, achieving an order-of-magnitude speedup over state-of-the-art methods, and training a fully differentiable sparse image coder from scratch.

Comment: Provides an exact-budget differentiable top-k primitive with million-scale GPU execution.

Topic Match: Differentiable, scalable top-k selection is directly applicable to learned expert routing and sparse activation.

Relevance: 8 Novelty: 8


3. Path-Lock Expert: Separating Reasoning Mode in Hybrid Thinking via Architecture-Level Separation

ArXiv ID: 2604.27201

Primary Topic: MoE Training

Also Matches: Architecture and Training Dynamics

Authors: Shouren Wang, Wang Yang, Chuang Ma, Debargha Ganguly, Vikash Singh, Chaoda Song, Xinpeng Li, Xianxuan Long, Vipin Chaudhary, Xiaotian Han

Abstract: Hybrid-thinking language models expose explicit /think and /no_think modes, but current designs do not separate them cleanly. Even in /no_think mode, models often emit long and self-reflective responses, causing reasoning leakage. Existing work reduces this issue through better data curation and multi-stage training, yet leakage remains because both modes are still encoded in the same feed-forward parameters. We propose Path-Lock Expert (PLE), an architecture-level solution that replaces the single MLP in each decoder layer with two semantically locked experts, one for /think and one for /no_think, while keeping attention, embeddings, normalization, and the language-model head shared. A deterministic control-token router selects exactly one expert path for the entire sequence, so inference preserves the dense model's per-token computation pattern and each expert receives mode-pure updates during supervised fine-tuning. Across math and science reasoning benchmarks, PLE maintains strong /think performance while producing a substantially stronger mode separation, a /no_think mode with higher accuracy and far less reasoning leakage. On Qwen3-4B, for example, compared to the SFT-only baseline on AIME24, PLE generates 17x fewer reflective tokens (6.01 vs. 0.35 per answer) and 2x shorter outputs (8665 vs. 4101 tokens), and improves /no_think accuracy from 35.33% to 44.67%, while maintaining /think-mode performance (61.33% vs. 60.00%). These results suggest that controllable hybrid thinking is fundamentally an architectural problem, and separating mode-specific feed-forward pathways is a simple and effective solution.

Comment: Routes entire sequences through semantically locked think/no-think MLP experts with mode-pure updates.

Topic Match: The core mechanism is deterministic expert routing with explicit expert specialization during training.

Relevance: 8 Novelty: 7


4. EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inference

ArXiv ID: 2608.07964

Primary Topic: MoE Training

Also Matches: Large-Scale Training Systems and Efficiency, Efficiency, Compression, and Large-Scale Training

Authors: Yize Wu, Ke Gao, Ling Li, Yanjun Wu

Abstract: Load Balancing has emerged as a critical problem in expert-parallel distributed inference of Mixture-of-Experts (MoE) models. As routing distributions are typically skewed across experts, devices hosting lighter-loaded experts must idle to wait for the heaviest during expert computing, leading to inefficiency. Existing load-balancing approaches primarily rely on expert replication or migration within each layer, which introduce additional overhead and limit their flexibility and scalability. To address this problem, we propose EasyBalance, a cross-layer load balancing strategy that requires no modifications to the expert-device mapping, enabling instant adaptability and incurring essentially no additional overhead. Our key insights are that (1) experts of other layers can be viewed as naturally redundant for the current layer, and (2) cross-layer MoE workloads can be jointly executed to mitigate their individual imbalance. Based on these observations, EasyBalance greedily schedules a subset of cross-layer workloads to run at each MoE step and defers the remaining workloads for future balancing opportunities, effectively leveraging cross-layer imbalance mitigation. Extensive experiments across models, tasks, and configurations demonstrate that EasyBalance consistently accelerates distributed MoE inference, reducing GPU idling by mostly over 40%. Code is available at https://github.com/yize-wu/EasyInfra.

Comment: Balances expert-parallel workloads by interleaving skewed expert computation across MoE layers.

Topic Match: Cross-layer expert load balancing is directly relevant to MoE parallel execution, although evaluated for inference.

Relevance: 8 Novelty: 7


Large-Scale Training Systems and Efficiency (9)

1. ZeroLock: Concurrent Memory-Efficient LLM Training via Modular Update Decoupling

ArXiv ID: 2608.07974

Primary Topic: Large-Scale Training Systems and Efficiency

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Wentao Dai, Xuanran Li, Yuxiang Zhang, Ming Tang, Chao Huang

Abstract: Large language model (LLM) fine-tuning at the edge adapts the model to scenario-specific data while preserving privacy. Although existing studies proposed pipeline parallelism to address the limited memory and computing resources of edge devices, they commonly rely on backpropagation (BP) training, which has a fundamental limitation of update locking and could experience severe throughput and memory bottlenecks. In this work, we propose a BP-free algorithm, called ZeroLock, that decouples the model updates into independent chunk updates by local objective construction. It breaks the update locking of BP and hence can improve throughput at the algorithm level and lower memory usage by reducing activation storage. To the best of our knowledge, we provide the first theoretical framework for such local objective construction-based approaches under general model chunk division by mapping local objectives to the global objective. We prove that ZeroLock has a convergence rate of $\tilde{\mathcal{O}}(1/\sqrt{T})$, which differs from BP only by polylogarithmic factors. We design a system for ZeroLock and build real-world prototypes, incorporating techniques such as early forwarding and failure recovery for efficient and robust implementation. Experiments on the prototype show that compared to BP-based baselines, ZeroLock reduces the memory by 26.5% and improves throughput by 4.9%.

Comment: Decouples model chunks into concurrent local updates that avoid end-to-end backpropagation locking.

Topic Match: This is an algorithm-system co-design for lower-memory, higher-throughput LLM training.

Relevance: 9 Novelty: 8


2. Periodic Asynchrony: An On-Policy Approach for Accelerating LLM Reinforcement Learning

ArXiv ID: 2511.18871

Primary Topic: Large-Scale Training Systems and Efficiency

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Jian Lu, Yi Luo

Abstract: Since the introduction of the GRPO algorithm, reinforcement learning (RL) has attracted increasing attention for LLM post-training, yet training efficiency remains a critical challenge. In mainstream RL frameworks, inference and training are co-located on the same devices, and their synchronous execution prevents concurrent inference and training. In this work, we revisit the strategy of separating inference and training deployment, and propose a periodically asynchronous framework that transforms synchronous RL training into an asynchronous producer-consumer pipeline. By synchronising model weights at the beginning of each training iteration and generating all rollouts from the same policy, the proposed framework remains inherently on-policy -- without any modification to standard RL algorithms -- thereby avoiding the off-policy bias introduced by existing asynchronous approaches. We further introduce a unified tri-model architecture and a shared-prompt attention mechanism to support efficient asynchronous execution and reduce redundant computation. Experiments on NPU platforms show approximately 2x throughput improvement from asynchronous execution, with additional gains from system-level optimisations, substantially outperforming mainstream RL frameworks in end-to-end throughput, with speedups of up to 3x on GPU platforms, further confirming cross-architecture generalisability while maintaining comparable accuracy. The proposed framework thus offers a practical, algorithm-agnostic solution for scalable RL post-training without sacrificing on-policy correctness. Code available at: https://github.com/janelu9/EasyLLM

Comment: Pipelines rollout generation and optimization asynchronously while synchronizing once per iteration to remain on-policy.

Topic Match: The central contribution is a distributed execution schedule that materially increases large-model training throughput.

Relevance: 8 Novelty: 7


3. Optimization as a Dynamical System: Generative Schedules from Latent ODEs

ArXiv ID: 2509.23052

Primary Topic: Large-Scale Training Systems and Efficiency

Also Matches: Architecture and Training Dynamics

Authors: Matt L. Sampson, Peter Melchior

Abstract: We present a new meta-learning method to determine the optimal learning rate schedule for gradient descent. It leverages training runs from a hyperparameter search to learn a latent representation of the training process, which is modeled as a dynamical system. Given current training metrics, it predicts the future learning rate schedule with the best long-term validation performance. Our scheduler generalizes beyond previously observed training dynamics and creates specialized schedules that deviate noticeably from even the best-performing parametric functions. It outperforms all baselines we compare to on results for image classification with CNN and ResNet models as well as for next-token prediction with a transformer model. The trained models are located in flatter regions of the loss landscape and thus provide better generalization than those trained with other schedules. Our method is computationally efficient, optimizer-agnostic, and can easily be layered on top of ML experiment-tracking platforms to streamline training of neural networks from scratch.

Comment: Learns state-conditioned learning-rate schedules by modeling training trajectories with a latent ODE.

Topic Match: The core contribution is optimizer-agnostic learning-rate control informed by training dynamics; evidence at large pretraining scales remains limited.

Relevance: 8 Novelty: 7


4. To Memorize or to Retrieve: Scaling the Interaction Between Pretraining and Retrieval

ArXiv ID: 2604.00715

Primary Topic: Large-Scale Training Systems and Efficiency

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Karan Singh, Michael Yu, Varun Gangal, Zhuofu Tao, Sachin Kumar, Emmy Liu, Steven Y. Feng

Abstract: Retrieval-augmented generation (RAG) improves language model (LM) performance by providing relevant context at test time for knowledge-intensive situations. In this work, we systematically study the trade-off between pretraining and retrieval by training OLMo-2-based LMs ranging from 30M to 3B parameters on up to 100B DCLM tokens, while varying pretraining data scale, retrieval store size, and retrieval store source (pretraining vs. new data) across reasoning, scientific QA, and open-domain QA benchmarks. We find that retrieval gains depend on model capacity and pretraining exposure and are strongly front-loaded, with a median 91% of the largest observed improvement realized by one retrieval token per model parameter. However, the interaction is objective-dependent: smaller models gain more in gold-answer perplexity, whereas larger, more-pretrained models gain more in accuracy. Retrieval from previously seen data also preserves most of the held-out retrieval gain. Retrieval is therefore a task-, regime-, and metric-dependent complement to parametric learning whose value also depends on datastore size and information novelty. Overall, this motivates the explicit partitioning of data between internalization and external access for LM design.

Comment: Maps how model size, pretraining exposure, and retrieval budget jointly determine performance gains.

Topic Match: The scaling analysis informs how pretraining data and external storage should be budgeted.

Relevance: 6 Novelty: 7


5. Capacity Confounds and Coverage Guarantees in Adaptive Sub-model Federated Learning

ArXiv ID: 2608.07157

Primary Topic: Large-Scale Training Systems and Efficiency

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Alireza Moayedikia, Alicia Troncoso Lora

Abstract: Sub-model federated learning lets resource-constrained clients train width-reduced versions of a global model, but existing methods allocate capacity by device resources alone. A natural next step, allocating capacity by each client's data heterogeneity as estimated from the updates the server already observes, has been repeatedly suggested. We ask whether that step is possible, using HAS-FL, an adaptive capacity-allocation framework, as a test case. Our findings are threefold. First, validated against ground-truth label-distribution divergence on reproducible partitions, update-divergence estimates of client heterogeneity are dominated by capacity rather than data: across two corrected estimators, multiple datasets, and all seeds, the estimates correlate strongly and negatively with device capacity, and no data signal remains once capacity is controlled for. This previously undocumented confound affects any method estimating client statistics from sub-model updates. Second, adaptive allocation has a hidden failure mode: when every client is capped below full width, the uncovered parameters stay at random initialization and progressively corrupt the global model. A simple coverage guarantee removes the failure and explains why uniform allocation collapses. Third, a matched-budget control settles what adaptivity contributes: random allocation to the same average budget performs no differently on both image benchmarks, and on the naturally partitioned text benchmark the adaptive policy is the weakest of the three strategies while consuming the most capacity. Sub-model training remains valuable because it admits constrained clients at quadratically reduced cost, but what protects accuracy is parameter coverage rather than allocation intelligence. Its apparent benefits come from capacity budgeting and coverage, and future designs need heterogeneity signals separable from capacity effects.

Comment: Identifies capacity-confounded heterogeneity estimates and establishes a parameter-coverage guarantee for sub-model training.

Topic Match: Its main relevance is a distributed-training failure analysis and corrective coverage condition under constrained capacity.

Relevance: 6 Novelty: 7


6. SSTQ:Privacy-Preserving Vector Quantization via Subsampled Stochastic TurboQuant

ArXiv ID: 2608.05127

Primary Topic: Large-Scale Training Systems and Efficiency

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Adel Javanmard, David P. Woodruff, Vahab Mirrokni

Abstract: Achieving local differential privacy in distributed optimization while maintaining low communication cost remains challenging. Existing vector quantization methods, such as vqSGD, use high-dimensional geometric constructions but incur unfavorable dimension-dependent variance. In this work, we propose Subsampled Stochastic TurboQuant (SSTQ), a framework that combines overcomplete equal-norm tight frames, coordinate subsampling, and privacy-aware one-dimensional quantization. SSTQ includes two variants: a Flat Randomized Response version and a Metric-Aware Laplace version, the latter being better suited to higher codebook bit-width regimes. We show that SSTQ achieves optimal mean squared error scaling while using only $\lceil \log_2 N \rceil + b$ bits per client, where $N = Θ(d)$ is the frame size. We also derive a surrogate privacy-aware codebook objective that reduces the codebook-dependent MSE scaling from $O(4^b)$ to $O(2^b)$. Finally, we empirically evaluate SSTQ against established baselines on federated learning tasks using CIFAR-10 and Fashion-MNIST, demonstrating favorable utility and communication efficiency. Some of the analytical derivations were first obtained using a fully automated Gemini-based agentic system developed internally at Google. The authors have verified those derivations and edited them for clarity of presentation.

Comment: Combines tight-frame quantization and coordinate subsampling for communication-efficient private distributed optimization.

Topic Match: Its strongest fit is a communication algorithm for distributed learning, with quantization as the enabling mechanism.

Relevance: 6 Novelty: 7


7. LGNNIC: Acceleration of Large-Scale GNN Training using SmartNICs

ArXiv ID: 2608.07733

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Liad Gerstman, Aditya Dhakal, Dejan Milojicic, Avi Mendelson

Abstract: Graph Neural Networks (GNNs) are widely used across domains such as natural sciences, social network analysis, chip design, and recommendation systems. However, as graph sizes grow, storing and processing them entirely on a single-node CPU-GPU system becomes increasingly impractical. A promising approach is to distribute the graph across multiple remote memory nodes, though this introduces a major bottleneck: inter-node network congestion during training. To address this, we propose LGNNIC, a novel inter-node system architecture that leverages SmartNICs co-located with remote memory nodes-a configuration already available in modern systems-to reduce communication overhead in distributed GNN training. LGNNIC offloads key preprocessing tasks to SmartNICs, reducing the volume of data transferred to computational (training) nodes and alleviating network congestion. We introduce two complementary techniques executed on the SmartNICs during the preprocessing phase: Neighbor Sampling, which performs mini-batch sampling, and Quantization of the sampled batches. To evaluate LGNNIC under different communication infrastructures, we designed both an optimized low-overhead DMA-based synchronization mechanism and a high-overhead socket-based alternative used as a benchmark. We evaluate the core SmartNIC offloading mechanisms across standard GNN workloads and sampling hyperparameters using a proof-of-concept (PoC) system comprising one remote-memory node with an NVIDIA BlueField-2 SmartNIC and one compute node with an A100 GPU. Both Neighbor Sampling and Quantization on the remote node demonstrated substantial training speedups in most configurations. Neighbor Sampling achieved up to 62.4x and 17.5x speedups with Sockets and DOCA-DMA, respectively, primarily due to reduced data transaction time. Quantization provided additional speedups of up to 3.6x and 1.3x, respectively, by reducing data transfer.

Comment: Offloads neighbor sampling and quantization to SmartNICs to reduce distributed GNN training traffic.

Topic Match: The core contribution is a communication-reducing architecture for distributed training.

Relevance: 6 Novelty: 6


8. Automated Numerical Stability Analysis of Deep Learning Operators

ArXiv ID: 2607.25494

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Xinye Chen

Abstract: Finite-precision arithmetic unavoidably introduces numerical approximation errors. Numerical computations may use insufficient precision or an improper formulation, which leads to numerical instability. In this paper, we introduce a unified software tool for stochastic numerical validation of deep-learning operators. The tool follows CESTAC on supported exposed operations and uses an operator-level data-perturbation approximation for GEMM-like kernels. Our developed software not only enables numerical validation with a single computation pass but also detects the sources of numerical instability and provides numerical stability monitoring during deep learning training and inference. We verified its effectiveness on the detection of polluted operators with injected numerical instabilities across various tasks. We believe that our developed method and tools provide valuable insights into developing numerically stable computing kernels, which are particularly critical for numerically stable and efficient deep learning training and inference.

Comment: Automates operator-level numerical-stability detection during neural-network training and inference.

Topic Match: The tool targets numerical reliability of training operators and GEMM-like kernels.

Relevance: 6 Novelty: 6


9. Stream Learning: Partition-Fair Gossip Learning Without Tokens

ArXiv ID: 2608.06946

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Fabien Mathieu, Alexandre Pham, Maria Gradinariu Potop-Butucaru, S{é}bastien Tixeuil

Abstract: In gossip learning, a network of nodes trains a shared model collaboratively, without a central coordinator, by repeatedly exchanging parts of their local models. The state-of-the-art protocol, Partitioned Token Gossip Learning (PTGL) of Heged{ü}s et al., splits the weight matrix into S fixed partitions and disseminates them using a token-based fairness mechanism coupled with per-neighbor metadata exchange. We revisit partition scheduling by analogy with peer-to-peer live streaming, where model partitions act as video chunks and partition age acts as chunk scarcity. The analogy yields a design space of two-stage selection strategies (partition first, or neighbor first), from which we instantiate ten concrete protocols collectively called Stream Learning. Our main finding is that the simplest of these protocols, which transmits the locally least-trained partition to a uniformly random neighbor (Ri), matches PTGL on fault-free workloads while requiring neither token counters nor metadata exchange. Under an adversarial 30% permanent crash of the best-performing nodes, Ri matches or outperforms PTGL across all complete-graph configurations tested, with the gap reaching 5.53% on HAR and 5.41% on MNIST in the most heterogeneous regime (Dirichlet $β$ = 0.1). In our experiments, partition fairness, captured by a single local rule on partition age, accounts for the gap; token-based rate control and utility maximization do not improve over this rule and, under heterogeneity, sit below it.

Comment: Schedules decentralized model partitions using local partition age without token counters or neighbor metadata.

Topic Match: Its core is a communication-light distributed learning protocol, albeit outside large-model pretraining.

Relevance: 6 Novelty: 6


Architecture and Training Dynamics (35)

1. TEMPER: Tensorized Efficient Manifold-constrained Parameterization for Expressive Residual Routing

ArXiv ID: 2608.07851

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Yuxuan Gu, Wuyang Zhou, Huijun Xing, Danilo Mandic

Abstract: Residual connections rely on a static residual pathway, and are essential for training deep neural networks. Hyper-connections (HC) increase the expressivity of residual routing by incorporating multiple residual streams and learning dynamic information flow, while manifold-constrained (mHC) variants stabilize training through doubly stochastic residual mixing. However, a generator-level bottleneck remains in existing methods: they use dense, unstructured generators for pre-branch aggregation, residual mixing, and post-branch redistribution, which results in parameter count growing rapidly with the number of streams. To address this issue, we propose \underline{\textbf{T}}ensorized \underline{\textbf{E}}fficient \underline{\textbf{M}}anifold-constrained \underline{\textbf{P}}arameterization for \underline{\textbf{E}}xpressive Residual \underline{\textbf{R}}outing (\textbf{TEMPER}), which represents these generators as multi-way tensors over the input-stream, feature, and output-stream modes, and parameterizes them using tensor networks. Such a structured low-rank formulation is shown to preserve token-dependent manifold-constrained routing interface while substantially reducing parameter growth. It also promotes interpretability and intuition, as: i) tensor ranks control the dimensionality of the learned routing subspace, with full ranks recovering dense routing; while ii) the generator approximation errors bound differences in routing logits and, consequently, in the routed-block outputs. Comprehensive experiments show that TEMPER matches or outperforms existing methods across language modeling and commonsense reasoning tasks, while requiring substantially fewer additional parameters. At eight residual streams, TEMPER achieves the best CORE score while using about $84\%$ fewer additional parameters than mHC, thus showing a stronger performance-parameter efficiency trade-off.

Comment: Tensor-factorizes manifold-constrained residual-routing generators to reduce parameter growth as residual streams increase.

Topic Match: Residual-stream parameterization is the core architectural contribution, with tensor compression providing the efficiency benefit.

Relevance: 10 Novelty: 7


2. Skip a Layer or Loop It? Learning Program-of-Layers in LLMs

ArXiv ID: 2606.06574

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Ziyue Li, Yang Li, Tianyi Zhou

Abstract: Large language models (LLMs) perform inference by following a fixed depth and order, non-recurrent execution of all layers. We reveal the wide existence of training-free, flexible, dynamic program-of-layers (PoLar), where pretrained layers can be packed as modules and then skipped or looped to form a customized program for each input. For most inputs, substantially shorter program executions can achieve the same or better accuracy, while incorrect predictions of the original LLM can be corrected by alternative programs with fewer layers. These observations indicate that inference admits multiple valid latent computations beyond the standard forward pass. To efficiently achieve PoLar in practice, we propose a lightweight PoLar prediction network, which learns to generate execution programs that dynamically skip or repeat pretrained layers for each input. Experiments on mathematical reasoning benchmarks demonstrate that PoLar consistently improves accuracy over standard inference and prior dynamic-depth methods, often while executing fewer layers, and that these gains persist under out-of-distribution evaluation. Our results suggest that fixed-depth execution captures only a narrow subset of an LLM's latent reasoning capacity.

Comment: Learns input-conditioned programs that skip or repeat pretrained transformer layers.

Topic Match: Dynamic layer execution is primarily a new computational architecture, with inference savings as a secondary benefit.

Relevance: 9 Novelty: 8


3. Modular TTT: Rethinking Test-Time Training as Composable Modules

ArXiv ID: 2608.07110

Primary Topic: Architecture and Training Dynamics

Authors: Bohao Tang, Zhen Qin, Yuqi Pan, Zheng Li, Pengfei Liu, Ya Zhang

Abstract: Test-time training (TTT) views sequence modeling as an online learning problem in which fast weights are updated by an internal learning rule. Despite the growing number of TTT variants, existing approaches typically hard-code each variant separately, which makes it difficult to design new TTT methods and to isolate the role of each component. To address this, we propose Modular TTT, a framework that represents the inner learner as a directed acyclic graph and exposes the fast-weight network, loss function, learning rate, weight decay, and normalization as explicit design dimensions. Modular TTT automatically composes primitive-level train-view forward, train-view backward, and causal query-view rules into the full graph-level TTT computation, including the fast-weight state transition. Using Modular TTT, we systematically ablate the components of TTT and find that small learning-rate initialization, weight decay, and a single-layer nonlinearity improve performance, while MSE and inner-product losses perform similarly. Deeper fast-weight networks and normalization tend to hurt performance because they induce excessively large activations, while residual connections and gating provide little measurable benefit. Guided by these findings, we train the best resulting variant as 410M- and 1.45B-parameter models on 100B tokens, and observe training loss and benchmark performance comparable to Gated DeltaNet.

Comment: Factorizes test-time training into composable fast-weight, loss, optimization, and normalization modules.

Topic Match: The paper systematically redesigns and analyzes the inner-learning mechanism of a recurrent sequence architecture.

Relevance: 9 Novelty: 8


4. Faster Query-Key Learning Sharpens Attention in Self-Attention Models

ArXiv ID: 2608.06776

Primary Topic: Architecture and Training Dynamics

Authors: Rahul Vashisht, Harish G. Ramaswamy

Abstract: A standard self-attention layer consists of two interacting circuits: the query-key circuit that governs attention allocation, and the output-value circuit that maps attended representations to predictions. Collapsed and factorized parameterizations of the query-key and output-value circuits lead to qualitatively different attention patterns. In particular, some parameterizations give sharper attention to task-relevant tokens, at a similar training loss. We analyze how the parameterizations of these circuits shape the parameter trajectories in single-layer self-attention models trained for next-token prediction. Through gradient-flow analysis, we show that factorization induces implicit rescaling of the two circuits' learning rates. We derive closed-form dynamics showing that output-value and query-key parameters move along a line, with relative speeds determined by their learning rates. Faster query-key learning relative to output-value learning thus produces sharper attention, as the model compensates for slower output-value learning by increasing attention mass on relevant tokens. Experiments show that differences in the relative learning rates of the two circuits govern attention concentration. This improves attention interpretability proxies while maintaining comparable predictive performance.

Comment: Derives how factorization implicitly changes query-key versus output-value learning rates and attention sharpness.

Topic Match: This is a direct mechanistic account of how attention parameterization controls training trajectories.

Relevance: 9 Novelty: 8


5. Barriers to Universal Reasoning With Transformers (And How to Overcome Them)

ArXiv ID: 2604.25800

Primary Topic: Architecture and Training Dynamics

Authors: Oliver Kraus, Yash Sarrof, Yuekun Yao, Alexander Koller, Michael Hahn

Abstract: Chain-of-Thought (CoT) has been shown to empirically improve Transformers' performance, and theoretically increase their expressivity to Turing completeness. However, whether Transformers can learn to generalize to CoT traces longer than those seen during training is understudied. We use recent theoretical frameworks for Transformer length generalization and find that -- under standard positional encodings and a finite alphabet -- Transformers with CoT cannot solve problems beyond $TC^0$, i.e. the expressivity benefits do not hold under the stricter requirement of length-generalizable learnability. However, if we allow the vocabulary to grow with problem size, we attain a length-generalizable simulation of Turing machines where the CoT trace length is linear in the simulated runtime up to a constant. Our construction overcomes two core obstacles to reliable length generalization: repeated copying and last-occurrence retrieval. We assign each tape position a unique signpost token, and log only value changes to enable recovery of the current tape symbol through counts circumventing both barriers. Further, we empirically show that the use of such signpost tokens and value change encodings provide actionable guidance to improve length generalization on hard problems.

Comment: Proves standard positional encodings block length-generalizable universal computation and constructs a signpost-token remedy.

Topic Match: It provides a mechanistic theory of transformer expressivity and an architectural encoding that overcomes the identified barriers.

Relevance: 8 Novelty: 9


6. Exact Network Surgery: Functional Invariance and Gradient Plasticity in Reactive Computational Graphs

ArXiv ID: 2607.16568

Primary Topic: Architecture and Training Dynamics

Authors: Abdallah Khemais

Abstract: Function-preserving network growth techniques such as Net2Net and progressive stacking expand a model's capacity without destroying its learned function, but existing formulations either tolerate numerical perturbations or require a full rebuild of the training program. We formalize Exact Network Surgery: the in-place insertion of a residual block into a live computational graph such that (i) the network function is preserved -- bit-exactly under explicit floating-point hypotheses -- and (ii) inserted parameters remain trainable immediately after insertion. We prove an identity-morphism theorem for gated residual blocks, a structural-locality theorem showing that a reactive invalidation engine recomputes exactly the downstream cone of the insertion point, leaving every other node's value and optimizer state untouched, and an escape-from-initialization proposition showing that the Gradient Shadowing gate alpha, initialized at zero over a randomly initialized branch, receives a generically non-zero gradient at insertion time. We identify a degenerate configuration -- zero-initialized output projections combined with a zero gate -- that is an exact saddle point gradient descent cannot escape. Every claim is validated on the reference implementation in NeuroDSL, a reactive graph engine in Julia: grafting is bit-exact on every logit tested (0 mismatches out of 1600); the gate escapes zero at the first optimizer step and unlocks branch gradients at the second, exactly as predicted; the degenerate configuration exhibits gradients identically zero for the entire 600-step run; surgery cost tracks downstream cone size with r = 0.9992 while graft-plus-invalidation bookkeeping is constant (about 0.75 ms) across insertion depths; and training resumes bit-identically across a real process restart. A flagged preliminary appendix reports first single-seed observations on post-insertion gate dynamics.

Comment: Provides bit-exact, function-preserving insertion of trainable residual blocks into a live network.

Topic Match: It develops an architectural growth mechanism and analyzes the gradients that make inserted capacity trainable.

Relevance: 8 Novelty: 8


7. Beyond Structural Symmetries: Linear Mode Connectivity via Neuron Identifiability

ArXiv ID: 2606.04754

Primary Topic: Architecture and Training Dynamics

Authors: Vincent Bürgin, Daniel Herbst, Ya-Wei Eileen Lin, Stefanie Jegelka

Abstract: Many striking phenomena in deep learning, such as linear mode connectivity and the structured behavior of training dynamics, are closely tied to parameter symmetries: transformations that leave the realized function unchanged. Despite growing attention to parameter symmetries, the exact interplay between parameters, data, and representations remains underexplored. To investigate this, we develop a theoretical framework of effective function classes, i.e., the set of functions a neuron can realize on its input support, and the norm cost of realizing them. We then formalize effective symmetry breaking via neuron identifiability across independent training runs. Our analysis shows that neural networks can admit large families of approximately equivalent solutions even in structurally asymmetric models. We further show that neuron identifiability enables representation merging without prior alignment, and characterize when such merging admits a linear low-loss path. These findings highlight the role of effective function classes in affecting the loss landscape.

Comment: Links neuron identifiability to unaligned representation merging and linear low-loss paths.

Topic Match: The paper directly analyzes parameter symmetries, solution geometry, and linear mode connectivity in trained networks.

Relevance: 8 Novelty: 8


8. Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers

ArXiv ID: 2608.07436

Primary Topic: Architecture and Training Dynamics

Also Matches: Large-Scale Training Systems and Efficiency

Authors: Ali Janati, Kaoutar El Maghraoui, Andrei Kanavalau, Anass Belfatmi

Abstract: Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head. Muon groks modular addition faster, but its solutions do not hold. All nine configurations on $(a+b) \bmod 113$ grok and later lose generalization. Across five seeds the selected AdamW reference falls below threshold on four, reaching 27.59%. Instability persists across two moduli, two widths, two training fractions, subtraction, and depth. The failure arises at the representation-readout interface, identified only jointly up to an invertible map unselected by the loss. After solving the training set, the gradient falls to order $10^{-6}$ and the optimizers respond differently: step-size elasticity is -0.03 for Muon versus +1.5 for AdamW, and the Muon group moves 8.0 times faster per parameter. From bit-identical states, freezing either group prevents failure. Freezing embeddings/readout removes it in five runs over 451,400 post-grokking steps and five paired seeds: unfrozen arms record 137-321 sub-threshold evaluations, frozen arms none. Removing Muon's normalization and orthogonalization is no substitute: it collapses representation from 326 effective conjugate pairs to 4, shows no recurrent collapse, and fails terminally. Fourier filtering separates circuit failure from masking. Across 43 checkpoints over five seeds and three regimes, the task-aligned family reaches exactly 100% alone. In circuit failure it no longer solves the task; in masking it remains perfect while the full model reaches 45.85%, giving a positive margin on every example, including errors, but being outvoted by a near-equal adversarial remainder. Rescaling it restores 99.9%; grokking is the same condition resolving upward. The task selects the family, swapping $(k,k)$ for $(k,-k)$ under subtraction. Across an abrupt collapse, standard Fourier support is unchanged and the power-distribution cosine remains 0.9899.

Comment: Diagnoses post-grokking collapse as optimizer mismatch at the representation-readout interface.

Topic Match: The central result explains a concrete training instability involving Muon and AdamW parameter groups.

Relevance: 8 Novelty: 8


9. Pay Attention to Attention Distribution: A New Local Lipschitz Bound for Transformers

ArXiv ID: 2507.07814

Primary Topic: Architecture and Training Dynamics

Authors: Nikolay Yudin, Sergei Kudriashov, Alexander Gaponov, Maxim Rakhuba

Abstract: We introduce a novel upper bound on the local Lipschitz constant of the dot-product self-attention block showing its dependence on the attention map distributions. The proposed bound is not only tighter than the prior art, but for the first time, reveals how the distribution of attention probabilities shapes the local Lipschitz constant of the self-attention block. The theoretical basis of the proposed upper bound lies in the refined closed-form upper bounds on singular values of the Jacobian of softmax function. Leveraging these theoretical insights, we introduce JaSMin (Jacobian Softmax norm Minimization), a lightweight regularizer that directly controls the local Lipschitz constant of each block and, consequently, the entire model. Additionally, we discuss how the nature of the attention map distribution contributes to the gradient dynamics and, consequently, transformer training stability.

Comment: Derives an attention-distribution-dependent Lipschitz bound and a regularizer for transformer stability.

Topic Match: The paper directly links attention geometry to gradients and transformer training stability.

Relevance: 8 Novelty: 8


10. Clustered Attractor Manifolds and Dynamical Condensation in Self-Attention

ArXiv ID: 2608.08922

Primary Topic: Architecture and Training Dynamics

Authors: Qucheng Gao, Zuyi Yang, Xiao Chen

Abstract: Transformer layers generate state-dependent interaction networks: token representations determine the attention matrix, which in turn updates the representations. We study this feedback in a minimal normalized self-attention dynamics and identify the overlap gap as the central quantity governing its attractor structure in the thermodynamic limit. When tokens form internally aligned clusters and their similarity to members of the same cluster exceeds that to every other cluster by a nonvanishing amount, inter-cluster attention is exponentially suppressed as the dimension increases. This mechanism produces a high-dimensional manifold of clustered fixed points, ranging from a few macroscopic clusters to extensive microscopic fragmentation, and also controls their stability against perturbations. Starting from an unstructured Gaussian state, we find that clustered states nucleate from the diffuse background only above a finite threshold in attention sharpness, giving rise to a dynamical attention-condensation transition.

Comment: Derives clustered attractor manifolds and a sharpness-driven condensation transition in normalized self-attention.

Topic Match: The paper directly analyzes the state-dependent dynamics and stability of self-attention.

Relevance: 8 Novelty: 8


11. Shattered Compositionality: Counterintuitive Learning Dynamics of Transformers for Arithmetic

ArXiv ID: 2601.22510

Primary Topic: Architecture and Training Dynamics

Authors: Xingyu Zhao, Darsh Sharma, Rheeya Uppaal, Yiqiao Zhong

Abstract: Large language models (LLMs) often achieve strong benchmark accuracy yet remain brittle under small distribution shifts. While recent mechanistic studies reveal the discrepancy between LLMs and humans in skill compositions, the learning dynamics of skill acquisition and the role of data distributions remain elusive. In this study, we train transformers on synthetic arithmetic tasks with black-box model-agnostic metrics for analyzing non-human skill compositions. We discover that transformers often acquire skills for arithmetic in reverse order or in parallel instead of human-like sequential rules--a phenomenon we refer to as shattered compositionality. To explain these behaviors, we provide evidence that correlational matching to the training data, rather than causal or procedural composition, shapes learning dynamics. As a consequence, this non-human acquisition creates competition between partially learned skills, producing characteristic mixing errors and weaker robustness under controlled distribution shifts. We further show that the same qualitative behavior persists in modern LLMs and is not mitigated by pure model scaling or scratchpad supervision. Our results highlight a mismatch between training-time skill acquisition and the human-like hierarchical compositions, with implications for reasoning reliability and out-of-distribution robustness.

Comment: Finds that arithmetic skills emerge in non-hierarchical orders through correlational matching and competing partial skills.

Topic Match: It directly analyzes transformer skill-acquisition dynamics under controlled training distributions.

Relevance: 8 Novelty: 8


12. Transformer Circuits Can Realize Clustering Algorithms

ArXiv ID: 2506.19125

Primary Topic: Architecture and Training Dynamics

Authors: Kenneth L. Clarkson, Lior Horesh, Takuya Ito, Charlotte Park, Parikshit Ram

Abstract: Although transformers are most commonly optimized as statistical sequence models, it is unclear to what extent they can implement and learn exact algorithmic computations. Here, we specify a transformer implementation from first principles that executes a fundamental and widely used method for $k$-means clustering: Lloyd's algorithm. We theoretically prove and empirically demonstrate that this implementation of a transformer architecture, which we term the $k$-means transformer, exactly implements Lloyd's algorithm for $k$-means clustering using the standard circuit mechanisms of modern transformers: attention block, residual connections, and feed-forward block. In learning experiments, we find that training this base architecture on $k$-means clustering yields a generalizable clustering algorithm that surpasses Lloyd's algorithm in terms of clustering quality. Finally, we demonstrate that interpretable alterations (e.g., inclusion of layer normalizations) to this architecture yields diverse and novel variants of clustering algorithms, including soft $k$-means, spherical $k$-means, trimmed $k$-means. Overall, our results show that transformer circuit mechanisms can instantiate exact algorithmic routines for clustering, while simultaneously providing an effective learnable model.

Comment: Constructs exact transformer circuits for Lloyd's k-means and studies how normalization alters the realized algorithm.

Topic Match: The core contribution analyzes the algorithmic capabilities of attention, residual, and feed-forward mechanisms.

Relevance: 7 Novelty: 8


13. Rethinking Reasoning with MDLMs: Early Exits, Post-hoc Reasoning, and Beyond

ArXiv ID: 2510.19990

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Zachary Horvitz, Raghav Singhal, Hao Zou, Carles Domingo-Enrich, Zhou Yu, Rajesh Ranganath, Kathleen McKeown

Abstract: The reasoning paradigm, where language models reason before answering, has enabled breakthroughs on tasks such as mathematical problem-solving. While current tooling for reasoning is built around next-token prediction trained models, recent works introduce an alternative choice: masked diffusion language models (MDLMs). MDLMs are trained to infill positions in randomly masked sequences. We introduce reasoning-as-infilling, a prompting technique that pre-fills tokens to explicitly delimit reasoning and answer regions, unlocking a unified set of capabilities for MDLM reasoning. Because answer positions are explicitly designated, the model's conditional distributions over answer tokens are directly accessible during generation. This enables early exits when the model is certain of its answer. The same framework supports post-hoc reasoning: given question-answer pairs, MDLMs can sample high-quality reasoning traces from their posterior, a distribution that is intractable for autoregressive models. On GSM8k, fine-tuning LLaDA-8B-Base on these posterior traces improves accuracy by +14.9%, matching gains from human-written traces (+13.4%). Finally, given a reference answer, the answer region distributions enable scoring partial reasoning traces at intermediate steps, providing intermediate rewards that are more strongly correlated with correctness than scores from a specialized process reward model. At intermediate steps, answer-likelihood scores from autoregressive models are significantly less predictive of correctness than those from MDLMs. Our results demonstrate that the MDLM training objective provides promising benefits for reasoning.

Comment: Uses masked-LM infilling structure for early exits, posterior reasoning traces, and intermediate rewards.

Topic Match: The main insight derives new computation and training capabilities from the masked diffusion language-model objective.

Relevance: 7 Novelty: 8


14. Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation

ArXiv ID: 2608.08469

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Kaichen Zhang, Wei Huang, Keming Wu, Bo Li, Xiaojuan Qi

Abstract: Existing streaming multimodal models process observations incrementally but still follow a turn-based prefill-then-decode pattern, making them non-duplex: new observations cannot naturally enter an active generation stream. Proactive alternatives use micro-turn polling or external response gates, which fragment continuous interaction, decouple response timing from language generation, and complicate KV-cache-friendly serving. We introduce Aero Realtime, a 4B streaming multimodal model with a duplex architecture for realtime generation. Aero Realtime aligns video, audio, and textual output on a shared temporal grid, where each approximately 80-ms audio slot predicts either a lexical token or a silence token. This allows input and output to advance together, enabling one autoregressive objective to learn both when to respond and what to generate. During inference, Aero Realtime appends only the newest multimodal slot, carries forward the previous output state, and reuses the KV cache for efficient incremental execution. We further provide a complete training and serving recipe, including realtime QA construction, slot-aligned supervision, hardware-aware distributed training, and resumable inference. On four NVIDIA A6000 workstation GPUs, Aero Realtime maintains 84-ms median and 173-ms P95 processing lag over 20 minutes of a continuously streamed video, remaining within 200~ms of the source timeline. These results demonstrate the feasibility of fully aligned input-output modeling for duplex, proactive, and hardware-aligned multimodal interaction.

Comment: Introduces a duplex autoregressive architecture that aligns streaming inputs and outputs on a shared temporal grid.

Topic Match: The central contribution is a new streaming architecture, with incremental KV-cache reuse as a secondary efficiency benefit.

Relevance: 7 Novelty: 8


15. Second Order Drifting Models

ArXiv ID: 2608.07924

Primary Topic: Architecture and Training Dynamics

Authors: Drake Brown, Yuhao Huang, Shih-Hsin Wang, Bao Wang

Abstract: Drifting models are a recent class of one-step generative models that evolve the model distribution during training using a predefined sample-based drift field. Although they avoid iterative inference, their kernel-based drift fields induce frequency-dependent training dynamics: In the linearized regime, each Fourier mode of the density residual decays at a rate determined by the kernel spectrum, leading to slow recovery of fine-scale structure. We propose Second-Order Drifting Models, which lift drifting dynamics into phase space by augmenting generated samples with artificial velocity variables. We show that the resulting density perturbations obey accelerated second-order dynamics in Fourier space, connecting drifting models to the celebrated Nesterov acceleration from optimization theory. This provides a principled mechanism for mitigating the spectral stiffness of first-order drifting while preserving one-step inference. We derive a practical semi-implicit training algorithm and evaluate it on synthetic distribution matching, sequential data generation, and robotic control. Across these settings, the second-order drifting model improves convergence behavior and achieves competitive or superior performance over first-order drifting baselines.

Comment: Introduces phase-space training dynamics that accelerate slow Fourier modes in one-step generative models.

Topic Match: Its main advance is a second-order architectural and optimization mechanism with an explicit convergence explanation.

Relevance: 7 Novelty: 8


16. Two-Layer Linear Auto-Regressive Models Estimate Latent States

ArXiv ID: 2606.12691

Primary Topic: Architecture and Training Dynamics

Authors: Yahya Sattar, Sunmook Choi, Leo Maynard-Zhang, Yassir Jedra, Maryam Fazel, Sarah Dean

Abstract: Auto-regressive models have emerged as powerful tools for sequential data, from language to video. Understanding how and why these models learn latent representations remains an open theoretical question. In this work, we demonstrate that when trained by empirical risk minimization on data from partially observed linear dynamical systems, two-layer linear auto-regressive models naturally learn to approximate Kalman filtering. In particular, we show that the learned hidden representation coincides, up to a similarity transformation, with the state estimates produced by the optimal (Kalman) filter, even though the model has no explicit knowledge of the underlying dynamics or state. The result follows from three main insights. First, we establish that the Kalman filter is well approximated by an auto-regressive model with bounded truncation error. Second, we show that despite non-convexity, the two-layer optimization landscape is benign, i.e., all stationary points are either strict saddles or global minima. Finally, as our main contributions, we provide finite-sample guarantees on prediction error, parameter estimation error, and latent state recovery. Numerical simulations support the theoretical results and demonstrate that the latent representations of auto-regressive models recover state estimates.

Comment: Proves that trained two-layer autoregressive models recover Kalman-filter latent states under controlled dynamics.

Topic Match: It provides mechanistic training-dynamics and representation-recovery guarantees for autoregressive models.

Relevance: 7 Novelty: 8


17. Tokenisation over Bounded Alphabets is Hard

ArXiv ID: 2511.15709

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Violeta Kastreva, Philip Whittington, Dennis Komm, Tiago Pimentel

Abstract: Recent works have shown that tokenisation is NP-complete. However, these works assume tokenisation is applied to inputs with unboundedly large alphabets -- an unrealistic assumption, given that in practice tokenisers operate over fixed-size alphabets, such as bytes or Unicode-characters. We close this gap by analysing tokenisation over bounded alphabets, considering two natural variants: bottomup tokenisation and direct tokenisation, where we must, respectively, select a sequence of merge operations or a vocabulary whose application optimally compresses a dataset. We prove that even with binary alphabets, both variants are not only NP-complete, but also APX-hard and thus admit no polynomial-time approximation scheme (unless P=NP). We further show that direct tokenisation remains NP-complete even when applied to unary alphabets. These results establish that the computational intractability of tokenisation is not an artifact of large alphabets or complex constructions, but a fundamental barrier. Overall, our results explain why current practical algorithms such as BPE and UnigramLM are heuristic, and point toward approximation algorithms being an important path going forward for tokenisation research.

Comment: Introduces continuous-time attention with sparse gated projections and subquadratic top-k interactions.

Topic Match: The novel attention mechanism is primary, while its sparse interaction scheme supplies an efficiency contribution.

Relevance: 7 Novelty: 8


18. Symmetry Reveals Layerwise Dynamics: How Transformers Perform In-Context Classification

ArXiv ID: 2604.11613

Primary Topic: Architecture and Training Dynamics

Authors: Patrick Lutz, Themistoklis Haris, Arjun Chandra, Aditya Gangrade, Venkatesh Saligrama

Abstract: Transformers can perform in-context classification from a few labeled examples, yet the inference-time algorithm remains opaque. We study multi-class linear classification in the hard no-margin regime and make the computation identifiable by enforcing feature- and label-permutation equivariance at every layer. This enables interpretability while maintaining functional equivalence and yields highly structured weights. From these models we extract an explicit depth-indexed recursion: an end-to-end identified, emergent update rule inside a softmax transformer, to our knowledge the first of its kind. Attention matrices formed from mixed feature-label Gram structure drive coupled updates of training points, labels, and the test probe. The resulting dynamics implement a geometry-driven algorithmic motif, which can provably amplify class separation and yields robust expected class alignment.

Comment: Extracts an explicit layerwise recursion showing how transformer attention performs in-context classification.

Topic Match: This is a mechanistic analysis of the computation implemented by transformer layers.

Relevance: 7 Novelty: 8


19. How Many Different Outputs Can a Transformer Generate?

ArXiv ID: 2605.22223

Primary Topic: Architecture and Training Dynamics

Authors: Maxime Meyer, Mario Michelessa, Caroline Chaux, Vincent Y. F. Tan

Abstract: We study how we can leverage only a handful of characteristics of a transformer's architecture to closely predict the number of different sequences it can output, both qualitatively and quantitatively. We provide an upper bound depending on the length of the prompt, which we show empirically to be tight up to a factor less than 10, across architectures and model sizes. Our analysis also provides a theoretical explanation for previously observed empirical failures of transformers on simple sequence tasks, such as copying and cramming. Formally, we prove that (i) the maximal length of accessible sequences (those that the transformer can output for some prompt) grows linearly with the prompt length, (ii) beyond a critical threshold, the proportion of accessible sequences decays exponentially with sequence length, and (iii) the linear coefficient relating prompt length to accessible sequence length admits a theoretical upper bound. Notably, these results hold even with unbounded context and computation time.

Comment: Bounds the number and length of sequences accessible to transformers from architectural characteristics.

Topic Match: The paper provides foundational analysis of a transformer's architectural output capacity.

Relevance: 7 Novelty: 8


20. On the Infinite Width and Depth Limits of Predictive Coding Networks

ArXiv ID: 2602.07697

Primary Topic: Architecture and Training Dynamics

Authors: Francesco Innocenti, El Mehdi Achour, Rafal Bogacz

Abstract: Predictive coding (PC) is a biologically plausible alternative to standard backpropagation (BP) that minimises an energy function with respect to network activities before updating weights. Recent work has improved the training stability of deep PC networks (PCNs) by leveraging some BP-inspired reparameterisations, but the scalability and theoretical basis of these methods remain unclear. To address this gap, we study the infinite width and depth limits of PCNs. For linear networks, we derive stable and "non-lazy" parameterisations when scaling both the model width and depth, revealing that the output of standard PCNs explodes with width during training. Moreover, under stable parameterisations, we show that the gradients computed by PC at activity equilibrium converge to the BP gradients for networks that are much wider than deep ($depth/width\to0$). Experiments show high gradient alignment between PC and BP at large width for different nonlinear models, including convolutional networks and transformers. Overall, this work constrains the parameterisations that are scalable with PC, while suggesting how BP could be implemented using only local updates in much wider than deep networks like the brain.

Comment: Derives stable width-depth parameterizations for predictive-coding networks and their convergence toward backpropagation gradients.

Topic Match: The paper directly studies scalable parameterization and training stability for an alternative neural learning mechanism.

Relevance: 7 Novelty: 8


21. The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought

ArXiv ID: 2605.18079

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Moritz Brösamle, Stephan Eckstein

Abstract: Existing expressivity results for transformers typically rely on hardmax attention, high precision, and other architectural modifications that disconnect them from the models used in practice. We bridge this gap by analyzing standard transformer decoders with softmax attention and rounding of activations and attention weights, while allowing depth and width to grow logarithmically with the context length. As an intermediate step, we construct hardmax transformers with ternary activations and well-separated attention scores that simulate Turing machines using Chain-of-Thought (CoT). This lets us convert the constructions to equivalent softmax transformers without the unrealistic parameter magnitudes or activation precision that prior approaches would require. Using the same technique, we analyze a recently proposed summarized CoT paradigm and show that it simulates Turing machines more efficiently, with model size scaling logarithmically in a space bound rather than a time bound. We empirically test predictions made by our results on a Sudoku reasoning task and find better alignment with learnability than for prior high-precision results. Our code is available at https://github.com/moritzbroe/transformer-expressivity.

Comment: Establishes computational expressivity bounds for softmax Transformers with rounded activations and attention weights.

Topic Match: Analyzes the computational consequences of attention precision and CoT organization, including model-size bounds; training implications are indirect.

Relevance: 7 Novelty: 8


22. Is SwiGLU's Open Positive Tail Necessary? Evidence from Closed-Tail Gating with MemGLU

ArXiv ID: 2608.07323

Primary Topic: Architecture and Training Dynamics

Authors: Yuting Ge, Pengju Yang, Mingkai Nie

Abstract: We test whether decoder-only language-model FFNs require SwiGLU's open positive tail. We introduce MemGLU as a closed-tail comparator derived from a memristive branch geometry. Across paired 9M and 30M pretraining runs with three seeds, MemGLU remains within about 0.1% of SwiGLU in validation NLL. Trained SwiGLU checkpoints are sensitive to positive-tail suppression, while mechanism diagnostics show that the two models use their gates differently despite similar losses. These results suggest that models adapt to the gate geometry available during pretraining. At the tested scales, SwiGLU's open positive tail is not necessary for decoder-only language-model FFNs.

Comment: Tests a closed-tail FFN gate and shows pretrained models adapt to gate geometry without needing SwiGLU's open tail.

Topic Match: The work directly probes an architectural gating mechanism through controlled pretraining and mechanistic diagnostics.

Relevance: 8 Novelty: 6


23. Think Deep, Speak Once: Relit, A Recursive Latent Implicit Transformer Framework

ArXiv ID: 2608.08113

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Abhishek Panwar, Maheep Singh, Saksham Bansal

Abstract: Chain-of-Thought (CoT) prompting has become the dominant paradigm for eliciting reasoning in Large Language Models (LLMs), yet it creates substantial computational overhead by forcing models to externalize intermediate reasoning steps as discrete tokens. Recent latent reasoning approaches attempt to internalize this process within continuous hidden states. One of the latest advancements in the field of latent reasoning, Tiny Recursive Models (TRMs) excel at symbolic reasoning but struggle to preserve semantic coherence in natural language settings. To bridge this gap, we introduce ReLIT (Recursive Latent Implicit Transformer), a hybrid framework that grounds deep recursive reasoning within the rich semantic representations of a foundational model. ReLIT augments a frozen LLM backbone (TinyLlama-1.1B) with a lightweight, trainable recursive block that iteratively refines its latent thinking (z) before committing to a final output, structurally solving linguistic intuition from algorithmic processing and enabling "deep thinking" via gradient-isolated recurrent loops without the latency of explicit token generation. Empirically, ReLIT achieves high parameter efficiency on the GLoRE logical reasoning benchmark, matching or outperforming significantly larger models on challenging tasks such as ProofWriter and RuleTaker despite minimal supervision. These results demonstrate that reasoning capability can be scaled efficiently through recurrent depth rather than parameter width, offering a principled framework for semantically grounded implicit reasoning.

Comment: Adds a trainable latent recurrent block with gradient-isolated loops to increase computation through recursive depth.

Topic Match: The recursive block changes computation and gradient flow, with parameter-efficiency benefits demonstrated through reasoning-task adaptation of a frozen backbone.

Relevance: 8 Novelty: 6


24. State Propagation Also Satisfies: A Complex-Valued State-Space Model for Deterministic State Tracking

ArXiv ID: 2608.03425

Primary Topic: Architecture and Training Dynamics

Authors: Xiaohe Li, Yang Lu

Abstract: Transformer-based architectures have dominated sequence modeling, largely due to the expressive power of attention mechanisms. However, for a class of deterministic state tracking tasks---such as parity checking, modular counting, and parenthesis matching---attention may be overkill. In this paper, we show that \textbf{state propagation alone is sufficient}. We propose the \textbf{Complex State Propagator (CSP)}, a minimalistic recurrent architecture that \textbf{only propagates hidden states} across layers without output projections at intermediate steps. The state is represented as a complex-valued vector, updated via input-dependent rotations in the complex domain. To enable deep propagation without gradient vanishing or degradation, we introduce a \textbf{block-level skip connection} alongside element-wise complex normalization and SiLU activation at sequence boundaries. Applied with Focal Loss, CSP achieves \textbf{100\% accuracy} with perfect F1 scores across canonical tasks.

Comment: Introduces a complex-valued recurrent state-propagation architecture with normalization and block-level skips.

Topic Match: The core contribution is a new recurrent sequence mechanism designed for stable deep propagation.

Relevance: 7 Novelty: 7


25. Neuronal Attention Circuit (NAC) for Representation Learning

ArXiv ID: 2512.10282

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Waleed Razzaq, Yun-Bo Zhao

Abstract: Attention improves representation learning over RNNs, but its discrete nature limits continuous-time (CT) modeling. We introduce Neuronal Attention Circuit (NAC), a novel, biologically inspired CT-attention mechanism that reformulates attention logit computation as the solution to a linear first-order ODE with nonlinear interlinked gates derived from repurposing the wiring of C. elegans Neuronal Circuit Policies (NCPs). NAC replaces dense projections with sparse sensory gates for query-key projections and introduces a sparse backbone network with two heads for computing content-target and learnable time-constant gates, enabling efficient adaptive dynamics. To improve efficiency and memory consumption, we implement an adaptable, sparse, subquadratic Top-K pairwise concatenation mechanism that selectively curates query-key interactions. We provide rigorous theoretical guarantees, including state stability and bounded approximation errors. Empirically, we implement NAC in diverse domains, including irregular time-series, long-range forecasting, lane-keeping for autonomous vehicles, and industrial prognostics. We observe that NAC matches or outperforms in accuracy against several CT state-of-the-art baselines, while being interpretable at the neuron cell level.

Comment: Reformulates attention-logit computation as a gated continuous-time ODE with stability and approximation guarantees.

Topic Match: The new attention operator is the primary contribution; sparse projections and subquadratic interaction selection add efficiency relevance, although large-model training remains untested.

Relevance: 7 Novelty: 7


26. PRISM: Principled Reference Identification for Schrodinger Bridge Model

ArXiv ID: 2608.06893

Primary Topic: Architecture and Training Dynamics

Authors: Forouzan Fallah, Yezhou Yang

Abstract: Schrödinger bridge models restore a clean signal from a degraded observation by following the conditional bridges of a reference process, yet this reference is chosen heuristically, typically white noise with a hand-tuned schedule. We develop PRISM, a theory of bridge reference design. We characterize the time-varying Gaussian references that remain exactly tractable with per-mode schedules: precisely those whose instantaneous covariances commute. We then prove an invisibility principle: with the exact drift and unlimited solver steps, every admissible reference recovers the true posterior. The choice of reference therefore matters only under finite computational resources. For a fixed step budget, we derive the finite-step objective in closed form and prove that every optimal noise spectrum is proportional to Pk, the spectrum of information destroyed by the sensor, with a mode-independent constant x*(T) = (2 ln T)^-1/2 (1 + o(1)). The analysis shows that noise color and temporal scheduling are interchangeable, and regularization provably shifts the optimal reference toward white noise. Experiments in Gaussian settings confirm the predicted orderings and the closed-form loss floors. On FFHQ, the distortion-- perception trade-off and spectral localization transfer, but white noise outperforms the matched reference; a pre-registered study that changes the training regime refutes ridge whitening as the explanation. A 2x2 mechanism study then traces the inversion to the non-Gaussian per-mode statistics of real images. PRISM turns reference design from a hyperparameter sweep into a calculation in the Gaussian regime, and locates exactly where real images break it.

Comment: Derives compute-budget-optimal reference processes for Schrödinger bridge training and explains finite-step behavior.

Topic Match: The work centers on mechanistic training dynamics and principled reference-process design.

Relevance: 6 Novelty: 8


27. Kolmogorov-Arnold Energy Models: Fast, Interpretable Generative Modeling

ArXiv ID: 2506.14167

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Prithvi Raj

Abstract: Generative models typically rely on either simple latent priors (e.g., Variational Autoencoders, VAEs), which are efficient but limited, or expressive iterative samplers (e.g., Diffusion and Energy-based Models), which are costly and opaque. We introduce a new unsupervised model, the Kolmogorov-Arnold Energy Model (KAEM), to bridge this trade-off and provide new opportunities for interpretability. Based on a novel adaptation of the Kolmogorov-Arnold Representation Theorem, KAEM imposes a univariate latent prior, enabling fast and exact inference via the inverse transform method. On small datasets, we show that importance sampling becomes a tractable, unbiased, and single-pass posterior inference method. For settings requiring exploration, we propose a population-based strategy that decomposes the posterior into a sequence of annealed distributions, serving as a new remedy for poor mixing in Energy-based Models. KAEM attains competitive Fréchet Inception Distance among latent-prior models on SVHN, CIFAR10, and CelebA while sampling in a single forward pass at lower cost than iterative EBMs, and exposing an interpretable prior built from 1D densities.

Comment: Introduces an energy model with a univariate learned prior enabling exact single-pass sampling.

Topic Match: The new generative architecture is primary, with non-iterative inference providing an efficiency match.

Relevance: 6 Novelty: 8


28. No Unique Minimizer, No Problem: On the Consistency of Robust Neural Classifiers

ArXiv ID: 2608.08489

Primary Topic: Architecture and Training Dynamics

Authors: Subhabrata Majumdar, Anand Deo, Partha Pratim Saha, Abhik Ghosh

Abstract: Neural network classifiers trained by cross-entropy minimization are highly sensitive to label noise and adversarial contamination. While robust alternatives offer bounded influence and resistance to corruption, their statistical foundations in the deep learning setting are insufficient due to a fundamental difficulty: neural parameterizations are non-identifiable, so the population loss minimizer is an equivalence class of parameters, not a unique point. We develop a consistency theory for robust neural classifiers based on the S-divergence family that requires no identifiability assumption. Casting training as stochastic optimization over a non-identifiable parameter space, we prove that empirical S-divergence minimizers converge to the population-optimal equivalence class under mild regularity conditions, and verify these conditions for three architecture choices. We further establish that limit points of the robust training algorithm are stationary points of the empirical objective. Experiments on vision and language benchmark datasets confirm that S-divergence training maintains clean-data accuracy while exhibiting performance competitive with existing robust methods.

Comment: Proves consistency of robust neural objectives without requiring parameter identifiability.

Topic Match: The contribution is optimization theory for non-identifiable neural parameterizations and robust training losses.

Relevance: 6 Novelty: 7


29. Understanding Reasoning from Pretraining to Post-Training

ArXiv ID: 2607.16097

Primary Topic: Architecture and Training Dynamics

Authors: Jingyan Shen, Ang Li, Salman Rahman, Yifan Sun, Micah Goldblum, Matus Telgarsky, Pavel Izmailov

Abstract: Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic questions remain open: (1) how do pretraining choices (model size, data) shape the returns to RL compute, and (2) what does RL actually do to the model? These questions are difficult to study in the standard LLM setting: pretraining corpora are vast and uncontrolled, making it hard to attribute behaviors to pretraining versus RL, and systematic compute sweeps across both stages are prohibitively expensive. To address these challenges, we use chess as a controlled testbed for studying reasoning across the full pretraining-to-post-training pipeline. We follow the standard LLM training pipeline by pretraining language models from 5M to 1B parameters on human chess games, supervised fine-tuning on synthetic reasoning traces, and running RL on chess puzzles with verifiable rewards. Using this framework, we find that the post-RL performance at given RL compute level is well-predicted from the pretraining loss, and slope of the RL reward curves improves approximately linearly with the pretraining tokens. Beyond scaling, we find that RL does not simply sharpen the SFT policy: on easy puzzles it amplifies correct moves the SFT policy already preferred, while on hard puzzles it surfaces correct moves that were nearly absent under SFT. We further test whether our findings transfer beyond chess by training a 1B language model on math-domain text, where the same predictive pattern emerges: longer-pretrained checkpoints reach higher post-RL performance and improve faster under RL. In sum, we provide a quantitative account of the pretraining-to-RL interface and a controlled testbed for studying the science of reasoning across the full pretraining-to-post-training pipeline.

Comment: Quantifies how pretraining loss and token exposure predict subsequent reinforcement-learning gains.

Topic Match: The most relevant contribution is its analysis of training dynamics across pretraining and post-training stages.

Relevance: 6 Novelty: 7


30. Semantic Adapter Routing with Fine-Tuning Task Embeddings

ArXiv ID: 2606.19079

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Enrico Cassano, Michał Brzozowski, Paolo Mandica, Zuzanna Dubanowska, Neo Christopher Chung

Abstract: Parameter-efficient fine-tuning (PEFT) has led to model ecosystems in which a single backbone is paired with many task-specialized adapters. Given such a library, routing aims to select the most appropriate adapter for a user query. While existing adapter routers typically require access to adapter weights or supervised training, we develop training-free semantic adapter routing methods using task embeddings. In ARIADNE, we reframe adapter selection as a classification problem, where PEFT adapters are represented by task embeddings and an unlabeled query is routed to the nearest adapter in the encoder's latent space. Evaluated on 23 tasks, ARIADNE recovers 97.4% of Oracle task performance and scales to 44 adapters at 89.7% selection accuracy, without touching a single adapter parameter. However, training data needed for ARIADNE may not be available when adapters come from public hubs or third-party providers. To overcome this limitation, we introduce GRACE, which recovers an adapter's fine-tuning data from its output logits alone via a modified contrastive decoding diffing (CDD) procedure. Synthetic data generated from CDD-UM is then used to construct task embeddings. Across three backbones (Llama-3.2-1B, Qwen2.5-3B, Qwen2.5-32B), GRACE recovers 72--100\% of Oracle task accuracy and matches or exceeds ARROW on 48 of 69 task/backbone combinations, while requiring neither training data nor model weights. Overall, we demonstrate that fine-tuning task embeddings provide an accurate and efficient path to semantic adapter routing.

Comment: Routes queries among task adapters using training-free semantic task embeddings.

Topic Match: This is primarily a modular-computation routing mechanism, with parameter-efficient adapter use as a secondary match.

Relevance: 6 Novelty: 7


31. Early Data Exposure Improves Robustness to Subsequent Fine-Tuning

ArXiv ID: 2605.12705

Primary Topic: Architecture and Training Dynamics

Authors: Lawrence Feng, Gaurav R. Ghosal, Jacob Mitchell Springer, Ziqian Zhong, Aditi Raghunathan

Abstract: How can we train models whose post-trained capabilities survive subsequent fine-tuning? Rather than focusing on downstream interventions to mitigate forgetting of upstream capabilities, we study how upstream training choices - that is, the manner in which a capability is acquired - shape how robustly that capability is retained. We investigate this question in a controlled three-stage language-model pipeline: pretraining, post-training to acquire a target capability, and downstream fine-tuning on a new objective. Across 135M and 1B models, two post-training domains, and two downstream fine-tuning tasks, we find that immediate post-training performance does not reliably predict retention after subsequent fine-tuning: training recipes that look equivalent immediately after post-training can retain the target capability very differently after subsequent fine-tuning. In particular, early exposure - mixing post-training data into pretraining - consistently improves the frontier between retained upstream performance and downstream performance. In compute-matched experiments, where the target data must be allocated between pretraining and post-training, we find that the optimum lies at neither extreme. Together with our other empirical and theoretical findings, this supports the view that post-training drives immediate specialization while early exposure improves robustness to later forgetting. Replay and dropout, typically used to mitigate forgetting as it occurs during fine-tuning, provide complementary gains to early exposure when applied during post-training. Our findings suggest that robustness to subsequent fine-tuning should be treated as a first-class objective of upstream training, addressed preventatively through choices like early exposure rather than reactively during fine-tuning itself.

Comment: Shows that mixing capability data into pretraining improves resistance to forgetting during later fine-tuning.

Topic Match: The strongest match is a training-dynamics result about how data timing changes capability retention.

Relevance: 6 Novelty: 7


32. Characteristic Learning for Provable One Step Generation

ArXiv ID: 2405.05512

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Zhao Ding, Chenguang Duan, Yuling Jiao, Ruoxuan Li, Jerry Zhijian Yang, Pingwen Zhang

Abstract: We propose the characteristic generator, an one-step generative model that combines the sampling efficiency of generative adversarial networks (GANs) with the training stability of flow-based models. The proposed model is based on characteristics along which probability-density transport is governed by ordinary differential equations (ODEs). Specifically, we first estimate the underlying velocity field and numerically solve the probability-flow ODE using an exponential integrator, thereby obtaining discrete approximations of the characteristic trajectories. We then train a deep neural network to approximate these trajectories, yielding a one-step transport map from a Gaussian prior to the target distribution. Theoretically, we analyze the errors introduced by velocity-field estimation, numerical discretization, and characteristic approximation, and establish a nonasymptotic convergence rate in the Wasserstein-2 distance under mild assumptions on the data distribution. Moreover, under a low-dimensional linear-subspace assumption, we show that the convergence rate depends on the intrinsic data dimension rather than the potentially much larger ambient dimension, demonstrating the model's ability to alleviate the curse of dimensionality. Experiments on synthetic and real-world datasets show that the characteristic generator produces high-quality, high-resolution samples using only one or a few neural-network evaluations.

Comment: Learns characteristic trajectories to collapse probability-flow generation into one or a few network evaluations.

Topic Match: The central contribution is a new generative computation and training mechanism, with sampling efficiency as a secondary benefit.

Relevance: 6 Novelty: 7


33. Intelligence Foundation Model: A New Perspective to Approach Artificial General Intelligence

ArXiv ID: 2511.10119

Primary Topic: Architecture and Training Dynamics

Authors: Borui Cai, Yao Zhao

Abstract: We propose a new perspective for approaching artificial general intelligence (AGI) through an intelligence foundation model (IFM). Unlike existing foundation models (FMs), which specialize in pattern learning within specific domains such as language, vision, or time series, IFM aims to acquire the underlying mechanisms of intelligence by learning directly from diverse intelligent behaviors. Vision, language, and other cognitive abilities are manifestations of intelligent behavior; learning from this broad range of behaviors enables the system to internalize the general principles of intelligence. Based on the fact that intelligent behaviors emerge from the collective dynamics of biological neural systems, IFM consists of two core components: a novel network architecture, termed the state neural network, which captures neuron-like dynamic processes, and a new learning objective, neuron output prediction, which trains the system to predict neuronal outputs from collective dynamics. The state neural network emulates the temporal dynamics of biological neurons, allowing the system to store, integrate, and process information over time, while the neuron output prediction objective provides a unified computational principle for learning these structural dynamics from intelligent behaviors. Together, these innovations establish a biologically grounded and computationally scalable foundation for building systems capable of generalization, reasoning, and adaptive learning across domains, representing a step toward truly AGI.

Comment: Proposes a state-neural-network architecture and neuronal-output prediction objective for temporal neural dynamics.

Topic Match: A new recurrent computational architecture and learning objective form the paper's core proposal.

Relevance: 6 Novelty: 7


34. Error Bound Analysis for the Regularized Loss of Deep Linear Neural Networks

ArXiv ID: 2502.11152

Primary Topic: Architecture and Training Dynamics

Authors: Po Chen, Rujun Jiang, Peng Wang

Abstract: The optimization foundations of deep linear networks have recently received significant attention. However, due to their inherent non-convexity and hierarchical structure, analyzing the loss functions of deep linear networks remains a challenging task. In this work, we study the local geometry of the regularized squared loss of deep linear networks around each critical point. Specifically, we obtain a closed-form characterization of the critical point set building on existing results and establish an error bound for the regularized loss under mild conditions on network width and regularization parameters. Notably, this error bound quantifies the distance from a point to the critical point set in terms of the current gradient norm, which can be used to derive linear convergence of first-order methods. To support our theoretical findings, we conduct numerical experiments and demonstrate that gradient descent converges linearly to a critical point when optimizing the regularized loss of deep linear networks.

Comment: Establishes local error bounds and linear first-order convergence around deep-linear-network critical sets.

Topic Match: Optimization geometry and convergence dynamics are central, although the analysis uses deep linear networks.

Relevance: 6 Novelty: 6


35. LF${}^{2}$AR: Accounting for Layerwise Dynamics to Improve Multimodal Adaptation of Language Models

ArXiv ID: 2503.06211

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Santiago Cuervo, Adel Moumen, Yanis Labrak, Sameer Khurana, Antoine Laurent, Mickael Rouvier, Phil Woodland, Ricard Marxer

Abstract: Text-pretrained language models (LMs) encode rich world knowledge, but adapting them to process and generate perceptual modalities such as audio and images while effectively leveraging that knowledge remains challenging. Perceptual modalities are finer-grained and less semantically dense than text, making it unclear how functions learned during text pretraining can be reused. We study this problem through the lens of a layerwise abstraction-refinement dynamic observed in transformer LMs: representations first become more abstract and compositional, then are refined into representations predictive of fine-grained structure. This perspective suggests that adapting an LM to finer-grained modalities requires: (i) allocating additional fine-to-coarse processing at the input and coarse-to-fine processing at the output, consistent with late modality fusion and an output-side analogue we term late fission; and (ii) allowing the output predictor to preserve input-dependent selective access to both high-level semantic structure and low-level perceptual detail, motivating our use of attention residuals in fission. We instantiate this view in LF${}^{2}$AR, a simple architecture combining these mechanisms, and study it on text-as-images and speech, two modalities for which semantic correspondence to text can be controlled. Across models ranging from 135M to 2B parameters and adapted to these modalities, we find that these components increase feature abstraction, strengthen cross-modal alignment, enable such alignment to emerge at smaller compute budgets, improve preservation of text-like predictive structure, and yield better performance on text-as-images and speech versions of language understanding and reasoning benchmarks. Additionally, attention residuals induce sparse, interpretable use of deep backbone layers, enabling early-exit decoding with a 1.9$\times$ generation speedup.

Comment: Late fusion plus late fission with attention residuals, derived from a layerwise abstraction-then-refinement view of transformer representations; residuals induce sparse deep-layer use and 1.9x early-exit decoding.

Topic Match: An architectural mechanism (residual/fission design) motivated by a layerwise training-dynamics analysis, though evaluated through multimodal adaptation.

Relevance: 6 Novelty: 6


Efficiency, Compression, and Large-Scale Training (45)

1. LegoLM: Structured Weight Sharing for Large Language Models

ArXiv ID: 2608.08652

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Joseph Bingham

Abstract: We present \LegoLM{}, a structured weight-sharing compression framework for large language models grounded in a systematic study of why global weight sharing fails and how to fix it. We identify two distinct failure modes. Distributional mismatch: for vector blocks of dimension d <= 2, transformer layers with heterogeneous weight scales impose a scale-mismatch penalty that grows linearly with d and cannot be resolved by increasing K, producing perplexity in the millions.Outlier dominance: for scalar blocks, a fraction ~1/K of weights lies beyond the outermost Lloyd-Max decision threshold and cannot be represented by any centroid; their misrepresentation accumulates across layers, causing catastrophic quality loss. \LegoLM{} resolves both failure modes via three data-free adaptations: 1 scalar-block encoding to eliminate the $d$-linear mismatch component, 2 percentile-selective replacement that identifies and preserves outlier weights verbatim, and 3 boundary-layer protection for the first and last transformer blocks. Across GPT-2 small (124M) and Mistral-7B, \LegoLM{} achieves +0.03% PPL degradation at 4.41X compression on Mistral-7B - outperforming PTQ-8bit in both quality and compression ratio - and -0.02% at 2.67X. Downstream evaluation on LAMBADA and HellaSwag confirms that \LegoLM{} at K=64, p=99% preserves accuracy within noise at 5.12 X compression, exceeding PTQ-8bit's compression ratio while matching its accuracy. We further discover that outlier dominance grows with model scale: full replacement at K=128 degrades GPT-2 small by only +23% but catastrophically degrades Mistral-7B by +1,134,279%, while selective replacement at p=99% rescues both models to under +15%. A controlled ablation confirms that selective replacement is the dominant mechanism: adding it to per-layer K-means also yields near-lossless quality, matching \LegoLM{} within 0.02%.

Comment: Compresses LLM weights through scalar block sharing, explicit outlier preservation, and boundary-layer protection.

Topic Match: Structured weight sharing and outlier-aware encoding directly target large-model compression.

Relevance: 9 Novelty: 8


2. Every Cache Entry Earns Its Place: Global Allocation of Resolution and Coverage for KV Cache Compression

ArXiv ID: 2608.07001

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Haolin Tian, Yuzhe Liu, Tonghan Wang

Abstract: As large language models (LLMs) process increasingly long contexts, KV cache storage and repeated access have become a major bottleneck. Existing KV cache compression methods rely on predefined, fixed compression rules and are typically developed around either token eviction or merging. As a result, cache resources can neither flow freely across layers, heads, and context slots, nor be jointly allocated to balance local resolution and information coverage. Therefore, we propose GraceKV, a global approach for the allocation of resolution and coverage in KV cache compression, and formulate the compression process as a global resource allocation problem under a fixed cache budget. GraceKV treats each layer-KV head-slot combination as an atomic unit and builds a prototype tree. Leaf nodes correspond to token-level KV entries, while each internal node uses a single prototype to compress the KV space covered by its children. A set of non-overlapping nodes in the tree forms the representation of an atomic unit. Adding the root of a new tree expands information coverage, whereas splitting a selected node improves local resolution. All candidate actions compete globally for a shared cache budget. Finally, the nodes retained across all trees form the compressed KV cache. This process adaptively determines the allocation of cache resources among atomic units globally and the balance between resolution and coverage. GraceKV requires no additional training, and the entire compression and inference process is performed on the GPU. Systematic experiments across diverse long-context tasks and compression ratios show that GraceKV ranks first in 24 of 32 settings and remains robust up to 128-fold compression. These results validate the effectiveness of global budget allocation in coordinating information coverage and local resolution.

Comment: Globally allocates KV-cache budget across layers, heads, slots, information coverage, and local resolution.

Topic Match: The global resource-allocation mechanism is directly centered on KV-cache compression.

Relevance: 9 Novelty: 8


3. Statistically-Lossless Quantization of Large Language Models

ArXiv ID: 2605.02404

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Michael Helcig, Eldar Kurtic, Dan Alistarh

Abstract: Model quantization has become essential for efficient large language model deployment, yet existing approaches present clear trade-offs: methods such as GPTQ and AWQ achieve practical compression but are lossy, while lossless techniques preserve fidelity but lack inference acceleration. This paper explores the middle ground of statistically-lossless compression, examining three complementary aspects of what losslessness means for quantized LLMs. First, task-lossless compression preserves zero-shot benchmark accuracy within natural sampling variance and is achievable at aggressive bitwidths. Second, we formalize the stricter notion of distribution-lossless compression, requiring the quantized model's next-token distribution to be practically indistinguishable from the original, and propose the Expected Acceptance Rate (EAR), the maximum token-agreement probability under optimal coupling, as a directly interpretable fidelity metric. For example, EAR >= 0.99 means 99% agreement. Third, we prove a gamma-squared variance law showing that symmetric quantization inflates noise variance by gamma^2 relative to asymmetric quantization, making asymmetric quantization a prerequisite for distribution-lossless fidelity but not for task-level preservation. Through SLQ, a layer-wise non-uniform method with asymmetric quantization and wide bitwidth search, we obtain task-lossless compression at well below 4 bits per parameter, as low as 3.3 bits depending on the model, distribution-lossless compression at 5-6 bits per parameter on average, and inference speedups of 1.7-3.7x compared to FP16 using optimized kernels. Source code is available at https://github.com/IST-DASLab/SLQ.

Comment: Introduces asymmetric, non-uniform LLM quantization with formal distribution-fidelity criteria and optimized low-bit kernels.

Topic Match: Quantization that preserves task or token-distribution behavior while accelerating inference is directly within compression efficiency.

Relevance: 9 Novelty: 8


4. Autonomy-of-Heads: Data-Free Sparse Attention from Frozen Query-Key Geometry

ArXiv ID: 2608.06849

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Yehan Yang, Junyuan Shang, Yang Li, Guanqun Zhao, Shuohuan Wang, Dianhai Yu

Abstract: Long-context LLM inference is bottlenecked by quadratic attention computation and growing KV-cache costs. Existing sparse attention and KV-compression methods typically decide which tokens or heads to preserve from runtime attention scores, observation windows, calibration prompts, or learned gates, making head diagnosis input-dependent and costly to deploy. We propose Autonomy-of-Heads (AoH), a data-free method that identifies retrieval and streaming heads from the spectral geometry of query-key projections. AoH defines the kernel attention operator $M_h = W_K^{h\top}W_Q^h$ and uses its effective-rank as a weight-space measure of head function: concentrated spectra indicate a small number of dominant query-key matching directions and are associated with retrieval heads, whereas diffuse spectra indicate the absence of a dominant global matching direction and are associated with streaming heads. We further derive an efficient $d_\text{head}$-dimensional computation that avoids constructing the full $d_\text{model}\times d_\text{model}$ matrix. We conducted extensive experiments across models demonstrating that at 50\% sparsity, AoH retains 96.5\% of Full Attention performance on average while reducing prefill and decode latency by up to 41.4\% and 66.0\%, respectively, and KV-cache memory by 50.0\% at 256K tokens.

Comment: Derives data-free sparse-attention head roles from frozen query-key spectral geometry.

Topic Match: Its primary result reduces attention and KV-cache costs through a new weight-space sparsification mechanism.

Relevance: 9 Novelty: 8


5. Shape Mutating Expert Compression:LorExperts and BTExperts

ArXiv ID: 2608.07814

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: MoE Training

Authors: Inesh Chakrabarti, Sourjya Roy, Bowen Bao, Thiago Crepaldi, Spandan Tiwari, Ashish Sirasao

Abstract: Mixture-of-Experts (MoE) language models deliver high capacity at low per-token compute, but deploying them cheaply requires compressing their many expert weight matrices. Expert pruning (e.g., REAP) and merging reduce cost but sacrifice accuracy and require retraining the router; low-rank delta decomposition of experts (e.g., D^2-MoE) preserves all experts and the router, but degrades sharply as the expert count grows because a single shared component cannot approximate many near-orthogonal experts. Because MoE expert weights are near-orthogonal, a single shared component (as in prior delta decomposition) scales poorly with the expert count; we show that experts nonetheless organize into functional co-activation communities that are decoupled from weight similarity. Building on this, we introduce LorExperts, a router-preserving compression method that clusters experts, keeps one full-precision dominant per cluster, and represents the remaining members as low-rank corrections to their local dominant. LorExperts retains all experts and the original router (no router retraining). At ~50% expert compression on Qwen3-30B-A3B and Gemma-4-26B-A4B, LorExperts preserves downstream accuracy and perplexity better than the baselines on most of the tasks; the margin over D^2-MoE grows with expert count E. We further give a reconstruction fine-tuning procedure for LorExperts, and BTExperts, a tree organization of dominants and corrections that enables inference-time amortization of shared computation.

Comment: Compresses MoE experts as cluster-local low-rank corrections while retaining every expert and the original router.

Topic Match: Router-preserving expert compression is the central mechanism, with MoE expert structure providing the design basis.

Relevance: 9 Novelty: 8


6. RotaryQuant: Fitting 120B MoE Models on Consumer Hardware via Fused Compressed-Space Attention

ArXiv ID: 2608.08081

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: MoE Training

Authors: Anthony. Lui, Mohamed. Elsaied, N. P. Savani

Abstract: Large mixture-of-experts (MoE) language models with 26--120 billion parameters exceed the memory capacity of consumer devices through three simultaneous pressures: resident weight matrices, key-value (KV) cache state that grows linearly with context, and dozens of expert sublayers that must be paged on demand. We present RotaryQuant, a three-axis compression system that addresses all three. Mixed-precision weight quantization assigns bit-widths by architectural role: 4-bit for dense layers, 2-bit for routed experts, and 8-bit for the shared expert whose high activation kurtosis resists aggressive compression. LRU expert offloading pages non-resident experts to disk under genuine memory pressure. The novel axis is IsoQuant, a KV cache compression method that applies a Walsh--Hadamard transform followed by block-diagonal SO(4) rotations to isotropize activation distributions before 3-bit scalar quantization, requiring $O(d \log d)$ operations and 256 stored parameters per head versus $O(d^2)$ and 16{,}384 for dense rotation methods. A fused four-kernel Metal GPU pipeline performs attention directly on packed 3-bit tensors without materializing full-precision KV state---a different execution model, not just a quantization scheme. The combined system fits Gemma 4-26B-A4B and Qwen3-30B-A3B within a 16\,GB budget and Nemotron-H 120B within 32\,GB, running interactively at 9--19 tok/s with near-zero perplexity degradation ($Δ$PPL $\leq +0.0012$) and 100\% retrieval accuracy at 32K context.

Comment: Jointly compresses MoE weights, routed experts, and KV state while computing attention directly on packed 3-bit caches.

Topic Match: The central advance is multi-axis quantization and compressed-space execution that sharply reduces MoE inference memory.

Relevance: 9 Novelty: 8


7. Quantization Degradation in Large Language Models: A Signal-Noise Perspective

ArXiv ID: 2608.08188

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Chenxi Zhou, Pengfei Cao, Jinyu Ye, Bohan Yu, Haida Yu, Jiang Li, Jun Zhao, Kang Liu

Abstract: Post-training quantization reduces the deployment cost of large language models, yet how severely a quantized model degrades is not determined by bit-width alone. We systematically study weight-only post-training quantization across bit-widths, quantization methods, model scales and downstream tasks on multiple model families. We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation, and at 3-bit, degradation becomes apparent but varies markedly with task type, quantization method and model scale. To explain this variability, we use the signal-to-noise ratio (SNR) to measure how strongly quantization perturbs full-precision representations. We trace degradation back to two linked processes: how quantization errors arise within individual modules, and how they accumulate across layers. First, a source SNR decomposition shows that newly introduced errors depend on three factors: the magnitude of the weight error, the strength of the task-specific signal, and how strongly the quantization error aligns with task-specific activations. Different factors affect these components in distinct ways. Second, a cross-layer propagation analysis shows that these errors can be attenuated, preserved, or amplified as they pass across layers, and that larger models benefit from weaker error amplification. Together, these results establish that quantization degradation is governed by how errors are introduced at the source and how they accumulate across the network.

Comment: Decomposes quantization noise at its module source and tracks its attenuation or amplification across transformer layers.

Topic Match: The paper directly explains when weight quantization degrades LLMs, with cross-layer dynamics supporting the compression analysis.

Relevance: 9 Novelty: 8


8. Attn-QAT: 4-Bit Attention With Quantization-Aware Training

ArXiv ID: 2603.00040

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Large-Scale Training Systems and Efficiency, Architecture and Training Dynamics

Authors: Peiyuan Zhang, Matthew Noto, Wenxuan Tan, Chengquan Jiang, Will Lin, Wei Zhou, Hao Zhang

Abstract: Achieving reliable 4-bit attention is a prerequisite for end-to-end FP4 computation on emerging FP4-capable GPUs, yet attention remains the main obstacle due to FP4's tiny dynamic range and attention's heavy-tailed activations. This paper presents the first systematic study of 4-bit quantization-aware training (QAT) for attention. We find ``drop-in'' QAT -- naively combining an FP4 forward pass with high-precision Flash Attention (FA)-style backward pass -- leads to training instability. We identify two key principles for stable FP4 attention: (1) matching low-precision recomputation of attention scores in the backward pass and (2) resolving implicit precision assumptions in FA's gradient calculation. Based on these insights, we propose Attn-QAT and implement fused Triton kernels for training and FP4 inference kernels. Across diffusion and language models, Attn-QAT recovers the quality drop from FP4 attention without explicit outlier-mitigation heuristics used in prior FP4 attention, and delivers up to a 1.5x speedup on an RTX 5090 over SageAttention3 and up to a 1.74x speedup over FA4 on a GB300. Video demos can be found at https://drive.google.com/drive/folders/190F6xbBDUF2kGQYIcXBt3ehSYij5jlim?usp=sharing. Code can be found at https://haoailab.com/FastVideo/training/attn_qat/.

Comment: First systematic FP4 quantization-aware training for attention: matches low-precision score recomputation in the backward pass and fixes FlashAttention's implicit precision assumptions, with fused Triton kernels.

Topic Match: A new quantization mechanism plus training kernels that change what an FP4 training run costs and whether it stays stable — squarely efficiency and compression.

Relevance: 9 Novelty: 8


9. SPECTRA: Pushing the KV Cache Beyond the 2-Bit Cliff via Spectral Transform Coding

ArXiv ID: 2608.07915

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Jiamu Zhang, Liang Wu, Kelly Wan, Hanjie Chen, Liangjie Hong

Abstract: Large language models (LLMs) increasingly read long inputs in the agentic era, from whole documents and codebases to conversations across many turns. Their inference memory is then dominated by the key-value (KV) cache, the stored attention keys and values of every token the model has read and generated. Because the cache grows with context length and is re-read in full at every generated token, a longer context means more GPU memory. To reduce this cost, most existing methods compress the KV cache by lowering every stored value to the same low precision, a technique known as quantization. They can push this to nearly two bits per value, but rarely further, because quality drops sharply at this 2-bit cliff: four levels are too few for the cache's outlier-heavy values, where a few large entries consume the levels and collapse the rest into noise. A natural remedy is to spend more bits on the channels (feature dimensions) that matter and fewer on the rest, but the raw cache offers no handle: its channels are strongly correlated, so none stands out as more important. Our analysis shows that this handle appears once the cache is rotated into a coordinate system computed from its own statistics, removing these correlations. There, a small fraction of channels carries almost all the information, and spending the budget on those few is far more accurate than spreading it evenly. Guided by this analysis, we develop SPECTRA, a training-free, drop-in codec that re-encodes the cache into this coordinate system and concentrates the bit budget on the channels that carry the signal. On Llama-3.1-8B and Qwen2.5-7B over long-context benchmarks, SPECTRA is near-lossless at 4x compression, competitive at 8x where uniform quantization has collapsed, and reaches up to 12x, pushing usable compression past the 2-bit cliff so the same GPU holds longer contexts and larger batches.

Comment: A statistics-derived spectral transform concentrates KV-cache quantization bits on high-signal channels, enabling compression beyond uniform two-bit quantization.

Topic Match: The core contribution is a KV-cache compression mechanism that directly reduces long-context LLM memory requirements.

Relevance: 9 Novelty: 8


10. KronQ: LLM Quantization via Kronecker-Factored Hessian

ArXiv ID: 2607.07964

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Donghyun Lee, Yuhang Li, Ruokai Yin, Priyadarshini Panda

Abstract: Post-training quantization (PTQ) is a widely adopted technique for compressing large language models (LLMs) without retraining. Most existing second-order PTQ methods, including GPTQ, construct quantization objectives from input activation statistics, effectively assuming that all output channels contribute equally to the layer-wise reconstruction objective. We propose KronQ, a PTQ framework that challenges this assumption by introducing the gradient covariance into the quantization pipeline. Under the Kronecker-factored Hessian approximation, the quantization loss depends jointly on both the activation and gradient covariances, and KronQ exploits this at two complementary levels. (1) KronQ introduces bidirectional incoherence processing, extending the existing input-side random rotation to the output dimension using the gradient covariance, reducing weight magnitude variance across both input and output dimensions. (2) KronQ derives a new sensitivity metric for inter-layer mixed-precision allocation, driven by the gradient and activation Hessian traces. Notably, in the case of 2-bit weight-only quantization on LLaMA-3-70B, while GPTQ and GPTAQ diverge or produce degenerate quantizations (>2000 perplexity on WikiText-2), \KronQ{} achieves 7.93 perplexity.

Comment: Uses activation and gradient covariances in a Kronecker-factored quantization objective to guide rotations and mixed-precision allocation.

Topic Match: The primary contribution is second-order LLM weight quantization that accounts for output sensitivity and enables effective two-bit compression.

Relevance: 9 Novelty: 8


11. ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantization

ArXiv ID: 2608.07019

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Yongge Ma, Guoan Wang, Feiyu Wang, Yaoming Li, Qian Zhang, Zihan Yan, Yinjun Han, Tong Yang

Abstract: Post-training quantization (PTQ) is widely used to reduce the memory and computational cost of large language models. Existing PTQ methods typically obtain an initial quantized model through heuristic rules or greedy optimization, and once quantization is completed the resulting integer assignments are usually treated as final. This observation motivates a complementary optimization stage within PTQ that keeps quantized weights improvable after an executable quantized model has been produced, while preserving the quantized format. We introduce ReQuant, a backpropagation-free fixed-grid refinement procedure for this stage. Agnostic to the PTQ initializer, ReQuant takes an existing quantized model as a feasible starting point and iteratively revisits its discrete weight assignments on the fixed quantization grid. Accepted updates strictly reduce the mean squared reconstruction error and remain on the original grid. In this way, ReQuant turns the initially fixed PTQ output into an iteratively optimizable discrete solution and serves as a plug-and-play post-processing stage for existing PTQ pipelines. Experiments across diverse model families, bit-widths, and downstream tasks show that ReQuant consistently improves quantized models from heterogeneous PTQ initializers, with especially large gains on simple initializers and lower bit-widths. Notably, ReQuant can refine a simple round-to-nearest initialization across multiple sweeps until it approaches or surpasses GPTAQ under the same quantization format. These results establish ReQuant as a practical complementary stage for further improving existing PTQ pipelines.

Comment: Refines PTQ integer assignments through backpropagation-free optimization on the original quantization grid.

Topic Match: The fixed-grid refinement procedure is directly centered on improving low-bit model compression.

Relevance: 9 Novelty: 7


12. CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit Weights

ArXiv ID: 2608.06763

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Xuetian Gao

Abstract: Weight quantization for large-language-model inference must balance adaptive reconstruction levels with representations regular enough for efficient GPU execution. Uniform integers constrain each group to a linear grid. Low-bit floating-point formats use a fixed exponent-mantissa structure, while learned codebooks gain flexibility at the cost of irregular decoding and additional metadata. We introduce CubicQuant, a parametric non-uniform scalar format that preserves a dense integer code stream while adapting reconstruction levels within each weight group. A monotonic cubic curve, specified by two shape parameters and one scale, maps uniformly spaced magnitude codes to non-uniform levels. The family spans 1-8-bit weight payloads, contains symmetric uniform integer quantization as an exact special case, and has effective width B + 64/G bits per weight for payload width B and group size G. We derive population distortion under Uniform, Gaussian, and Laplace distributions, formulate continuous and Dynamic-A8-carrier-aware fitting objectives, and describe direct packed-weight GPU execution. For finite groups of G=128 with 15,360 samples per distribution, W4 CubicQuant reduced reconstruction RMSE relative to optimally clipped four-bit uniform integer quantization by 3.90% on Uniform, 13.49% on Gaussian, and 28.14% on Laplace samples. Relative to the best enumerated four-bit finite floating-point format, the reductions were 3.90%, 9.44%, and 6.27%. Preliminary H200 kernel measurements show a workload-dependent crossover: model-dtype execution is faster for narrow GEMV, while Dynamic A8 becomes favorable as row count grows. The results establish the format's representational promise and direct executability; downstream model quality and cross-device end-to-end performance remain open evaluation questions.

Comment: Introduces a parametric non-uniform quantization codebook that retains a regular packed integer stream for GPU execution.

Topic Match: The new low-bit weight format jointly addresses reconstruction quality and executable kernel structure.

Relevance: 9 Novelty: 7


13. Length-MAX Tokenizer for Language Models

ArXiv ID: 2511.20849

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Large-Scale Training Systems and Efficiency

Authors: Dong Dong, Weijie Su

Abstract: We introduce a new tokenizer for language models that minimizes the average tokens per character, thereby reducing the number of tokens needed to represent text during training and to generate text during inference. Our method, which we refer to as the Length-MAX tokenizer, obtains its vocabulary by casting a length-weighted objective maximization as a graph partitioning problem and developing a greedy approximation algorithm. On FineWeb and diverse domains, it yields 14--18\% fewer tokens than Byte Pair Encoding (BPE) across vocabulary sizes from 10K to 50K, and the reduction is 13.0\% when the size is 64K. Training GPT-2 models at 124M, 355M, and 1.3B parameters from scratch with five runs each shows 18.5\%, 17.2\%, and 18.5\% fewer steps, respectively, to reach a fixed validation loss, and 13.7\%, 12.7\%, and 13.7\% lower inference latency, together with a 16\% throughput gain at 124M, while consistently improving on downstream tasks including reducing LAMBADA perplexity by 11.7\% and enhancing HellaSwag accuracy by 4.3\%. Moreover, the Length-MAX tokenizer achieves 99.62\% vocabulary coverage and the out-of-vocabulary rate remains low at 0.12\% on test sets. These results demonstrate that optimizing for average token length, rather than frequency alone, offers an effective approach to more efficient language modeling without sacrificing -- and often improving -- downstream performance. The tokenizer is compatible with production systems and reduces embedding and KV-cache memory by 18\% at inference.

Comment: Optimizes vocabulary for average token length, reducing both pretraining steps and inference sequence length.

Topic Match: The tokenizer materially changes training and inference cost by reducing the number of processed tokens.

Relevance: 9 Novelty: 7


14. LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment

ArXiv ID: 2608.03020

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Large-Scale Training Systems and Efficiency

Authors: Linhan Xia, Rui Liu, Zhaofeng Zhang, Yihao Wang, Binrui Shen, Shengxin Zhu

Abstract: Parameter-efficient post-training reduces the number of trainable parameters, but still requires repeated end-to-end backpropagation through the frozen backbone. Every adaptation step therefore needs backward-capable hardware and must store or recompute activations. We ask whether this repeated backward chain can be replaced by a one-time calibration. We introduce Local Credit Assignment (LoCA), a two-stage method for small-shift adaptation. One probe backward pass fits a low-rank map at each transformer block from the final prediction error to a local hidden-state correction. LoCA then reuses these maps to form blockwise regression targets from forward activations and fits low-rank adapters with closed-form ridge solves. No further backbone backward pass is required. We evaluate LoCA on five discriminative benchmarks with Qwen2.5 models from 0.5B to 14B. In 16 of 25 reported task--scale comparisons, LoCA yields lower evaluation cross-entropy than the corresponding LoRA run. Its measured full-run GPU peak, including calibration, is 26--29\% lower than LoRA's. After calibration, its CPU steady-state memory is 36--52\% lower and its per-pass time is 43--48\% lower. A shared scale-normalized candidate set is reused across all tested Qwen2.5 sizes and on SmolLM2-1.7B. LoCA thus amortizes global credit assignment into one calibration and enables later forward-only tuning when repeated backpropagation is impractical. The code associated with this paper is available \href{https://github.com/Xia12121/LoCA}{here}.

Comment: Replaces repeated backbone backpropagation with one calibration pass and forward-only local adapter fitting.

Topic Match: The method materially reduces adaptation memory and computation by changing the tuning algorithm.

Relevance: 8 Novelty: 8


15. Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation

ArXiv ID: 2606.09864

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Bruce Changlong Xu, Adarsh Kumarappan, Mu Zhou

Abstract: Key-value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely focus on measuring perplexity and accuracy without assessing the safety impact. In this study, we explore alignment preservation under KV cache quantization. Across eleven instruction-tuned models (3.8B-72B) and five benchmarks (1,894 prompts), we find that low-bit quantization can silently destroy safety alignment: Mistral-7B loses 15.2% of its refusals at only 1.03x perplexity, and no universal safe bit-width exists, with sharp model-specific phase transitions invisible to standard metrics. We identify that the root cause is geometric: safety features occupy a low-dimensional activation subspace 10^2-10^3x more vulnerable to quantization noise than the full representation space perplexity averages over. Inspired by this observation, we propose Per-Channel Reduction (PCR), a diagnostic that classifies each model into one of three mechanistic failure modes: outlier-crushes-safety, where safety lives in non-outlier channels collaterally damaged by outlier-driven scale factors; outlier-as-safety, where safety overlaps outlier channels and finer granularity cannot rescue it; and multi-layer dilution, where safety is distributed across many layers and per-layer fixes fail. PCR predicts the correct mitigation direction on all nine primary models and one held-out model from an independent family using 20 calibration prompts. PCR generalizes across unseen prompts, models, and production quantizers, including KIVI with up to 97.2% recovery, succeeding where attention-based allocation methods fail. The resulting training-free protocol, requiring approximately 35 GPU-minutes, recovers up to 97% of lost alignment at minimal memory overhead, addressing vulnerabilities confirmed in production vLLM serving with FP8 KV cache on NVIDIA GPUs.

Comment: Diagnoses safety-sensitive KV quantization failure modes through channel geometry and selects model-specific mitigations.

Topic Match: KV-cache quantization behavior and its mechanistic failure analysis are the central technical contributions.

Relevance: 8 Novelty: 8


16. Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Models

ArXiv ID: 2608.07890

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Ali Janati, Kaoutar El Maghraoui, Xinyi Luo, Wenyuan Shen, Owen Zou, Yankai Mao

Abstract: Mixture-of-Experts (MoE) models decouple total parameters from per-token compute, but deployment still requires storing every expert. Recent theory shows that pruning experts with the smallest router-norm changes during fine-tuning can preserve accuracy, but assumes full fine-tuning. We test whether lightweight adaptation can recover this signal. We briefly fine-tune with a parameter-efficient adapter, rank experts by the induced $\ell_2$ router change, and prune the least-changed experts in one shot. On Mixtral-8$\times$7B-Instruct (44.83% MMLU-Pro), router-only LoRA trains 0.002% of parameters and outperforms all-module LoRA at matched rank with half the experts removed (27.54% vs. 24.42%); signal quality declines as adaptation spreads to attention and expert weights. Accuracy improves monotonically with LoRA rank, reaching 28.76%. IA3, which leaves router weights frozen, matches direct router adaptation, whereas unconstrained additive adapters degrade the signal. Router-guided MMLU-Pro accuracy decays quasi-linearly rather than collapsing, remains nearly 1.8 times that of magnitude-based or random pruning at maximal compression, and reduces memory by 49% and per-token latency by 37%. At 25% compression, retention is competitive with methods using full activation statistics. The criterion also transfers to Qwen1.5-MoE fine-tuned for mathematics, retaining 49.7% mean accuracy over eleven benchmarks with half the experts removed while random pruning falls to single digits. Router sensitivity under lightweight fine-tuning therefore makes provably motivated expert pruning practical at scale.

Comment: Uses router sensitivity under lightweight adaptation to prune experts and reduce MoE memory and latency.

Topic Match: The central contribution is practical expert compression using a cheaper pruning signal; it extends an existing router-sensitivity criterion rather than introducing a new routing mechanism.

Relevance: 9 Novelty: 6


17. RippleKV: Cross-Layer KV Cache Allocation via Perturbation Propagation

ArXiv ID: 2608.08684

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Dongjie Xu, Kai Qian, Julius, Weijie Shi, Yuxuan Sun, Minghua Tang, Fenglei Jin, Hanchi Dong, Jiajie Xu

Abstract: Long-context LLM inference is bottlenecked by KV cache memory, yet distributing a limited cache budget across layers remains challenging. Existing methods rely on proxies such as layer depth, attention statistics, or representation change. These proxies do not measure how perturbations at each layer propagate to the output and may therefore cause sensitive layers to be underallocated while tolerant layers are overallocated. To address this issue, we propose RippleKV, which allocates cache across layers by estimating how perturbations to each layer's value cache affect the final predictive distribution. RippleKV independently injects norm-adaptive perturbations into each layer's value cache and measures the induced KL divergence at the model output over a small calibration set. Averaging these responses yields a sensitivity profile specific to the model that need not vary monotonically with depth. RippleKV then converts the sensitivity profile into layer budget multipliers by normalizing the sensitivity scores and applying an exponential mapping. A ratio parameter controls the allocation disparity between sensitive and tolerant layers, while a final normalization preserves the KV cache budget. Experiments on LongBench demonstrate that RippleKV achieves the highest average performance among the evaluated KV cache compression methods under matched cache budgets.

Comment: Allocates per-layer KV-cache budgets using output-distribution sensitivity to calibrated value-cache perturbations.

Topic Match: Sensitivity-based cache allocation directly addresses LLM memory efficiency under a fixed compression budget.

Relevance: 9 Novelty: 6


18. SimSD: Simple Speculative Decoding in Diffusion Language Models

ArXiv ID: 2606.02544

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Junxia Cui, Haotian Ye, Runchu Tian, Hongcan Guo, Jinya Jiang, Haoru Li, Chaojie Ren, Yiming Huang, Kaijie Zhu, Zhongkai Yu, Kun Zhou, Jingbo Shang

Abstract: Diffusion large language models (dLLMs) have recently emerged as a promising alternative to autoregressive (AR) LLMs, offering faster inference through parallel or blockwise decoding. However, their masked language modeling formulation remains incompatible with standard token-level speculative decoding, one of the most effective acceleration techniques for AR models. In AR decoding, the causal mask preserves temporally valid token-level contexts, enabling a target model to verify multiple drafted tokens in a single forward pass. In contrast, dLLMs rely on mask tokens and bidirectional attention, causing the effective context to change across denoising steps and preventing direct token-level speculative verification. To bridge this gap, we propose a simple but effective speculative decoding algorithm for diffusion language models, named SimSD, which mainly adopts a plug-and-play masking strategy that equips dLLMs with temporally valid token-level contexts for speculative decoding. Our method explicitly introduces reference tokens from draft-model predictions and designs an attention mask that regulates their interaction with current-step tokens, allowing dLLMs to compute valid logits for drafted tokens in a single forward pass. This restores the key verification ability provided by causal masking in AR models while preserving the parallel decoding advantages of dLLMs. The proposed method is training-free and can be flexibly integrated with other acceleration techniques such as KV cache and blockwise decoding. Experiments on SDAR-family dLLMs across four benchmarks show that our method achieves up to 7.46x higher decoding throughput while maintaining and even improving average generation quality.

Comment: Enables speculative verification in diffusion LMs with reference tokens and a custom attention mask.

Topic Match: Decoding acceleration is the primary result, enabled by a new attention-context construction for diffusion language models.

Relevance: 8 Novelty: 7


19. DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM Inference

ArXiv ID: 2608.08878

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Asaad Althoubi

Abstract: Transformer-based large language models (LLMs) achieve strong performance across many tasks, but their Key-Value (KV) cache grows linearly with sequence length, creating a severe memory bottleneck for long-context inference. Existing heuristic eviction methods (e.g., H$_2$O and SnapKV) rely on static attention or positional signals that often fail to capture a token's future predictive influence. We propose DistillCache, a reinforcement learning framework that formulates KV-cache eviction as a sequential decision problem. DistillCache learns a lightweight policy network using rich internal model signals (attention statistics, value norms, entropy, and position) and trains it with REINFORCE via a per-step KL-divergence reward to preserve the full-cache output distribution. On a 7B-parameter instruction-tuned Transformer (Mistral-7B-Instruct-v0.3), DistillCache retains 94.2% of full-cache accuracy on LongBench at a 25% cache budget, outperforming both strong heuristic baselines (H$_2$O, SnapKV) by up to 2.7 absolute points and, under our re-implementations, concurrent RL-based methods (ForesightKV, RLKV) by up to 1.4 points on long-context tasks. On reasoning benchmarks, DistillCache is competitive with the best concurrent method and surpasses it under aggressive compression. It also delivers up to 2.1x full-cache throughput while maintaining competitive practical efficiency. These results highlight the effectiveness of learned, distribution-aware policies for memory-efficient long-context LLM inference.

Comment: Learns KV eviction policies from per-step KL rewards to preserve full-cache distributions.

Topic Match: Adaptive KV-cache compression directly reduces long-context memory and inference cost.

Relevance: 8 Novelty: 7


20. Tied Trit-Planes: Constraining PTQTP to a Uniform Nine-Level Quantizer, with a Persistent Folded Format for Disk-Streamed Mixture-of-Experts Serving

ArXiv ID: 2608.08910

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Matteo Grella

Abstract: PTQTP decomposes LLM weight matrices into two ternary (trit) planes with two free per-group scales. Tying the scales to a fixed ratio of three collapses the decomposition into a single uniform nine-level quantizer, a known balanced-ternary identity. To our knowledge, at the time of writing, this work is the first to impose that identity as a constraint inside PTQTP's solver. The two trit planes then fold losslessly into one 4-bit code plane that we make the persistent serving representation: disk bytes, expert-cache bytes, and kernel input are the same 4.0625-bits/weight blocks, consumed in one integer dot pass. For this conjunction (ratio-3 nine-level code, CPU-SIMD kernels, SSD expert streaming, identical persistent bytes) we likewise found no precedent. We apply this to the routed experts of DeepSeek-V4-Flash-0731, a 284B-A13B mixture-of-experts model, quantizing in one shot from the released MXFP4 expert weights and streaming experts from SSD on a 64 GB laptop. Against a 4.5-bit Q4_K baseline, measured one process per fixture with an expert-lossless anchor arm as reference control, the tied model matches the official serving API on 5/5 fixtures at step 0 (Q4_K: 4/5) and 12/14 captured continuation steps (11/14), scores 86 vs. 84 on a 100-item MMLU subset, decodes 6.7% faster in decode phase, and ships 9% smaller files: no detected fidelity difference at these small evaluation sizes, and every fixture-level difference between the arms traces to a single measured near-tie cell. The tied fit nevertheless shows higher weight-reconstruction error and worse perplexity, a measured dissociation between proxy metrics and reference fidelity. A cumulative trunk-ternarization ladder and bitwise-pinned aarch64/x86-64 kernels complete the report. All code, formats, and evaluation artifacts are open source in the fucina inference stack.

Comment: Constrains routed-expert weights to a foldable nine-level quantizer used directly for SSD-streamed serving.

Topic Match: The central contribution is a compact persistent quantization format and matching execution path for a large MoE model.

Relevance: 8 Novelty: 7


21. Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference

ArXiv ID: 2604.26968

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Sanjeev Rao Ganjihal

Abstract: Key-value (KV) cache memory management is the primary bottleneck limiting throughput and cost-efficiency in large-scale GPU inference serving. Current systems suffer from three compounding inefficiencies: (1) the absence of unified KV cache sizing across all attention architectures--particularly multi-head latent attention (MLA), which is unsupported in general-purpose frameworks, resulting in up to 57x memory over-provisioning; (2) confinement of KV cache to a single memory tier (GPU HBM) despite the availability of a rich hierarchy spanning CPU DRAM, CXL-attached memory, NVMe via GPUDirect Storage, RDMA fabric, and parallel filesystems; and (3) reactive eviction policies that discard reusable state, forcing redundant recomputation. We present a unified system addressing all three. Our architecture-variant-aware sizing engine computes exact memory requirements per attention type; the resulting batch size gain reaches 7.4x for the one MLA model we evaluate (DeepSeek-V3), while the three GQA models see 1.0x, 1.0x, and 0.7x, so the GQA benefit is fleet-wide unified sizing rather than larger per-model batches. A six-tier memory hierarchy extends effective KV cache capacity from 40 GB to over 38 TB per node while maintaining sub-millisecond time-to-first-token (TTFT) for hot entries. A Bayesian reuse predictor with Beta conjugate priors over 16 (block-type, transition-type) pairs drives EMA-scored head-granular eviction and RoPE-aware prefetching. Component-level validation on trace replay using ShareGPT, LMSYS-Chat-1M, and agentic workloads demonstrates 70-84% cache hit rates. Analytical projections combining validated component behavior with published hardware specifications indicate TTFT reductions of 1.4x to 2.1x, throughput improvements of 1.7x to 2.9x, and 47% cost reduction relative to published baselines; these cluster-scale projections are analytical and carry no error bars.

Comment: Combines exact architecture-aware KV sizing with predictive allocation across six memory tiers.

Topic Match: Its unified KV-cache allocation and reuse mechanism directly targets large-model memory and inference cost.

Relevance: 8 Novelty: 7


22. IFCLoRA: Topology-Aware Rank Allocation for Parameter-Efficient Fine-Tuning

ArXiv ID: 2607.22251

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Wei Zhang, Xinwu Liu, Yihang Cheng

Abstract: Low-Rank Adaptation (LoRA) is a widely used approach to parameter-efficient fine-tuning (PEFT) of LLMs whose effectiveness depends on rank allocation. Existing adaptive LoRA methods derive ranks from local gradient, activation, or matrix statistics collected before or during fine-tuning; training-time variants add overhead, and local signals reveal little about each module's structural role in information propagation, giving weak global grounding for scarce-capacity allocation. We propose IFCLoRA, a topology-aware method for pre-fine-tuning rank allocation and adapter initialization. Using a small calibration set, IFCLoRA performs intervention tracing on the frozen model and constructs a sparse task-conditioned interaction graph over LoRA target modules. From this graph it extracts a global information-flow topology prior and fuses it with each node's local gradient sensitivity to form a topology-dominant Information-Flow Centrality (IFC) score, measuring participation in task-conditioned multi-hop propagation. The IFC scores then serve as module-level routing signals for one-shot discrete rank allocation under a rank-budget constraint. Reusing response vectors from tracing, IFCLoRA constructs a function-preserving flow-response subspace initialization, giving adapters task-relevant output subspaces. Across all settings, IFCLoRA achieves higher mean scores than standard LoRA with comparable fine-tuning time and peak memory; it requires a one-time offline calibration stage. On GSM8K, IFCLoRA attains the highest mean accuracy among compared PEFT methods on both base models, exceeding standard LoRA by 4.75 percentage points on LLaMA-3.1-8B. Resulting rank allocations are non-uniform and vary across tasks and base models, suggesting that task-conditioned global information-flow topology can serve as a useful structural prior for rank allocation in low-budget PEFT.

Comment: Allocates LoRA ranks from a task-conditioned global information-flow graph and local gradient sensitivity.

Topic Match: Topology-aware rank budgeting and initialization constitute a substantive parameter-efficient adaptation method.

Relevance: 8 Novelty: 7


23. Unified Static-Dynamic Pruning for Efficient LLM Inference

ArXiv ID: 2607.21985

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Jinhyeok Kim, Yejoon Lee, Jaeyoung Do

Abstract: The increasing deployment of large language models (LLMs) has magnified the computational and memory bottlenecks of autoregressive decoding, where low compute intensity and bandwidth-bound kernels dominate inference cost. Weight pruning offers a promising remedy, but existing methods remain confined to either static pruning (SP), which permanently removes redundant weights but lacks adaptivity, or dynamic pruning (DP), which adapts to input sparsity but introduces runtime irregularity. This paper presents SPDP, a unified sparse-inference framework that integrates unstructured SP with input-adaptive DP for efficient LLM inference on GPUs. SPDP co-designs a new Tiled-Column-wise Bitmap Compressed (Tiled-CBC) format and two complementary GPU kernels: (1) a CUDA-core spMspV kernel featuring Hybrid Activation-aware Dynamic Shared-Memory Bitmap Decoding (HAD-SMBD) for fine-grained, runtime activation skipping, and (2) a Tensor-Core SpMM kernel optimized for prefill computation. This joint format-kernel design harmonizes static and dynamic sparsity, maintaining bandwidth-efficient memory access and high compute intensity under both phases of LLM inference. Comprehensive evaluations on inference-optimized GPUs demonstrate that SPDP achieves 1.24x-1.37x average speedup (up to 2.51x) over state-of-the-art sparse frameworks such as SpInfer, while matching perplexity with up to 25% higher sparsity. SPDP advances the inference efficiency-quality Pareto frontier, showing that unified static-dynamic pruning can deliver substantial throughput and performance-per-watt improvements in large-scale LLM serving.

Comment: Co-designs a compressed sparse format and GPU kernels for combined static and input-adaptive pruning.

Topic Match: The contribution directly advances sparse LLM execution through joint compression-format and kernel design.

Relevance: 8 Novelty: 7


24. Understanding Calibration and Truncation Error Propagation in Training-Free Low-Rank Compression for LLMs

ArXiv ID: 2608.08506

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Mohanad Odema, Gabrielle De Micheli, Dayin Gou, Nilesh Malpeddi, Prathamesh Vaste, Jacob Song

Abstract: Training-free low-rank compression frameworks have been gaining prominence for LLM compression given their effectiveness in reducing model parameter count while maintaining task-level accuracy. However, existing SOTA frameworks share two key limitations: (1) residual errors in calibration data activations accumulate across layers during compression, causing misalignment between representations simulated at compression time and those experienced at inference; (2) the assumption that layer importance distribution is preserved post-compression does not hold. Together, these two effects introduce misalignment in the compression process in relation to the deployed model. We study these effects and propose a simple, training-free methodology compatible with existing frameworks to mitigate them, comprising: (1) Layer-by-Layer Compression with Calibration Correction; (2) Iterative Compression with Rank Allocation Correction. Implemented atop an existing SOTA decomposition framework, and evaluated on Llama and Qwen3 models across various benchmarks and compression rates, our approach demonstrates up to ~1-2.5 accuracy point improvements over per-weight and joint decomposition baselines on zero-shot tasks.

Comment: Corrects layerwise calibration drift and reallocates ranks iteratively during training-free low-rank compression.

Topic Match: The paper directly analyzes and mitigates error propagation in low-rank LLM compression.

Relevance: 8 Novelty: 7


25. Dropping the Anchor: Statistical Context Summarization for Distributed Systems via Pulsar Attention

ArXiv ID: 2607.20457

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Aryan Sood, Shantanu Acharya, Gaurav Kumar Nayak

Abstract: Inference with large language models (LLMs) on long sequences is computationally expensive due to the quadratic complexity of self-attention. Distributed blockwise methods such as Star Attention reduce this cost by sharding context across hosts, but rely on prepending a static, content-blind copy of the first block to every host. We propose Pulsar Attention, which replaces the static anchor with two lightweight, content-aware components: a small attention-sink prefix that stabilizes softmax, and compact cross-block summaries built via a Max-IDF heuristic that selects chunks containing globally rare tokens. This reduces the Phase 1 per-GPU FLOPs by up to 3.3x over Star Attention while retaining an identical KV cache footprint. On RULER with Llama-3.1-8B-Instruct, Pulsar Attention outperforms Star Attention at sequence lengths up to 128K tokens and remains competitive with dense attention across most tasks, with task-dependent absolute gains of up to 4.7% over the dense baseline.

Comment: Replaces replicated static context anchors with attention sinks and compact content-aware cross-block summaries.

Topic Match: Its primary contribution is an efficient distributed long-context attention mechanism with reduced computation.

Relevance: 8 Novelty: 7


26. CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents

ArXiv ID: 2608.07855

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Weizhong Huang, Jinchao Zhang, Xiawu Zheng

Abstract: Multi-turn Reasoning-and-Acting (ReAct) agents accumulate growing trajectories of reasoning, tool calls, and observations. Their key-value (KV) caches grow accordingly, increasing memory use and attention cost during model inference. Existing KV cache compression methods reduce these costs by evicting states with low attention scores. However, low attention in the current turn does not imply future irrelevance, as temporarily inactive information may become important later. Snapshot-based eviction methods therefore do not explicitly distinguish temporarily dormant information from information that appears to have completed its role. In this paper, we present CommitKV, which identifies KV lifecycles through commit transitions. Specifically, CommitKV first divides completed agent events into token pages and compares each eligible page's deletion effect before a tool-call commit and after the commit's returned observation has been incorporated. Based on these paired measurements, CommitKV distinguishes dormant pages from high-to-low completion candidates. It then applies a greedy joint test, accepting candidates for retirement only when their combined post-commit effect remains bounded. Finally, at a later compression checkpoint, accepted pages are excluded, a bounded set of pages awaiting post-commit measurement is protected, and the remaining KV states are retained within the cache budget using the same token indices for keys, values, and absolute positions. These mechanisms ensure that CommitKV can distinguish dormant information from information that has completed its observed role and can be safely removed. Experiments on various benchmarks show that CommitKV reduces agent memory use, accelerates end-to-end inference, and achieves higher accuracy than existing KV cache compression methods.

Comment: Prunes KV pages only after paired pre- and post-commit tests indicate that their agent-event role has completed.

Topic Match: Lifecycle-aware KV-cache compression is the primary mechanism and directly reduces memory and attention cost.

Relevance: 8 Novelty: 7


27. CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG

ArXiv ID: 2608.07458

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Gyuwan Kim, Cheoneum Park, Tao Yang

Abstract: Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoid processing long retrieved contexts for higher efficiency, while significant information redundancy and noise still remain in the coarse-grained chunks. This paper optimizes the Pareto frontier under low prefill latency constraints while maximizing accuracy by proposing CoinRAG (Contextualized Information Nugget KV Cache Reuse for Long-Context RAG). The name metaphorically reflects our core mechanism: much like assembling small tokens (or "coins") to accumulate a larger value, CoinRAG compositionally reuses offline-computed, fine-grained nugget caches to form a learned contextual representation efficiently in a more semantically relevant but compact manner. Specifically, instead of full-chunk encoding, CoinRAG identifies query-relevant semantic units within retrieved chunks through two-stage retrieval and seamlessly assembles their sliced KV representations with a chunk-level context. Extensive evaluations on LongBench multi-hop question answering tasks demonstrate that CoinRAG significantly reduces operational costs and outperforms the other baselines with a new Pareto frontier and an average 5.3% relative improvement in answer quality (F1) under a standard fast prefill latency budget.

Comment: Composes fine-grained, query-relevant KV-cache slices to reduce long-context prefill computation.

Topic Match: The contribution is a finer-grained cache-reuse mechanism that changes the prefill cost-quality tradeoff, making efficiency the primary fit despite its RAG-specific workload.

Relevance: 8 Novelty: 7


28. Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models

ArXiv ID: 2608.06901

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Minseok Kang, Hyunwoo Kim, Chanyoung Kim, Minwoo Kim, Jaekoo Lee, Dahuin Jung

Abstract: Vision-language models (VLMs) have achieved remarkable generalization across diverse multimodal tasks through large-scale pre-training, yet their rapidly increasing computational and memory requirements pose significant challenges for deployment in constrained environments. Existing pruning strategies often depend on task-specific criteria or LLM-oriented importance measures, making them unsuitable for task-agnostic pruning, where no task-specific samples are available at pruning time and the pruned model remains broadly applicable. We introduce a retraining-free VLM pruning framework called PORTA that derives a task- and modality-agnostic importance formulation based on activation variation, estimated from generic calibration data, which reliably captures feature-level representation utility across modalities. PORTA further incorporates an adaptive sparsity allocation mechanism that assigns layer-wise pruning ratios based on output feature variability, avoiding the limitations of uniform sparsity and reducing performance degradation at high compression levels. Extensive experiments across VLM architectures, such as CLIP, BLIP, and Qwen2-VL, demonstrate that PORTA achieves competitive downstream performance under high sparsity without requiring any retraining, supporting efficient VLM compression. Code is available at https://github.com/cau-hai-lab/PORTA.git.

Comment: Allocates VLM pruning sparsity from activation variation without task data or retraining.

Topic Match: Task-agnostic structural pruning directly targets model compute and memory requirements.

Relevance: 8 Novelty: 6


29. High-Layer Attention Pruning with Rescaling

ArXiv ID: 2507.01900

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Songtao Liu, Peng Liu

Abstract: Pruning is a highly effective approach for compressing large language models (LLMs), significantly reducing inference latency. However, conventional training-free structured pruning methods often employ a heuristic metric that indiscriminately removes some attention heads across all pruning layers, without considering their positions within the network architecture. In this work, we propose a novel pruning algorithm that strategically prunes attention heads in the model's higher layers. Since the removal of attention heads can alter the magnitude of token representations, we introduce an adaptive rescaling parameter that calibrates the representation scale post-pruning to counteract this effect. We conduct comprehensive experiments on a wide range of LLMs, including LLaMA3.1-8B, Mistral-7B-v0.3, Qwen2-7B, and Gemma2-9B. Our evaluation includes both generation and discriminative tasks across 27 datasets. The results consistently demonstrate that our method outperforms existing structured pruning methods. This improvement is particularly notable in generation tasks, where our approach significantly outperforms existing baselines. Code is available at https://github.com/SongtaoLiu0823/HARP.

Comment: Prunes attention heads preferentially in higher layers and rescales representations after removal.

Topic Match: The central contribution is a structured, training-free LLM pruning method.

Relevance: 8 Novelty: 6


30. Diving into Kronecker Adapters: Component Design Matters

ArXiv ID: 2602.01267

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Jiayu Bai, Danchen Yu, Zhenyu Liao, TianQi Hou, Feng Zhou, Robert C. Qiu, Zenan Ling

Abstract: Kronecker adapters have emerged as a promising approach for fine-tuning large-scale models, enabling high-rank updates through tunable component structures. However, existing work largely treats the component structure as a fixed or heuristic design choice, leaving the dimensions and number of Kronecker components underexplored. In this paper, we identify component structure as a key factor governing the capacity of Kronecker adapters. We perform a fine-grained analysis of both the dimensions and number of Kronecker components. In particular, we show that the alignment between Kronecker adapters and full fine-tuning depends on component configurations. Guided by these insights, we propose Component Designed Kronecker Adapters (CDKA). We further provide parameter-budget-aware configuration guidelines and a tailored training stabilization strategy for practical deployment. Experiments across various architectures and modalities demonstrate the effectiveness of CDKA. Code is available at https://github.com/rainstonee/CDKA.

Comment: Designs Kronecker-adapter component dimensions and counts under explicit parameter budgets.

Topic Match: Parameter-efficient low-rank adaptation and its stabilization are the paper's central contributions.

Relevance: 8 Novelty: 6


31. OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference

ArXiv ID: 2510.07651

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Yuzhe Gu, Xiyu Liang, Jiaojiao Zhao, Enmao Diao

Abstract: Large language models (LLMs) with extended context windows enable powerful applications but impose significant memory overhead, as caching all key-value (KV) states scales linearly with sequence length and batch size. Existing cache eviction methods address this by exploiting attention sparsity, yet they typically rank tokens heuristically using accumulated attention weights without considering their true impact on attention outputs. We propose Optimal Brain Cache (OBCache), a principled framework that formulates cache eviction as a layer-wise structured pruning problem. Building upon the Optimal Brain Damage (OBD) theory, OBCache quantifies token saliency by measuring the perturbation in attention outputs induced by pruning tokens, with closed-form scores derived for isolated keys, isolated values, and joint key-value pairs. Our scores account not only for attention weights but also for information from value states and attention outputs, thereby enhancing existing eviction strategies with output-aware signals. Experiments on LLaMA and Qwen models demonstrate that replacing the heuristic scores in existing works, which estimate token saliency across different query positions, with OBCache's output-aware scores consistently improves long-context accuracy. Code is available at https://github.com/DreamSoul-AI/OBCache.

Comment: Casts KV eviction as layer-wise structured pruning under Optimal Brain Damage, giving closed-form saliency scores for isolated keys, isolated values, and joint KV pairs instead of accumulated-attention heuristics.

Topic Match: KV-cache memory efficiency with a genuinely new output-aware scoring mechanism rather than a tuned heuristic.

Relevance: 8 Novelty: 6


32. VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference

ArXiv ID: 2608.08569

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Wenxu Jia, Dongjie Fu, Xize Cheng, Fangming Feng, Linjun Li, Wenshi Chen, Yingming Li, Zhou Zhao, Tao Jin

Abstract: Recent advancements in Speech Large Language Models have demonstrated remarkable capabilities in understanding complex audio tasks. Despite this progress, their long-context inference remains severely bottlenecked by prohibitive KV cache memory demands. Existing text-centric compression methods struggle here, often disrupting speech continuity or discarding crucial semantic cues. To address this, we propose VoxZip, a train-free, two-stage semantic-anchored KV cache compression framework. The first stage uses automatic speech recognition (ASR) transcriptions as explicit semantic anchors to temporally align, compress, and fuse audio tokens, significantly reducing the initial KV cache while elevating token information density. To further improve the compression ratio, the second stage employs a dynamic filtering strategy based on temporally decayed accumulated attention to evict non-essential tokens while mitigating early-token bias. Comprehensive evaluations on Qwen3-Omni across six diverse audio benchmarks demonstrate the superiority of our approach. VoxZip excels in long-audio reasoning and consistently maintains high-fidelity perception on short-form tasks. Notably, it sustains over 90\% of the uncompressed baseline performance even under an aggressive 20x KV cache compression in long-context scenarios. Furthermore, at a 4x compression ratio, VoxZip yields a 1.9x increase in inference throughput alongside a 3.3x reduction in peak memory overhead. Code and models will be available at https://github.com/MM-Speech/VoxZip.

Comment: Semantic-anchored temporal KV-cache compression reduces speech-LLM inference memory through aligned token fusion and time-aware eviction.

Topic Match: The core contribution is an audio-specialized LLM cache-compression algorithm with measured memory and throughput improvements.

Relevance: 8 Novelty: 6


33. Hybrid Policy Distillation for LLMs

ArXiv ID: 2604.20244

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Wenhong Zhu, Ruobing Xie, Rui Wang, Pengfei Liu

Abstract: Knowledge distillation (KD) is a powerful paradigm for compressing large language models (LLMs), whose effectiveness depends on intertwined choices of divergence direction, optimization strategy, and data regime. We break down the design of existing KD methods and present a unified view that establishes connections between them, reformulating KD as a reweighted log-likelihood objective at the token level. We further propose Hybrid Policy Distillation (HPD), which integrates the complementary advantages of forward and reverse KL to balance mode coverage and mode-seeking, and combines off-policy data with lightweight, approximate on-policy sampling. We validate HPD on long-generation math reasoning as well as short-generation dialogue and code tasks, demonstrating improved optimization stability, computational efficiency, and final performance across diverse model families and scales. The code related to this work is available at https://github.com/zwhong714/Hybrid-Policy-Distillation.

Comment: Unifies token-level distillation and mixes forward/reverse KL with approximate on-policy sampling.

Topic Match: Knowledge distillation is the primary compression mechanism, while the KL analysis also informs optimization behavior.

Relevance: 7 Novelty: 7


34. PTQ4SNN: Membrane-Aware Post-Training Quantization for Spiking Neural Networks

ArXiv ID: 2608.07066

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Hui Xie, Tong Shi, Haotong Qin, Aishan Liu, Xiaode Liu, Jinyang Guo

Abstract: Spiking neural networks (SNNs) enable sparse and event-driven computation, but their low-bit deployment remains incomplete because recurrent membrane states are commonly retained in floating point even after weight quantization. Quantizing these states is challenging because their distributions differ across channels and from the preceding weights, while small perturbations near the firing threshold may alter spike decisions and accumulate over time. We propose PTQ4SNN, a membrane-aware post-training quantization framework that jointly quantizes weights and recurrent membrane states using only a small calibration set. First, a channel-wise Unified Scale Bridge constrains the membrane scale as s_mem,c = s_w,c * 2^k_c, adapting to membrane distributions while enabling shift-compatible scale conversion. Second, Mixed-Precision Bit Allocation assigns 2/4/8-bit precision to membrane channels according to firing activity and quantization sensitivity under an average-bit budget. The framework operates on reusable projection-LIF pairs and supports both convolutional SNNs and spike-driven Transformers without backbone retraining. Experiments on static and event-based classification and semantic segmentation show that PTQ4SNN effectively preserves model accuracy under W4 quantization and approximately 4-bit membrane precision.

Comment: Jointly quantizes SNN weights and recurrent membrane states with scale bridging and mixed precision.

Topic Match: The core contribution is a new post-training quantization mechanism covering both parameters and recurrent state.

Relevance: 7 Novelty: 7


35. LibraSpec: Dynamic Diffusion-Based Speculative Decoding via Marginal-Gain-Driven Optimization

ArXiv ID: 2608.08721

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Zexun Lin, Yuan Feng, Junlin Lv, Kevin S. Zhou, Xike Xie

Abstract: Speculative decoding accelerates large language model inference by drafting multiple tokens for parallel verification, with efficiency critically determined by the speculative length selected at each decoding round. Existing dynamic speculation methods select the speculation length by estimating how many tokens will be accepted, which is reasonable for autoregressive drafters that generates tokens sequentially. The recent wave of diffusion-based drafters, however, generates candidate blocks in parallel at substantially lower drafting cost, shifting the key question from how many tokens to generate to how many generated tokens are worth verifying. We therefore reformulate dynamic speculative-length selection as expected-speedup optimization and derive a marginal criterion that extends the speculative sequence only when its acceptance gain outweighs the additional verification cost. Building on this criterion, we develop \textit{LibraSpec}, a training-free and plug-and-play algorithm that iteratively determines the speculative length using drafter confidence scores. Theoretically, we prove that LibraSpec monotonically converges toward the optimal speculative length. Experiments across six target models, three diffusion-based speculative decoding methods, and math, coding, and chat benchmarks show consistent improvements under both greedy and sampling settings, achieving a further $0.5\sim1.5\times$ improvement over baselines and up to $8.49\times$ speedup over autoregressive decoding.

Comment: Optimizes diffusion-drafter speculative length using the marginal acceptance gain relative to verification cost.

Topic Match: The contribution is a new training-free decoding algorithm that materially reduces inference cost.

Relevance: 7 Novelty: 7


36. Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs

ArXiv ID: 2608.08794

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Kyeongyoon Lee, Hongyeob Kim, Youngeun Kim, Sungeun Hong

Abstract: Omni-modal LLMs jointly process audio, video, and text, but long multimodal sequences incur substantial prefill and KV-cache costs. Existing omni-modal compression methods primarily focus on pre-LLM token reduction, leaving modality-specific compression across the LLM boundary underexplored. We propose A-PACK, a two-stage framework that defers audio pruning until query-conditioned multimodal interactions emerge. Our analysis shows that audio exhibits higher task-relevant information density and representational diversity per token than video. We further find that local audio-visual dynamics provide a more effective cue for visual selection than token-wise matching. We therefore preserve audio and compress video with local dynamics before the LLM, then progressively prune low-relevance audio and visual tokens and their KV-cache entries inside the LLM. Across four benchmarks on Qwen2.5-Omni-7B/3B, A-PACK achieves the strongest average performance among the evaluated prior methods while reducing prefill FLOPs by up to 78% and improving decoding throughput by up to 2.21x.

Comment: Progressively prunes multimodal tokens and their KV entries after query-conditioned audio-visual interactions emerge.

Topic Match: The core contribution is a staged token and KV-cache pruning mechanism that materially lowers inference cost.

Relevance: 7 Novelty: 7


37. Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills

ArXiv ID: 2608.07885

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Agamdeep Singh, Srishti Gautam, Priyanshu Gupta, Nikita Mehrotra, Tanmay Bakshi, Sumit Gulwani

Abstract: Reasoning modes of language models outperform their non-reasoning counterparts on multi-step agentic tasks, but pay a 3-6x premium in output tokens on every episode -- much of it spent re-deriving procedures that are shared across episodes of the same domain. We show this recurring cost can be amortized: a coding agent analyses a small corpus of existing trajectories from a training split and compiles a compact natural-language skill that is injected into the non-reasoning model's system prompt. Across four agentic benchmarks (ALFWorld, tau$^2$-bench telecom and retail, and SpreadsheetBench-Verified), skills recover 55%-100%+ of the reasoning gap for GPT-5.4-mini on held-out tasks -- exceeding the reasoning mode outright on two of four -- while emitting 2.7-6x fewer output tokens and zero reasoning tokens. Notably, reasoning traces are not a prerequisite: skills distilled from non-reasoning trajectories alone remain competitive with skills distilled from paired reasoning/non-reasoning corpora, with domain-dependent differences between the two sources. We interpret these results through a search lens: test-time reasoning is deep search inside a single episode, re-paid at every deployment, while corpus distillation is wide search across episodes, paid once. The two recover overlapping procedural knowledge, and width over cheap trajectories is often the better buy -- with the residual gap on some domains (telecom, SpreadsheetBench) delineating where genuinely per-instance deep search remains necessary.

Comment: Amortizes agent reasoning into reusable prompt skills, sharply reducing inference tokens.

Topic Match: Its strongest foundational connection is reducing repeated inference computation through cross-episode distillation.

Relevance: 6 Novelty: 8


38. RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs

ArXiv ID: 2608.07088

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Qiyanhui Lu, Han Wu, Rongjian Xu, Tingzhang Luo, Cheng Fan, Xinghao Chen, Minjing Dong, Jufeng Yang, Jianyuan Guo

Abstract: Multimodal large language models (MLLMs) encode images as long visual token sequences, making prefilling and KV-cache storage expensive. Existing training-free pruning methods select tokens by importance, diversity, or spatial coverage, but treat retained tokens as interchangeable and do not explicitly track which object-related regions are already covered. We present RoRA, a training-free framework that casts visual token pruning as role-oriented regional evidence allocation. Given a fixed budget, RoRA partitions tokens into a protected semantic core, complementary context, and fine-grained detail. It first calibrates text-conditioned attention with a positional prior and a prompt-calibrated object prior, then builds Attention-Anchored Regions (AARs) from high-confidence anchors as lightweight proxies for covered object support. Context is explored mainly outside AARs, while a small AAR-guided budget restores local detail; pairwise similarity is used only for context-stage redundancy filtering. Under matched budgets, RoRA consistently outperforms strong training-free baselines across LLaVA and Qwen-VL families, retaining most of the unpruned accuracy even at aggressive pruning ratios, e.g., 96.5% of full performance at 88.9% pruning on LLaVA-1.5, and improving over D2Pruner by about 5% on Qwen3-VL at 75-90% pruning. At a 66.7% pruning ratio, RoRA requires only 0.7 ms for token selection and reduces end-to-end inference time by 24.6%, corresponding to a 1.33x speedup over unpruned inference on an NVIDIA H800.

Comment: Prunes visual tokens via role-specific regional coverage to reduce prefilling and KV-cache cost.

Topic Match: The central contribution is a new token-pruning allocation mechanism for lower inference compute and memory.

Relevance: 7 Novelty: 6


39. MiCoPro: End-to-End Mixed Precision HW/SW Co-design with HW-aware Proxy Model

ArXiv ID: 2608.06916

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Zijun Jiang, Yangdi Lyu

Abstract: Quantized Neural Networks~(QNN) with low-bitwidth data have proven promising in efficient storage and computation on edge devices. To mitigate accuracy degradation while maximizing speedup, layer-wise mixed-precision quantization~(MPQ) becomes a popular solution. However, existing algorithms for exploring MPQ schemes are limited in flexibility and efficiency. Comprehending the complex impacts of different MPQ schemes on post-training quantization and quantization-aware training results is a challenge for conventional methods. Furthermore, an end-to-end framework for the optimization and deployment of MPQ models is missing in existing work. To address these challenges, we propose the MiCo framework, a holistic MPQ exploration and deployment framework for edge AI applications. The framework adopts a novel optimization algorithm to search for accuracy-optimal quantization configurations under strict latency constraints. We further extended the framework to MiCoPro, which introduces a robust Hardware-Aware Proxy (HAP) model to enhance prediction accuracy and hardware versatility. By leveraging target-specific latency modeling, MiCoPro enables rapid exploration and direct deployment from PyTorch models to bare-metal C code. We demonstrate the versatility of our framework on both the BitFusion accelerator and SIMD-extended RISC-V processors, achieving up to 40\% of latency reduction with less than 3\% of accuracy drop.

Comment: Searches hardware-aware mixed-precision assignments under latency constraints and deploys them to edge targets.

Topic Match: The paper directly addresses mixed-precision quantization and hardware-aware compute reduction.

Relevance: 7 Novelty: 6


40. LoCA: Spatially-Aware Low-Rank Convolutional Adaptation of Vision Foundation Models

ArXiv ID: 2607.06918

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Sojung An, Junha Lee, Sujeong You, Nam Ik Cho, Donghyun Kim

Abstract: Pre-trained Vision Foundation Models (VFMs) provide strong visual representations for diverse downstream tasks. The key challenge of VFM adaptation stems from the prohibitive costs of full fine-tuning and catastrophic forgetting. To address this, Low-Rank Adaptation (LoRA) has emerged as the prevailing paradigm for Parameter-Efficient Fine-Tuning (PEFT). However, LoRA is typically designed for transformer self-attention layers parameterized by 2D matrices. Since convolutional kernels inherently couple spatial and channel information within a 4D tensor, forcing them into a monolithic 2D matrix disrupts the inherent spatial topology. In this paper, we propose Low-Rank Convolutional Adaptation (LoCA), a convolution-aware PEFT framework that addresses spatial-channel entanglement by decoupling channel and spatial adaptation. LoCA introduces a low-rank channel adaptation for dense cross-channel mixing and refines spatial bases extracted from pre-trained kernels via Singular Value Decomposition (SVD). Experimental results show that LoCA preserves pre-trained spatial priors and achieves competitive or state-of-the-art performance across fine-grained classification, domain-generalized semantic segmentation, and generative benchmarks.

Comment: Factorizes convolutional adaptation into low-rank channel updates and refinements of pretrained spatial bases.

Topic Match: The paper's core contribution is a new low-rank parameter-efficient adaptation mechanism.

Relevance: 7 Novelty: 6


41. Domain-Aware Pruning: Sparsity and Domain Generalization via Regularized Probabilistic Masking

ArXiv ID: 2608.08624

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Parham Sazdar, Mostafa Tavassolipour, Reshad Hosseini

Abstract: Domain generalization (DG) and neural network pruning are conventionally treated as distinct objectives, targeting out-of-distribution (OOD) robustness and model efficiency, respectively. In this work, we bridge this gap by introducing Domain-Aware Pruning (DAP), a framework that leverages network sparsity as a mechanism to implicitly enhance generalization to unseen domains. Diverging from standard binary mask optimization, DAP learns a continuous parameter retention probability $p \in [0, 1]$, framing network compression as a continuous probabilistic masking problem. By introducing a regularization objective that actively penalizes the retention of domain-sensitive weights during the mask training, DAP identifies a domain-invariant subnetwork. Empirical results across five DG benchmark datasets demonstrate that DAP achieves significant sparsity while consistently matching or exceeding the OOD performance of its dense counterparts. Crucially, DAP is an algorithm-agnostic framework that integrates seamlessly with existing DG pipelines without necessitating post-hoc fine-tuning. Beyond efficiency and generalization, we show that DAP natively provides increased robustness to adversarial perturbations and yields highly interpretable models, where the retained weights reliably encapsulate the most domain-invariant and task-critical representations.

Comment: Learns probabilistic pruning masks that preferentially remove domain-sensitive parameters.

Topic Match: The core method is a new pruning objective that creates sparse subnetworks, despite its domain-generalization motivation.

Relevance: 7 Novelty: 6


42. LoRSA: Toward Generalizable Parameter-Efficient Fine-Tuning for Biomedical Downstream Tasks

ArXiv ID: 2608.07749

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Saed Moradi, Benyamin Ghojogh, M. Hadi Sepanj, Yimin Yang, Ashirbani Saha

Abstract: Parameter-efficient fine-tuning enables the adaptation of vision foundation models to biomedical tasks under limited computational resources, but a single low-rank update can constrain all task-specific changes to one narrow parameter subspace. This restriction may prevent the model from simultaneously representing globally shared task structure and localized residual directions required for generalization to unseen imaging domains. We introduce LoRSA, a global--residual adaptation framework that jointly learns a dense low-rank component and a dynamically structured-sparse low-rank component. The dense component captures globally coordinated task adaptation, while the structured component provides complementary residual corrections whose support evolves during training. We characterize the representational capacity, approximation properties, rank structure, and singular-subspace complementarity of this decomposition. We evaluate LoRSA for four-class breast-density classification using DINOv3-Base, with VinDr-Mammo as the source domain and MammosighTR and RSNA as unseen external domains. LoRSA remains competitive on the internal validation set and achieves the best external macro-F1 on both target datasets, improving upon the strongest competing method by 2.15 percentage points on MammosighTR and 3.09 percentage points on RSNA. Weight-matrix analysis further shows that approximately $92\%$ of the energy of each adaptation component lies outside the bilateral singular subspace of the other, indicating that the two components learn largely complementary update directions. These results suggest that organizing adaptation capacity into distinct global and residual paths can improve the external-domain generalization of parameter-efficiently adapted biomedical vision models.

Comment: Combines dense low-rank updates with dynamically structured-sparse low-rank corrections to expand adaptation capacity.

Topic Match: The contribution is a new parameter-efficient adaptation decomposition, although demonstrated benefits chiefly concern biomedical domain generalization.

Relevance: 7 Novelty: 6


43. Addressable Memory for Video World Models

ArXiv ID: 2608.07408

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Xindi Wu, Sven Elflein, James Lucas, Olga Russakovsky, Laura Leal-Taixé, Despoina Paschalidou, Jonathan Lorraine, Aljoša Ošep

Abstract: We study visual persistence in interactive video world models. These models rely on a Key-Value (KV) cache as a growing visual memory to carry forward previously generated frames. However, we find that models can no longer reliably address stored content once rollouts extend beyond the training horizon, because temporal Rotary Positional Embeddings (RoPE) offsets then fall outside the range seen during training and the model struggles to retrieve the relevant visual information through attention. Moreover, naively compressing the cache in the RoPE-rotated space corrupts memory by averaging together incompatible positional phases. To address this, we propose WorldTrace, a training-free memory framework for long-horizon visual persistence. WorldTrace keeps compressed memory addressable by assigning each summary slot a distinct, in-distribution virtual position. Within this addressable cache, we study two memory compression approaches: WorldTrace-Field compresses history for temporal coherence, while WorldTrace-Landmark stores verbatim scene traces at detected transitions for episodic recall. We further introduce LoopBench, a benchmark evaluating whether a compressed cache can reconstruct a previously visited scene after a long detour. WorldTrace-Field improves temporal consistency by +15.5%, and WorldTrace-Landmark improves episodic recall by +19.5% on LoopBench, extending visually persistent generation without retraining.

Comment: Keeps compressed KV memory addressable by assigning summaries in-distribution virtual RoPE positions.

Topic Match: The primary mechanism is position-aware KV-cache compression, with a secondary architectural insight about RoPE addressing.

Relevance: 6 Novelty: 7


44. HEART: Exploiting Head Heterogeneity in Sparse Attention for Video Diffusion

ArXiv ID: 2605.14513

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Xuzhe Zheng, Yuexiao Ma, Jing Xu, Xiawu Zheng, Rongrong Ji, Fei Chao

Abstract: Sparse attention accelerates video diffusion by allowing each attention head to focus on only a small subset of interactions. Existing methods already construct head-specific sparse patterns conditioned on the input. However, we find that these heads also differ in two less obvious but practically important ways. First, for some heads, sparse attention masks remain stable across many denoising steps, whereas others change rapidly. Second, heads differ substantially in their sensitivity to sparsification: applying the same threshold can induce markedly different errors in the final denoising velocity. Ignoring these differences leads to redundant mask prediction and suboptimal threshold calibration across heads under a global sparsity budget. We present HEART, short for Heterogeneity-Exploiting Adaptive Refresh and Thresholding, a training-free framework that exploits both forms of head heterogeneity. First, Temporal Mask Reuse (TMR) uses a lightweight per-head query-key drift signal to determine whether a cached sparse mask remains reliable across denoising steps, refreshing it only when the drift exceeds a prescribed threshold. Second, Error-guided Budgeted Calibration (EBC) evaluates candidate thresholds offline using a frequency-weighted denoising-velocity error, and assigns each head an appropriate threshold under a global sparsity budget. HEART requires no retraining, weight modification, or sparse-kernel changes, and can be integrated directly into existing sparse-attention pipelines. Across Wan2.1-1.3B, Wan2.1-14B, and HunyuanVideo-13B, HEART consistently pushes the quality--efficiency frontier of advanced sparse attention methods such as XAttention and SVG2 outward.

Comment: Cuts sparse-attention inference cost through per-head mask reuse and error-budgeted thresholds.

Topic Match: The main mechanism reduces attention computation by exploiting head-specific temporal stability and error sensitivity.

Relevance: 6 Novelty: 6


45. CHIME: A Case for Efficient Long-Context Attention-FC Disaggregated Inference with DIMM-PIM

ArXiv ID: 2504.17584

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Large-Scale Training Systems and Efficiency

Authors: Qingyuan Liu, Liyan Chen, Haocheng Wang, Yanning Yang, Dong Du, Zhigang Mao, Naifeng Jing, Yubin Xia, Haibo Chen

Abstract: Attention-FC Disaggregated (AFD) LLM inference systems offload memory-bound Attention operations to memory-rich accelerators (e.g., CPUs, HBM-PIM) while retaining compute-bound Fully-Connected (FC) operations on GPUs. In this paper, we first design a Disaggregated Roofline Model (DRM) to characterize AFD performance, revealing that system throughput is constrained by the accelerator's limiting factor: either memory bandwidth or capacity. We observe that prior AFD systems often overlook these constraints and fail to balance them, leading to resource underutilization or constrained throughput. Therefore, we propose CHIME, the first AFD system integrating DIMM-PIM, which is a case of the new accelerator that strikes the balance with scalable capacity and bandwidth. To address the synchronization challenges inherent to the distributed cooperating DRAM chips in DIMM-PIM, CHIME employs bubble-free pipelining and hybrid-grained re-layout for efficient attention computation. Furthermore, it maximizes cross-device resource utilization via rankset-granular communication-computation overlapping and alignment-predicting scheduling. Evaluations show CHIME achieves up to 5.15$\times$ speedup over state-of-the-art HBM-PIM solutions.

Comment: Pipelines disaggregated attention on DIMM-PIM while overlapping rank-level communication and GPU FC computation.

Topic Match: The primary contribution is an accelerator and scheduling design that reduces long-context inference cost.

Relevance: 6 Novelty: 6


Paper Selection Prompt

System Prompt

You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.

User Prompt

Instructions

Respond in JSONL. Output exactly one JSON object per paper, one per line:

{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}

Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only: daily_hot, new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return []. - daily_hot means the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. - new_frontier means the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.

Scoring Criteria

Relevance and Novelty are independent axes. Score both from 1 to 10.

Relevance Scoring

  • 9-10: directly centered on the target foundational topics; highest when the core contribution is clearly within them.
  • 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
  • 5-6: touches the target topics, but the main contribution is elsewhere.
  • 3-4: largely outside the target topics, often application-focused or domain-specific.
  • 1-2: unrelated.

Important: Broad frontier relevance, major launch status, or daily buzz is not enough for a high Relevance score here. Those cases belong in the hotspot digest unless the paper strongly matches the specialized paper topics.

Novelty Scoring

  • 9-10: new paradigm, theory, or major methodological breakthrough.
  • 7-8: substantial methodological advance or strong new insight.
  • 5-6: meaningful but incremental extension or refinement.
  • 3-4: minor, narrow, or mostly engineering or domain-specific improvement.
  • 1-2: little originality; mainly standard application of existing methods.

Topic Registry

Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - moe_training: MoE Training - Mixture-of-Experts training: routing, load balancing, expert granularity and shared experts, training stability, MoE parallelism and communication, and MoE kernels. - training_systems: Large-Scale Training Systems and Efficiency - Large-scale training systems and efficiency: distributed training algorithms, optimizers, and communication. - architecture_training: Architecture and Training Dynamics - Core architectural or computational mechanisms, dynamic computation, and training-stability dynamics. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, memory or cache efficiency, and large-scale training systems that materially change cost or behavior.

Papers

[PAPER LIST HERE]

Relevant Topics

This is a training-side feed. The centre of gravity is Mixture-of-Experts training and the large-scale training systems it sits on; architecture and efficiency work is kept when it bears on how a large model is trained or what it costs to train.

Keep a paper when its CORE CONTRIBUTION falls in one of the four topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.

  1. MoE Training (primary) - Keep: expert routing and gating mechanisms; load balancing, with or without auxiliary losses; expert granularity, fine-grained and shared experts; MoE training stability (loss spikes, router collapse, z-loss and related fixes); expert and hybrid parallelism; all-to-all dispatch/combine and the communication schedules MoE needs; MoE kernels and grouped-GEMM style implementations; expert capacity, token dropping, and upcycling dense checkpoints into MoE. - Filter: papers that merely train on top of a MoE model without a new routing, balancing, stability, parallelism, or kernel contribution; "mixture of experts" in the classical ensemble or recommender sense.

  2. Large-Scale Training Systems - Keep: distributed training algorithms (data, tensor, pipeline, sequence, and expert parallelism, FSDP and ZeRO-style sharding); optimizers and preconditioners for large-scale pretraining; communication algorithms, overlap, and collective design; throughput and MFU work that materially changes what a training run costs or whether it converges; scaling-law work that informs how such runs are configured. - Filter: routine infrastructure tuning with no new algorithmic or systems idea.

  3. Architecture and Training Dynamics - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, dynamic or modular computation); optimisation and training-dynamics analysis that explains why large models train the way they do. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight.

  4. Efficiency and Compression - Keep: quantization, sparsity, pruning, low-rank adaptation, KV-cache and memory-efficient designs, and other work that materially changes the cost of running or training a large model, especially when the mechanism is new rather than a tuned variant. - Filter: straightforward application of standard efficiency methods, and deployment or serving engineering with no new idea.

Leave these to the hotspot digest unless the core contribution clearly falls in one of the four topics above: - major model or product releases - agent, tooling, and RAG launches - benchmark, leaderboard, or evaluation-only papers - representation-learning theory, interpretability, and feature-formation analysis - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning) - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains