Previous Day 2026-08-29
Monthly Overview 2026-08
Next Day 2026-09-01

This is a remedial run for missed papers from 08/29/2026 to 08/30/2026.

Results generated on 09/14/2026.

Personalized Daily ArXiv Papers 2026-08-31

Model Metric Usage Papers
Prompt Completion Total Total arXiv Scanned Relevant
local agent CLI Tokens not reported not reported not reported 624 624 36
Cost not reported not reported not reported

Token counts are not reported for this run. 27 of 27 model calls succeeded, 3,050s of model wall clock.

Topic Coverage:

TopicPapers
MoE Training2
Large-Scale Training Systems and Efficiency4
Architecture and Training Dynamics9
Efficiency, Compression, and Large-Scale Training21

Table of contents by topic:

MoE Training (2)

  1. Structure Aware Neural Architecture Search for Mixture of Experts Authors: Petr Babkin, Oleg Bakhteev

  2. Evolutionary Soups: Evolving Mixture-of-Experts for Multi-Objective LLM Alignment Authors: Lingxiao Kong, Steffen Staab, Cong Yang, Oya Beyan, Zeyd Boukhers

Large-Scale Training Systems and Efficiency (4)

  1. When Do Larger Batches Help Scale LLM Reinforcement Learning? Authors: Ziniu Li, Jinbo Wang, Guanhua Huang, Feiyuan Zhang, Pengbo Li, Alex Chen

  2. A Unified Framework for Fair and Personalized Decentralized Learning under Communication Constraints Authors: Krishnendu S. Tharakan, Carlo Fischione

  3. GRASP: Geometry-aware Residual Alignment for Scalable Pretraining Data Attribution Authors: Yue Min, Ruining Chen, Yujun Li

  4. SS-ESOAP: Self-Scaled Adaptive Preconditioning for Physics-Informed Learning Authors: Guangyuan Wang, Mads Toftrup, Sebastian Loeschcke, Yixuan Wang, Anima Anandkumar

Architecture and Training Dynamics (9)

  1. Don't Read Everything: A Curvature-Conditioned Query for Linear Attention Authors: Dong Le, Thong Nguyen, Cong-Duy Nguyen, Anh Tuan Luu

  2. Branch Scaling Manifests as Implicit Architectural Regularization for Improving Generalization in Overparameterized ResNets Authors: Zixiong Yu, Guhan Chen, Jianfa Lai, Bohan Li, Songtao Tian

  3. Higher-Dimensional Rotary Position Embedding Authors: Yixing Li, Ruobing Xie, Yudong Zhang, Yushi Bai, Samm Sun, Yu Cheng

  4. Understanding Deep Learning via Notions of Rank Authors: Noam Razin

  5. Beyond Dense States: Sparse Transcoders as Causally Testable Operators for LLM Latent Reasoning Authors: Yadong Wang, Haodong Chen, Yu Tian, Chuanxing Geng, Dong Liang, Xiang Chen

  6. Every Layer Counts: An Exponential $L_2$ Depth Hierarchy for ReLU Networks Authors: Itay Safran

  7. SHAKE-GNN: Scalable Hierarchical Kirchhoff-Forest Graph Neural Network Authors: Zhipu Cui, Johannes Lutzeyer

  8. Training-Free Hidden-State Refinement for Flow-Matching Image Generators Authors: Yuanyi Yan, Xinzhe Rao, Canyu Shen, Yang Chen, Yunlu Chen, Meng Tang, Teng Long, Vincent Tao Hu

  9. Federated Personalization of Early-Exit Networks Authors: Boyi Liu, Zimu Zhou, Cheng Fang, Yongxin Tong

Efficiency, Compression, and Large-Scale Training (21)

  1. OCGQuant: Outlier-Companion Grouping for NVFP4 Quantization Authors: Yishan Yao, Binjun Li, Hanling Yi, Pengyu Li, Xiaoqing Liu, Zihan Yang, Xiaotian Yu, Zhiwen Yu

  2. TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning Authors: Wonpyo Park, Seung-won Hwang

  3. An Efficient Sparse Fine-Tuning with Low Quantization Error via Neural Network Pruning Authors: Cen-Jhih Li, Aditya Bhaskara

  4. Locality-Aware Redundancy Pruning for LLM Depth Compression Authors: Vincent-Daniel Yun, Youngrae Kim, Woosang Lim, YoungJin Heo, Minkyu Kim, Sunwoo Lee

  5. DASH: Dual-Branch Score Distillation for Guidance-Calibrated Compact Diffusion Models Authors: Abdullah Al Shafi, Kazi Saeed Alam, Sk Imran Hossain, Engelbert Mephu Nguifo

  6. WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource Hardware Authors: Jiamu Zhang, Liang Wu, Mayank Darbari, Liangjie Hong

  7. A Deterministic Constant-Competitive Algorithm for Dynamic Mixture-of-Experts Serving Authors: Ian D'Ambrosio

  8. Towards Provable and Scalable Training of Quantized Neural Networks with Ising Optimization Authors: Wenxin Li, Chuan Wang, Hongdong Zhu, Qi Gao, Yin Ma, Hai Wei, Kai Wen

  9. SupraTok: Cross-Boundary Tokenization for Enhanced Language Model Performance Authors: Andrei-Valentin Tănase, Elena Pelican

  10. Approximate Speculative Decoding Authors: Yuannuo Feng, Zegang Peng, Yuxin Xie, Yubing Ye, Yizhe Chen, Wenshuai Yao, Wenyong Zhou, Wang Kang

  11. AutoCRAT: Within-trajectory Joint Control of Stochasticity and Compute for LLM Reasoning Authors: Hanjun Luo, Qiushi Liu, Jingya Zhang, Haihong Pang, Jiaheng Wen, Yifei Ma, Yu Yao, Chengxi Zhang, Hanrong Zhang, Yankai Chen, Hanan Salam

  12. SpecPV: Improving Self-Speculative Decoding for Long-Context Generation via Partial Verification Authors: Zhendong Tan, Xingjun Zhang, Chaoyi Hu, Junjie Peng, Kun Xia

  13. EpaCache: Error-Propagation-Aware Caching for Accelerating Diffusion-Based Visual Generation Authors: Yuhan Liu, Zongwei Hong, Jinglun Li, Linze Li, Shen Zhang, Yao Tang

  14. DIP: Dynamic In-Context Planner For Diffusion Language Models Authors: Yang Li, Han Meng, Chenan Wang, Zhenyu Bi, Xuan Wang, Haipeng Chen

  15. Prefix-Adaptive Block Diffusion for Efficient Document Recognition Authors: Mingxu Chai, Ziyu Shen, Chenyu Liu, Jihua Kang, Tao Gui, Qi Zhang

  16. Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning Authors: Yu Xia, Zhouhang Xie, Xin Xu, Byungkyu Kang, Prarit Lamba, Xiang Gao, Julian McAuley

  17. Spectral Analysis for Sparse Matrix Computation: Insights and Potential Authors: Ruifeng Zhang, Xipeng Shen

  18. Masked Distillation: Internalizing the Chain-of-Thought in Language Models Authors: Durgesh Kalwar, Vardhan Palod, Subbarao Kambhampati

  19. AdaVLA: Adaptive Step Flow Matching for Training-free Acceleration of Vision-Language-Action Models Authors: Sunghwan Han, Youngtae Han, Youngmin Yi

  20. Entropy-Aware Token Rejection for Improving Speculative Decoding Authors: Tiancheng Su, Meicong Zhang, Guoxiu He

  21. PruneShift: A Framework for Evaluating Decision Reliability in Structured Pruning Authors: Hao Ye, Gaopeng Zhang


MoE Training (2)

1. Structure Aware Neural Architecture Search for Mixture of Experts

ArXiv ID: 2608.29817

Primary Topic: MoE Training

Authors: Petr Babkin, Oleg Bakhteev

Abstract: Neural Architecture Search (NAS) has so far rarely been applied to Mixture-of-Experts (MoE) models, and existing MoE designs leave the alignment between experts and the structure of the data to emerge on its own. We propose an architecture search framework that makes this alignment an explicit search variable: the assignment of data clusters to experts is optimised jointly with the per-expert architectures. We cast the joint problem as a cluster-aware likelihood maximisation, show that it coincides with the incomplete-data maximum likelihood of a latent-variable mixture, and solve it by a generalised Expectation-Maximisation procedure whose otherwise intractable expert-quality term is supplied by an adaptively refined surrogate. We prove that the iterates converge whenever the surrogate errors are summable, and that at every limit point no candidate the search produces improves the true objective. On a heterogeneous image-classification mixture the method recovers the underlying domain partition on 95% of clusters without ever observing domain labels, and on that benchmark and a four-domain time-series forecasting one alike it outperforms the MoE and NAS baselines that likewise use no label information.

Comment: Jointly searches cluster-to-expert assignments and per-expert architectures using surrogate-guided generalized expectation-maximization.

Topic Match: Expert assignment makes moe_training the closest label, but the described cluster-level latent-mixture search appears closer to classical mixture design than sparse large-model training.

Relevance: 6 Novelty: 7


2. Evolutionary Soups: Evolving Mixture-of-Experts for Multi-Objective LLM Alignment

ArXiv ID: 2608.29978

Primary Topic: MoE Training

Authors: Lingxiao Kong, Steffen Staab, Cong Yang, Oya Beyan, Zeyd Boukhers

Abstract: Large language models are increasingly required to generate responses that satisfy multiple competing objectives. Since optimal trade-offs depend on both user preferences and input prompts, controllable multi-objective generation must dynamically adapt models at inference time without retraining. To address this, we propose Evolutionary Soups, a mixture-of-experts framework for fine-grained generation control, with gating networks trained via an evolutionary algorithm. The per-layer gating networks dynamically produce expert-merging coefficients from hidden-state representations, while the evolutionary algorithm incorporates greedy hypervolume contribution for effective evolution of these gating networks, achieving consistent improvements on large and noisy training datasets and broader coverage of the non-convex Pareto front. Experiments across three tasks demonstrate the effectiveness of Evolutionary Soups over baselines: it achieves the best hypervolume, linear utility, and Tchebyshev utility (~20% improvement) among controllable methods on all tasks.

Comment: Uses hypervolume-guided evolution to train hidden-state-dependent, per-layer expert-merging gates.

Topic Match: Expert-gate learning provides a substantive MoE connection, but the central contribution is Pareto-controlled alignment through model merging.

Relevance: 6 Novelty: 6


Large-Scale Training Systems and Efficiency (4)

1. When Do Larger Batches Help Scale LLM Reinforcement Learning?

ArXiv ID: 2608.29296

Primary Topic: Large-Scale Training Systems and Efficiency

Also Matches: Architecture and Training Dynamics

Authors: Ziniu Li, Jinbo Wang, Guanhua Huang, Feiyuan Zhang, Pengbo Li, Alex Chen

Abstract: Larger batches reduce the variance of stochastic gradients per update and are therefore often expected to accelerate training. Yet whether this statistical benefit translates into lower wall-clock time-to-target remains unclear, because each update consumes more samples and may take longer to execute. We study this tradeoff in reinforcement learning for large language models. We separate its algorithmic and systems effects by comparing learning and execution along their natural axes. At the algorithmic level, we compare configurations at equal cumulative sample counts while retuning batch-dependent hyperparameters. Over a bounded range of batch sizes, this procedure yields an approximately batch-size-invariant family whose members follow similar sample-indexed learning trajectories. At the systems level, we exploit the computational asymmetry between rollout generation and training: autoregressive generation is often memory-bandwidth-bound at low concurrency, whereas training work scales approximately with the number of processed tokens. Combining these two views yields a direct decision rule: a larger-batch configuration reduces time-to-target only when its throughput gain exceeds its samples-to-target penalty. Experiments with GRPO and PPO support both sides of this decomposition. At the algorithmic level, square-root learning-rate scaling with Adam produces approximately batch-size-invariant learning curves over a bounded range of batch sizes. At the systems level, larger batches improve generation throughput by up to 2.29x on fixed hardware. In GRPO, combining higher throughput with learning-rate retuning reduces time-to-target by up to 29%, whereas increasing the batch without retuning is slower despite its higher throughput.

Comment: Determines when larger-batch throughput gains outweigh samples-to-target penalties after learning-rate retuning.

Topic Match: The core contribution separates execution throughput from learning dynamics to configure training for lower wall-clock cost, qualifying despite its RL post-training setting.

Relevance: 8 Novelty: 6


2. A Unified Framework for Fair and Personalized Decentralized Learning under Communication Constraints

ArXiv ID: 2608.26493

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Krishnendu S. Tharakan, Carlo Fischione

Abstract: Decentralized learning systems aim to collaboratively train models across multiple clients without relying on a central coordinator. While decentralization improves scalability, privacy, and robustness, it also exacerbates three fundamental challenges: statistical heterogeneity across clients, fairness in client-level performance, and stringent communication constraints. This raises a natural question: \emph{how fair can decentralized learning be under limited communication?} We address this question by presenting a unified framework for decentralized learning under communication constraints, bringing together graph-based personalization, agnostic fairness, and compressed event-triggered communication. Specifically, we propose a new algorithm DMFL-SQ, a decentralized multi-task learning algorithm that couples personalized model training over a communication graph with an agnostic mixture fairness objective, while reducing communication through sparsification, quantization, and event-triggered synchronization. We establish convergence guarantees for general non-convex objectives and show that DMFL-SQ achieves an $\mathcal{O}(T^{-1/2})$ rate in expected squared Moreau-envelope stationarity despite sparse, quantized, and event-triggered communication. We further derive PAC-Bayes generalization guarantees for the fairness-aware mixture objective. Experiments on CIFAR-10 and the real heterogeneous MUSMET EEG dataset demonstrate that DMFL-SQ substantially reduces communication while maintaining predictive performance and improving fairness across clients. Together, our theoretical and empirical results show that personalization, fairness, and communication efficiency can be jointly achieved in decentralized learning while preserving the dominant convergence rate.

Comment: Sparse, quantized, event-triggered synchronization preserves convergence guarantees for decentralized training.

Topic Match: The distributed optimization and synchronization algorithm fits training_systems, with narrower relevance because it targets heterogeneous-client fairness and personalization.

Relevance: 7 Novelty: 6


3. GRASP: Geometry-aware Residual Alignment for Scalable Pretraining Data Attribution

ArXiv ID: 2606.06892

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Yue Min, Ruining Chen, Yujun Li

Abstract: Scalable data attribution methods typically assign isolated utility scores to individual training examples. This prevalent additive assumption fundamentally fails to capture critical subset dynamics, including data redundancy and complementary coverage. In this work, we reframe attribution as subset-level counterfactual utility prediction and introduce GRASP, an interaction-aware surrogate. Grounded in a theoretical smoothness lower bound, GRASP explicitly models subset interactions through a quadratic geometric penalty. To achieve pretraining-scale efficiency without relying on hidden oracle tuning, we couple low-dimensional feature sketches with a strictly finite lower-confidence bound selection protocol. Extensive subset-retraining evaluations demonstrate that GRASP decisively outperforms existing scalable baselines. It more than doubles the task-level rank correlation for counterfactual subset fidelity while reducing upfront artifact construction costs by nearly an order of magnitude. Downstream diagnostics further show that this scoring mechanism transfers to language model curation and cross-domain vision selection, establishing a robust foundation for optimizing massive pretraining corpora.

Comment: Interaction-aware subset scoring supports pretraining-data selection using low-dimensional feature sketches.

Topic Match: Pretraining-data selection is adjacent to run configuration; the demonstrated efficiency gain concerns attribution artifacts rather than training runs.

Relevance: 6 Novelty: 7


4. SS-ESOAP: Self-Scaled Adaptive Preconditioning for Physics-Informed Learning

ArXiv ID: 2608.29448

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Guangyuan Wang, Mads Toftrup, Sebastian Loeschcke, Yixuan Wang, Anima Anandkumar

Abstract: Physics-informed neural networks (PINNs) often face ill-conditioned objectives that limit high-accuracy training. Dense quasi-Newton methods improve local conditioning but require expensive optimizer state, while Kronecker-factored methods such as SOAP scale to larger networks but rely on periodic basis updates. We introduce \method, which augments SOAP-style preconditioning with a scalar secant-energy correction adapted to Kronecker geometry and an adaptive basis update followed by variance-state downscaling. We characterize the directional secant matching induced by the scalar correction and give a bound on variance-state mismatch across basis changes. Across eight PDE benchmarks, \method attains the lowest final residual on six, including Burgers and Boussinesq, while SOAP-family baselines perform better on Gray-Scott and Ginzburg-Landau. On Boussinesq, \method reaches a residual of $10^{-5}$ in 4.1 hours with 9.2 GB peak VRAM, while Adam does not reach this target within 14 hours. Three-seed $L^2$ and $H^1$ errors on four representative PDEs support the link between lower residuals and improved solution accuracy. These results position \method as a scalable option for stiff, high-accuracy physics-informed training, rather than a uniform replacement for existing optimizers.

Comment: Adds secant-energy rescaling and adaptive basis updates to SOAP-style preconditioning, with analysis of variance-state mismatch.

Topic Match: The new preconditioning mechanism connects directly to optimizer design, although its development and validation target physics-informed networks.

Relevance: 6 Novelty: 6


Architecture and Training Dynamics (9)

1. Don't Read Everything: A Curvature-Conditioned Query for Linear Attention

ArXiv ID: 2606.01294

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Dong Le, Thong Nguyen, Cong-Duy Nguyen, Anh Tuan Luu

Abstract: Linear attention reduces the quadratic cost of softmax attention by maintaining a recurrent fast-weight state, but it consistently lags on in-context retrieval and long-context tasks. Existing remedies act on the write side of memory through gating, delta updates, or kernel feature maps, but the read step is left unchanged: every past key contributes additively to the output, so useful targets are diluted by the bulk of stored vectors. We borrow one specific piece of softmax's geometry to construct a cheap read-time contraction of the query. A second-order Taylor expansion of the softmax log-partition at the isotropic-attention point gives a local quadratic model whose curvature coincides with the running key covariance, a quantity that can be maintained with the same recurrent/chunkwise mechanism as the linear-attention state. The associated linear operator contracts the query along the high-variance directions of memory before it reads the state. We call this mechanism Curvature-Conditioned Query (CCQ). CCQ modifies only the read step and is composable with any linear-attention backbone. Attached to GLA and Gated DeltaNet, it improves perplexity, zero-shot downstream accuracy, S-NIAH retrieval at and beyond the training context, length-extrapolation perplexity from 4K to 20K, and LongBench accuracy.

Comment: Contracts linear-attention queries along high-variance key directions using running covariance to reduce retrieval dilution.

Topic Match: A new attention read operator changes the core sequence-model mechanism while retaining recurrent and chunkwise execution.

Relevance: 9 Novelty: 8


2. Branch Scaling Manifests as Implicit Architectural Regularization for Improving Generalization in Overparameterized ResNets

ArXiv ID: 2403.04545

Primary Topic: Architecture and Training Dynamics

Authors: Zixiong Yu, Guhan Chen, Jianfa Lai, Bohan Li, Songtao Tian

Abstract: Scaling factors in residual branches have emerged as a prevalent method for boosting neural network performance, especially in normalization-free architectures. While prior work has primarily examined scaling effects from an optimization perspective, this paper investigates their role in residual architectures through the lens of generalization theory. Specifically, we establish that wide residual networks (ResNets) with constant scaling factors become asymptotically unlearnable as depth increases. In contrast, when the scaling factor exhibits rapid depth-wise decay combined with early stopping, over-parameterized ResNets achieve minimax-optimal generalization rates. To establish this, we demonstrate that the generalization capability of wide ResNets can be approximated by kernel regression associated with the Neural Tangent Kernel (NTK). Our theoretical findings are validated through experiments on synthetic data and real-world classification tasks, including MNIST and CIFAR-100.

Comment: Depth-dependent residual branch scaling and early stopping control learnability and generalization in wide ResNets.

Topic Match: Directly analyzes a residual-design mechanism and explains how its scaling with depth changes learning behavior.

Relevance: 9 Novelty: 7


3. Higher-Dimensional Rotary Position Embedding

ArXiv ID: 2608.29715

Primary Topic: Architecture and Training Dynamics

Authors: Yixing Li, Ruobing Xie, Yudong Zhang, Yushi Bai, Samm Sun, Yu Cheng

Abstract: Transformers rely on position embedding mechanisms in long context modeling in most cases. Rotary Position Embedding (RoPE) embeds positional information with independent 2D rotations, forming relative position terms in self-attention. However, its pairwise, block-based, and decoupled structure limits deep mixing and robustness across channels. We propose HD-RoPE, which extends RoPE from independent 2D rotations to higher-dimensional rotations and introduces a Paley-I orthogonal basis to obtain balanced, isotropic, and dense phase mixing within each rotation subspace. This significantly enhances channel coupling and rotational degrees of freedom while maintaining orthogonal stability and the relative position closure property. Furthermore, HD-RoPE is easily optimized for engineering efficiency without introducing additional trainable parameters. We have conducted extensive evaluation results demonstrating that HD-RoPE achieves significant performance improvements over standard RoPE across various popular benchmarks and in both long and short contexts.

Comment: Introduces higher-dimensional orthogonal positional rotations that couple attention channels while preserving relative-position structure.

Topic Match: The core contribution changes the positional geometry of attention without adding trainable parameters.

Relevance: 9 Novelty: 7


4. Understanding Deep Learning via Notions of Rank

ArXiv ID: 2408.02111

Primary Topic: Architecture and Training Dynamics

Authors: Noam Razin

Abstract: Despite the extreme popularity of deep learning in science and industry, its formal understanding is limited. This thesis puts forth notions of rank as key for developing a theory of deep learning, focusing on the fundamental aspects of generalization and expressiveness. In particular, we establish that gradient-based training can induce an implicit regularization towards low rank for several neural network architectures, and demonstrate empirically that this phenomenon may facilitate an explanation of generalization over natural data (e.g., audio, images, and text). Then, we characterize the ability of graph neural networks to model interactions via a notion of rank, which is commonly used for quantifying entanglement in quantum physics. A central tool underlying these results is a connection between neural networks and tensor factorizations. Practical implications of our theory for designing explicit regularization schemes and data preprocessing algorithms are presented.

Comment: Establishes that gradient-based training induces implicit low-rank regularization across several neural architectures.

Topic Match: Optimization-induced rank bias directly informs training dynamics, although the broader generalization and expressiveness focus makes this peripheral to large-scale pretraining.

Relevance: 7 Novelty: 7


5. Beyond Dense States: Sparse Transcoders as Causally Testable Operators for LLM Latent Reasoning

ArXiv ID: 2602.01695

Primary Topic: Architecture and Training Dynamics

Authors: Yadong Wang, Haodong Chen, Yu Tian, Chuanxing Geng, Dong Liang, Xiang Chen

Abstract: Latent reasoning reduces the token-generation cost of chain-of-thought reasoning by replacing explicit intermediate tokens with continuous latent transitions. However, existing latent reasoning methods usually rely on dense and entangled transitions, making their reasoning trajectories difficult to inspect or intervene on. We introduce LSTR (Latent Sparse Transcoder Reasoning), a framework that turns sparse transcoders from post-hoc diagnostic tools into in-loop, intervenable transition components for latent reasoning. At each latent step, a Latent Transition Transcoder (LTT) combines a linear skip path with a Top-k sparse innovation path, exposing a small set of active sparse features. Under matched compression settings, LSTR offers a mechanistically inspectable alternative to dense latent reasoning. On GSM8K-Aug, ablating only a few top-active sparse features reduces accuracy by up to 16.5%, whereas analogous interventions have much smaller effects in dense latent baselines. These results indicate that the active sparse features are causally involved in the latent transition process, rather than merely post-hoc descriptors. Additional experiments on mathematical benchmarks and StrategyQA suggest that sparse latent transitions can preserve the compression benefits of latent reasoning while making the resulting trajectories more inspectable and intervenable.

Comment: Introduces a latent transition operator combining a linear skip path with sparse Top-k innovations.

Topic Match: Sparse transcoders become computational components inside the reasoning loop, making this an architectural contribution despite its emphasis on interpretability.

Relevance: 7 Novelty: 7


6. Every Layer Counts: An Exponential $L_2$ Depth Hierarchy for ReLU Networks

ArXiv ID: 2608.23877

Primary Topic: Architecture and Training Dynamics

Authors: Itay Safran

Abstract: We prove a depth hierarchy for ReLU neural networks in which every additional ReLU layer can save exponentially many neurons. For all $k\geq2$, we construct a globally $[0,1]$-valued, $1$-Lipschitz function realized by a depth-$(k+1)$ network of width $\mathcal{O}(d^4)$, whereas any depth-$k$ network with unrestricted weights and width at most $\frac{2^d}{2d(k-1)}$ has squared $L_2$ error at least $1/24$ under an absolutely continuous distribution supported at exponential distance from the origin. To the best of our knowledge, this is the first exponential hierarchy across all adjacent fixed depths, and the first exponential separation for ReLU networks between two fixed depths whose shallower network has depth at least $3$. The lower bound also immediately yields the corresponding hierarchy for exact computation. Moreover, the case $k=2$ gives a compactly supported separation between depths $3$ and $2$ with unrestricted shallow-network weights, answering a question raised by Safran, Eldan, and Shamir (2019). The distribution used in our construction nevertheless has all its mass at exponential radius, placing the hierarchy outside the regularity regime in which such a separation would imply major threshold-circuit lower bounds. We also prove an exact separation for a more regular target, which is globally $[0,1]$-valued and $\mathcal{O}(\sqrt d)$-Lipschitz and maps the unit hypercube onto $[0,1]$. It is computed by a polynomial-width depth-$4$ network, whereas any depth-$3$ network agreeing with it on the unit hypercube requires exponentially many first-layer neurons, even with unrestricted weights.

Comment: Proves exponential width savings between every pair of adjacent ReLU depths, allowing unrestricted weights.

Topic Match: Depth-width expressivity is the closest architectural topic; these approximation-theoretic separations do not establish large-model training behavior or achievable training-cost savings.

Relevance: 6 Novelty: 8


7. SHAKE-GNN: Scalable Hierarchical Kirchhoff-Forest Graph Neural Network

ArXiv ID: 2509.22100

Primary Topic: Architecture and Training Dynamics

Authors: Zhipu Cui, Johannes Lutzeyer

Abstract: Graph Neural Networks (GNNs) have achieved remarkable success across a range of learning tasks. However, scaling GNNs to large graphs remains a significant challenge, especially for graph-level tasks. In this work, we introduce SHAKE-GNN, a novel scalable graph-level GNN framework based on a hierarchy of Kirchhoff Forests, a class of random spanning forests used to construct stochastic multi-resolution decompositions of graphs. SHAKE-GNN produces multi-scale representations, enabling flexible trade-offs between efficiency and performance. We introduce an improved, data-driven strategy for selecting the trade-off parameter and analyse the time-complexity of SHAKE-GNN. Experimental results on multiple large-scale graph classification benchmarks demonstrate that SHAKE-GNN achieves competitive performance while offering improved scalability.

Comment: Hierarchical Kirchhoff-forest decomposition enables multiscale graph computation with controllable cost.

Topic Match: The core contribution is a hierarchical architectural mechanism, with narrower relevance through graph-level GNN scalability.

Relevance: 7 Novelty: 6


8. Training-Free Hidden-State Refinement for Flow-Matching Image Generators

ArXiv ID: 2608.29160

Primary Topic: Architecture and Training Dynamics

Authors: Yuanyi Yan, Xinzhe Rao, Canyu Shen, Yang Chen, Yunlu Chen, Meng Tang, Teng Long, Vincent Tao Hu

Abstract: We aim to improve frozen flow-matching image generators by adding inference computation inside the denoiser, without changing model weights or the outer sampler. Existing generators usually spend extra test-time computation by increasing the number of sampling steps, which repeatedly evaluates the entire denoiser and couples quality gains to sampler cost. A key challenge is how to use extra computation inside a frozen transformer denoiser: the method must decide which tokens, layers, and sampling times receive repeated updates while preserving the original generation pipeline. We introduce a training-free looping framework that repeatedly applies selected transformer layers inside each denoising call. Dense and Sparse Token Loop vary the token scope; Sampling-Progress Gating and the loop layer range specify when and where looping is active; loop count and strength control the repeated updates; and Loop Guidance combines ordinary and looped vector-field predictions. Across two Scale-RAE model scales, loop variants improve primary and auxiliary quality metrics with competitive quality--efficiency trade-offs. Loop Guidance further improves both primary metrics across all three tested models; on Scale-RAE DiT2.4B, it raises GenEval from 0.4471 to 0.5691 and DPG-Bench from 0.7656 to 0.8053. Code will be released.

Comment: Selective repetition of transformer layers allocates extra computation across tokens and denoising steps.

Topic Match: Reusing transformer blocks is a dynamic-computation mechanism, with evidence limited to inference refinement in frozen image generators.

Relevance: 7 Novelty: 6


9. Federated Personalization of Early-Exit Networks

ArXiv ID: 2601.10015

Primary Topic: Architecture and Training Dynamics

Authors: Boyi Liu, Zimu Zhou, Cheng Fang, Yongxin Tong

Abstract: Personalized Federated Learning (PFL) excels at tailoring client-specific models, which is particularly critical for decentralized and heterogeneous data environments, yet existing methods produce static models with a fixed tradeoff between accuracy and efficiency. This inherent static nature limits their ability to adapt to inference demands that vary with context and resource availability, posing a challenge for real-world deployment. Early-exit networks (EENs), which enable adaptive inference via intermediate classifiers, offer a promising solution. However, integrating EENs into PFL introduces two intertwined conflicts: client-wise heterogeneity across clients and depth-wise interference arising from conflicting exit objectives. Prior studies fail to resolve both conflicts simultaneously, leading to suboptimal performance. In this paper, we propose X-FED, a novel Conflict-Aware Cross-Client Federated Exit Distillation framework that jointly addresses both client- and depth-wise conflicts while extending PFL to early-exit networks. At its core, X-FED employs a progressive, depth-prioritized student coordination mechanism that mitigates interference among shallow and deep exits while enabling effective personalized knowledge transfer across clients. Furthermore, we introduce a client-decoupled formulation that reduces communication overhead with theoretical soundness. Extensive evaluations on various datasets show that compared to state-of-the-art PFL and PFL-EE methods, X-FED achieves higher accuracy while reducing inference costs by 30.79%-46.86%.

Comment: Uses depth-prioritized training coordination to reduce conflicting objectives across early exits.

Topic Match: Training interference in adaptive-depth networks is the nearest foundational connection; the main advance is personalized federated distillation, with large-model training applicability unestablished.

Relevance: 6 Novelty: 6


Efficiency, Compression, and Large-Scale Training (21)

1. OCGQuant: Outlier-Companion Grouping for NVFP4 Quantization

ArXiv ID: 2609.00066

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Yishan Yao, Binjun Li, Hanling Yi, Pengyu Li, Xiaoqing Liu, Zihan Yang, Xiaotian Yu, Zhiwen Yu

Abstract: NVFP4 is an efficient microscaling format for low-bit inference, but activation outliers can still degrade quantization accuracy within NVFP4 blocks. Within each quantization block, large activations can dominate the block scale, increasing the quantization error of the remaining values sharing the same scale. Existing post-training quantization (PTQ) methods mitigate outlier errors through strategies such as mixed precision, rotation, or residual compensation, but these approaches are either not specifically tailored to NVFP4 or introduce additional computation. In this work, we revisit NVFP4 from a channel-grouping perspective and define the reducible error incurred by remaining block values under the scale set by the block maximum as Collateral Quantization Error. Based on this insight, we propose OCGQuant, a post-training quantization method centered on Outlier-Companion Grouping (OCG), which adaptively pairs outlier channels with low-magnitude companion channels to improve NVFP4 activation block composition. Experiments on Llama3 and Qwen3 show that OCGQuant achieves the lowest WikiText-2 perplexity and highest average downstream accuracy among evaluated PTQ methods, while maintaining prefill speedup close to RTN and matching its peak decoding memory. Code is available at https://github.com/Eshamont/OCGQuant.

Comment: Groups activation outliers with low-magnitude companion channels to reduce shared-scale NVFP4 quantization error.

Topic Match: The contribution is a new channel-grouping mechanism for low-bit quantization that preserves practical inference speed and memory benefits.

Relevance: 9 Novelty: 7


2. TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning

ArXiv ID: 2608.03276

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Wonpyo Park, Seung-won Hwang

Abstract: Long-context inference with large language models is constrained by the linear growth of the key-value cache to sequence length. While pruning offers mitigation, prevailing methods determine query-specific token importance that cannot be reused across unseen queries. In contrast, we introduce TaskPress, a framework for task-guided, query-agnostic KV cache eviction. Instead of optimizing the cache for a single query, TaskPress constructs a reusable memory representation conditioned on a high-level task guide. The guide functions as a meta-query during prefill to filter irrelevant tokens before downstream queries are issued. In addition, TaskPress leverages quantization scale factors as a zero-cost signal for detecting influential representation outliers, providing an efficient proxy for token importance. Experiments on conducted on various tasks with long context input demonstrate that TaskPress efficiently creates a compact, reusable cache across diverse queries.

Comment: Combines task-guided, query-agnostic KV eviction with quantization-scale signals to construct reusable compressed caches.

Topic Match: The core mechanism directly reduces long-context KV memory while allowing the compressed cache to support multiple unseen queries.

Relevance: 9 Novelty: 7


3. An Efficient Sparse Fine-Tuning with Low Quantization Error via Neural Network Pruning

ArXiv ID: 2502.11439

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Cen-Jhih Li, Aditya Bhaskara

Abstract: Fine-tuning is an important step in adapting foundation models such as large language models to downstream tasks. To make this step more accessible to users with limited computational budgets, it is crucial to develop fine-tuning methods that are memory and computationally efficient. Sparse Fine-tuning (SpFT) and Low-rank adaptation (LoRA) are two frameworks that have emerged for addressing this problem and have been adopted widely in practice. In this work, we develop a new SpFT framework, based on ideas from neural network pruning. At a high level, we first identify "important" neurons/nodes using feature importance metrics from network pruning (specifically, we use the structural pruning method), and then perform fine-tuning by restricting to weights involving these neurons. Experiments on common language tasks show our method improves SpFT's memory efficiency by 20-50\% while matching the accuracy of state-of-the-art methods like LoRA's variants. Code available at: https://github.com/CenjhihLi/sparsity_finetuning

Comment: Pruning-derived neuron selection restricts trainable weights and reduces sparse fine-tuning memory by 20-50%.

Topic Match: The core contribution changes which foundation-model weights are updated to improve adaptation memory efficiency while preserving accuracy.

Relevance: 9 Novelty: 6


4. Locality-Aware Redundancy Pruning for LLM Depth Compression

ArXiv ID: 2605.27786

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Vincent-Daniel Yun, Youngrae Kim, Woosang Lim, YoungJin Heo, Minkyu Kim, Sunwoo Lee

Abstract: Large language models are known to contain representational redundancy across network depth, making depth pruning an effective approach for improving inference efficiency. Existing one-shot pruning methods rely on local layer importance or fixed redundancy assumptions across architectures. We propose Locality-Aware Redundancy Pruning (LoRP), a training-free one-shot depth pruning framework guided by representation locality. We show that inter-layer redundancy can be either localized or globally distributed depending on the LLM architecture. To characterize this phenomenon, we introduce Representation Locality Score (RLS), derived from global inter-layer hidden-state similarity. Using a small calibration set, LoRP computes pairwise layer similarity, clusters layers by representational similarity, and allocates pruning according to residual intra-cluster redundancy. Experiments across diverse LLM families show improvements in both perplexity and downstream task accuracy. Official github repository: https://github.com/daniel-eai/LoRP-Locality-Aware-Redundancy-Pruning/

Comment: Prunes LLM depth by clustering inter-layer representations and allocating removals according to residual redundancy.

Topic Match: Architecture-dependent redundancy directly guides a training-free LLM compression mechanism.

Relevance: 9 Novelty: 6


5. DASH: Dual-Branch Score Distillation for Guidance-Calibrated Compact Diffusion Models

ArXiv ID: 2606.00798

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Abdullah Al Shafi, Kazi Saeed Alam, Sk Imran Hossain, Engelbert Mephu Nguifo

Abstract: Parameter compression of class-conditional diffusion models exposes a structural limitation in output-level distillation: supervising only the guided output leaves the two score branches non-identifiable, so the classifier-free guidance gap is not determined in the student. The zero-loss set admits degenerate solutions in which both branches collapse toward identical predictions, and guidance loses its effect at inference despite low training loss. For each fixed input the objective is exactly flat in a branch-output direction at every residual value, so the ambiguity is a property of the output-level objective rather than a poor local minimum. This paper introduces DASH, which supervises the conditional and unconditional branches independently. An anchor term regularises the conditional prediction toward ground-truth noise, and the teacher's final learned per-timestep curriculum transfers into the student as a frozen prior. Across CIFAR-10, CIFAR-100, and class-conditional ImageNet-64, a more than 5x compressed student stays within four FID points of its teacher and recovers at least 89% of the teacher's guidance-gap magnitude, where no other two-branch baseline with defined calibration exceeds 82%. Ablation isolates unconditional supervision as the term that separates this formulation from output-level distillation. The account is quantitative: the null space fixes the ratio between a composite student's two branch errors, and the measured ratios approximately match that prediction on all three datasets.

Comment: Independent conditional and unconditional score supervision prevents guidance collapse during diffusion-model compression.

Topic Match: Preserving guidance in a student over five times smaller makes compression primary; the objective-null-space analysis also provides training-dynamics insight.

Relevance: 8 Novelty: 7


6. WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource Hardware

ArXiv ID: 2606.21868

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Jiamu Zhang, Liang Wu, Mayank Darbari, Liangjie Hong

Abstract: Modern local and agentic workloads often need large-model capacity at low concurrency, but run on GPUs that cannot keep a frontier-scale model resident. Mixture-of-Experts (MoE) models are a natural fit because they activate only a small subset of experts per token, but their sparsity saves computation, not residency: the full expert pool still has to be stored, and any expert used by a layer must be in GPU memory when that layer runs. Static layer-level CPU offload makes such models fit, but transfers the expert layer in bulk on every forward pass, losing much of the sparsity advantage. We view low-resource MoE serving as a working-set problem on the GPU. Routed expert weights and the KV cache are two memory-demand streams competing for the same limited VRAM. We implement this view in WiSP (Working-Set Paging), a routing-aware expert pager that plugs into an unmodified serving engine and preserves byte-identical outputs. On a real 24 GiB RTX 3090, WiSP achieves up to 2.0x the decode throughput of static offload at the same memory budget when the model does not fit. A natural next step is to predict future experts and prefetch them. We find that this does not help in single-stream decode: the bottleneck is PCIe bandwidth, not prediction quality, so speculative transfers compete with demand transfers instead of hiding them. This shifts the design question from prefetching to allocation: how should one VRAM budget be divided between resident experts and the KV cache? We answer with MV-WSA (Marginal-Value Working-Set Allocation), which splits memory by marginal latency benefit per byte while enforcing a KV-admission floor. As a startup configurator, MV-WSA is the only policy we test that stays near-best on both prefill and decode; as a live controller, it resizes both pools while serving and reduces end-to-end time by up to 1.19x over a fixed offline split, without changing model outputs.

Comment: Jointly allocates VRAM between routed expert weights and the KV cache according to marginal latency benefit per byte.

Topic Match: Routing-aware paging and joint expert/KV allocation introduce concrete memory-efficiency mechanisms for MoE inference under tight VRAM budgets.

Relevance: 8 Novelty: 7


7. A Deterministic Constant-Competitive Algorithm for Dynamic Mixture-of-Experts Serving

ArXiv ID: 2608.16947

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Ian D'Ambrosio

Abstract: Dynamic Mixture-of-Experts Serving allocates k replica GPUs among m experts as workloads change. At each round, the online algorithm sees the current workload, chooses integral replica counts, and pays bottleneck service cost plus replica movement. It does not know future workloads. Huang, Lou, and Xiao gave an O(sqrt(log k))-competitive randomized algorithm for this problem. We prove a deterministic O(1)-competitive algorithm. For every number of experts and every k>=1, the algorithm satisfies ALG_det <= 10 C_PB OPT + (5 C_PB + 8) k + 16, where C_PB is the absolute constant from Chasing Positive Bodies at resource augmentation one and covering sparsity two. Consequently, CR_det(k)<=10 C_PB for every k>=1, so CR_det(k)=Theta(1). The multiplicative factor does not depend on the number of experts, replica budget, horizon, or workload values. Thus randomization is not needed for the asymptotic guarantee. The proof has two layers. A finite tangent envelope, summable positive resets, and a nonexpansive balanced projection reduce reciprocal-max service costs to a deterministic exact-budget fractional path. A new deterministic rounding theorem converts every such path to integral allocations with service distortion three and movement bounded by the fractional movement plus 6k. The complete reduction, rounding theorem, causal composition, and quantified main theorem are machine-checked in Lean 4 relative to the positive-body result as the sole scientific source premise. The theorem concerns the allocation model above. It does not include network topology, shared-edge congestion, or routing decisions.

Comment: Proves a deterministic constant-competitive expert-replica allocation algorithm accounting for service and replica-movement costs.

Topic Match: The contribution is a new serving-efficiency algorithm and rounding theorem within an abstract resource-allocation model.

Relevance: 7 Novelty: 8


8. Towards Provable and Scalable Training of Quantized Neural Networks with Ising Optimization

ArXiv ID: 2506.18240

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Wenxin Li, Chuan Wang, Hongdong Zhu, Qi Gao, Yin Ma, Hai Wei, Kai Wen

Abstract: Training quantized neural networks remains fundamentally challenging due to non-convex loss landscapes and discrete parameter spaces. We introduce an exact Quadratic Constrained Binary Optimization (QCBO) framework with provable guarantees. We first characterize the stratified topology of network zero-loss level sets: generic interior strata are smooth, yet globally optimal components can remain disconnected even under overparameterization. To address this non-convex obstruction, we compile finite-depth architectures with parameter codebooks and Forward Interval Propagation (FIP)-bounded states into bounded QCBOs, yielding an exact completely positive convex formulation that preserves the global discrete optimum with zero relaxation gap. To overcome monolithic sample scaling, we formulate sample-wise Decomposed Lower-Bound Optimization (DLBO) to reduce each Ising call from dataset to single-sample scale. The DLBO moment hierarchy also forms a Hamiltonian-locality hierarchy, with order two giving an auxiliary-free pairwise QUBO oracle and higher orders trading interaction locality for tighter bounds. Strictly feasible discrete parameters are recovered via Spectral--ADMM and randomized rounding. Experiments on a coherent Ising machine achieve $94.95\%$ accuracy on binary Fashion-MNIST (coats vs. sandals) at 1.1-bit precision, demonstrating resilience against low-bit representational collapse. Multi-class DLBO evaluations on 3-class Fashion-MNIST, 3-class Wine, and 3-class Digits further validate scalable convergence.

Comment: Exact discrete formulations and sample-wise Ising decomposition provide an alternative method for training quantized neural networks.

Topic Match: Direct optimization of low-bit weights clearly matches quantized training; practical relevance to large models remains uncertain given the small experimental settings.

Relevance: 7 Novelty: 8


9. SupraTok: Cross-Boundary Tokenization for Enhanced Language Model Performance

ArXiv ID: 2508.11857

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Andrei-Valentin Tănase, Elena Pelican

Abstract: Tokenization remains a persistent bottleneck in language modeling, especially when vocabulary learning is limited by whitespace boundaries. We present SupraTok, a tokenizer that crosses whitespace boundaries using three modular components: optional entropy-based data curation, staged curriculum training with PMI-guided candidate search, and multilingual script handling. At 100k vocabulary on the same unfiltered training data, SupraTok improves compression over standard BPE by 17.5% and over the official SuperBPE implementation by 1.8%, while training 2.1x faster than SuperBPE. Across 50k-300k vocabularies in the same matched setting, SupraTok remains ahead of SuperBPE by 1.8%-8.6%. We evaluate entropy filtering separately as a pipeline step: at 100k vocabulary it raises SupraTok from 5.78 to 5.99 C/T, while matched controls show a smaller gain for SuperBPE and almost no change for SP-BPE-CrossBoundary. On FLORES-200 across 14 languages, SupraTok yields a macro-averaged 34.9% relative gain over the BPE baseline. In separate downstream experiments with matched compute and fixed token budgets, using 12L-768d and 24L-1024d GPT-2-style backbones with 256k vocabularies, SupraTok improves HellaSwag and MMLU. Overall, these results show that crossing whitespace boundaries gives consistent compression gains under controlled public-data comparisons, while optional entropy filtering provides a separate pipeline benefit.

Comment: Cross-boundary tokenization improves text compression, reducing token requirements for a fixed amount of training text.

Topic Match: Token compression directly affects pretraining sequence budgets, with matched-compute downstream experiments supporting its practical relevance.

Relevance: 8 Novelty: 6


10. Approximate Speculative Decoding

ArXiv ID: 2608.03447

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Yuannuo Feng, Zegang Peng, Yuxin Xie, Yubing Ye, Yizhe Chen, Wenshuai Yao, Wenyong Zhou, Wang Kang

Abstract: Speculative decoding accelerates autoregressive generation by verifying a draft block with a target model in parallel. Under standard greedy verification, decoding stops at the first draft token that differs from the target argmax, discarding the remaining target-scored suffix. Although accepting such a mismatch changes the decoding trajectory, it can make a contiguous suffix reusable when its tokens remain target-greedy under the realized prefix. In this paper, we introduce \textbf{Approximate Speculative Decoding (ASD)}, a training-free verifier that replaces binary first-mismatch truncation with budgeted longest-prefix selection. ASD accepts selected mismatches subject to a local target-logit regret gate, a per-block exception cap, and a persistent request-level regret budget, then reuses the contiguous target-greedy suffix without additional approximate decisions or target-model forward passes. ASD requires neither a new draft model nor fine-tuning, and exactly reduces to standard greedy verification when the budget is zero. Experiments show that ASD improves fixed-workload throughput by $3.05\%$--$15.26\%$ over matched strict verification and averages a $7.78\%$ gain across seven Qwen3-14B + DSpark-14B tasks. On DeepSeek-V4-Flash (284B) with DSpark it also raises verifier-side acceptance by roughly $10\%$--$16\%$ on GSM8K and MATH-500 in an FP4-to-FP8 compatibility setting. The source code is publicly available at: https://github.com/Kissmetothemoon/ASD

Comment: Regret-budgeted speculative verification accepts selected mismatches and reuses draft suffixes to improve decoding throughput.

Topic Match: A new verification rule directly reduces LLM inference overhead, making it a substantive efficiency contribution despite its inference-only scope.

Relevance: 8 Novelty: 6


11. AutoCRAT: Within-trajectory Joint Control of Stochasticity and Compute for LLM Reasoning

ArXiv ID: 2608.29988

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Hanjun Luo, Qiushi Liu, Jingya Zhang, Haihong Pang, Jiaheng Wen, Yifei Ma, Yu Yao, Chengxi Zhang, Hanrong Zhang, Yankai Chen, Hanan Salam

Abstract: Large language models (LLMs) achieve strong reasoning performance, which depends critically on inference-time decisions. Yet these decisions are commonly handled by static, one-size-fits-all policies, limiting adaptation to diverse tasks and reasoning stages. Recent adaptive methods partially address this limitation, but they primarily adapt either decoding stochasticity (how the model explores) or reasoning compute (how long the model reasons) in isolation, leaving their interaction within a single reasoning trajectory unmodeled. To address this challenge, we shift toward a within-trajectory joint control view, and instantiate it in AutoCRAT, a decoder-side controller for frozen backbones. Using only signals available during decoding, AutoCRAT jointly adjusts sampling stochasticity and reasoning budget during generation. AutoCRAT operates over a discrete action space and updates control decisions only at semantic boundaries, improving stability while remaining responsive to the evolving reasoning process. Comprehensive evaluation across 6 benchmarks demonstrates that AutoCRAT (I) uses 13.8-52.7% fewer inference tokens on average than recommended static configurations, (II) surpasses recommended static and adaptive baselines by 1.5-4.5% in relative accuracy, and (III) enjoys strong cross-backbone transferability.

Comment: Joint adaptation of decoding stochasticity and reasoning length reduces inference-token use by 13.8–52.7%.

Topic Match: The central mechanism adaptively allocates inference compute within a trajectory using a controller over frozen LLMs.

Relevance: 8 Novelty: 6


12. SpecPV: Improving Self-Speculative Decoding for Long-Context Generation via Partial Verification

ArXiv ID: 2512.02337

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Zhendong Tan, Xingjun Zhang, Chaoyi Hu, Junjie Peng, Kun Xia

Abstract: Growing demands from tasks like code generation, deep reasoning, and long-document understanding have made long-context generation a crucial capability for large language models (LLMs). Speculative decoding is one of the most direct and effective approaches for accelerating generation. It follows a draft-verify paradigm, where a lightweight draft model proposes several candidate tokens and the target model verifies them. However, we find that as the context length grows, verification becomes the dominant bottleneck. To further accelerate speculative decoding in long-context generation, we introduce SpecPV, a self-speculative decoding approach that performs fast verification using partial key-value states (KV) and periodically applies full verification to eliminate accumulated errors. We validate SpecPV across multiple long-context benchmarks and models, including LLaMA-3.1-8B-Instruct and Qwen3-series. Experimental results show that SpecPV achieves up to 6x decoding speedup over standard autoregressive decoding with minor degradation.

Comment: Reduces speculative verification cost using partial KV states and periodic full verification, with a reported small quality tradeoff.

Topic Match: Its central mechanism reduces long-context verification computation and KV access, directly addressing inference efficiency.

Relevance: 8 Novelty: 6


13. EpaCache: Error-Propagation-Aware Caching for Accelerating Diffusion-Based Visual Generation

ArXiv ID: 2608.29264

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Yuhan Liu, Zongwei Hong, Jinglun Li, Linze Li, Shen Zhang, Yao Tang

Abstract: Diffusion-based visual generative models deliver strong image and video synthesis quality but incur high inference costs because sequential samplers repeatedly evaluate large networks. Caching-based methods reduce inference latency by reusing intermediate computations across adjacent timesteps. However, existing cache controllers rely primarily on local temporal variation and overlook the trajectory-level consequences of cache reuse. We introduce Error-Propagation-Aware Cache (EpaCache), a training-free caching policy that adaptively allocates the reuse budget on timesteps with lower downstream impact. Experiments on image and video synthesis models demonstrate that EpaCache consistently improves the latency--fidelity trade-off over existing caching methods. On FLUX.1-dev, EpaCache outperforms the prior state-of-the-art caching method in both latency and fidelity, reducing inference time from $11.7$ s to $11.3$ s while improving PSNR from $21.4$ to $22.8$. On HunyuanVideo, EpaCache achieves a $2.63\times$ speedup over uncached inference and improves SSIM from $0.891$ to $0.905$ over the prior state-of-the-art method at matched latency.

Comment: Allocates diffusion cache reuse according to downstream error propagation across the sampling trajectory.

Topic Match: The core contribution is a cache controller that reduces repeated computation in large diffusion models while accounting for accumulated output error.

Relevance: 8 Novelty: 6


14. DIP: Dynamic In-Context Planner For Diffusion Language Models

ArXiv ID: 2601.03199

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Yang Li, Han Meng, Chenan Wang, Zhenyu Bi, Xuan Wang, Haipeng Chen

Abstract: Diffusion language models (DLMs) have shown strong potential for general natural language tasks with in-context examples. Existing In-Context Learning (ICL) approaches largely inherit the practice of autoregressive language models (ARLMs), incorporating all examples into a fixed prompt. However, applying this rigid, static-prompt paradigm to DLMs incurs substantial computational overhead, as the model must evaluate the maximum context length at every step. We address this inefficiency with a key discovery: the block-wise KV-cache mechanism inherent to DLM inference enables the \textit{low-cost dynamic adjustment of the context}. Following this intuition, our core idea is to start generation with a minimal prompt and progressively insert additional examples on the fly only when the generated tokens are of low confidence. Through rigorous empirical evaluations, we observe that average verified token confidence correlates strongly with generation accuracy, making it a reliable and computationally efficient signal of token quality. Formally, we propose \textbf{D}ynamic \textbf{I}n-Context \textbf{P}lanner (DIP), a context-optimization algorithm based on average verified confidence that dynamically ranks and inserts in-context examples during generation, rather than providing all examples up front. Experimental results on math and coding benchmarks with LLaDA-1.5 and LLaDA-8B-Instruct show that DIP achieves up to $1.59\times$ and $1.36\times$ speedups, respectively, while largely preserving the generation quality of the fixed-prompt baseline. Code: https://github.com/wmd3i/DIP

Comment: Confidence-triggered insertion of in-context examples reduces repeated context computation in diffusion language models.

Topic Match: Cache-aware context scheduling directly reduces LLM inference computation, with reported speedups while largely preserving generation quality.

Relevance: 8 Novelty: 6


15. Prefix-Adaptive Block Diffusion for Efficient Document Recognition

ArXiv ID: 2605.16861

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Mingxu Chai, Ziyu Shen, Chenyu Liu, Jihua Kang, Tao Gui, Qi Zhang

Abstract: Block Diffusion Models (BDMs) support parallel generation, flexible-length output, and KV caching, making them promising for efficient document parsing. However, existing BDMs bind denoising and cache commitment to fixed block boundaries: parallelism shrinks during intra-block denoising, while generated tokens cannot be cached until the whole block is completed. Moreover, intra-block bidirectional denoising conflicts with inter-block autoregression, creating inconsistent information flow that can challenge structure-sensitive recognition. We propose the Prefix-Adaptive Block Diffusion Model (PA-BDM), which replaces intra-block bidirectional denoising with causal denoising from prefix to suffix and treats the block size as a maximum candidate range rather than a fixed commitment unit. PA-BDM uses Confidence-gated Structural Loss (CSL) to build low-entropy prefixes before extending training to longer continuations. During inference, Progressive Prefix Commitment (PPC) then dynamically commits the longest reliable prefix into the KV cache and resets the next candidate range from the updated prefix, restoring a large parallel decoding space at each step. Experiments show that the 3B PA-BDM achieves higher recognition scores on several benchmarks and improves inference throughput by 71.6\% over the 2.5B MinerU-Diffusion.

Comment: Dynamically commits reliable denoised prefixes to the KV cache to restore parallel decoding capacity.

Topic Match: Adaptive prefix commitment directly changes decoding efficiency, supported by causal denoising and a prefix-focused training loss; the mechanism qualifies despite document-specific evaluation.

Relevance: 7 Novelty: 7


16. Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning

ArXiv ID: 2606.03965

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Yu Xia, Zhouhang Xie, Xin Xu, Byungkyu Kang, Prarit Lamba, Xiang Gao, Julian McAuley

Abstract: Large language models improve final-answer accuracy through extended chain-of-thought reasoning, but often spend tokens inefficiently and offer little inference-time control. Existing efficient reasoning methods control thinking length by shortening, early-stopping, or compressing traces, leaving how the model thinks implicit. In this paper, we propose Agentic Chain-of-Thought Steering (ACTS), which formulates reasoning steering as a Markov decision process where a controller agent adaptively steers a frozen reasoner during inference. At each step, the controller observes the reasoning trace and remaining thinking budget, then issues a steering action consisting of a reasoning strategy and a steering phrase that initiates the next reasoner step. This enables budget-aware strategy control for efficient reasoning while preserving the reasoner's generation continuity. We initialize the controller agent from our constructed synthetic steering trajectories with multi-budget augmentation, and further optimize it via reinforcement learning with budget-conditioned reward shaping. Experiments across multiple benchmarks show that ACTS achieves competitive accuracy with substantial token savings, and enables controllable accuracy-token trade-offs across different reasoners and tasks. The code is available at https://github.com/Andree-9/ACTS.

Comment: A budget-aware controller steers reasoning strategies to reduce a frozen LLM's generated tokens.

Topic Match: Adaptive inference-token allocation directly targets runtime efficiency, although its connection to pretraining is peripheral and controller overhead remains unquantified in the abstract.

Relevance: 7 Novelty: 7


17. Spectral Analysis for Sparse Matrix Computation: Insights and Potential

ArXiv ID: 2608.29362

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Ruifeng Zhang, Xipeng Shen

Abstract: Sparse computations are fundamental to scientific computing, graph analytics, and machine learning, yet their performance is highly sensitive to the diverse sparsity and patterns. This is because cache reuse, memory coalescing, and load balancing depend critically on the sparsity patterns. This work gives the first known exploration of the connections between sparse matrix computation and spectral analysis by treating sparse matrices as two-dimensional signals and analyzing their frequency-domain representations through Fast Fourier Transform. We show that spectral signatures uncover global structural characteristics that are not sufficiently captured by conventional spatial statistics and provide complementary information for understanding sparse computation performance. Experiments on incorporating spectral features into machine-learning-based SpMV format selection demonstrate the usefulness of such spectral analysis over a state-of-the-art spatial-only model. By uncovering the principled connections between spectral characteristics and sparse matrix computations, this work introduces a novel analytical perspective into sparse computation, and provides a new approach to enhancing the current sparse structure characterization and optimization. On pruned LLM decoding, adding spectral features improves kernel selection and yields 1.035--1.245$\times$ kernel speedups.

Comment: Uses Fourier features of sparsity patterns to improve kernel selection for pruned LLM decoding.

Topic Match: A new characterization of sparse matrix structure improves sparse computation efficiency, with measured benefits for pruned-model decoding.

Relevance: 7 Novelty: 7


18. Masked Distillation: Internalizing the Chain-of-Thought in Language Models

ArXiv ID: 2607.22629

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Durgesh Kalwar, Vardhan Palod, Subbarao Kambhampati

Abstract: Large Reasoning Models (LRMs) produce long, explicit chains of intermediate steps before generating a final answer at inference time. These intermediate traces dominate latency, memory usage, and serving cost, even though the final answer correctness is not causally related to the trace correctness and the trace length is not a reliable indicator of the problem complexity. This raises a natural question: can the computation expressed in these intermediate tokens be internalized into the parameters of a language model, enabling it to produce answers directly (or with much shorter intermediate traces)? We introduce \textit{masked distillation}, a knowledge-distillation framework in which a student LLM is trained to predict only the solution tokens conditioned on the question, while a reasoning teacher provides feedback on the student's responses after conditioning on the question and its own CoT trace. We instantiate this framework in two settings: (i) a \textit{self-distillation} setting, in which the same model serves as the teacher in thinking mode and as the student in non-thinking mode, and (ii) a \textit{dual-model} setting, in which a larger reasoning teacher supervises a separate smaller non-thinking student over the solution tokens. By treating intermediate tokens as a scaffold which reasoning models use to fit over the solution tokens, We additionally vary the length of intermediate-token scaffolding the student is supervised on, interpolating between full internalization (the student emits only the solution) and no internalization (the student emits the full trace before the answer). We evaluate the framework through controlled experiments on two reasoning domains: GSM8K (grade-school arithmetic) and Countdown (a number-puzzle search task).

Comment: Distills solution-token predictions from a CoT-conditioned teacher so students can shorten or omit emitted reasoning traces.

Topic Match: The distillation mechanism targets inference cost by compressing explicit reasoning traces, with experiments limited to two reasoning domains.

Relevance: 7 Novelty: 6


19. AdaVLA: Adaptive Step Flow Matching for Training-free Acceleration of Vision-Language-Action Models

ArXiv ID: 2608.29208

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Sunghwan Han, Youngtae Han, Youngmin Yi

Abstract: Vision-Language-Action (VLA) models, built upon Vision-Language Models (VLMs), have significantly enhanced robotic capabilities by leveraging internet-scale knowledge and multimodal reasoning. However, the intensive computational overhead of VLAs constrains on-device deployment, hindering real-time responses to environmental changes. While various acceleration techniques have been proposed, they often rely on fine-tuning or access to training datasets, which are frequently unavailable due to privacy and proprietary concerns. Moreover, although flow-matching-based VLAs have emerged as efficient alternatives to standard diffusion models, current acceleration efforts largely target VLM inference costs, failing to address the iterative ODE solving process inherent in flow matching inference. To address these limitations, we propose AdaVLA, an online, training-free adaptive framework for fast yet accurate flow-matching-based Vision-Language-Action models. We introduce a novel metric derived from the flow matching trajectory curvature to quantify action generation confidence during inference. This metric enables the dynamic reduction of inference steps and the adaptive adjustment of MLP pruning ratios through an efficiently computed importance evaluation, requiring no access to training data. Experimental results on the LIBERO benchmark using a Jetson AGX Orin device demonstrate that our method achieves $1.87\times$ and $2.24\times$ speedups for $π_{0.5}$ and X-VLA, respectively, with negligible degradation in success rates. Furthermore, we validate the robustness of our approach on real-world robotic tasks using SmolVLA.

Comment: Uses flow-trajectory curvature to adapt inference step counts and MLP pruning without retraining.

Topic Match: Adaptive computation and pruning constitute the core acceleration mechanism, with demonstrated scope limited to flow-based vision-language-action models.

Relevance: 7 Novelty: 6


20. Entropy-Aware Token Rejection for Improving Speculative Decoding

ArXiv ID: 2512.23765

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Tiancheng Su, Meicong Zhang, Guoxiu He

Abstract: Speculative decoding (SD) accelerates large language model (LLM) inference by using a lightweight draft model to propose tokens and a stronger target model to verify them. However, standard SD is mainly designed for acceleration, and its output quality is typically constrained by the target model. In this work, we propose Entropy-Aware Speculative Decoding (EASD), a lightweight and training-free extension of SD that improves reasoning quality through token-level entropy-guided rejection. EASD detects cases where both draft and target models exhibit high uncertainty while strongly overlapping in their top predictions. In such uncertain-agreement cases, EASD rejects the aligned token and resamples from the target distribution, preventing low-confidence errors from propagating. Experiments on challenging reasoning benchmarks show that EASD consistently improves accuracy over standard SD and reward-guided variants while maintaining comparable inference efficiency. Notably, EASD can surpass the standalone performance of the target model, suggesting that speculative decoding can serve not only as an acceleration method but also as an effective mechanism for improving reasoning quality. The code is available at https://github.com/ECNU-Text-Computing/EASD.

Comment: Entropy-guided rejection modifies speculative decoding to improve reasoning accuracy at comparable inference cost.

Topic Match: The contribution changes speculative token verification, directly connecting to efficient LLM decoding, although its primary gains concern output quality.

Relevance: 7 Novelty: 6


21. PruneShift: A Framework for Evaluating Decision Reliability in Structured Pruning

ArXiv ID: 2608.29765

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Hao Ye, Gaopeng Zhang

Abstract: Structured pruning uses surrogate objectives because direct task evaluation over every feasible mask is too expensive. Most evaluations report average surrogate error or rank correlation on broadly sampled masks. These summaries do not directly test the mask chosen by the surrogate. We introduce PruneShift, an evaluation framework that separates broad predictive fidelity, fidelity near selector outputs, and the quality of the selected pruning decision. We first prove that Spearman and Kendall agreement can approach one while normalized selection regret remains maximal. We then derive sufficient conditions based on uniform error, selector suboptimality, decision margin, density ratio, and comparison mass. The analysis also yields a finite pool certificate with an explicit excess cost bound. Four studies test different links in this argument. External TextbookQA confirmation is heterogeneous: 7 of 20 simultaneous intervals favor the surrogate-selected mask, 6 favor its fixed comparator, and 7 cross zero. On a fixed Natural Questions pool, strict improvement holds in one of four settings. A controlled QQP experiment supports the proposed coverage mechanism in all 16 prespecified endpoints, although the sufficient bounds are conservative. Finally, a restricted OSSCAR reconstruction study on OPT-125M shows better local than broad fidelity in 68 of 75 primary endpoints. Independent fixed-mask confirmation is inconclusive in 24 of 25 endpoints and favors the comparator in one. These results show why predictive fit, decision reliability, and pruning method quality require separate evidence.

Comment: Shows that near-perfect surrogate rank agreement can coexist with maximal pruning-selection regret.

Topic Match: Structured pruning is directly topical; the core contribution is evaluation methodology and reliability certification.

Relevance: 6 Novelty: 7


Paper Selection Prompt

System Prompt

You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.

User Prompt

Instructions

Respond in JSONL. Output exactly one JSON object per paper, one per line:

{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}

Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only: daily_hot, new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return []. - daily_hot means the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. - new_frontier means the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.

Scoring Criteria

Relevance and Novelty are independent axes. Score both from 1 to 10.

Relevance Scoring

  • 9-10: directly centered on the target foundational topics; highest when the core contribution is clearly within them.
  • 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
  • 5-6: touches the target topics, but the main contribution is elsewhere.
  • 3-4: largely outside the target topics, often application-focused or domain-specific.
  • 1-2: unrelated.

Important: Broad frontier relevance, major launch status, or daily buzz is not enough for a high Relevance score here. Those cases belong in the hotspot digest unless the paper strongly matches the specialized paper topics.

Novelty Scoring

  • 9-10: new paradigm, theory, or major methodological breakthrough.
  • 7-8: substantial methodological advance or strong new insight.
  • 5-6: meaningful but incremental extension or refinement.
  • 3-4: minor, narrow, or mostly engineering or domain-specific improvement.
  • 1-2: little originality; mainly standard application of existing methods.

Topic Registry

Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - moe_training: MoE Training - Mixture-of-Experts training: routing, load balancing, expert granularity and shared experts, training stability, MoE parallelism and communication, and MoE kernels. - training_systems: Large-Scale Training Systems and Efficiency - Large-scale training systems and efficiency: distributed training algorithms, optimizers, and communication. - architecture_training: Architecture and Training Dynamics - Core architectural or computational mechanisms, dynamic computation, and training-stability dynamics. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, memory or cache efficiency, and large-scale training systems that materially change cost or behavior.

Papers

[PAPER LIST HERE]

Relevant Topics

This is a training-side feed. The centre of gravity is Mixture-of-Experts training and the large-scale training systems it sits on; architecture and efficiency work is kept when it bears on how a large model is trained or what it costs to train.

Keep a paper when its CORE CONTRIBUTION falls in one of the four topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.

  1. MoE Training (primary) - Keep: expert routing and gating mechanisms; load balancing, with or without auxiliary losses; expert granularity, fine-grained and shared experts; MoE training stability (loss spikes, router collapse, z-loss and related fixes); expert and hybrid parallelism; all-to-all dispatch/combine and the communication schedules MoE needs; MoE kernels and grouped-GEMM style implementations; expert capacity, token dropping, and upcycling dense checkpoints into MoE. - Filter: papers that merely train on top of a MoE model without a new routing, balancing, stability, parallelism, or kernel contribution; "mixture of experts" in the classical ensemble or recommender sense.

  2. Large-Scale Training Systems - Keep: distributed training algorithms (data, tensor, pipeline, sequence, and expert parallelism, FSDP and ZeRO-style sharding); optimizers and preconditioners for large-scale pretraining; communication algorithms, overlap, and collective design; throughput and MFU work that materially changes what a training run costs or whether it converges; scaling-law work that informs how such runs are configured. - Filter: routine infrastructure tuning with no new algorithmic or systems idea.

  3. Architecture and Training Dynamics - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, dynamic or modular computation); optimisation and training-dynamics analysis that explains why large models train the way they do. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight.

  4. Efficiency and Compression - Keep: quantization, sparsity, pruning, low-rank adaptation, KV-cache and memory-efficient designs, and other work that materially changes the cost of running or training a large model, especially when the mechanism is new rather than a tuned variant. - Filter: straightforward application of standard efficiency methods, and deployment or serving engineering with no new idea.

Leave these to the hotspot digest unless the core contribution clearly falls in one of the four topics above: - major model or product releases - agent, tooling, and RAG launches - benchmark, leaderboard, or evaluation-only papers - representation-learning theory, interpretability, and feature-formation analysis - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning) - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains