This is a remedial run for missed papers from 09/08/2026 to 09/08/2026.
Results generated on 09/14/2026.
Personalized Daily ArXiv Papers 2026-09-09
| Model | Metric | Usage | Papers | ||||
|---|---|---|---|---|---|---|---|
| Prompt | Completion | Total | Total arXiv | Scanned | Relevant | ||
local agent CLI |
Tokens | not reported | not reported | not reported | 523 | 523 | 35 |
| Cost | not reported | not reported | not reported | ||||
Token counts are not reported for this run. 24 of 24 model calls succeeded, 5,740s of model wall clock.
Topic Coverage:
| Topic | Papers |
|---|---|
| MoE Training | 1 |
| Architecture and Training Dynamics | 21 |
| Efficiency, Compression, and Large-Scale Training | 13 |
Table of contents by topic:
MoE Training (1)
- Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning Authors: Haoze He, Xingyuan Ding, Xuan Jiang, Xinkai Zou, Alex Cheng, Yibo Zhao, Juncheng Billy Li, Heather Miller
Architecture and Training Dynamics (21)
-
Learning Length-Extrapolatable Recurrent Models Authors: Hanwen Jiang
-
Do Dynamic Routers Need Memory? HeRo: History-Aware Routing for Efficient LLM Inference Authors: Hongjin Lin, Wentao Wan, Keze Wang
-
HAARES Half-Split Residual Basis Routing for Deep Transformers Authors: Kehan Wang
-
Critical initialization destabilizes higher input derivatives in wide scalar-input networks Authors: Prashant Singh, Pranav Singh
-
Equivariance Breaks the Learning Rate Authors: Andrei Manolache, Mathias Niepert
-
Manifold-Aligned Generative Transport Authors: Xinyu Tian, Xiaotong Shen
-
The Dynamics of Generalization in Deep Learning Authors: Rubing Yang, Pratik Chaudhari
-
Do New Attention Mechanisms Actually Fix Attention Sinks at Million-Token Context? Authors: Sara Rizwan, Samaanah Abdus Salam
-
Noise in Diffusion Models Is a Learnable Input Authors: Shengzhi Deng, Chenqi Ye, Yanze Guo
-
Revisiting Spectral Representations in Generative Diffusion Models Authors: Yuehao Wang, Peihao Wang, Hanwen Jiang, Ziyi Yang, Qixing Huang, Zhangyang Wang
-
The Exact Time-Uniform Rate Frontier for Stochastic Gradient Descent on Smooth Convex Objectives Authors: Ruijie Li, Kang Chen, Tianyu Wang
-
Adaptive Anisotropic Attention for Axis-Structured Signals Authors: Mahir Jain, Parshva Runwal, Aditya Ray Mishra, Arvasu Kulkarni, Sandeep Singh, Siddharth Panwar
-
Curriculum Learning as Transport: Understanding Curricula with Wasserstein Geodesics Authors: Changho Shin, David Alvarez-Melis
-
Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack Authors: Sohir Maskey, Philipp Scholl, Jonas Knupp, Pit Neitemeier, Sascha Wirges
-
CDFlow: Building Invertible Layers with Circulant and Diagonal Matrices Authors: Xuchen Feng, Siyu Liao
-
Explaining f-Divergence-Based Regularization via Local Curvature and Sharpness-Aware Minimization Authors: Nour Jamoussi, Marios Kountouris
-
Target-Independent Micro-Interventions for Predicting Training Response Across Language-Model Families Authors: Zhongxuan Liu, Sicheng Zhou, Hongzhi Wang
-
Constant-Stepsize Stochastic Approximation: Finite-Time Convergence, Gaussian Approximation, and Tail Bounds Authors: Zedong Wang, Yuyang Wang, Ijay Narang, Felix Wang, Yuzhou Wang, Siva Theja Maguluri
-
DiffLUT-Net: Differentiable Training of FPGA LUT Networks with Learnable Connectivity Authors: Jiaqi Ye, Xinrui Gong, Jingcun Wang, Olga Kondrateva, Bing Li, Grace Li Zhang
-
Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training Authors: Yunpeng Xu, Kun Zheng
-
On the Effectiveness of the z-Transform Method in Quadratic Optimization Authors: Francis Bach
Efficiency, Compression, and Large-Scale Training (13)
-
KBBQ: A Predictive Noise Law and the Limits of Spectrum Flattening in FP4 Quantization Authors: Lexington Whalen, Yuki Ito, Ryo Sakamoto
-
Less is MoE: Trimming Experts in Domain-Specialist Language Models Authors: Haoze He, Xinkai Zou, Xuan Jiang, Xingyuan Ding, Ao Qu, Juncheng Billy Li, Heather Miller
-
Activation- and Influence-Aware Ranks (AIR): Function-Preserving SVD Compression for LLMs Authors: Nico Harder, Daniel Becking, Karsten Mueller, Wojciech Samek
-
Squeeze10-LLM: Squeezing LLMs' Weights by 10 Times via a Staged Mixed-Precision Quantization Method Authors: Qingcheng Zhu, Yangyang Ren, Linlin Yang, Yanjing Li, Sheng Xu, Haodong Zhu, Juan Zhang, Runqi Wang, Baochang Zhang
-
Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts Authors: Dohyeon Kim, Bedionita Soro, Sung Ju Hwang
-
Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live Authors: Hanchen Li, Runyuan He, Qiuyang Mang, Qizheng Zhang, Huanzhi Mao, Xiaokun Chen, Hangrui Zhou, Huanchen Zhang, Alvin Cheung, Joseph Gonzalez, Ion Stoica
-
Compressing Sequences in the Latent Embedding Space: $K$-Token Merging for Large Language Models Authors: Zihao Xu, John Harvill, Ziwei Fan, Yizhou Sun, Hao Ding, Hao Wang
-
ActionSplice: In-Flight Action Editing for Interactive World Models Authors: Pardis Taghavi, Tingyu Guo, Jonas Lossner, Gaurav Pandey, Reza Langari
-
SequenceO1: End-to-End Ultra-Long (100K) Sequence Modeling in Recommendation with Low-Rank Caching Authors: Lin Guan, Jia-Qi Yang, Zhishan Zhao, Jiaqi Huang, Hangyu Wang, Longbin Li, Beichuan Zhang, Haonan Jiang, Jinan Ni, Xiangyu Fan, Xiaowen Li, Ziyao Ren, Yuhang Qi, Xiaolong Zhu, Xuanyuan Luo, Qiwei Chen, Yi Cheng, Lele Yu
-
More Bang for the Buck: Improving the Inference of Large Language Models at a Fixed Budget using Reset and Discard (ReD) Authors: Sagi Meir, Tommer D. Keidar, Noam Levi, Shlomi Reuveni, Barak Hirshberg
-
CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs Authors: Nhat-Tan Bui, Varshini Elangovan, Arun Reddy Anugu, Sreyas Mohan, Wei Ye, Dilin Wang, JQ Huang, Rakesh Ranjan, Aviral Chharia, Fernando De la Torre
-
UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models Authors: Lei Xin, Bin Gu, Peize Li, Zitong Wang, Jianbo Zhao, Changjiang Jiang, Chao Huang, Xuyang Zhao, Zunhai Su, Fanhu Zeng, Zhenglun Kong
-
Encrypt What Matters: When Selective Homomorphic Inference Is Efficient Authors: Ali Backour, Juan Reyes, Jaime Punyed, Ana Onoprishvili
MoE Training (1)
1. Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning
ArXiv ID: 2604.23036
Primary Topic: MoE Training
Authors: Haoze He, Xingyuan Ding, Xuan Jiang, Xinkai Zou, Alex Cheng, Yibo Zhao, Juncheng Billy Li, Heather Miller
Abstract: Despite MoE models leading many benchmarks, supervised fine-tuning (SFT) for the MoE architectures remains difficult because its router layers are fragile. Methods such as DenseMixer and ESFT mitigate router collapse with dense mixing or auxiliary load-balancing losses, but these introduce noisy gradients that often degrade performance. In preliminary experiments, we systematically pruned experts and observed that while certain super experts are activated far more frequently, discarding less used experts still leads to notable performance degradation. This suggests that even rarely activated experts encode non-trivial knowledge useful for downstream tasks. Motivated by this, we propose an auxiliary-loss-free MoE SFT framework that combines bias-driven sparsification with always-active gated condenser experts. Rather than enforcing balanced activation across all experts, our method encourages task-relevant experts to remain active while pushing long-tailed experts toward inactivity. The condenser experts provide a persistent, learnable pathway that alleviates gradient starvation and facilitates consolidation of information that would otherwise remain fragmented across sparsely activated experts. Analysis further suggest that this design better preserves long-tailed expert information under sparse routing. Experiments on large-scale MoE models demonstrate that our approach outperforms state-of-the-art SFT baselines such as DenseMixer and ESFT, achieving average gain of 2.5%+ on both mathematical reasoning and commonsenseQA benchmarks.
Comment: Auxiliary-loss-free routing with always-active condenser experts addresses sparse-expert gradient starvation.
Topic Match: The central contribution changes MoE routing and expert learning, directly qualifying despite its supervised fine-tuning setting.
Relevance: 10 Novelty: 7
Architecture and Training Dynamics (21)
1. Learning Length-Extrapolatable Recurrent Models
ArXiv ID: 2609.09157
Primary Topic: Architecture and Training Dynamics
Authors: Hanwen Jiang
Abstract: Recurrent models provide a natural path to long-context modeling, yet models trained with backpropagation through time (BPTT) often fail beyond their training horizon. Classical analyses emphasize gradients that vanish or explode along temporal paths. However, dense per-token losses can still train a shared recurrent rule despite severe decay, showing that decay alone does not determine whether learning fails. We instead study state credit: the signal through which future losses reach earlier recurrent states before contributing to parameter updates. Accordingly, we intervene directly on state credit and propose Credit Stabilization through Time (CST). During backward propagation, CST locally rescales the state-credit signal to stabilize its norm without rotating the component being corrected, while leaving the forward computation unchanged. Because controlled synthetic tasks and real data exhibit different credit dynamics, we specialize CST to each regime. In both settings, CST improves performance beyond the training horizon, with gains observed at up to 128x the training length.
Comment: Locally rescales backward state-credit signals to stabilize recurrent learning and improve extrapolation beyond the training horizon.
Topic Match: Directly analyzes and modifies recurrent-model training dynamics, distinguishing state credit from ordinary temporal gradient decay while preserving forward computation.
Relevance: 9 Novelty: 8
2. Do Dynamic Routers Need Memory? HeRo: History-Aware Routing for Efficient LLM Inference
ArXiv ID: 2609.08189
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Hongjin Lin, Wentao Wan, Keze Wang
Abstract: Dynamic layer routing reduces the inference cost of Large Language Models (LLMs) by learning to skip layers for individual tokens. Existing methods, however, treat each routing decision as a local operation conditioned solely on the current hidden state which is a formulation that overlooks the sequential, path-dependent nature of routing across depth: earlier decisions shape the representations seen by downstream routers, and the layer-usage objective couples all decisions jointly. We propose History-Aware Routing (HeRo), a dynamic routing framework that resolves this mismatch by introducing a router memory mechanism to maintain an explicit routing state across model depth. The memory is constructed via linear attention, incrementally aggregating preceding routing scores and their induced residual updates into a compact history representation. At each routed layer, the router conditions jointly on this accumulated state and the current hidden representation to select the executed branch. Instantiated for token-wise FFN routing, HeRo trains only lightweight routers and adapters on a frozen backbone, requiring no modification to pretrained parameters. Across Llama 3.1-8B, Llama 2-7B, and Llama 2-13B, HeRo consistently achieves the highest aggregate performance retention among ten baselines. On Llama 3.1-8B, it bypasses 26.87% of model parameters while achieving 100.24% of dense model performance across seven benchmarks, and retains 97.01% while bypassing 38.82% of model parameters under a tighter computation budget. Ablation studies confirm that removing routing history consistently degrades performance, most notably on multistep reasoning and code generation, validating that explicit routing memory enables more accurate and adaptive dynamic routing than solely conditioning on hidden state.
Comment: Conditions token-wise FFN skipping on accumulated routing history to improve dynamic computation.
Topic Match: The core architectural mechanism coordinates layer-execution decisions across depth, with reduced active computation providing a second efficiency match.
Relevance: 9 Novelty: 7
3. HAARES Half-Split Residual Basis Routing for Deep Transformers
ArXiv ID: 2606.06564
Primary Topic: Architecture and Training Dynamics
Authors: Kehan Wang
Abstract: Block-level residual routing makes learned residual aggregation practical by routing over block summaries, but each summary compresses an ordered sequence of attention and MLP updates into one cumulative vector. We propose \method{}, a lightweight residual basis router that keeps the cumulative block source and adds one half-split detail basis, computed as the difference between first-half and second-half residual updates. The detail basis is RMS-matched and updated online, exposing coarse intra-block trajectory information without dense sublayer-level routing. Across OpenWebText, cross-domain character-level benchmarks, and BPE-tokenized OpenWebText, the empirical pattern is depth-dependent: gains are small or mixed at shallow depth and most reliable in 48-layer models. In the 201M 48-layer setting, \method{} improves over Block AttnRes across all three seeds, while a 453M two-seed probe shows the same direction. Ablations rule out source duplication, random signed details, fixed detail-source biases, or block-count changes alone. Cost analysis shows that the method is FLOP-light but not wall-clock-free: it adds memory and routing overhead, yet its relative arithmetic cost is amortized as width grows and earlier convergence can reduce time-to-target.
Comment: Adds an RMS-matched half-block detail basis to learned residual routing in deep Transformers.
Topic Match: Residual-path design is the central mechanism, supported by depth-dependent convergence evidence and controlled ablations.
Relevance: 9 Novelty: 6
4. Critical initialization destabilizes higher input derivatives in wide scalar-input networks
ArXiv ID: 2609.09244
Primary Topic: Architecture and Training Dynamics
Authors: Prashant Singh, Pranav Singh
Abstract: The edge-of-chaos condition preserves first-order input perturbations in wide randomly initialized networks, but physics-informed losses, score matching and derivative regularization depend on higher input derivatives. For smooth scalar-input fully connected networks, using a joint Gaussianity of the finite derivative jet that holds in the infinite-width limit at each fixed depth, we derive mean-field recursions through third order that are exact at the variance fixed point, with finite-depth corrections that decay geometrically. At criticality, the first-derivative variance is depth-invariant, whereas the second-derivative variance grows linearly whenever the activation has nonzero curvature. The resulting third-order system closes on mean-field susceptibilities. For residual networks with branch scale L^{-1/2}, we prove that every fixed finite derivative order has uniformly bounded variance under explicit regularity assumptions. Simulations verify the critical growth laws, the residual bound, and the closed recursion. The results concern initialization, not trained-network performance.
Comment: Derives higher-derivative instability at critical initialization and proves variance bounds under depth-scaled residual branches.
Topic Match: Directly analyzes initialization stability and residual scaling, although its guarantees concern scalar-input networks before training.
Relevance: 8 Novelty: 7
5. Equivariance Breaks the Learning Rate
ArXiv ID: 2609.08381
Primary Topic: Architecture and Training Dynamics
Authors: Andrei Manolache, Mathias Niepert
Abstract: Equivariant networks are commonly trained with Adam, yet recent work reports that matrix-structured optimizers such as Muon can perform better on these architectures without explaining why. We identify one source of this difference inside equivariant linear layers. Each irrep block learns a channel-mixing matrix $W_l$ shared across its $2l+1$ components, giving the expanded map $W_l \otimes I_{2l+1}$. For a single application of the layer, the gradient of $W_l$ sums $2l+1$ outer product contributions and has rank at most $2l+1$. Adam rescales stored weights individually without using the irrep boundaries, so one learning rate can produce different spectral step sizes across blocks within a layer. We address this mismatch by normalizing each block update separately, without introducing a new hyperparameter. This changes only the scale of the update, leaving Adam's moment estimates and its direction within each block unchanged. We evaluate the mechanism in a controlled $\mathrm{SO}(3)$-equivariant model with a matched dense control and in an e3nn interatomic potential model trained on rMD17 and MD22. The toy setup isolates a mismatch that grows with width while the dense control shows no corresponding growth. In the interatomic potential model, block normalization and tuning Adam's momentum coefficients independently improve performance, but neither alone matches Muon. Combined, they make Adam competitive with Muon on all datasets, indicating that blockwise step control and momentum accumulation account for much of Muon's advantage.
Comment: Irrep-wise update normalization corrects Adam's unequal spectral step sizes across equivariant blocks.
Topic Match: Mechanistically connects equivariant parameter sharing to optimizer behavior, isolating the effects of blockwise update scale and momentum.
Relevance: 8 Novelty: 7
6. Manifold-Aligned Generative Transport
ArXiv ID: 2602.19600
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Xinyu Tian, Xiaotong Shen
Abstract: Many high-dimensional datasets concentrate near a low-dimensional structure embedded in the ambient space. Generative models for such data must control off-support mass while remaining computationally practical. Diffusion models use iterative denoising at inference, whereas standard normalizing flows require invertible, dimension-preserving maps. We propose MAGT (Manifold-Aligned Generative Transport), a direct transport from a low-dimensional base distribution to the data space. Its core objective compares the data and generator-induced scores at a selected Gaussian smoothing level. A posterior identity expresses this score through a latent conditional mean, which is approximated by self-normalized importance sampling over a finite anchor set. After training, generation requires one evaluation of the transport, whose image also carries an intrinsic density with respect to manifold volume. We establish a minimax-optimal Wasserstein convergence rate for an explicitly constructed localized spline-RePU coordinate transport estimator, and treat finite-anchor approximation separately. Experiments on synthetic, image, and tabular benchmarks compare fidelity, support alignment, and sampling cost with diffusion, flow-matching, and adversarial baselines.
Comment: A score-based objective trains a direct low-dimensional generative transport requiring a single evaluation at inference.
Topic Match: The new generative training mechanism is the primary fit; single-pass sampling provides a secondary efficiency contribution, although large-model scaling remains unestablished.
Relevance: 7 Novelty: 8
7. The Dynamics of Generalization in Deep Learning
ArXiv ID: 2504.16450
Primary Topic: Architecture and Training Dynamics
Authors: Rubing Yang, Pratik Chaudhari
Abstract: We derive a differential equation that governs the evolution of the generalization gap when a model is trained by gradient descent-based methods. This differential equation is driven by two key quantities, a contraction factor that brings together trajectories corresponding to slightly different datasets, and a perturbation factor that accounts for them training on different datasets. The coupled decay of contraction and perturbation guarantees a controlled accumulation of generalization gap during training. We analyze this differential equation to show that the generalization gap is given by a quadratic form that consists of an ``effective Gram matrix'' that depends upon the training trajectory and a certain residual of the predictor at initialization. Our framework is applicable to general deep networks and smooth loss functions. In numerical experiments on different neural network architectures, datasets and sample sizes, we show that this quadratic form accurately captures the actual generalization gap. We also show how to instantiate our framework in a number of examples via analytical calculations. For example, for high-dimensional linear regression, our framework matches existing calculations of generalization gap in the literature exactly in under-parameterized, over-parameterized and critical regimes.
Comment: Derives generalization-gap dynamics from trajectory contraction and dataset perturbation during gradient-based training.
Topic Match: The central contribution explains generalization along optimization trajectories, although its implications for large-model training remain indirect.
Relevance: 7 Novelty: 8
8. Do New Attention Mechanisms Actually Fix Attention Sinks at Million-Token Context?
ArXiv ID: 2609.08574
Primary Topic: Architecture and Training Dynamics
Authors: Sara Rizwan, Samaanah Abdus Salam
Abstract: Long context language models now advertise windows of one million tokens, but two habits limit how much of that window is used. Attention heads with nothing useful to read still spend their budget on the first token, which is called the attention sink, and where a fact sits in the context changes whether the model finds it. Gated attention cut first token attention from 46.7 percent to 4.8 percent at NeurIPS 2025, and Kimi K3 pairs that idea with Kimi Delta Attention and Attention Residuals behind a one million token window, eight times past the range where these diagnostics have been reported. This paper asks whether the fix survives that jump. We build SinkProbe, a suite that measures sink mass, massive activation, position resolved recall and the recency gap, and apply it to four small models that differ only in how they mix tokens and depth. Three results follow. The training objective produces the sink, not the architecture. Gating did not reproduce its published effect at our scale. Sink mass, activations and position bias moved independently. Code, data and the measurement protocol are released at https://github.com/sararizwan7/Attention-Mechanisms-in-1M-Context-Window
Comment: Tests whether training objectives or architectural token-mixing mechanisms determine attention-sink formation.
Topic Match: Controlled architectural comparisons provide training-dynamics insight into attention sinks beyond reporting benchmark performance; the evidence is limited to small models.
Relevance: 8 Novelty: 6
9. Noise in Diffusion Models Is a Learnable Input
ArXiv ID: 2608.02575
Primary Topic: Architecture and Training Dynamics
Authors: Shengzhi Deng, Chenqi Ye, Yanze Guo
Abstract: Stochastic learning objectives are typically written as expectations over abstract random variables. Actual training, however, uses concrete random inputs that enter both the realized loss and its gradient. Structure in these inputs that is accessible to the learning system can therefore be learned and exploited. Much prior structured-noise work asks how noise should be distributed or designed; we instead ask what structure in the concrete realized randomness becomes exploitable by the learner. We develop this general view and analyze its mechanism in diffusion models: in noise prediction, clean data and realized noise jointly form the noisy input, so the model can improve prediction by learning clean-data regularities or exploiting structure introduced through the noise, and the two routes can interact. Using pseudorandom streams as controlled, reproducible instances of structured noise, we provide mechanistic evidence on MNIST and CIFAR-10: random-role ablations localize the dominant effect to diffusion noise; in a diffusion probe, structured-noise training can reduce prediction loss below the IID reference, but replacing the test noise with IID reverses this advantage; and shuffling the same values largely removes the source-dependent loss reduction. This learned dependence can also affect generation. The same view offers a unified interpretation of data-dependent noise assignment, noise-based backdoors, and temporally correlated noise in video diffusion: although these methods introduce different structures, all alter what the model can exploit through noise and its interaction with clean-data learning. Our results indicate that diffusion noise is not merely a passive stochastic perturbation, but a learnable---and therefore potentially designable---input dimension.
Comment: Identifies how structure in realized diffusion noise becomes exploitable during optimization.
Topic Match: Controlled training ablations explain how stochastic-input structure changes learned behavior and loss, although the experiments use small models and datasets.
Relevance: 7 Novelty: 7
10. Revisiting Spectral Representations in Generative Diffusion Models
ArXiv ID: 2609.08253
Primary Topic: Architecture and Training Dynamics
Authors: Yuehao Wang, Peihao Wang, Hanwen Jiang, Ziyi Yang, Qixing Huang, Zhangyang Wang
Abstract: Diffusion models have shown remarkable performance on diverse generation tasks. Recent work finds that imposing representation alignment on the hidden states of diffusion networks can both facilitate training convergence and enhance sampling quality, yet the mechanism driving this synergy remains insufficiently understood. In this paper, we investigate the connection between self-supervised spectral representation learning and diffusion generative models through a shared perspective on perturbation kernels. On the diffusion side, samples (e.g., images, videos) are produced by reversing a stochastic noise-injection process specified by Gaussian kernels; on the spectral representation side, spectral embeddings emerge from contrasting positive and negative relations induced by random perturbation kernels. Motivated by this, we propose a self-supervised spectral representation alignment method to facilitate diffusion model training. In addition, we clarify how joint spectral learning can benefit diffusion training from a geometric perspective. Furthermore, we find that the optimization of the spectral alignment objective is in an equivalent form of diffusion score distillation in the representation space. Building on these findings, we integrate a spectral regularizer into diffusion training objectives to improve the performance of diffusion models on multiple datasets. Experiments across images and 3D point clouds show consistent gains in generation quality. Code is released at https://github.com/yuehaowang/spectral-reg-diffusion.
Comment: A spectral regularizer connects diffusion-training representation alignment to representation-space score distillation.
Topic Match: The core combines a diffusion training-objective change with mechanistic analysis of representation alignment, supporting a focused training-dynamics match.
Relevance: 7 Novelty: 7
11. The Exact Time-Uniform Rate Frontier for Stochastic Gradient Descent on Smooth Convex Objectives
ArXiv ID: 2609.08537
Primary Topic: Architecture and Training Dynamics
Authors: Ruijie Li, Kang Chen, Tianyu Wang
Abstract: We study the time-uniform convergence of the raw iterate of standard stochastic gradient descent (SGD) for unconstrained smooth convex objectives. We prove that, under standard noise assumptions, the time-uniform convergence rate gets arbitrarily close to $\sqrt{\log n / n}$ but never reaches it. More specifically, we prove that for every positive, eventually nondecreasing sequence $h$ satisfying $h(n) = o(\sqrt{n})$, a bound of order $h(n)/\sqrt{n}$, holding simultaneously for all $n$ with probability at least $1-α$ and uniformly over the problem class, is achievable if and only if [ \sum_{j = 1}^{\infty} \frac{1}{h(2^j)^2} < \infty. ] The constructive sufficiency result follows from a dyadic horizon-free schedule together with an additive conditional-restart inequality. The necessity counterpart applies to every deterministic nonnegative schedule and holds even for a one-dimensional analytic smooth convex objective with Gaussian noise.
Comment: Establishes an exact time-uniform SGD convergence frontier and constructs a horizon-free learning-rate schedule achieving the admissible rates.
Topic Match: Convergence limits and learning-rate scheduling fit optimization dynamics, with direct large-model relevance limited by the smooth-convex assumptions.
Relevance: 6 Novelty: 8
12. Adaptive Anisotropic Attention for Axis-Structured Signals
ArXiv ID: 2609.08788
Primary Topic: Architecture and Training Dynamics
Authors: Mahir Jain, Parshva Runwal, Aditya Ray Mishra, Arvasu Kulkarni, Sandeep Singh, Siddharth Panwar
Abstract: Dense self-attention treats all token pairs as equally plausible before learning, an interaction-isotropic prior that can be mismatched to structured signals. For structured, low signal-to-noise ratio (SNR) signals such as EEG, dependencies are organized along the electrode and time axes, and this uniform prior exposes each token to many irrelevant interactions. We introduce Adaptive Anisotropic Attention (AAA), which splits attention into two paths: a temporal path, where each token attends to the tokens of its own electrode across time, and a spatial path, where it attends to the tokens of the other electrodes at the same time step. A small gate predicts, for every token, a convex combination of the two path outputs: two non-negative weights that sum to one. On six EEG downstream tasks, the resulting model, AXON (AXis-factorized Operator Network), improves mean balanced accuracy over a dense baseline under both linear probing and full fine-tuning. We show that both paths (temporal and spatial) are necessary and that the weighted sum beats a hard choice of one path; most of the benefit comes from the gate learning a different temporal/spatial balance at each layer of the network. Controlled audio spectrogram experiments show that axis factorization transfers beyond EEG. These results suggest that aligning attention with the natural axes of structured signals provides a useful inductive bias.
Comment: Learned gating between temporal and spatial attention paths introduces an axis-factorized attention mechanism.
Topic Match: The attention mechanism itself is the core contribution, with ablations explaining layer-dependent path mixing; its demonstrated scope remains structured EEG and audio signals.
Relevance: 7 Novelty: 6
13. Curriculum Learning as Transport: Understanding Curricula with Wasserstein Geodesics
ArXiv ID: 2609.09099
Primary Topic: Architecture and Training Dynamics
Authors: Changho Shin, David Alvarez-Melis
Abstract: Curriculum learning is governed by several coupled design choices---how difficulty is defined, how examples are ordered, how much exposure each level receives, and how quickly training moves across levels---making it hard to isolate what actually helps. We present Wasserstein curriculum paths, a simple transport-based framework that decouples these factors by representing curricula as trajectories of training distributions over discrete difficulty levels. Across a calibrated synthetic suite with 12 tasks and 33 difficulty axes, we use this framework to isolate the effects of ordering, matched exposure, endpoint smoothness, and pacing under fixed training budgets. We find that curriculum effects are strongly context-dependent: no single strategy dominates across tasks, difficulty axes, and budgets, and curricula mainly change where a fixed budget is spent most effectively. Within this framework, easy-to-hard ordering improves hard-level performance relative to exposure-matched static sampling, showing that the benefit is not explained by cumulative exposure alone. We further show that endpoint smoothness and pacing substantially affect where along the difficulty spectrum a curriculum is effective. Finally, we show that the same transport view naturally supports extensions to learned pacing through geometry and to structured difficulty spaces beyond one-dimensional orderings.
Comment: Exposure-matched experiments isolate how curriculum ordering and pacing affect learning under fixed training budgets.
Topic Match: The core contribution analyzes training dynamics through controlled curriculum paths, with relevance limited by evidence from synthetic tasks rather than large-scale pretraining.
Relevance: 7 Novelty: 6
14. Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack
ArXiv ID: 2609.08966
Primary Topic: Architecture and Training Dynamics
Authors: Sohir Maskey, Philipp Scholl, Jonas Knupp, Pit Neitemeier, Sascha Wirges
Abstract: Language-model checkpoints are commonly selected by pretraining loss or benchmark scores, assuming that the highest-scoring checkpoint will remain the best starting point for subsequent training. We show that this assumption can fail in a full 30B mixture-of-experts training pipeline. The checkpoints that perform better after the full downstream training stack also have higher solution density, i.e., retain downstream performance under local weight perturbations.
Comment: Links pretraining checkpoint quality after subsequent training to robustness under local weight perturbations.
Topic Match: Checkpoint geometry and its association with subsequent training outcomes fit training dynamics; the 30B MoE supplies the experimental setting.
Relevance: 7 Novelty: 6
15. CDFlow: Building Invertible Layers with Circulant and Diagonal Matrices
ArXiv ID: 2510.25323
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Xuchen Feng, Siyu Liao
Abstract: Normalizing flows are deep generative models that enable efficient likelihood estimation and sampling through invertible transformations. A key challenge is to design linear layers that enhance expressiveness while maintaining efficient computation of the Jacobian determinant and inverse. We introduce a novel invertible linear layer based on the product of circulant and diagonal matrices. This decomposition reduces parameter complexity from $\mathcal{O}(n^2)$ to $\mathcal{O}(mn)$ using $m$ diagonal matrices and $m-1$ circulant matrices while still approximating general linear transformations. By leveraging the Fast Fourier Transform, our approach reduces the time complexity of matrix inversion from $\mathcal{O}(n^3)$ to $\mathcal{O}(mn\log n)$ and that of computing the log-determinant from $\mathcal{O}(n^3)$ to $\mathcal{O}(mn)$, where $n$ is the input dimension. We build upon this layer to develop Circulant-Diagonal Flow (CDFlow), which achieves strong density estimation on natural image datasets and effectively models data with inherent periodic structure. Furthermore, CDFlow significantly accelerates key operations in normalizing flows, providing practical benefits for scalable generative modeling.
Comment: Circulant and diagonal factorizations create invertible layers with cheaper inverses and log-determinants.
Topic Match: The core contribution is a structured invertible architectural layer, with direct parameter and computational savings.
Relevance: 7 Novelty: 6
16. Explaining f-Divergence-Based Regularization via Local Curvature and Sharpness-Aware Minimization
ArXiv ID: 2609.09367
Primary Topic: Architecture and Training Dynamics
Authors: Nour Jamoussi, Marios Kountouris
Abstract: Divergence-based regularization and Sharpness-Aware Minimization (SAM) are two prominent approaches for improving generalization in deep learning, both motivated by robustness to perturbations. However, their relationship has remained largely unexplored. Building on classical second-order expansions of $f$-divergences, we show that the two methods are locally consistent under parameter-space perturbations: both induce curvature-sensitive penalties, with divergence regularization yielding a Fisher-weighted quadratic form and SAM penalizing sharpness through the dominant Hessian eigenvalue. For negative log-likelihood objectives with exponential-family output distributions, this correspondence becomes especially transparent, since the Fisher and Gauss-Newton matrices coincide. We further show that the same local geometric perspective extends to input-space perturbations, where divergence-based regularization is defined through transformations of the input. In this setting, the regularizer induces a pullback quadratic form on the input space, providing a more general perturbation framework than standard SAM while preserving the same local sensitivity interpretation. To validate the analysis empirically, we use the asymmetric $α$-skew Jensen-Shannon divergence (JSD) family as a controlled testbed. Its local curvature coefficient scales as $α(1-α)$ and is maximized at the symmetric point $α=\tfrac12$, which recovers the standard JSD. Loss-landscape visualizations in the input-perturbation regime show that stronger induced curvature penalization is associated with flatter local minima. Experiments on four benchmark datasets further demonstrate that both accuracy and negative log-likelihood are consistently best near this regime of maximal curvature penalization.
Comment: Connects divergence regularization to SAM through local Fisher and Hessian curvature penalties.
Topic Match: The main contribution explains optimization geometry and regularization behavior, providing a focused training-dynamics insight.
Relevance: 7 Novelty: 6
17. Target-Independent Micro-Interventions for Predicting Training Response Across Language-Model Families
ArXiv ID: 2609.08618
Primary Topic: Architecture and Training Dynamics
Authors: Zhongxuan Liu, Sicheng Zhou, Hongzhi Wang
Abstract: Benchmark scores describe what a checkpoint can do now, but they do not determine how it will respond to the next training episode. We measure this missing state by branching four short, standardized, target-independent micro-interventions from the same checkpoint and recording their effects in a common capability space. Together with current capability, these responses form L-State; its pulse block supports a flexible direct readout and a structure-preserving operator readout. Under smooth local dynamics, the operator construction admits an end-to-end cross-family bound with explicit source- and target-family coordinate heterogeneity. In three-family leave-one-family-out development, both pulse readouts reduce source-standardized MSE by 39.4% relative to capability alone, while separating the best response and direction estimates. On sealed GLM-4-9B, the direct and operator readouts reduce MSE by 71.8% and 78.3%, respectively, and the operator readout raises sign balanced accuracy from 0.366 to 0.754. On sealed Granite-3.1-8B, the direct readout reaches RMSE 0.544 and a development-fitted action-wise selector reaches 0.554, compared with 1.172 for capability alone. A five-family audit finds that the operator coordinate varies by action and family, and that modeling these deviations improves retrospective held-trajectory prediction. Target-independent interventions therefore expose training-response information that current capability misses, with direct and structured readouts covering complementary transfer regimes.
Comment: Uses standardized short training interventions to predict how checkpoints respond to subsequent training.
Topic Match: Checkpoint-dependent training response is adjacent to training dynamics, but the contribution primarily measures and forecasts responses across model families rather than explaining or changing a training mechanism.
Relevance: 6 Novelty: 7
18. Constant-Stepsize Stochastic Approximation: Finite-Time Convergence, Gaussian Approximation, and Tail Bounds
ArXiv ID: 2602.13960
Primary Topic: Architecture and Training Dynamics
Authors: Zedong Wang, Yuyang Wang, Ijay Narang, Felix Wang, Yuzhou Wang, Siva Theja Maguluri
Abstract: Constant-stepsize stochastic approximation (SA) is widely used in learning for computational efficiency, yet the distribution of the iterates is typically intractable. Classical asymptotics results give $X_k^{(α)} \approx X^{(α)} \approx x^\star+\sqrtαY$, where $X^{(α)}$ is the steady state and $Y$ is an appropriate Gaussian limit, by progressively taking the time $k\uparrow\infty$ and stepsize $α\downarrow0$. Such limit results, however, do not quantify finite-time, finite-stepsize errors. We develop an explicit pre-limit characterization for SA with i.i.d.\ and Markovian noise. We establish existence and uniqueness of the stationary law, a geometric Wasserstein convergence to stationarity, and almost-sure and $L^3$ convergence of the steady state to the root $x^\star$, identifying the scale $\sqrtα$ as first-order fluctuation. At this scale, we derive a higher-order quantitative Gaussian approximation with a Wasserstein error, using Stein's method and Poisson equation techniques. We further obtain non-uniform Berry--Esseen-type tail bounds, incorporating both steady-state approximation and finite-time convergence errors. We instantiate the theory for strongly convex smooth SGD, linear SA, and nonlinear contractive SA. Beyond strong convexity, for general convex SGD, we identify a Gibbs limiting law and prove a pre-limit Wasserstein approximation error under stability and Stein-equation hypothesis, which are validated numerically.
Comment: Quantifies constant-step SGD convergence, stationary fluctuations, and finite-time tail probabilities.
Topic Match: The connection is optimization dynamics under convexity or contraction assumptions; implications for large-model training remain indirect.
Relevance: 6 Novelty: 7
19. DiffLUT-Net: Differentiable Training of FPGA LUT Networks with Learnable Connectivity
ArXiv ID: 2609.09254
Primary Topic: Architecture and Training Dynamics
Authors: Jiaqi Ye, Xinrui Gong, Jingcun Wang, Olga Kondrateva, Bing Li, Grace Li Zhang
Abstract: Field-programmable gate arrays (FPGAs) enable efficient neural-network inference, but most deployment flows either accelerate multiply-accumulate operations or convert pretrained quantized models into lookup tables (LUTs). We present DiffLUT-Net, an FPGA-native network connected by six-input LUTs that are trained from scratch. We jointly learn the 64 truth-table entries of a LUT and the source to each of its six input ports using a differentiable LUT function relaxation and hardware source selection. After training, the truth tables and connections are discretized, unused logic can be pruned, and the network is exported directly as synthesizable Verilog. Across five benchmarks, DiffLUT-Net achieves favorable accuracy-resource trade-offs. These results demonstrate the effectiveness of jointly learning LUT functions and sparse connectivity for compact FPGA-native inference. The code is available at https://github.com/TUDa-HWAI/DiffLUT-Network.
Comment: Jointly learns LUT truth tables and sparse connectivity through differentiable relaxation, introducing a hardware-native computational architecture.
Topic Match: Learning both computational primitives and wiring constitutes an architectural mechanism, although its demonstrated benefits concern compact FPGA inference and its large-model training relevance remains unestablished.
Relevance: 6 Novelty: 7
20. Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training
ArXiv ID: 2609.09081
Primary Topic: Architecture and Training Dynamics
Authors: Yunpeng Xu, Kun Zheng
Abstract: Mid-training, the stage between pre-training and alignment, is where a model's per-domain data composition is typically set by data availability rather than principled design. We ask what that decision buys, and whether a later alignment pass can undo it. In a controlled logical-reasoning setting (Qwen3-8B-Base, with a 4B replication; five semantically rule-disjoint KOR-Bench domains) we train 30 allocations spanning the five-domain simplex, 24 sweep configurations plus six withheld from the fit, at five seeds each. Three findings emerge. First, every domain has an interior coverage optimum: the moderate band ($10\%$-$40\%$) is best for all five domains, and a calibrated permutation test for quadratic interiority gives $P\approx0.010$; the fitted mid-training-only curves, with 8B peaks between $9.9\%$ and $35.1\%$, reproduce for curve shape but not peak location. Second, the gaps survive a fixed-budget alignment pass: compensatory SFT raises 116/120 cells (mean $+4.32\%$) yet bridges $0/240$ pairs at a $5\%$ threshold and $30/240$ at a $10\%$ ratio, an equal-budget uniform control behaves almost identically, and a permutation null would bridge $13.8\pm3.3$ and $77.9\pm8.5$ pairs ($P<0.001$). Third, zero coverage collapses mid-training-only accuracy, though a FineWeb-Edu-only control shows the collapse is commingled with generic drift. An exploratory $θ^*$ allocation attains the largest full-pipeline gain ($+4.36\%$ vs. $+0.80\%$/$+0.64\%$\,pp) but is marginal under Welch test.
Comment: Controlled data-mixture sweeps identify per-domain mid-training coverage optima and persistent effects of allocation choices.
Topic Match: Sensitivity to training-data composition provides a limited training-dynamics match; the findings are confined to a controlled logical-reasoning setting.
Relevance: 6 Novelty: 6
21. On the Effectiveness of the z-Transform Method in Quadratic Optimization
ArXiv ID: 2507.03404
Primary Topic: Architecture and Training Dynamics
Authors: Francis Bach
Abstract: The z-transform of a sequence is a classical tool used within signal processing, control theory, computer science, and electrical engineering. It allows for studying sequences from their generating functions, with many operations that can be equivalently defined on the original sequence and its $z$-transform. In particular, the z-transform method focuses on asymptotic behaviors and allows the use of Taylor expansions. We present a sequence of results of increasing significance and difficulty for linear models and optimization algorithms, demonstrating the effectiveness and versatility of the z-transform method in deriving new asymptotic results. Starting from the simplest gradient descent iterations in an infinite-dimensional Hilbert space, we show how the spectral dimension characterizes the convergence behavior. We then extend the analysis to Nesterov acceleration, averaging techniques, and stochastic gradient descent.
Comment: Uses spectral dimension and z-transforms to characterize convergence of gradient descent, Nesterov acceleration, and SGD.
Topic Match: Convergence analysis provides a training-dynamics connection, although the results concern quadratic optimization and linear models.
Relevance: 6 Novelty: 6
Efficiency, Compression, and Large-Scale Training (13)
1. KBBQ: A Predictive Noise Law and the Limits of Spectrum Flattening in FP4 Quantization
ArXiv ID: 2609.08135
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Lexington Whalen, Yuki Ito, Ryo Sakamoto
Abstract: We develop a second-order theory of quantization noise in matrix multiplication in which the quantization format is characterized by the variance it assigns to each element. The constant variance profile of integer quantization recovers existing integer-noise theory, while the multiplicative profile of floating-point rounding reduces the data dependence to a scalar, the participation factor $κ$, yielding a closed-form signal-to-noise-ratio law. The resulting functional also admits a closed-form upper bound $κ^{*}$ that no function-preserving linear transform can exceed and that is attained by a recent state-of-the-art method. Building on this analysis, we introduce KBBQ (\textbf{K}appa-\textbf{B}raked \textbf{B}lockwise \textbf{Q}uantization), which parameterizes the extent to which a transform approaches this ceiling. At W4A4, across four base models and two FP4 formats, KBBQ outperforms the prior state of the art without additional deployment-time computation.
Comment: A predictive FP4 quantization-noise law guides W4A4 transforms without adding deployment-time computation.
Topic Match: Quantization-error theory, a bound on function-preserving transforms, and the resulting low-bit quantization method directly match compression and computational efficiency.
Relevance: 9 Novelty: 8
2. Less is MoE: Trimming Experts in Domain-Specialist Language Models
ArXiv ID: 2606.05538
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Haoze He, Xinkai Zou, Xuan Jiang, Xingyuan Ding, Ao Qu, Juncheng Billy Li, Heather Miller
Abstract: Mixture-of-Experts (MoE) models achieve strong performance through conditional computation, but their large parameter footprint poses deployment challenges. Prior MoE compression approaches catastrophically fail when evaluated on general-purpose benchmarks beyond commonsense reasoning. We trace this failure to the granularity of compression: important capabilities are distributed across experts but concentrated in FFN sparse intermediate dimensions. To identify these dimensions, we use Fisher importance which outperforms activation-, router-score-, and magnitude-based alternatives, and identifies tiny sets of task-critical dimensions: in Qwen1.5-MoE, removing as few as 12 of 1.35M routed-FFN intermediate dimensions collapses GSM8K accuracy while largely preserving factual-knowledge performance. Building on this, we propose Fisher-MoE, which operates within FFN to remove intermediate dimensions ranked by Fisher importance. At the same 50% MoE compression ratio, Fisher-MoE preserves model capability, while reducing weight memory by ~45% and improving inference throughput by 21%. These findings suggest intermediate dimension granularity is an effective unit for both compression and ranking where capability concentrates in MoE models.
Comment: Fisher-ranked pruning of MoE FFN intermediate dimensions reportedly reduces weight memory by about 45% while preserving capabilities.
Topic Match: The core contribution is fine-grained MoE compression with measured memory and throughput benefits; the appropriate category is efficiency.
Relevance: 9 Novelty: 7
3. Activation- and Influence-Aware Ranks (AIR): Function-Preserving SVD Compression for LLMs
ArXiv ID: 2606.19993
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Nico Harder, Daniel Becking, Karsten Mueller, Wojciech Samek
Abstract: We present Activation- and Influence-Aware Ranks (AIR), an SVD-based LLM compression framework that guides each weight matrix's low-rank approximation with a backward-signal influence metric. Starting from the activation-aware optimum of SVD-LLM(W), AIR runs a single closed-form alternating least squares (ALS) sweep that integrates influence element-wise under a monotone-descent guarantee. AIR is layer-local and composes orthogonally with end-to-end methods: alone it exceeds ACIP, and AIR+LoRA outperforms it further. AIR improves perplexity over SVD-LLM(W) by >18% at <=60% parameter retention, matches its quality with ~90% less calibration data, and turns parameter savings into FLOP, peak-memory, and per-token latency gains.
Comment: Adds backward-signal influence to activation-aware SVD compression through one ALS sweep with a monotone-descent guarantee.
Topic Match: Directly improves LLM weight compression and translates parameter reduction into FLOP, peak-memory, and per-token latency savings.
Relevance: 9 Novelty: 7
4. Squeeze10-LLM: Squeezing LLMs' Weights by 10 Times via a Staged Mixed-Precision Quantization Method
ArXiv ID: 2507.18073
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Qingcheng Zhu, Yangyang Ren, Linlin Yang, Yanjing Li, Sheng Xu, Haodong Zhu, Juan Zhang, Runqi Wang, Baochang Zhang
Abstract: Deploying large language models (LLMs) is challenging due to their massive parameters and high computational costs. Ultra low-bit quantization can significantly reduce storage and accelerate inference, but extreme compression (i.e., mean bit-width <= 2) often leads to severe performance degradation. To address this, we propose Squeeze10-LLM, effectively "squeezing" 16-bit LLMs' weights by 10 times. Specifically, Squeeze10-LLM is a staged mixed-precision post-training quantization (PTQ) framework and achieves an average of 1.6 bits per weight by quantizing 80% of the weights to 1 bit and 20% to 4 bits. We introduce Squeeze10LLM with two key innovations: Post-Binarization Activation Robustness (PBAR) and Full Information Activation Supervision (FIAS). PBAR is a refined weight significance metric that accounts for the impact of quantization on activations, improving accuracy in low-bit settings. FIAS is a strategy that preserves full activation information during quantization to mitigate cumulative error propagation across layers. Experiments on LLaMA and LLaMA2 show that Squeeze10-LLM achieves state-of-the-art performance for sub-2bit weight-only quantization, improving average accuracy from 43% to 56% on six zero-shot classification tasks--a significant boost over existing PTQ methods.
Comment: Activation-aware mixed-precision quantization reaches 1.6 bits per weight while addressing cumulative activation error.
Topic Match: Ultra-low-bit weight compression is the central contribution, supported by new weight-significance and activation-supervision refinements.
Relevance: 9 Novelty: 6
5. Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts
ArXiv ID: 2609.09241
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Dohyeon Kim, Bedionita Soro, Sung Ju Hwang
Abstract: Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling model capacity while preserving efficient inference in large foundation models. However, most MoE models use a fixed top-$k$ expert selection policy, assigning the same expert budget to every token even when fewer experts may be sufficient. Inference-time dynamic top-$k$ routing can reduce computation without retraining, but existing methods often overlook the distributional shift caused by deviating from the training-time routing configuration. We show that reducing the number of activated experts consistently increases the RMS scale and variance of SMoE outputs, inducing a representation mismatch that contributes to downstream performance degradation in addition to the loss of expert capacity. To address this correctable component, we propose Layer-wise Distribution Alignment (LDA), a lightweight inference-time correction that uses layer-wise calibration statistics to align reduced-routing representations with the default configuration. Across multiple SMoE LLMs, benchmarks, and routing strategies, LDA recovers much of the performance lost induced by the distributional shift under reduced routing while preserving sparse-inference efficiency with negligible overhead.
Comment: Calibrates layer outputs to correct the scale shift introduced by reducing active MoE experts.
Topic Match: The correction preserves quality under cheaper expert routing at inference with negligible overhead, making efficiency the primary contribution.
Relevance: 8 Novelty: 7
6. Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
ArXiv ID: 2511.02230
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Hanchen Li, Runyuan He, Qiuyang Mang, Qizheng Zhang, Huanzhi Mao, Xiaokun Chen, Hangrui Zhou, Huanchen Zhang, Alvin Cheung, Joseph Gonzalez, Ion Stoica
Abstract: KV cache management is essential for efficient LLM inference. To maximize utilization, existing inference engines evict finished requests' KV cache if new requests are waiting. This policy breaks for agentic workloads, which interleave LLM calls with tools, introducing pauses that prevent effective KV reuse across turns. Since many tool calls have much shorter durations than human response multi-turn chatbot, it would be promising to retain the KV cache in during these tools. However, many challenges remain. First, we need to consider both the potential cost of recomputation or reloading (if offloading enabled) as well as the increasing queueing delays after eviction from GPU. Second, due to the internal variance of tool call durations, the method needs to remain robust under limited predictability of tool call durations. We present Continnum, a serving system to optimize job completion time for multi-turn agent workloads by introducing time-to-live mechanism for KV cache retention. For requests that generate tool calls, Continnum selectively pins the KV cache in GPU memory with a time-to-live value determined by the reload cost and potential queueing delay induced by eviction. When the TTL expires, the KV cache can be automatically evicted to free up GPU memory, providing robust performance under edge cases. When combined with program-level first-come-first-serve, Continnum preserves multi-turn continuity, and reduces delay for agentic workflows. Evaluations on real-world agents (SWE-Bench, BFCL, OpenHand) with Llama-3.1 8B/70B, Gemma-3 12B, and GLM-4.5 355B shows that Continnum improves the average job completion times by over 8x while improving throughput.
Comment: Introduces cost-aware KV-cache retention deadlines that account for reloading and eviction-induced queueing.
Topic Match: The core contribution is a cache-retention scheduling mechanism that reduces recomputation and completion time under uncertain tool-call durations.
Relevance: 8 Novelty: 6
7. Compressing Sequences in the Latent Embedding Space: $K$-Token Merging for Large Language Models
ArXiv ID: 2604.15153
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Zihao Xu, John Harvill, Ziwei Fan, Yizhou Sun, Hao Ding, Hao Wang
Abstract: Large Language Models (LLMs) incur significant computational and memory costs when processing long prompts, as full self-attention scales quadratically with input length. Token compression aims to address this challenge by reducing the number of tokens representing inputs. However, existing prompt-compression approaches primarily operate in token space and overlook inefficiencies in the latent embedding space. In this paper, we propose K-Token Merging, a latent-space compression framework that merges each contiguous block of K token embeddings into a single embedding via a lightweight encoder. The compressed sequence is processed by a LoRA-adapted LLM, while generation remains in the original vocabulary. Experiments on structural reasoning (Textualized Tree), sentiment classification (Amazon Reviews), and code editing (CommitPackFT) show that K-Token Merging lies on the Pareto frontier of performance vs. compression, achieving up to 75% input length reduction with minimal performance degradation. Code is available at https://github.com/shsjxzh/K-Token-Merging.
Comment: Merges contiguous token embeddings into learned latent tokens to shorten the sequence processed by an LLM.
Topic Match: Latent token compression directly reduces sequence-dependent attention and memory costs.
Relevance: 8 Novelty: 6
8. ActionSplice: In-Flight Action Editing for Interactive World Models
ArXiv ID: 2609.08230
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Pardis Taghavi, Tingyu Guo, Jonas Lossner, Gaurav Pandey, Reza Langari
Abstract: Chunk-autoregressive video world models typically condition each generated chunk on one action. An action received during sampling must therefore wait for the next chunk, condition future solver evaluations on a state produced under the previous action, or trigger rollback that repeats completed evaluations. We introduce ActionSplice, an inference framework that formulates this problem as Counterfactual State Transport (CST). A lightweight corrector transports the interrupted backbone-native representation toward the matched state induced by the revised action at the same solver step. The world model and sampler remain frozen, and sampling resumes without replaying completed evaluations. The retargeting variant $\mathrm{CST}{R}$ updates the entire active chunk, while the temporal-splicing variant $\mathrm{CST}}$ preserves a temporal prefix and updates only the suffix. Across minWM-Wan Action2V and HY-WM1.5, $\mathrm{CST{R}$ reduces rollback-relative LPIPS by 61.5% and 75.9% relative to direct condition swapping. $\mathrm{CST}$ obtains a PSNR of 25.66 dB, an SSIM of 0.6902, and an LPIPS of 0.1337 against the original rollout.}$ reduces suffix LPIPS by 56.1% and 77.5%, respectively, while providing $2.73\times$ and $1.69\times$ pixel-ready speedups over waiting. Under the HY-WorldPlay protocol, $\mathrm{CST}_{R
Comment: Transports an interrupted diffusion state to accommodate revised actions while reusing completed solver evaluations.
Topic Match: A new mechanism for avoiding recomputation directly improves inference efficiency in large video generators, although its usefulness is specific to interactive action changes.
Relevance: 7 Novelty: 7
9. SequenceO1: End-to-End Ultra-Long (100K) Sequence Modeling in Recommendation with Low-Rank Caching
ArXiv ID: 2609.08443
Primary Topic: Efficiency, Compression, and Large-Scale Training
Also Matches: Architecture and Training Dynamics
Authors: Lin Guan, Jia-Qi Yang, Zhishan Zhao, Jiaqi Huang, Hangyu Wang, Longbin Li, Beichuan Zhang, Haonan Jiang, Jinan Ni, Xiangyu Fan, Xiaowen Li, Ziyao Ren, Yuhang Qi, Xiaolong Zhu, Xuanyuan Luo, Qiwei Chen, Yi Cheng, Lele Yu
Abstract: Modeling long-term user behavior is central to sequential recommendation and billion-scale industrial recommender systems, yet production ranking models operate under strict latency, memory, communication, and training-throughput constraints. At the 100K scale, the challenge extends beyond attention complexity: raw sequence features must be stored, transferred, and repeatedly processed during training and online serving. Existing approaches based on history truncation, multi-stage behavior retrieval, compressed lifelong histories, or train-short/infer-long extrapolation either weaken end-to-end optimization or retain substantial length-dependent cost. We present SequenceO1, an end-to-end framework for ultra-long user behavior sequence modeling, deployed at full traffic on Douyin with histories of up to 100K interactions. SequenceO1 follows a compress-then-reason design. Its Sketch Attention (SA) uses learnable prototypes and prototype-wise normalization to compress the raw history into a fixed-size, target-agnostic user representation. Target-conditioned Stacked Target-to-History Cross Attention (STCA) then models complementary time scales: a recent 10K suffix for short-term interests and the compact sketch for long-term preferences. To make training and inference practical, SequenceO1 combines low-rank user representation caching, multi-request user-level batching, pipeline lift, and a fused FlashSA kernel to amortize feature storage, communication, and computation across targets, training instances, and consecutive requests. Production experiments show consistent offline and online gains, while the compact cached sketch retains most of the benefit of directly scaling end-to-end sequence ranking to 100K. These results provide a practical model-system approach to efficient attention, sequence compression, and scalable long-sequence and long-context recommendation systems.
Comment: Low-rank caching of prototype-attention sketches amortizes training and inference over 100K-interaction histories.
Topic Match: Sequence compression and cached computation drive the efficiency match, with sketch attention adding an architectural contribution; the reuse strategy is recommendation-specific.
Relevance: 7 Novelty: 7
10. More Bang for the Buck: Improving the Inference of Large Language Models at a Fixed Budget using Reset and Discard (ReD)
ArXiv ID: 2601.21522
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Sagi Meir, Tommer D. Keidar, Noam Levi, Shlomi Reuveni, Barak Hirshberg
Abstract: The performance of large language models (LLMs) on verifiable tasks is usually measured by pass@k, the probability of answering a question correctly at least once in k trials. At a fixed budget, a more suitable metric is coverage@cost, the average number of unique questions answered as a function of the total number of attempts. We connect the two metrics and show that the empirically-observed power-law behavior in pass@k leads to a sublinear growth of the coverage@cost (diminishing returns). To solve this problem, we propose Reset-and-Discard (ReD), a query method of LLMs that increases coverage@cost for a given budget, regardless of the pass@k form. Moreover, given a pass@k, we can quantitatively predict the savings in the total number of attempts using ReD. If pass@k is not available for the model, ReD can infer its power-law exponent. Experiments on three LLMs across coding (HumanEval), math (GSM8K), and reasoning (MMLU-Pro) benchmarks demonstrate that ReD substantially reduces the required attempts, tokens, and USD cost to reach a desired coverage, while also offering an efficient way to measure inference power-laws. ReD's advantage is maintained for imperfect verifiers and outperforms the tested allocation baselines.
Comment: Reduces repeated-sampling inference costs through reset-and-discard query allocation with analytically predicted savings.
Topic Match: The core algorithm improves inference-budget efficiency across questions, a substantive but peripheral match to this training-focused feed.
Relevance: 7 Novelty: 7
11. CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs
ArXiv ID: 2609.08345
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Nhat-Tan Bui, Varshini Elangovan, Arun Reddy Anugu, Sreyas Mohan, Wei Ye, Dilin Wang, JQ Huang, Rakesh Ranjan, Aviral Chharia, Fernando De la Torre
Abstract: Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-training, sidestepping the scarcity of annotated 3D data. However, it produces thousands of redundant visual tokens whose cost grows with every view. Existing visual token pruners fall into two families, each limited in the 3D multi-view setting. Learned importance methods rank tokens by attention or encoder features; because redundancy here is fundamentally spatial, they keep near-duplicate tokens from a few prominent regions and leave most of the scene unrepresented. Voxelization methods improve spatial coverage but cannot enforce an exact token budget and saturate as multi-view observations overlap in 3D, capping retention well below the target. We show that spatial coverage is associated with 3D reasoning performance and introduce CoVeR, a deterministic, training-free selector that uses only token coordinates, with no learned signals. CoVeR selects tokens that collectively cover every region of the scene, and solves the limitations of both families: it enforces an exact per-scene budget, breaks the voxelization saturation plateau, and avoids the near-duplicate selections of learned importance. Extensive experiments show CoVeR outperforms prior SOTAs on all three 3D reasoning benchmarks and generalizes as a plug-and-play module tested across four VLMs. Notably, with only $\approx$8% of visual tokens, it preserves 93.5% of full-token performance, surpassing SOTA by 3.9 percentage points on average across benchmarks.
Comment: Introduces coordinate-based coverage selection that enforces exact visual-token budgets.
Topic Match: The core contribution is a token-pruning mechanism that reduces VLM computation, although its applicability depends on multi-view 3D geometry.
Relevance: 7 Novelty: 6
12. UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models
ArXiv ID: 2608.08627
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Lei Xin, Bin Gu, Peize Li, Zitong Wang, Jianbo Zhao, Changjiang Jiang, Chao Huang, Xuyang Zhao, Zunhai Su, Fanhu Zeng, Zhenglun Kong
Abstract: Sparse mixture-of-experts (MoE) layers expand recommendation capacity through conditional computation, yet a trained checkpoint still stores and routes over its full expert bank. We study a deployment problem: convert that checkpoint to a smaller standard MoE under an explicit expert budget, without adding a compression-specific online module. To address this, we introduce UniMoMo, a post-training compression framework formulated as a constrained graph coarsening problem. Rather than relying on parameter distance, UniMoMo groups experts based on their functional similarity, using an unlabeled calibration set to measure how similarly experts respond to shared recommendation states. To prevent performance degradation, we introduce a layer-adaptive protection mechanism that restricts the merging of high-traffic experts based on their routing exposure. Across Amazon Beauty, KuaiRec, and TenRec with 2, 4, and 6 MoE blocks, the final four-expert checkpoints obtain source-relative five-run mean NDCG@10 ratios of 99.92%--102.30% and measured A100 speedups of 1.28$\times$--1.63$\times$. An aggressive two-expert, top-1 operating point obtains ratios of 98.36%--104.24% and speedups of 1.47$\times$--2.21$\times$. These endpoint results evaluate the complete conversion-and-adaptation workflow and show that a trained recommendation MoE can be exported at multiple serving budgets.
Comment: Compresses expert banks through functional-similarity graph coarsening while protecting experts with high routing exposure.
Topic Match: The core contribution is model compression through expert merging; recommendation-only validation and its post-training deployment focus narrow its relevance to this feed.
Relevance: 7 Novelty: 6
13. Encrypt What Matters: When Selective Homomorphic Inference Is Efficient
ArXiv ID: 2609.09357
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Ali Backour, Juan Reyes, Jaime Punyed, Ana Onoprishvili
Abstract: Fully homomorphic encryption (FHE) enables inference on private data without revealing it to the server, but evaluating an entire input under FHE is expensive. We study \emph{selective homomorphic inference}, where only a sensitive region of interest (ROI) is encrypted, and computations independent of that region are performed in plaintext. Selective evaluation produces the same output as full FHE on the same model, without retraining. Its efficiency depends on how quickly encrypted dependencies spread through the network. For small encrypted ROIs, locality-preserving architectures can achieve order-of-magnitude homomorphic-evaluation speedups, whereas architectures with early global mixing provide essentially no speedup. These results identify locality as the key architectural property governing the benefit of selective homomorphic inference.
Comment: Selective encrypted computation preserves full-FHE outputs while exploiting locality to reduce inference cost.
Topic Match: The core contribution is an inference-efficiency mechanism, although its benefits are specific to encrypted computation and diminish with early global mixing.
Relevance: 6 Novelty: 7
Paper Selection Prompt
System Prompt
You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.
User Prompt
Instructions
Respond in JSONL. Output exactly one JSON object per paper, one per line:
{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}
Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only:
daily_hot,new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return[]. -daily_hotmeans the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. -new_frontiermeans the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.Scoring Criteria
Relevance and Novelty are independent axes. Score both from 1 to 10.
Relevance Scoring
- 9-10: directly centered on the target foundational topics; highest when the core contribution is clearly within them.
- 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
- 5-6: touches the target topics, but the main contribution is elsewhere.
- 3-4: largely outside the target topics, often application-focused or domain-specific.
- 1-2: unrelated.
Important: Broad frontier relevance, major launch status, or daily buzz is not enough for a high Relevance score here. Those cases belong in the hotspot digest unless the paper strongly matches the specialized paper topics.
Novelty Scoring
- 9-10: new paradigm, theory, or major methodological breakthrough.
- 7-8: substantial methodological advance or strong new insight.
- 5-6: meaningful but incremental extension or refinement.
- 3-4: minor, narrow, or mostly engineering or domain-specific improvement.
- 1-2: little originality; mainly standard application of existing methods.
Topic Registry
Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - moe_training: MoE Training - Mixture-of-Experts training: routing, load balancing, expert granularity and shared experts, training stability, MoE parallelism and communication, and MoE kernels. - training_systems: Large-Scale Training Systems and Efficiency - Large-scale training systems and efficiency: distributed training algorithms, optimizers, and communication. - architecture_training: Architecture and Training Dynamics - Core architectural or computational mechanisms, dynamic computation, and training-stability dynamics. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, memory or cache efficiency, and large-scale training systems that materially change cost or behavior.
Papers
[PAPER LIST HERE]
Relevant Topics
This is a training-side feed. The centre of gravity is Mixture-of-Experts training and the large-scale training systems it sits on; architecture and efficiency work is kept when it bears on how a large model is trained or what it costs to train.
Keep a paper when its CORE CONTRIBUTION falls in one of the four topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.
MoE Training (primary) - Keep: expert routing and gating mechanisms; load balancing, with or without auxiliary losses; expert granularity, fine-grained and shared experts; MoE training stability (loss spikes, router collapse, z-loss and related fixes); expert and hybrid parallelism; all-to-all dispatch/combine and the communication schedules MoE needs; MoE kernels and grouped-GEMM style implementations; expert capacity, token dropping, and upcycling dense checkpoints into MoE. - Filter: papers that merely train on top of a MoE model without a new routing, balancing, stability, parallelism, or kernel contribution; "mixture of experts" in the classical ensemble or recommender sense.
Large-Scale Training Systems - Keep: distributed training algorithms (data, tensor, pipeline, sequence, and expert parallelism, FSDP and ZeRO-style sharding); optimizers and preconditioners for large-scale pretraining; communication algorithms, overlap, and collective design; throughput and MFU work that materially changes what a training run costs or whether it converges; scaling-law work that informs how such runs are configured. - Filter: routine infrastructure tuning with no new algorithmic or systems idea.
Architecture and Training Dynamics - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, dynamic or modular computation); optimisation and training-dynamics analysis that explains why large models train the way they do. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight.
Efficiency and Compression - Keep: quantization, sparsity, pruning, low-rank adaptation, KV-cache and memory-efficient designs, and other work that materially changes the cost of running or training a large model, especially when the mechanism is new rather than a tuned variant. - Filter: straightforward application of standard efficiency methods, and deployment or serving engineering with no new idea.
Leave these to the hotspot digest unless the core contribution clearly falls in one of the four topics above: - major model or product releases - agent, tooling, and RAG launches - benchmark, leaderboard, or evaluation-only papers - representation-learning theory, interpretability, and feature-formation analysis - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning) - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains