Previous Day 2026-09-16
Monthly Overview 2026-09
Next Day 2026-09-18

Personalized Daily ArXiv Papers 2026-09-17

Model Metric Usage Papers
Prompt Completion Total Total arXiv Scanned Relevant
local agent CLI Tokens not reported not reported not reported 629 423 27
Cost not reported not reported not reported

Topic Coverage:

TopicPapers
MoE Training1
Large-Scale Training Systems and Efficiency3
Architecture and Training Dynamics10
Efficiency, Compression, and Large-Scale Training13

Table of contents by topic:

MoE Training (1)

  1. Beyond the Previous Layer: Residual Predictive Structure in Sparse MoE Routing Authors: Hao Li, Yasuyuki Tahara, Yuichi Sei

Large-Scale Training Systems and Efficiency (3)

  1. COMPASS-ABS: Reducing Fragmentation in Shared GPU Clusters for Deep Learning Training Workloads Authors: Yukai Zhou, Hongfan Wu

  2. Gradient Descent with Stochastic Subspaces via Persistence of Memory Authors: Subhroshekhar Ghosh, Clement Z. Q. Ng, Pierre-Louis Poirion, Akiko Takeda

  3. FedPGT: Progressive Gradient Transmission for Vehicular Federated Learning over Time-Varying Channels Authors: Jintao Yan, Tan Chen, Yuxuan Sun, Sheng Zhou, Zhisheng Niu

Architecture and Training Dynamics (10)

  1. Reaching Every Position Without Searching: Rotating Sparse Wiring on the Hypercube as a Substitute for Attention Authors: Yoshiaki Takashita

  2. Beyond Random Couplings: Contrastive Noise Alignment in Generative Flows Authors: Lennart Wittke, Vinicius Azevedo

  3. The Attention Within: Consensus Dynamics in Selective State Space Models Authors: Jo\~ao Pedro Silvestre, \'Alvaro Rodr\'iguez Abella, Paulo Tabuada

  4. Double descent is the principle of least action Authors: Congzhou M Sha

  5. Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data Authors: Jinli Hu, Ross M. Clarke, Yichuan Zhang, Jos\'e Miguel Hern\'andez-Lobato

  6. SAM-on-the-Curve: Sharpness-Aware Mode Connectivity for Robust Weight-Space Interpolation Authors: Alejandro Calatrava, Xu Zhang, Ren Wang

  7. QiT: Quantum-Inspired Transformer for Visual Recognition Task Authors: Badri N. Patro, Vijay Agneeswaran

  8. Objective vs. Search: Decomposing What Makes a Good Tokeniser Authors: Ahmetcan Yavuz, Clara Meister, Tiago Pimentel

  9. Beyond Embedding Transfer: Component Roles in Grokking Transfer and Stability Authors: Zeyu Jia

  10. Stability-Constrained Approximation in Spline KANs: Exact Layer Balancing and Budget-Compatible Saturation Authors: Aleksander Tankman

Efficiency, Compression, and Large-Scale Training (13)

  1. Higher-order pruning of experts in mixture-of-experts language models Authors: Alex M. Tseng, Prannay Kaul, Luca Zancato, Wei Xia, Stefano Soatto

  2. Fathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches Authors: Vivek Kalyanarangan

  3. GroupKV: Hierarchical KV Cache Management for Long-Context Diffusion LLM Inference Authors: Jinhao Wang, Zhexin Hu, Kangjie Zhou, Xin Zhou, Fangfang Liu

  4. OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning Authors: Ha Lan Nguyen, Huy Hoang Tran, Trac-Duy Tran, Dung D. Le

  5. The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction Authors: Yu Lin, Yiming Wang, Runyuan Cai, Hanze Liu, Xiaodong Zeng

  6. Colla-Q: Toward Collaborative Experts in MoE Quantization via Minimax Precision Balancing Authors: Eunju Shin, Jongbin Ryu

  7. ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference Authors: Amir Ziashahabi, Hossein Entezari Zarch, Lei Gao, Murali Annavaram, Salman Avestimehr

  8. Contiguity, Not Importance: Budgeted Repair of Stale KV Caches After Document Edits Authors: Mingyang Mao, Wyatt Mackey, Xiaomin Lin

  9. Accelerating Diffusion Sampling via Speculative Draft Trees Authors: Marcello Bullo, Yanxiao Liu, \"Oyk\"u S{\i}la G\"uner, Arpan Mukherjee, Deniz G\"und\"uz

  10. Token Latency Fairness: Performance Isolation for Multi-Tenant LLM Serving Authors: Dev Bali, Soujanya Ponnapalli, Yichuan Wang, Natacha Crooks, Scott Shenker, Matei Zaharia

  11. Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It Authors: Yipeng Liu, Yingqiang Zhang, Feifei Li, Huanchen Zhang

  12. Not All Layers Need Tuning: Diagnosing and Directing Adaptation in Vision-Language-Action Models Authors: Shahram Najam Syed, Arthur Jakobsson, Prayuj Sachdev, Jeffrey Ichnowski

  13. rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference Authors: Kaijun Zhou, Zhiyang Li, Le Chen, Jinyu Gu


MoE Training (1)

1. Beyond the Previous Layer: Residual Predictive Structure in Sparse MoE Routing

ArXiv ID: 2609.17940

Primary Topic: MoE Training

Authors: Hao Li, Yasuyuki Tahara, Yuichi Sei

Abstract: Sparse mixture-of-experts models route each token through a sequence of expert selections. We ask whether the immediately preceding selection adequately summarizes this trajectory for predicting the next router. Using frozen OLMoE and JetMoE models, we measure the held-out predictive gain from earlier expert selections while retaining the most recent selection as a common baseline. In OLMoE, extending the history from one to eleven layers raises router-logit $R^2$ from 0.59879 to 0.66544. A preregistered JetMoE replication yields four-layer gains of 0.14275 and 0.20528 at two target depths, with paired bootstrap intervals above zero. These gains survive nonlinear decoding: adding history to a small multilayer perceptron improves $R^2$ by 0.17137 and 0.21861, whereas nonlinear decoding of the recent state alone adds 0.00139 and 0.00936 over a linear probe. Parameter-matched controls preserve the advantage, and cross-fitted history residuals predict target residuals with $R^2$ of 0.20549 and 0.23556. These findings identify residual predictive structure in expert-selection trajectories beyond adjacent-layer persistence.

Comment: Earlier expert selections improve router prediction beyond the information in the immediately preceding layer.

Topic Match: Directly studies MoE routing traces, but the contribution is predictive probing of frozen models without a new routing or training mechanism.

Relevance: 6 Novelty: 6


Large-Scale Training Systems and Efficiency (3)

1. COMPASS-ABS: Reducing Fragmentation in Shared GPU Clusters for Deep Learning Training Workloads

ArXiv ID: 2609.18519

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Yukai Zhou, Hongfan Wu

Abstract: With the rapid advancement of deep learning technology, shared GPU clusters receive an increasing number of deep learning training (DLT) jobs. Yet resource fragmentation make such clusters underutilized and forces the DLT jobs running on them to endure long turnaround times. Extensive research has been devoted to quantifying fragmentation and developing scheduling algorithms that alleviate its impact. However, existing fragmentation measures break down in the absence of workload distribution information, while current schedulers cannot continuously maintain resource fragmentation at a low level. To tackle these problems, we first introduce Scheduler-Induced Fragmentation (SIF), a metric built on the notion of partial-nodes that is independent of historical workload knowledge. We then propose COMPASS-ABS, which employs the COMPact-ASSured (COMPASS) algorithm to confine the cluster state within a tight Anchor-Based Space (ABS), whose construction fully leverages the topological alignment between dominant workload size and node capacity. Moreover. We also prove that it ensures SIF is bounded by $\frac{2}{N}$ under a workload composition condition that matches both theory and production. Evaluations implemented on a physical cluster and a simulated cluster demonstrate COMPASS-ABS effectiveness at improving resource utilization, reducing DLT job completion time by reducing fragmentation.

Comment: Topology-aware scheduling bounds GPU fragmentation to improve shared training-cluster utilization.

Topic Match: A new scheduling algorithm and conditional fragmentation bound address training-cluster resource efficiency, supported by physical-cluster and simulation evaluations.

Relevance: 8 Novelty: 7


2. Gradient Descent with Stochastic Subspaces via Persistence of Memory

ArXiv ID: 2609.18416

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Subhroshekhar Ghosh, Clement Z. Q. Ng, Pierre-Louis Poirion, Akiko Takeda

Abstract: Stochastic subspace methods have gained popularity as gradient descent based techniques for large scale optimisation problems, especially in distributed settings. In this paper, we introduce the technique of "persistence of memory" to greatly extend and improve the random subspace methods. To this end, we leverage a vector that is only weakly correlated with the gradient in order to provide a guiding structure to the generative process of the random subspace along which the descent is going to take place. This guidance vector may be fixed for a large number of iterations, only to be refreshed at wide intervals (on whose size we can provide guarantees in terms of problem parameters). In important machine learning settings, such as optimisation problems embodying sparsity or a minibatch structure, we show that the guidance vector can be obtained in an effective and computationally inexpensive manner by leveraging the structured properties of the problem. En route, we establish to our knowledge the first theoretical analysis of classical SSD methods for sparse functions. In a local neighbourhood of the optimum, we demonstrate an alignment phenomenon of our gradient estimates with a low-lying eigenvector of the Hessian, allowing a once-for-all computation of the guidance vector which renders the method computationally favourable even in scenarios with unstructured objectives.

Comment: Reuses weak gradient guidance to bias stochastic descent subspaces, with guarantees on guidance refresh intervals.

Topic Match: The core contribution is a computationally efficient optimizer for large-scale problems, although the abstract does not establish performance on large-model pretraining.

Relevance: 7 Novelty: 7


3. FedPGT: Progressive Gradient Transmission for Vehicular Federated Learning over Time-Varying Channels

ArXiv ID: 2609.18089

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Jintao Yan, Tan Chen, Yuxuan Sun, Sheng Zhou, Zhisheng Niu

Abstract: Vehicular federated learning (VFL) enables privacy-preserving collaborative model training for intelligent transportation systems, where communication resource allocation and gradient sparsification techniques have been explored to reduce communication overhead. However, vehicle mobility leads to rapidly varying channel conditions and transmission capacity, rendering predetermined resource allocation and sparsification decisions ineffective. In this paper, we propose FedPGT, a progressive gradient transmission scheme for VFL over time-varying channels, where vehicles progressively transmit high-magnitude gradient entries in response to instantaneous channel conditions. We establish a convergence bound that characterizes the impact of transmitted gradient entries and reveals diminishing-return behavior governed by a power-law decay. Motivated by this result, we formulate a stochastic optimization problem for online decision-making, where the main challenge lies in a cumulatively coupled, non-separable objective. To handle this challenge, we introduce per-slot surrogate transmission variables to decouple the long-term dependence across time slots and convert the original objective into an additive per-slot optimization problem, enabling a Lyapunov drift-plus-penalty approach for online scheduling. We further develop a low-complexity resource allocation algorithm for efficient online implementation. Experimental results demonstrate that the proposed scheme achieves a 3.65% accuracy improvement on the CIFAR-10 image classification task and a 12.66% reduction in average displacement error on the Argoverse trajectory prediction task compared with state-of-the-art baselines, demonstrating its applicability to diverse learning tasks under highly dynamic vehicular environments.

Comment: Progressive sparse-gradient transmission adapts distributed training communication to changing channel capacity.

Topic Match: The core contribution changes gradient communication and scheduling, although its scope is wireless vehicular federation rather than large-model training.

Relevance: 7 Novelty: 6


Architecture and Training Dynamics (10)

1. Reaching Every Position Without Searching: Rotating Sparse Wiring on the Hypercube as a Substitute for Attention

ArXiv ID: 2609.18145

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Yoshiaki Takashita

Abstract: Attention pays, at every layer and for every input, the cost of searching for whom to connect. We ask how far one can get with wiring that is fixed, sparse, and simply rotated from layer to layer. Treating the $n$ positions of a sequence as the vertices of a $\log_2 n$-dimensional hypercube and connecting each position, at layer $\ell$, to its neighbour along dimension $\ell \bmod \log_2 n$, information from every position reaches every other in $\log_2 n$ layers with $2n$ links per layer instead of $n^2$. On a synthetic task that is unsolvable unless all positions are reached, this rotation matches all-to-all wiring at $1/32$ of the links, while the same sparse pattern held fixed across layers fails; what matters is that every dimension is touched, not the order. On character-level language modelling of a public corpus (the first $12$M characters of enwik8), a hybrid that keeps two attention layers among sixteen sparse ones reaches $0.06$ bits-per-character lower held-out loss than a fully attentive model of the same width at the same step budget (three seeds each, no overlap), with $1/7$ of the links, $42\%$ fewer parameters, and $2.4\times$ less wall-clock time; the purely rotated schedule is level with the hybrid. The same ordering holds on a second corpus of mixed Japanese, English and code, where the gap widens to $0.16$. The usable learning-rate window is four to eight times wider than attention's on both. We also report what did not work - learned coordinates, and a "dynamics" variant whose apparent gains turned out to be an artefact of a saturated kernel - and the measurement discipline (frozen corpus, full-coverage evaluation, seed spread as the bar for ranking) that we found necessary to say anything at all at this scale.

Comment: Rotating hypercube connections provide global sequence mixing through linear-size sparse wiring per layer.

Topic Match: The core contribution is an attention replacement with an explicit connectivity mechanism and measured compute and learning-rate stability benefits.

Relevance: 9 Novelty: 7


2. Beyond Random Couplings: Contrastive Noise Alignment in Generative Flows

ArXiv ID: 2609.18488

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Lennart Wittke, Vinicius Azevedo

Abstract: Diffusion and flow-matching models are typically trained by corrupting data through independently sampled Gaussian noise. While simple and scalable, this forward process induces arbitrary data-noise couplings, forcing the network to learn high-curvature transports between unrelated endpoints. Existing optimal-transport methods reduce this burden by reassigning fixed noise samples to data, but the source noise distribution itself remains passive. To address this, we introduce Contrastive Noise Alignment (CNA), a training-time method that creates dynamic, contrastive couplings by optimizing the noise representations directly. By modeling the noise batch as an interacting particle system, CNA employs a cross-modal InfoNCE objective to align noise particles with their paired data targets. To prevent spatial collapse, this alignment is regularized using an angular entropy term and a radial norm penalty. We show theoretically that this equilibrium asymptotically preserves Gaussian structures, maintaining tractability during inference. Empirically, CNA improves the alignment between noise and data, reduces flow curvature, and provides better generation quality with fewer required sampling steps. For few-step, pixel-space generation (2-4 NFEs), CNA reduces FID by over 50\% compared to standard rectified flow, and by at least 24\% against Optimal Transport baselines.

Comment: Optimizing noise particles during training reshapes data-noise couplings to reduce flow curvature.

Topic Match: Changes the training mechanism of generative flows through contrastive noise alignment, with reduced sampling requirements providing a secondary efficiency contribution.

Relevance: 8 Novelty: 8


3. The Attention Within: Consensus Dynamics in Selective State Space Models

ArXiv ID: 2609.17997

Primary Topic: Architecture and Training Dynamics

Authors: Jo\~ao Pedro Silvestre, \'Alvaro Rodr\'iguez Abella, Paulo Tabuada

Abstract: Selective state space models (SSMs) have recently emerged as a compelling alternative to transformers, combining competitive performance with substantially improved inference efficiency. At each SSM layer, a sequence of hidden states are propagated by a recurrence, mixing information of different tokens. Despite using a different mechanism, this mixing plays a role analogous to attention in transformers. In fact, recent works have shown that the two architectures may be closer than they first appear, as this recurrence admits a formulation akin to linear attention. In transformers, attention is known to drive the tokens to cluster, i.e., to reach consensus, collapsing in the limit to a single direction. Thus, we ask: does the recurrence at the core of SSMs drive the tokens to consensus, as attention does in transformers? To answer this question, we take a dynamical systems perspective on SSMs, modeling the evolution of tokens across layers as an ordinary differential equation. By exploiting input-to-state stability arguments, we establish local exponential stability of the consensus equilibria and characterize their domain of attraction for time-varying weight matrices, a setting not addressed by previous results. We thereby show that the resemblance between SSMs and transformers does run deeper: the recurrence at the core of SSMs aggregates tokens just as attention does. Numerical experiments on a pretrained Mamba-2 model point to the output gate as the component that regulates the extent of this consensus, preventing the tokens from reaching it in full.

Comment: Dynamical analysis explains how selective-SSM recurrence drives token consensus and how output gating limits collapse.

Topic Match: The analysis directly examines recurrence and gating as architectural mechanisms controlling token mixing, including stability under time-varying weights.

Relevance: 8 Novelty: 7


4. Double descent is the principle of least action

ArXiv ID: 2609.19076

Primary Topic: Architecture and Training Dynamics

Authors: Congzhou M Sha

Abstract: The test error of a model plotted against its number of parameters $d$ falls, peaks when the model can just fit the training data, and falls again, exhibiting the double descent phenomenon. We explain the phenomenon with statistical mechanics. The training trajectory of a stochastic gradient-based method is a particle wandering over the energy landscape of the training loss at an induced temperature $T$, and a run that has equilibrated visits every parameter vector of a given training loss equally often, the fundamental postulate of statistical mechanics, with probability given by the Boltzmann distribution. Because training starts at an initial point and has only finite time to diffuse, it carries an effective weight decay, which makes every parameter a quadratic degree of freedom. The equipartition theorem then distributes the energy among the $d$ degrees of freedom in shares of $T/2$, so at a fixed training loss adding parameters lowers the temperature and drives the Boltzmann distribution toward the stationary path. Finally, adding parameters can only lower the $L^2$ norm of the stationary path, so a solution sampled at fixed loss is less likely to be large with increasing $d$, effectively increasing weight regularization.

Comment: Links finite-time SGD and overparameterization to implicit regularization through a statistical-mechanics model.

Topic Match: The proposed explanation directly concerns optimization-induced regularization and double-descent behavior.

Relevance: 8 Novelty: 6


5. Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data

ArXiv ID: 2609.18842

Primary Topic: Architecture and Training Dynamics

Authors: Jinli Hu, Ross M. Clarke, Yichuan Zhang, Jos\'e Miguel Hern\'andez-Lobato

Abstract: The scaling laws hold that a language model grows more capable with more parameters and more training data, and Mixture-of-Experts (MoE) architectures have ridden these laws to remarkable results, activating only a fraction of an enormous stored parameter bank for each token. That success is built on static pretraining data. A deployed model faces a different world, where much of the data that would make it more useful is not in its training set but in the live interaction it is currently handling, such as the facts a user supplies or the corrections they give. A conventional model cannot learn from this data, because its weights are frozen after training. Instead, the knowledge and behaviour supplied at run time are placed in the prompt, by retrieval or instruction, and re-read on every request only to be discarded once the request ends. We ask how an architecture could learn from live interaction by writing it into its weights. Taking inspiration from MoE, we propose the \textbf{Infinite-Parameter LLM}. A compact hypernetwork turns the data given at run time into a low-rank modulation of a shared base network, so the feed-forward weights are generated from live data rather than stored in a fixed bank. Where prior weight generators read the context once and freeze, we carry a Bayesian belief over the generator's latent code and update it online, so the effective weight is re-derived from that evolving belief as the session proceeds rather than fixed after one read. The stored footprint stays fixed, yet the weights the model can compile are effectively infinite. For the knowledge and behaviour supplied at run time, carrying them in the weights rather than the prompt is amortized in compute, frees the context window, persists across turns, and can generalise better than in-context use. We specify an evaluation protocol that tests exactly this against in-context learning and retrieval.

Comment: Combines hypernetwork-generated low-rank FFN weights with online Bayesian updates to their latent code.

Topic Match: Dynamic FFN weight generation provides a substantive architectural match, although the focus is live-session adaptation and the abstract reports a proposed evaluation protocol without results.

Relevance: 7 Novelty: 7


6. SAM-on-the-Curve: Sharpness-Aware Mode Connectivity for Robust Weight-Space Interpolation

ArXiv ID: 2609.17748

Primary Topic: Architecture and Training Dynamics

Authors: Alejandro Calatrava, Xu Zhang, Ren Wang

Abstract: Deep neural networks that are independently trained to similar performance can be connected by low-loss parametric curves in weight space, a phenomenon known as Mode Connectivity (MC). This geometric property underpins practical techniques such as weight averaging, model ensembling, and model merging. We argue that low-loss connectivity is an incomplete geometric criterion: it controls loss only along a one-dimensional trajectory while leaving the surrounding weight-space neighborhood unconstrained, so the optimized curve may traverse sharp ridges that become fragile under distribution shift. We therefore reformulate mode connectivity as a neighborhood-robust path optimization problem, seeking a curve whose entire local neighborhood maintains low loss. We propose Sharp Mode Connectivity (SMC), which applies a first-order sharpness-aware approximation to the resulting minimax functional, enforcing flatness along the entire curve rather than only on it. We derive a practical optimization algorithm for connectivity paths under this sharpness-aware objective. Under severe blur corruptions from CIFAR-10-C, SMC achieves up to 6.09\% absolute accuracy improvement over standard MC. Remarkably, SMC produces negative loss barriers, meaning that models obtained at interior points of the optimized path can outperform the average endpoint loss. These results, validated across ResNet-18, VGG16-BN, and ViT-Tiny on CIFAR-10 and ImageNet-100, establish path-wise flatness as a practical principle for robust weight-space interpolation.

Comment: Optimizes connectivity curves for low loss throughout their surrounding weight-space neighborhoods.

Topic Match: Path-wise sharpness contributes to optimization geometry, with practical relevance concentrated on robust weight-space interpolation.

Relevance: 7 Novelty: 6


7. QiT: Quantum-Inspired Transformer for Visual Recognition Task

ArXiv ID: 2609.17789

Primary Topic: Architecture and Training Dynamics

Authors: Badri N. Patro, Vijay Agneeswaran

Abstract: Quantum machine learning offers a compelling representational perspective: angle-encoded states inhabit Hilbert spaces in which periodic similarities and interactions can be expressed naturally. Realizing this perspective for visual recognition remains difficult, however, because present quantum neural networks are constrained by limited qubit counts, costly circuit simulation and measurement, noise, and unstable optimization on noisy intermediate-scale quantum devices. We investigate whether useful structural ideas from quantum models can instead be realized as scalable classical Transformer operations. We introduce QiT, a Quantum-inspired Transformer for vision tasks with three components: (i) angle-inspired encoding that maps image tokens to learned trigonometric Hilbert-space features analogous to quantum rotation-based state encoding; (ii) self-attention over these periodic features, inducing a classical cosine kernel approximated to quantum fidelity kernels; and (iii) gated multiplicative emulation, a trainable classical surrogate for interaction terms found in variational circuits. All components are differentiable tensor operations, so QiT claims neither quantum computation nor quantum speedup and retains the $\mathcal{O}(N^2D)$ attention complexity of a standard Vision Transformer. Across image-classification benchmarks, QiT is competitive with a matched classical Transformer while avoiding the severe runtime cost observed for a small simulated quantum Transformer. QiT-B reaches 78.3\% ImageNet-1K top-1 accuracy with 45.7M parameters and 11.5 GFLOPs. These results position QiT as a scalable baseline for isolating and evaluating quantum-motivated inductive biases in visual recognition.

Comment: Periodic token features induce a cosine attention kernel inside a classical Transformer.

Topic Match: Introduces internal attention and token-interaction mechanisms, although validation is confined to visual recognition and attention complexity remains unchanged.

Relevance: 7 Novelty: 6


8. Objective vs. Search: Decomposing What Makes a Good Tokeniser

ArXiv ID: 2609.19145

Primary Topic: Architecture and Training Dynamics

Authors: Ahmetcan Yavuz, Clara Meister, Tiago Pimentel

Abstract: Two dominant tokenisation algorithms are used by modern language models: byte-pair encoding (BPE) and UnigramLM. These differ along two orthogonal axes: their optimisation objective (compression vs. log-likelihood) and their search procedure (bottom-up merging vs. top-down pruning). Existing comparisons confound these axes, making it unclear whether their observed differences stem from what is being optimised vs. how it is being optimised. We disentangle the two by introducing two new tokenisation algorithms that complete this 2x2 design space: BottomUpLL, a bottom-up likelihood-based tokeniser, and TopDownComp, a top-down compression-based tokeniser. We train language models with tokenisers produced by each algorithm, varying: model size, vocabulary sizes, and domain (English-only vs. multilingual). Evaluating models on bits-per-byte, we find that the search procedure -- not the objective -- is the dominant factor: bottom-up tokenisers consistently achieve lower bits-per-byte in most settings. Evaluating models on the BLiMP task, however, shows no consistent relationship between design choice and performance. Overall, our results disentangle the effect of tokeniser design choices on language modelling performance, offering concrete guidance for their more principled construction.

Comment: Separates tokenizer search from objective, finding that bottom-up construction generally produces lower bits per byte.

Topic Match: New tokenizer construction mechanisms inform the input design used during pretraining, a peripheral computational mechanism within this topic.

Relevance: 7 Novelty: 6


9. Beyond Embedding Transfer: Component Roles in Grokking Transfer and Stability

ArXiv ID: 2609.18078

Primary Topic: Architecture and Training Dynamics

Authors: Zeyu Jia

Abstract: Warm-start transfer can make algorithmic tasks generalize rapidly, yet it is unclear which model components provide the gain and whether that gain remains stable under continued optimization. We study cross-operator transfer on modular arithmetic and separate efficacy (early velocity) from stability (post-reach drawdown). In a scale-matched 108-run battery across 12 seed blocks (96-run 2^3 factorial plus 12-run scale control), transferring internal attention/MLP weights (B) alongside token embeddings and readout (E+U) improves early accuracy by 5.46 pp (Holm p=0.0039) and cuts confirmation latency by 558 steps (Holm p=0.0088). While readout plus internal-block transfer satisfies the pre-specified +/-500-step latency equivalence criterion in 1-layer models (TOST p=0.0011, though Full is faster in 11/12 paired seeds), a prospective 2-layer replication confirms the internal-block advantage (12/12 seeds, +704.67 integral units, p=4.88x10^-4) while revealing an architectural boundary: omitting donor embeddings falls 4475.6 units below Full, outside the +/-250-unit margin. Continued target training frequently triggers severe post-grokking relapse. Freezing transferred representation carriers (E, U) nearly eliminates offline relapse (19.40% -> 0.07%, Holm p=0.005859). Online validation-triggered gating slashes True Max Drawdown from 22.06% to 0.60% on 2a+b (p=0.000488), with prospective confirmations extending protection across affine, nonlinear quadratic, and 2-layer targets (10.94-23.47 pp reductions), distinguishing continual stabilization from static early stopping. In non-abelian S_5, unshielded transfer surges transiently (95.4% peak), but a prospective shielding cohort yields no confirmed benefit (+0.15 +/- 1.14 pp). These results establish a component-level dissociation between transfer acceleration and trajectory stability, and expose the empirical boundaries of parameter shielding.

Comment: Suppresses post-grokking relapse during continued optimization by freezing transferred embeddings and readouts.

Topic Match: Component-level interventions investigate training stability, with evidence limited to small arithmetic models.

Relevance: 7 Novelty: 6


10. Stability-Constrained Approximation in Spline KANs: Exact Layer Balancing and Budget-Compatible Saturation

ArXiv ID: 2609.17619

Primary Topic: Architecture and Training Dynamics

Authors: Aleksander Tankman

Abstract: Deep spline superposition networks face a tension between approximation order and stability across depth. We study approximation under a hard layerwise Lipschitz budget, and organise it around two quantities: the factorisation stability complexity of a given deep factorisation, and the budget-compatible approximation complexity of a discretisation operator. First, we solve exactly the finite-depth diagonal balancing problem for a fixed chain of nonnegative envelope matrices: the optimal uniform layer budget equals $|M_{L-1}\cdots M_0|_{\infty\to\infty}^{1/L}$, attained by an explicit one-pass minimiser, for rectangular layers, with a complete treatment of degeneracies and non-attainment. The optimum can be arbitrarily larger than the Lipschitz constant of the network itself, because passing to envelopes destroys sign cancellation. Second, we give a constructive spline discretisation theorem preserving the budget up to a controlled slack, with an explicit grid threshold. Conversely, for linear spline-valued operators that preserve the budget exactly, we prove budget-compatible minimax lower bounds on classes constrained simultaneously in the first and third derivative norms -- a constraint pair that is forced by the problem and that rules out the usual scaling escapes. Finally, we show that the corresponding layer errors need not cancel under composition: for every operator of the class there is a stable depth-$L$ tower realising a constant fraction of the accumulated error, so the linear-in-depth accumulation of the upper bound is not a proof artefact.

Comment: Exact diagonal layer balancing under hard Lipschitz budgets in deep spline networks.

Topic Match: Closest to architectural stability, but the results concern constrained approximation rather than optimization or training dynamics.

Relevance: 6 Novelty: 7


Efficiency, Compression, and Large-Scale Training (13)

1. Higher-order pruning of experts in mixture-of-experts language models

ArXiv ID: 2609.18916

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Alex M. Tseng, Prannay Kaul, Luca Zancato, Wei Xia, Stefano Soatto

Abstract: Mixture-of-Experts (MoE) language models suffer from large parameter counts, which create a significant memory bottleneck. Expert pruning is the most direct approach for reducing this parameter count, yet existing methods make pruning decisions for each expert independently, and assume experts' contributions are purely additive. In reality, expert usage in MoEs is inherently cooperative. We derive HOPE (Higher-Order Pruning of Experts), a second-order pruning objective which provably minimizes an upper bound on the error resulting from pruning. We show that REAP (a state-of-the-art first-order pruning method) is a special case of HOPE where interaction terms are ignored. Across three frontier MoE models (up to 122B parameters), two distinct calibration sets, and multiple benchmarks (including math, instruction following, coding, and an agentic suite), we demonstrate that HOPE produces better pruning decisions than existing methods, and its advantage is most pronounced at high pruning rates and on challenging agentic workloads. At 50% pruning, HOPE outperforms all baselines and achieves an average rank of 1.58 out of 5 methods (versus 2.42 for the next-best method, REAP), with gains of up to +6.1% on agentic coding. Over all conditions, HOPE again achieves the best average rank and surpasses every other method in the majority of head-to-head comparisons. By preserving cooperative expert structure that first-order methods ignore, HOPE enables aggressive compression with minimal degradation, particularly on complex tasks where diverse expert combinations are invoked over long sequences.

Comment: Derives a second-order expert-pruning objective that accounts for cooperative expert interactions and bounds pruning error.

Topic Match: Interaction-aware expert selection reduces stored MoE parameters, making compression the strongest fit.

Relevance: 9 Novelty: 8


2. Fathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches

ArXiv ID: 2609.17652

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Vivek Kalyanarangan

Abstract: When agentic sessions run to a million tokens with many sessions resident at once, the KV cache and the index that ranks it live in host memory, and the scan that ranks all n keys for a top-k step becomes the traffic that bounds decoding. We present Fathom, a key scan in which each query decides how many bits of each key channel to read. The 4-bit K cache is stored channel-major as bit planes, so a prefix of t planes is exactly the channel's t-bit quantizer, and the query spends its bit budget by reverse water-filling over the variance-weighted importance of its channels. At one million tokens on Qwen3-8B a decode step is 1.67x faster in GPU time than with the 136-bit scans of Double Sparsity, Loki and SparQ r=32, and in the same GPU time as SparQ's 68-bit read (r=16) Fathom reads 18% fewer bytes with lower attention error on six of seven model and context settings. On RULER-style tasks every per-token scan matches exact top-k decoding, and on real coding-agent sessions Fathom reaches the step agreement of the most accurate 136-bit scan at 92 bits. The store is the 4-bit K copy a quantized serving stack already holds, and the method is not faster when the index is resident in GPU memory.

Comment: Query-dependent bit-plane allocation reduces key-scan bandwidth for sparse decoding over offloaded KV caches.

Topic Match: The new adaptive precision mechanism directly reduces large-model decoding traffic and cost when KV data reside in host memory.

Relevance: 9 Novelty: 8


3. GroupKV: Hierarchical KV Cache Management for Long-Context Diffusion LLM Inference

ArXiv ID: 2609.17573

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Jinhao Wang, Zhexin Hu, Kangjie Zhou, Xin Zhou, Fangfang Liu

Abstract: Diffusion large language models (dLLMs) are emerging as a promising generative paradigm that complements autoregressive decoding. In long-context settings, KV cache bloat and offloading transfer overhead have become primary bottlenecks in inference systems. Meanwhile, the periodic full-sequence recomputation and localized token updates in dLLMs make the KV lifecycle substantially more dynamic, complicating cache management and prefetch scheduling while making heavyweight token-level indexing or clustering schemes harder to amortize effectively during decoding. To address these challenges, we present \textsc{GroupKV}, a lightweight hierarchical KV cache management system for long-context dLLM inference. We observe that under block-wise decoding, tokens within the same generation block tend to access highly overlapping and spatially concentrated context regions, making group-level sparse selection effective. Building on this observation, \textsc{GroupKV} partitions the context into contiguous groups and performs coarse-to-fine sparse selection. \textsc{GroupKV} further exploits cross-layer consistency to enable predictive prefetching, and incorporates a staleness correction mechanism to maintain cache coherence under dynamic KV updates. Additionally, \textsc{GroupKV} adopts streaming prefill to reduce peak memory consumption during prefilling. Experiments show that \textsc{GroupKV} extends the maximum serviceable context length by up to $48.00\times$ under constrained GPU memory, improves end-to-end inference performance by up to $3.73\times$ in offload-based long-context settings, and maintains competitive task accuracy.

Comment: Introduces group-level sparse KV selection with predictive prefetching and staleness correction for diffusion LLMs.

Topic Match: The core contribution directly reduces KV memory and transfer costs through cache management designed for dynamically updated diffusion states.

Relevance: 9 Novelty: 7


4. OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning

ArXiv ID: 2609.17890

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Ha Lan Nguyen, Huy Hoang Tran, Trac-Duy Tran, Dung D. Le

Abstract: Large reasoning models (LRMs) generate long chain-of-thought traces before answering, creating significant inference overhead. Pruning can reduce this cost, but its effectiveness depends on the calibration data used to estimate parameter importance. Recent work calibrates on the model's own rollouts instead of generic dataset, but treats all reasoning tokens uniformly, regardless of whether they contribute to successful reasoning. As a result, pruning protects weights by statistical salience rather than by their contribution to correct reasoning, so weights behind erroneous computation survive as readily as those behind correct computation. These erroneous patterns then get carried into the pruned model, degrading reasoning quality, producing both lower accuracy and longer reasoning traces. We propose Outcome-Based Calibration for Large Reasoning Model Pruning (OBC-Prune) to close this gap. OBC first constructs difficulty-matched pairs of correct and incorrect rollouts from problems the model answers inconsistently. It then estimates the causal importance of each reasoning sentence through intervention-based analysis, quantifying how removing its influence affects subsequent predictions. These causal importance scores are converted into per-token weights that rescale the calibration activations used by one-shot pruning methods (SparseGPT, Wanda, ALPS), without modifying the underlying pruning algorithms. Experiments on DeepSeek-R1-Distill-Qwen 1.5B, 7B, and 14B models at 40\% and 50\% sparsity demonstrate consistent improvements over state-of-the-art calibration baselines across most model sizes and sparsity levels on MATH500, LiveCodeBench, and AIME 2025. These results indicate that preserving causally important reasoning circuits is a substantially more effective pruning objective than uniformly preserving observed activations.

Comment: Uses outcome-conditioned, intervention-derived token weights to improve one-shot pruning calibration.

Topic Match: The central contribution is a new calibration mechanism for pruning billion-parameter models, directly matching compression and sparsity.

Relevance: 9 Novelty: 7


5. The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction

ArXiv ID: 2609.18063

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: MoE Training

Authors: Yu Lin, Yiming Wang, Runyuan Cai, Hanze Liu, Xiaodong Zeng

Abstract: Mixture-of-experts (MoE) inference on consumer hardware is bounded by weight memory: a 35B-class model is 19.5GB at 4-bit, and sparsity shrinks the compute per token, not the bytes that must be held. Naive offloading to SSD does not help on its own, because layer N+1's experts must be chosen before layer N's output exists, so the reads cannot start early enough to hide behind compute. We present Edge0, a streaming MoE inference engine that closes the gap with a prerouter: a per-layer head predicts the next layer's routing one token ahead, and the prediction is consumed as the routing itself, so the staged expert set equals the routed set and nothing is dropped. An unmerged recovery LoRA, trained on the student path, pays back the quality lost to int4 quantization and routing replacement. On a single 24GB machine, Edge0 serves a 35B MoE at 20tok/s inside 3GiB of peak active memory, within a few points of its fp16 teacher on average across five public benchmarks. An 8B tier runs on the same framework, and the framework, checkpoints, and adapters are open source.

Comment: Learned next-layer routing commits expert choices early enough to overlap SSD weight reads with computation.

Topic Match: The central contribution reduces MoE inference memory requirements through SSD streaming; replacing routing with trained predictions also supplies a substantive routing mechanism.

Relevance: 8 Novelty: 8


6. Colla-Q: Toward Collaborative Experts in MoE Quantization via Minimax Precision Balancing

ArXiv ID: 2609.18131

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Eunju Shin, Jongbin Ryu

Abstract: In this paper, we present a Mixture-of-Experts (MoE) quantization method based on activation entropy. Although quantization reduces memory and computational costs, it can substantially degrade performance. In particular, performance decline is pronounced in quantized MoE models, where individual experts have a small number of parameters that are sensitive to low-bit representation. Considering that MoE operates as an ensemble model with collaborative contributions from routed experts, a significant performance decline of a particular expert due to quantization can harm model performance. Therefore, we propose Colla-Q, a bit-allocation framework to maintain balanced performance across experts through an activation-entropy-based bit-width allocation algorithm. This approach encourages each expert to operate collaboratively in the quantized model, thereby 1) improving the overall MoE performance and 2) reducing the dependence on the calibration dataset. Since uniformly adjusting each expert's performance facilitates robustness and stability of the MoE model, the proposed MoE quantization method can generalize more consistently across different calibration datasets. Our code is available at: https://github.com/mmai-laboratory/Colla_Q

Comment: Activation-entropy-guided bit allocation balances quantization damage across MoE experts.

Topic Match: The core mechanism is expert-wise mixed-precision compression, targeting prediction quality and robustness across calibration datasets.

Relevance: 9 Novelty: 6


7. ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference

ArXiv ID: 2609.17943

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Amir Ziashahabi, Hossein Entezari Zarch, Lei Gao, Murali Annavaram, Salman Avestimehr

Abstract: Long-context LLM inference is bottlenecked by attention, whose repeated KV-cache reads make decoding memory-bound. Self-speculative decoding alleviates this by drafting tokens with sparse attention and verifying them with full attention, but existing batched methods remain synchronized: all requests in a batch share a single draft-verify schedule, even though the optimal draft length varies widely across requests and changes dynamically within each request. We propose ASPIRE, a non-synchronized batched self-speculative decoding framework built on three components. First, a unified mixed forward allows drafting and verifying requests to coexist in the same batched forward pass, removing the need for global draft-verify phases. Second, a lightweight online speculation scheduler uses per-request acceptance-rate estimates and a batch-aware cost model to let each request independently choose when to verify. Third, an intra-draft refresh layer performs full attention at a single designated layer during drafting, updating the sparse context at every draft step to reduce staleness during drafting. Across three models and five reasoning and long-context benchmarks, ASPIRE achieves $1.70$-$4.58\times$ speedup in decoding throughput over autoregressive baselines and improves average speedup by approximately $27\%$ over the strongest prior self-speculative baselines.

Comment: A mixed draft/verify forward removes batch-wide synchronization in self-speculative decoding.

Topic Match: Introduces a substantive scheduling mechanism for reducing memory-bound LLM decoding cost, with an inference-focused scope.

Relevance: 8 Novelty: 7


8. Contiguity, Not Importance: Budgeted Repair of Stale KV Caches After Document Edits

ArXiv ID: 2609.17983

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Mingyang Mao, Wyatt Mackey, Xiaomin Lin

Abstract: KV-cache reuse can reduce inference cost in retrieval-augmented generation and agentic systems, but cached contexts may become stale when retrieved knowledge, working memory, or user state is edited. Under causal self-attention, even a local edit can affect downstream KV states. A full re-prefill reliably restores consistency but is costly, whereas refreshing only the edited span can leave downstream dependencies stale. We formulate in-place repair as budgeted recomputation and compare training-free position-selection policies on a factual RAG benchmark with matched direct and derived edits. Across three model families, all policies repair direct cases, but derived cases clearly separate them. At the primary budget, a contiguous edit-local window recovers at least 0.94 of the post-edit answer margin and substantially outperforms attention-based, KV-deviation, and structural selectors. Mechanistic analysis shows that position sets effective under clean-state transplantation can fail under actual recomputation because scattered positions inherit surrounding staleness. The edit-local advantage also depends on adjacency and largely disappears when the answer-bearing text moves downstream. Because answer-relevant edits almost always corrupt model behavior, failure severity is difficult to predict, and repair is 13-21 times faster than full re-prefill, our results support unconditional edit-local repair when the dependent text remains adjacent to the edit.

Comment: Contiguous selective recomputation repairs stale KV caches while avoiding the cost of full re-prefill.

Topic Match: Budgeted cache repair directly reduces model execution cost, with a mechanistic explanation of why scattered recomputation inherits stale dependencies.

Relevance: 8 Novelty: 7


9. Accelerating Diffusion Sampling via Speculative Draft Trees

ArXiv ID: 2609.17691

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Marcello Bullo, Yanxiao Liu, \"Oyk\"u S{\i}la G\"uner, Arpan Mukherjee, Deniz G\"und\"uz

Abstract: Speculative sampling accelerates diffusion model generation by drafting inexpensive candidate states and correcting them under a coupling that preserves the target distribution exactly, reducing the number of expensive target evaluations. Existing diffusion samplers, notably those based on reflection maximal coupling, are topologically constrained: their lookahead drafts form a chain graph, a single linear sequence, which inherently limits the acceptance rate per target evaluation. We connect speculative sampling in diffusion models to relative entropy coding (REC). This perspective shows the lookahead need not be linear and motivates our central contribution, draft trees, which enrich the candidates considered per round and lower the target function evaluations. We further adopt greedy rejection sampling, an REC algorithm, as the draft-target coupling, improving acceptance while guaranteeing exact target samples. Experiments across diverse target and draft models demonstrate up to 8.3% acceleration over the reflection coupling baseline in practical settings.

Comment: Tree-structured drafts and rejection coupling reduce expensive diffusion evaluations while preserving the target distribution.

Topic Match: Introduces a sampling-acceleration mechanism with up to 8.3% speedup over the stated baseline; its scope is diffusion inference efficiency.

Relevance: 7 Novelty: 7


10. Token Latency Fairness: Performance Isolation for Multi-Tenant LLM Serving

ArXiv ID: 2609.18112

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Dev Bali, Soujanya Ponnapalli, Yichuan Wang, Natacha Crooks, Scott Shenker, Matei Zaharia

Abstract: LLM serving is typically offered as a shared, multi-tenant service, where high-demand workloads from one client can cause latency SLO violations for others. Existing solutions for performance isolation equalize client throughput in the long run, for example through queueing and batching fairness. However, these approaches do not provide latency isolation guarantees; as a result, well-behaved clients can still experience significant degradation to their token-level latencies. In this paper, we present FairInference, which provides the novel {\delta}-token fairness guarantee: for a well-behaved client, if a token is generated in d time units in isolation, it will be generated within d + {\delta} time units in multi-tenant execution, providing strong latency isolation guarantees for LLM serving. To achieve this, FairInference addresses a key challenge of LLM serving: bounding delays from sharing GPU resources without support for fine-grained scheduling or resource allocation. In FairInference, the scheduler enforces per-token deadlines, while bounding the delays from GPU compute sharing and accounting for the additional delays introduced by the shared KV caching in GPU memory. We show that FairInference effectively bounds token-level latency spikes for well-behaved clients and improves overall throughput compared to state-of-the-art LLM serving systems.

Comment: Introduces per-token deadline scheduling that bounds interference from shared GPU compute and KV caches.

Topic Match: A new resource-sharing algorithm improves LLM serving throughput and latency isolation, qualifying as runtime efficiency while remaining peripheral to training.

Relevance: 7 Novelty: 7


11. Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It

ArXiv ID: 2609.18849

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Yipeng Liu, Yingqiang Zhang, Feifei Li, Huanchen Zhang

Abstract: An agentic request spends substantial wall-clock time waiting for tools, and its KV cache holds GPU memory the whole time. Serving systems decide whether that cache stays, leaves, or comes back by guessing how long the tool will run, from the tool's name, its history, a duration declared before the call, or the engine's own occupancy. We show that no estimate fixed before a call starts can know its duration, and such estimates may not even rank the calls. Meanwhile, the running tool already holds the answer, but the agent stack together with the tool silences it. We propose that tool calls report their progress explicitly while they run, and we measure what that takes. A census of four public agent corpora finds a readable signal in most tool time once it is revealed, in two strengths: a fraction of the work remaining, or an accurate signal that the end is near. A harness recovers it without changing what the agent sees, at no measurable cost to the agent's benchmark score. At the points where a KV cache decision is made, the reported progress is between several times and an order of magnitude more accurate than the best published predictors, and it stays accurate when the environment changes. Plugged into a production engine through a few small hints, it cuts the p90 time to first token (TTFT) after a tool call by 20.7% (HBM only) and 20.8% (HBM + DRAM) against LRU, close to an oracle. A serving system should not guess what its tools can tell it.

Comment: Live tool-progress signals guide KV-cache residency and prefetch decisions during agent execution.

Topic Match: The contribution introduces a new information interface for cache management that improves large-model serving efficiency, with inference-specific scope.

Relevance: 7 Novelty: 7


12. Not All Layers Need Tuning: Diagnosing and Directing Adaptation in Vision-Language-Action Models

ArXiv ID: 2609.18084

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Shahram Najam Syed, Arthur Jakobsson, Prayuj Sachdev, Jeffrey Ichnowski

Abstract: Fine-tuning a Vision-Language-Action (VLA) model for a new deployment environment is expensive, yet most methods apply uniform-capacity adapters to every network region as if every region requires equal adjustment. This paper tests that assumption on five architecturally diverse VLAs (OpenVLA-OFT, $\pi_0$, SmolVLA, DTP, Octo; 93M-7B parameters). Measuring per-region adaptation cost as normalized parameter displacement under region-isolated fine-tuning reveals an adaptation spectrum in which appearance shifts concentrate cost in the vision encoder, instruction shifts in the language backbone, and novel-object shifts in the vision encoder together with the action head, across all five architectures. To exploit this structure, we introduce a pipeline that observes, diagnoses, allocates, and adapts. From ten unlabeled target observations and without fine-tuning, the diagnostic estimates per-region cost by combining reference-free gradient and Monte Carlo Dropout signals with a Centered Kernel Alignment score against a cached source reference; the allocator converts the estimates into variable-rank LoRA adapters under a parameter budget and freezes well-calibrated regions; and standard LoRA fine-tuning trains the resulting adapters. The diagnostic ranks regions within each deployment at a median Spearman of 0.91, and the allocation matches or exceeds uniform LoRA at every budget we tested on LIBERO and CALVIN. On a physical xArm-7, the pipeline matches full fine-tuning under an instruction-wording shift with 0.04% of its trainable parameters, and on five held-out scenes evaluated without retraining it leads every baseline, with 11-23 successes of 30 rollouts against 8-18 for the strongest parameter-efficient baseline at equal or larger budgets and 2-11 for full fine-tuning. These results suggest that adaptation cost in VLAs is structured enough to measure before fine-tuning begins.

Comment: Diagnostics from ten unlabeled observations allocate LoRA rank across regions under a parameter budget.

Topic Match: The core method allocates adapter capacity using regional adaptation estimates, making efficiency central; experiments focus on VLA fine-tuning.

Relevance: 7 Novelty: 6


13. rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference

ArXiv ID: 2609.19104

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Kaijun Zhou, Zhiyang Li, Le Chen, Jinyu Gu

Abstract: Factory work is a promising early scenario for embodied AI: assigning repetitive manual jobs to robots has clear economic payoff, and a structured station keeps the jobs tractable for current policies. Vision-Language-Action (VLA) models now dominate as the policy paradigm for these robots. The inference latency of VLA models directly affects robot responsiveness and motion smoothness. However, existing VLA inference frameworks do not fully exploit the characteristics of embodied workloads or account for the distinct bottlenecks across different stages of VLA inference. In this paper, we first characterize embodied workloads and identify substantial task similarity across repeated robot executions. We further find that such similarity extends beyond observations and action trajectories to internal model states. Drawing on these observations, we present rMuscle, a real-time VLA inference framework inspired by human muscle memory. It exploits cross-execution similarity through a dual-phase muscle-memory cache. The Context Cache reuses visual-token outputs to reduce computation, while the Action Cache reuses neuron activation patterns to reduce weight accesses. We keep both the cache memory footprint and access overhead low through online cache recomputation, sliding-window cache retrieval, and mask sharing across consecutive denoising steps. rMuscle achieves 1.29-1.42X speedup on RTX 4090 and Jetson Thor across LIBERO, RoboTwin, and physical manipulation tasks, while maintaining the original success rates on real-world robots.

Comment: Caches visual representations and neuron activation patterns across executions to reduce computation and weight accesses.

Topic Match: The central contribution is a model-state reuse mechanism that reduces inference cost, with applicability narrowed by its dependence on repetitive robotic workloads.

Relevance: 7 Novelty: 6


Paper Selection Prompt

System Prompt

You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.

User Prompt

Relevant Topics

This is a training-side feed. The centre of gravity is Mixture-of-Experts training and the large-scale training systems it sits on; architecture and efficiency work is kept when it bears on how a large model is trained or what it costs to train.

Keep a paper when its CORE CONTRIBUTION falls in one of the four topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.

  1. MoE Training (primary) - Keep: expert routing and gating mechanisms; load balancing, with or without auxiliary losses; expert granularity, fine-grained and shared experts; MoE training stability (loss spikes, router collapse, z-loss and related fixes); expert and hybrid parallelism; all-to-all dispatch/combine and the communication schedules MoE needs; MoE kernels and grouped-GEMM style implementations; expert capacity, token dropping, and upcycling dense checkpoints into MoE. - Filter: papers that merely train on top of a MoE model without a new routing, balancing, stability, parallelism, or kernel contribution; "mixture of experts" in the classical ensemble or recommender sense.

  2. Large-Scale Training Systems - Keep: distributed training algorithms (data, tensor, pipeline, sequence, and expert parallelism, FSDP and ZeRO-style sharding); optimizers and preconditioners for large-scale pretraining; communication algorithms, overlap, and collective design; throughput and MFU work that materially changes what a training run costs or whether it converges; scaling-law work that informs how such runs are configured. - Filter: routine infrastructure tuning with no new algorithmic or systems idea.

  3. Architecture and Training Dynamics - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, dynamic or modular computation); optimisation and training-dynamics analysis that explains why large models train the way they do. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight.

  4. Efficiency and Compression - Keep: quantization, sparsity, pruning, low-rank adaptation, KV-cache and memory-efficient designs, and other work that materially changes the cost of running or training a large model, especially when the mechanism is new rather than a tuned variant. - Filter: straightforward application of standard efficiency methods, and deployment or serving engineering with no new idea.

Leave these to the hotspot digest unless the core contribution clearly falls in one of the four topics above: - major model or product releases - agent, tooling, and RAG launches - benchmark, leaderboard, or evaluation-only papers - representation-learning theory, interpretability, and feature-formation analysis - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning) - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains

Scoring Criteria

Relevance and Novelty are independent axes. Score both from 1 to 10.

Relevance Scoring

  • 9-10: directly centered on the target foundational topics; highest when the core contribution is clearly within them.
  • 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
  • 5-6: touches the target topics, but the main contribution is elsewhere.
  • 3-4: largely outside the target topics, often application-focused or domain-specific.
  • 1-2: unrelated.

Important: Broad frontier relevance, major launch status, or daily buzz is not enough for a high Relevance score here. Those cases belong in the hotspot digest unless the paper strongly matches the specialized paper topics.

Novelty Scoring

  • 9-10: new paradigm, theory, or major methodological breakthrough.
  • 7-8: substantial methodological advance or strong new insight.
  • 5-6: meaningful but incremental extension or refinement.
  • 3-4: minor, narrow, or mostly engineering or domain-specific improvement.
  • 1-2: little originality; mainly standard application of existing methods.

Topic Registry

Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - moe_training: MoE Training - Mixture-of-Experts training: routing, load balancing, expert granularity and shared experts, training stability, MoE parallelism and communication, and MoE kernels. - training_systems: Large-Scale Training Systems and Efficiency - Large-scale training systems and efficiency: distributed training algorithms, optimizers, and communication. - architecture_training: Architecture and Training Dynamics - Core architectural or computational mechanisms, dynamic computation, and training-stability dynamics. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, memory or cache efficiency, and large-scale training systems that materially change cost or behavior.

Papers

[PAPER LIST HERE]

Instructions

Respond in JSONL. Output exactly one JSON object per paper, one per line:

{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}

Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only: daily_hot, new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return []. - daily_hot means the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. - new_frontier means the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.