Previous Day 2026-08-28
Monthly Overview 2026-08
Next Day 2026-08-31

This is a remedial run for missed papers from 08/28/2026 to 08/28/2026.

Results generated on 09/14/2026.

Personalized Daily ArXiv Papers 2026-08-29

Model Metric Usage Papers
Prompt Completion Total Total arXiv Scanned Relevant
local agent CLI Tokens not reported not reported not reported 424 424 33
Cost not reported not reported not reported

Token counts are not reported for this run. 20 of 20 model calls succeeded, 2,509s of model wall clock.

Topic Coverage:

TopicPapers
Large-Scale Training Systems and Efficiency4
Architecture and Training Dynamics11
Efficiency, Compression, and Large-Scale Training18

Table of contents by topic:

Large-Scale Training Systems and Efficiency (4)

  1. HARTS: Efficient Agentic Reinforcement Learning for Hybrid-Attention Models over Arbitrary Rollout Trees Authors: Boyuan Meng, Peihua Bao, Hong Liu, Xiaowei Zhu, Chao Wang, Gen Li, Zhenxuan Pan

  2. Learning-Theoretic Foundation for General Coded Computing: The Straggler Setting Authors: Parsa Moradi, Behrooz Tahmasebi, Mohammad Ali Maddah-Ali

  3. SpecGradFilter: A Spectral Gradient Filtering Framework for Taming Federated Heterogeneity Authors: Liyang Yuan, Yibo Yang, Dandan Guo, Peter Richtarik, Zhouchen Lin

  4. Ampere: Communication-Efficient and High-Accuracy Split Federated Learning Authors: Zihan Zhang, Leon Wong, Blesson Varghese

Architecture and Training Dynamics (11)

  1. InfoMamba: An Attention-Free Hybrid Mamba-Transformer Model Authors: Youjin Wang, Jiaqiao Zhao, Rong Fu, Run Zhou, Ruizhe Zhang, Jiani Liang, Suisuai Cao, Feng Zhou

  2. Learned Relay Representations for Forward-Thinking Discrete Diffusion Models Authors: Benjamin Rozonoyer, Jacopo Minniti, Dhruvesh Patel, Neil Band, Avishek Joey Bose, Tim G. J. Rudner, Andrew McCallum

  3. Attention as Conditioning: What Classical Learning Theory Predicts About Linear Transformers Authors: Mu Qiao

  4. The Approximation Rank of Softmax Attention: Sharp Geometric Laws and Robust Interaction Dimension Authors: Yuhe Sui, Jianing Zhang

  5. On the Depth Scalability of Logic Gate Networks Authors: Taegun An, Dohun kim, Haebeom Lee, Changhee Joo

  6. ABC: Any-Subset Autoregression via Non-Markovian Diffusion Bridges in Continuous Time and Space Authors: Gabe Guo, Thanawat Sornwanee, Lutong Hao, Elon Litman, Stefano Ermon, Jose Blanchet

  7. Memristive-Friendly Hadamard Reservoir Computing: Structured, Multiplier-Free Recurrences at Scale Authors: Andrea Ceni, Gianluca Milano, Carlo Ricciardi, Claudio Gallicchio

  8. SpikeOPD: Stable On-Policy Distillation for Autoregressive Spiking Language Models Authors: Enqiao Lu, Xingrui Yu, Yiwei Fu, Zhenglin Wan, Pengfei Zhou, Wangbo Zhao, Muqing Jian, Xueyi Zhang, Yang You, Ivor Tsang

  9. Equivariant Sheaf Neural Networks: Learning Geometric Transport on Graphs Authors: Alessio Borgi, Mario Severino, Fabrizio Silvestri, Pietro Liò

  10. Transformer-Encoder Trees for Efficient Multilingual Machine Translation and Speech Translation Authors: Yiwen Guan, Jacob Whitehill

  11. PolicyLong: Towards On-Policy Context Extension Authors: Junlong Jia, Jiang Zhou, Ziyang Chen, Xing Wu, Chaochen Gao, TingHao Yu, Feng Zhang, Songlin Hu

Efficiency, Compression, and Large-Scale Training (18)

  1. A Probabilistic Interpretation of KV Cache Eviction Authors: Renato Geh, Alex Chen, Daniel Israel, Aditya Grover, Guy Van den Broeck

  2. MultiHashFormer: Hash-based Generative Language Models Authors: Huiyin Xue, Atsuki Yamaguchi, Nikolaos Aletras

  3. HyQuant: Hybrid-Precision Quantization for LLM Attention Authors: Jiatong Ding, Bingxin Xing, Yu Zhang, Dian Ding, Xiaodong Yi, Xianbin Ouyang, Feihu Zhou, Kun Zhang, Zhenyu Guo, Hao Pan, Guangtao Xue, Yiming Zhang

  4. SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM Inference Authors: Daeha Lee, Do-Hyung Kim, Jae-Hong Kim

  5. Trust the Mass: Forced Weights in KV-Cache Eviction Authors: Jack Shi, Jerry Gu

  6. SABER: Stability-Aware Early Exit for LLM Reasoning via Adversarial Branch Probing Authors: Wanli Cheng, Haiya Xiang, Juntao Li, Hongling Wang, Wenliang Chen

  7. Learning What Matters: Supervising Global Context Pruning with Causal Evidence Sets Authors: James E. Allchin

  8. BRACE: Taming Sharp Irregularities via Barycentric Rational Forecasting for Fast Diffusion Transformers Inference Authors: Jinlong Yang, Jinke Wu, Lizilin, Yao Zhou

  9. The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning Authors: Dylan Jayabahu, Tinuade Adeleke

  10. Parser States Already Know: Structure-Conditioned KV Persistence for Structured Generation Authors: Linze Wu, Xinrui Chen

  11. DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory Architectures Authors: Peiming Yang, Sankeerth Durvasula, Ivan Fernandez, Mohammad Sadrosadati, Onur Mutlu, Gennady Pekhimenko, Christina Giannoula

  12. Speculative Probing: LLM Monitoring at Speculative-Decoding Cost Authors: Collin Zhang, Tingwei Zhang, Vitaly Shmatikov

  13. The information geometry of product-reference discrete diffusion: Interaction growth complexity and optimal scheduling Authors: Martin J. Wainwright

  14. AERA: Adaptive Evidence Residual Allocation for Efficient Test-Time Reasoning Authors: Ziming Wang, Ivor Tsang, Hangwei Qian

  15. Pro-Router: Token-Aware Progressive Model Routing with Adaptive Edge-Cloud Collaboration for Efficient Multimodal LLM Inference Authors: Xinyuan Gui, Shaowen Wang, Sheng Sun, Zijian Wang, Zishu Yu, Zheming Yang

  16. Nested Byte-Level Vocabularies Are Cheap to Deploy and Expensive to Share: A Pre-Registered Negative Result Authors: Christos Koutsiaris

  17. CoRe-MoE: Compact Reusable MoE for Continual Multimodal Instruction Tuning Authors: Runze Liu, Naibin Gu, Mingxu Ai, Yuqing Li, Peng Fu, Zheng Lin, Weiping Wang

  18. Sliding-window beats linear attention Authors: Alexia Jolicoeur-Martineau, Rhea Sanjay Sukthanker, Pashmina Cameron, Emy Gervais


Large-Scale Training Systems and Efficiency (4)

1. HARTS: Efficient Agentic Reinforcement Learning for Hybrid-Attention Models over Arbitrary Rollout Trees

ArXiv ID: 2608.28158

Primary Topic: Large-Scale Training Systems and Efficiency

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Boyuan Meng, Peihua Bao, Hong Liu, Xiaowei Zhu, Chao Wang, Gen Li, Zhenxuan Pan

Abstract: Agentic reinforcement learning (RL) often produces irregular rollout trees with shared histories. Training root-to-leaf trajectories independently recomputes these shared prefixes. Existing systems primarily target full-attention models and lack dense, differentiable hybrid-attention execution compatible with activation recomputation. We present HARTS (Hybrid-Attention RL over Tree Structures). HARTS jointly plans microbatches, data-parallel (DP) replica assignments, and microbatch-slot schedules using non-replay compact-token work after prefix compression. For chunkwise linear attention, a linear-time algorithm coordinates chunk-boundary state recovery and replay and produces the minimum number of sequential linear-attention calls under our packed execution model. HARTS preserves the chunkwise state partitioning of trajectory-wise training: it does not repeat projections, MLP/MoE computation, or final outputs, and performs only bounded state replay for numerical alignment. Per round, HARTS batches all branches into one packed call, propagates gradients through differentiable state handoffs, supports activation recomputation, and restores per-token log-probabilities. For deterministic, no-token-drop top-$k$ MoE routing, semantic multiplicities restore MoE-objective token weights and load statistics. Existing RL objectives retain their interface. To our knowledge, HARTS is the first system to demonstrate arbitrary-rollout-tree prefix-sharing speedups on a real hybrid-attention model. On an Agentic RL workload generated from SWE-bench tasks, HARTS achieves $4.81$--$4.87\times$ forward/backward/gradient speedup with activation recomputation across multiple parallel configurations. Its numerical differences are comparable to baseline self-rerun variation, and its reward trend is similar to the baseline over the first 120 steps of $τ^3$-Bench training.

Comment: Prefix-sharing execution and data-parallel scheduling reduce hybrid-attention training work while supporting differentiable state replay and activation recomputation.

Topic Match: The core contribution is a training execution and scheduling algorithm that eliminates redundant forward/backward work, despite its RL workload.

Relevance: 9 Novelty: 8


2. Learning-Theoretic Foundation for General Coded Computing: The Straggler Setting

ArXiv ID: 2608.28910

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Parsa Moradi, Behrooz Tahmasebi, Mohammad Ali Maddah-Ali

Abstract: Coded computing has emerged as a powerful paradigm for mitigating the impact of straggling workers in distributed computing systems. However, existing coded-computing schemes are predominantly designed for the exact recovery of highly structured computations, such as polynomial evaluation and matrix multiplication, and typically rely on strict recovery thresholds. These assumptions significantly limit their applicability to modern machine-learning workloads, particularly deep neural networks (DNNs), whose computations generally lack rigid algebraic structure and, in many applications, require only accurate approximations rather than exact recovery. To address this gap, we revisit coded computing from a learning-theoretic perspective and introduce General Coded Computing (GCC). Rather than adopting existing algebraic tools, GCC formulates coded computing through a natural end-to-end mean-squared error loss that directly measures the discrepancy between the desired computations and their recovered estimates. By deriving suitable upper bounds and restricting the encoder and decoder to a reproducing kernel Hilbert space (RKHS) with mild smoothness constraints, we show that both the encoder and decoder admit specific representations as linear combinations of RKHS kernel functions. This representation allows the corresponding coefficients to be computed efficiently. Moreover, this framework enables us to establish theoretical performance guarantees for GCC under two complementary straggler regimes. In the worst-case setting with $N$ worker nodes, and at most $S$ stragglers, we show that the end-to-end loss decays at least at rate $O(S^3N^{-3})$ for standard configurations. We then study a probabilistic setting in which each worker independently straggles with probability $p$. We prove that the expected loss can still converge at rate $O(\log_{1/p}^3(N)N^{-3})$.

Comment: RKHS-based coded computation mitigates stragglers with worst-case and probabilistic recovery-error guarantees.

Topic Match: Straggler-tolerant distributed computation is a relevant systems mechanism, although integration with large-model training is not demonstrated.

Relevance: 7 Novelty: 8


3. SpecGradFilter: A Spectral Gradient Filtering Framework for Taming Federated Heterogeneity

ArXiv ID: 2607.04189

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Liyang Yuan, Yibo Yang, Dandan Guo, Peter Richtarik, Zhouchen Lin

Abstract: Federated Learning (FL) is fundamentally challenged by statistical heterogeneity, where non-identically distributed (non-IID) data induces client drift that severely hampers global convergence. While existing approaches attempt to mitigate this drift through spatial-domain gradient correction or regularization, they overlook the intrinsic spectral structure of optimization signals. In this work, we revisit client drift from a novel frequency-domain perspective and uncover a critical Spectral Bias of Drift: inter-client gradient divergence is predominantly concentrated in low-frequency components which encode client-specific distributional shifts, while high-frequency components representing fine-grained features remain relatively consistent. Motivated by this, we propose SpecGradFilter, a unified Spectral Gradient Filtering Framework that tames heterogeneity by suppressing discordant low-frequency signals. Crucially, we demonstrate that SpecGradFilter is a generalizable principle, effective not only via precise FFT-based truncation but also through spatial approximations like Gaussian detrending. Extensive experiments on benchmarks such as CIFAR-10/100 and Tiny-ImageNet demonstrate that SpecGradFilter significantly performs better performance in highly Non-IID settings with negligible communication overhead, establishing a new paradigm for robust federated optimization.

Comment: Suppresses low-frequency gradient components to mitigate client drift in federated optimization.

Topic Match: A new distributed optimization rule fits training_systems, although its evidence comes from federated image-classification workloads.

Relevance: 7 Novelty: 7


4. Ampere: Communication-Efficient and High-Accuracy Split Federated Learning

ArXiv ID: 2507.07130

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Zihan Zhang, Leon Wong, Blesson Varghese

Abstract: A Federated Learning (FL) system collaboratively trains neural networks across devices and a server but is limited by significant on-device computation costs. Split Federated Learning (SFL) systems mitigate this by offloading a block of layers of the network from the device to a server. However, in doing so, it introduces large communication overheads due to frequent exchanges of intermediate activations and gradients between devices and the server and reduces model accuracy for non-IID data. We propose Ampere, a novel collaborative training system that simultaneously minimizes on-device computation and device-server communication while improving model accuracy. Unlike SFL, which uses a global loss by iterative end-to-end training, Ampere develops unidirectional inter-block training to sequentially train the device and server blocks with a local loss, eliminating the transfer of gradients. A lightweight auxiliary network generation method decouples training between the device and server, reducing frequent intermediate exchanges to a single transfer, which significantly reduces the communication overhead. Ampere mitigates the impact of data heterogeneity by consolidating activations generated by the trained device block to train the server block, in contrast to SFL, which trains on device-specific, non-IID activations. Extensive experiments on multiple CNNs and Transformers show that, compared to state-of-the-art SFL baseline systems, Ampere (i) improves model accuracy by up to 11.70 percentage points while training up to 18.6x faster, (ii) incurs up to 911x lower device-server communication overhead and up to 14.5x lower on-device computation, and (iii) reduces standard deviation of accuracy by 71.13% for various non-IID degrees highlighting superior performance when faced with heterogeneous data. Ampere is available from https://github.com/blessonvar/Ampere.

Comment: Decouples device and server training with local losses, eliminating gradient transfers and repeated activation exchange.

Topic Match: Introduces a distributed training decomposition that directly reduces communication and computation; its split-federated setting makes it peripheral to large-scale pretraining.

Relevance: 7 Novelty: 7


Architecture and Training Dynamics (11)

1. InfoMamba: An Attention-Free Hybrid Mamba-Transformer Model

ArXiv ID: 2603.18031

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Youjin Wang, Jiaqiao Zhao, Rong Fu, Run Zhou, Ruizhe Zhang, Jiani Liang, Suisuai Cao, Feng Zhou

Abstract: Balancing fine-grained local modeling with long-range dependency capture under computational constraints remains a central challenge in sequence modeling. While Transformers provide strong token mixing, they suffer from quadratic complexity, whereas Mamba-style selective state-space models (SSMs) scale linearly but often struggle to capture high-rank and synchronous global interactions. We present a consistency boundary analysis that characterizes when diagonal short-memory SSMs can approximate causal attention and identifies structural gaps that remain. Motivated by this analysis, we propose InfoMamba, an attention-free hybrid architecture. InfoMamba replaces token-level self-attention with a concept bottleneck linear filtering layer that serves as a minimal-bandwidth global interface and integrates it with a selective recurrent stream through information-maximizing fusion (IMF). IMF dynamically injects global context into the SSM dynamics and encourages complementary information usage through a mutual-information-inspired objective. Extensive experiments on classification, dense prediction, and non-vision tasks show that InfoMamba consistently outperforms strong Transformer and SSM baselines, achieving competitive accuracy-efficiency trade-offs while maintaining near-linear scaling.

Comment: Analyzes SSM limits in approximating attention and introduces a concept-bottleneck global interface with information-maximizing recurrent fusion.

Topic Match: The core contribution explains and changes sequence-model computation by combining selective recurrence with an attention-free global mechanism that maintains near-linear scaling.

Relevance: 9 Novelty: 8


2. Learned Relay Representations for Forward-Thinking Discrete Diffusion Models

ArXiv ID: 2605.22967

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Benjamin Rozonoyer, Jacopo Minniti, Dhruvesh Patel, Neil Band, Avishek Joey Bose, Tim G. J. Rudner, Andrew McCallum

Abstract: When Masked Diffusion Models (MDMs) generate sequences through iterative refinement, the rich internal computation over masked positions is discarded, forcing every subsequent refinement step to recompute the valuable internal information stored as model representations. To avoid a hard reset between denoising rounds, we propose Learned Relay Representations (Relay), a method that allows MDMs to be forward-thinking when denoising by explicitly learning how to propagate latent information for the benefit of future denoising steps. Relay introduces a differentiable per-token channel that passes information between forward passes and is trained via truncated backpropagation through time (BPTT). We show that this framework can be scaled to state-of-the-art Diffusion Language Models (DLMs), and is seamlessly compatible with techniques like block diffusion and KV caching. We first provide a thorough justification of the design choices in Relay on a challenging Sudoku-based planning task. We then scale Relay to Fast-dLLM v2, a state-of-the-art DLM, outperforming standard supervised finetuning on coding tasks while reducing inference latency by up to 32%. Our empirical results demonstrate that state-of-the-art DLMs can be explicitly trained to relay latent information forward across decoding steps, advancing the performance-latency Pareto frontier. We provide code for all our experiments.

Comment: A differentiable per-token relay trained with truncated BPTT carries latent computation across diffusion denoising passes.

Topic Match: Trainable recurrence changes the denoising architecture itself, while reusing intermediate information improves inference efficiency.

Relevance: 9 Novelty: 8


3. Attention as Conditioning: What Classical Learning Theory Predicts About Linear Transformers

ArXiv ID: 2508.08289

Primary Topic: Architecture and Training Dynamics

Authors: Mu Qiao

Abstract: Attention is widely understood as an associative memory, but that description alone does not predict how the memory will behave. Predictive theories do exist, but in the literature on animal learning. We show that the state updates of the major linear-attention families are term-for-term identical with named models from a century of animal learning theory: linear attention implements Hebbian contiguity, DeltaNet implements Rescorla--Wagner error correction, and decay variants such as RetNet implement contiguity with a stimulus trace. This dictionary turns conditioning phenomena into testable statements about the in-context behavior of linear transformers, while distinguishing algebraic consequences from empirical measurements. Algebraically, it yields an exact closed form for Kamin blocking, verified in simulation to $<10^{-7}$ across five learning rates. Empirically, it predicts a dissociation that survives training on generic in-context association: error-correcting attention exhibits cue competition, whereas contiguity-based attention does not. A single state also has two capacity regimes, with measured scaling exponents of 1.22 for faithful retrieval and 1.89 for identification, consistent with linear and near-quadratic predictions. Across the full head grid, retrieval error is governed primarily by total state size rather than its partition across heads, indicating that heads provide capacity rather than redundant copies. We also prove no spontaneous recovery for the analyzed single-state recurrences under cue-orthogonal retention trials; with a never-presented-cue control and probes within the trained positional range, we likewise find no recovery in trained models. Finally, we introduce PH-attention, a Pearce--Hall-inspired rule with an explicit feature-indexed associability state that yields cue-dependent learning rates and is absent from the token-computed gates we compare.

Comment: Analyzes linear-attention update rules and introduces feature-specific associability through PH-attention.

Topic Match: The algebraic analysis and new update rule directly concern recurrent attention mechanisms and their computational behavior.

Relevance: 8 Novelty: 8


4. The Approximation Rank of Softmax Attention: Sharp Geometric Laws and Robust Interaction Dimension

ArXiv ID: 2608.28150

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Yuhe Sui, Jianing Zhang

Abstract: Which geometry controls the rank complexity of normalized softmax attention? We study maximum-row-$\ell_1$ approximation rank, exactly the least unrestricted rank preserving every bounded vector-valued output. Two sharp worst-case laws isolate support geometry: for fixed $d$ and error $\varepsilon$, spherical self-attention has rank $Θ{d,\varepsilon}(\min{n,(1+β)^{(d-1)/2}})$, while full-ball geometry adds one radial degree and, for $β\geβ_0(d,\varepsilon)$ and $n\ge C_d e^{β/8}$, gives $Θ)$. For a fixed head, row-softmax quotients out row-scalar logit directions: the remaining visible query--key interaction dimension $r$ yields an $r/2$ per-instance upper law, and bounded constructions show this exponent is minimax sharp. Approximate interaction subspaces incur an explicit residual output error and yield a tolerance-indexed SVD dimension. On an 84-head BERT-base calibration set, we observe modest effective-dimension reductions across many head--temperature settings, together with positive associations with finite constructive rank upper certificates. Together, these results separate support geometry, which sets worst-case temperature scaling, from softmax-visible interaction geometry, which controls per-head approximation complexity.}(β^{d/2

Comment: Sharp output-preserving rank bounds characterize when softmax attention admits low-rank approximation.

Topic Match: Analyzes the attention operator itself, deriving geometric complexity laws that also inform low-rank attention compression.

Relevance: 8 Novelty: 8


5. On the Depth Scalability of Logic Gate Networks

ArXiv ID: 2607.21633

Primary Topic: Architecture and Training Dynamics

Authors: Taegun An, Dohun kim, Haebeom Lee, Changhee Joo

Abstract: Logic Gate Networks (LGNs) compute through compositions of Boolean operations, yet existing LGNs do not reliably benefit from increased depth. We identify two causes: optimization collapse and topology-induced degradation of output-specific credit that persists even after skip-biased initialization and straight-through estimation stabilize training. We introduce Input-Anchored Logic Gate Networks (IALGNs), in which each gate combines a private hidden spine with a direct input anchor. This topology prevents output-path merging while retaining input access at every layer. Credit diagnostics show that random wiring dilutes or conflicts output-specific gradients, whereas IALGN maintains usable and coherent credit. Random-$k_x$ relaxation improves anchor selection without relaxing the spine. Across MNIST, CIFAR-10, and CIFAR-100, IALGN exhibits consistent fixed-width depth--accuracy scaling up to 150 layers, while alternative topologies saturate or degrade. Linear probes, topology ablations, and operation-aware analysis show that trained IALGNs preserve private states and apply sparse anchor-conditioned updates. These results indicate that scalable LGN depth requires both stable optimization and credit-preserving information access.

Comment: Introduces input-anchored gate topology that preserves output-specific gradient credit as network depth increases.

Topic Match: The contribution connects architectural topology to optimization stability and depth scaling, with scope limited to logic gate networks.

Relevance: 8 Novelty: 7


6. ABC: Any-Subset Autoregression via Non-Markovian Diffusion Bridges in Continuous Time and Space

ArXiv ID: 2604.27443

Primary Topic: Architecture and Training Dynamics

Authors: Gabe Guo, Thanawat Sornwanee, Lutong Hao, Elon Litman, Stefano Ermon, Jose Blanchet

Abstract: Generating continuous-time, continuous-space stochastic processes (e.g., videos, weather forecasts) conditioned on partial observations (e.g., first and last frames) is a fundamental challenge. Existing approaches, (e.g., diffusion models), suffer from key limitations: (1) noise-to-data evolution fails to capture structural similarity between states close in physical time and has unstable integration in low-step regimes; (2) random noise injected is insensitive to the physical process's time elapsed, resulting in incorrect dynamics; (3) they overlook conditioning on arbitrary subsets of states (e.g., irregularly sampled timesteps, future observations). We propose ABC: Any-Subset Autoregressive Models via Non-Markovian Diffusion Bridges in Continuous Time and Space. Crucially, we model the process with one continual SDE whose time variable and intermediate states track the real time and process states. This has provable advantages: (1) the starting point for generating future states is the already-close previous state, rather than uninformative noise; (2) random noise injection scales with physical time elapsed, encouraging physically plausible dynamics with similar time-adjacent states. We derive SDE dynamics via changes-of-measure on path space, yielding another advantage: (3) path-dependent conditioning on arbitrary subsets of the state history and/or future. To learn these dynamics, we derive a path- and time-dependent extension of denoising score matching. Our experiments show ABC's superiority to competing methods on multiple domains, including video generation and weather forecasting.

Comment: Introduces path-dependent diffusion-bridge autoregression whose learned SDE evolves in physical time and supports arbitrary-subset conditioning.

Topic Match: The contribution changes the generative sequence mechanism and its score-matching training objective; its continuous-process focus makes it a narrower architectural match.

Relevance: 7 Novelty: 8


7. Memristive-Friendly Hadamard Reservoir Computing: Structured, Multiplier-Free Recurrences at Scale

ArXiv ID: 2608.28295

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Andrea Ceni, Gianluca Milano, Carlo Ricciardi, Claudio Gallicchio

Abstract: Reservoir Computing (RC) designs Recurrent Neural Networks around a fixed, i.e., untrained, recurrent layer, and is a natural candidate for neuromorphic hardware. Memristive-friendly reservoirs derive the neuron dynamics from memristive-device kinetics, but still rely on dense recurrent matrices, which are expensive to realize physically. In this paper, we replace the dense matrix with a structured orthogonal operator, built from sign diagonals, a permutation, and a fast Walsh-Hadamard transform. The operator is multiplier-free, requires $O(N)$ parameters and $O(N\log N)$ operations per step, and is never materialized as a matrix. We instantiate it in a standard and in a memristive-friendly Echo State Network, with one binary input connection per unit. Our mathematical analysis shows that exact orthogonality yields an echo state condition that is tight in the recurrent scaling, and a noise response that is predictable at design time. Moreover, the operator mixes the whole state in a single application. Experiments on twenty classification and seven regression benchmarks, at reservoir sizes up to $N = 8192$, show that the structured models match dense orthogonal reservoirs, and achieve better mean performance than the cycle reservoir by a margin that widens with size. Furthermore, we time the recurrent step on three hardware platforms, where it is up to $50\times$ faster than a dense product and $10^4\times$ smaller in memory. Finally, we ablate the operator and measure the response to noise, quantization, device mismatch and discrete faults.

Comment: Structured orthogonal Hadamard recurrence replaces dense state mixing with O(N log N) computation.

Topic Match: The structured recurrent operator and its stability analysis fit architecture_training, with computation and memory savings demonstrated in fixed reservoir networks.

Relevance: 7 Novelty: 7


8. SpikeOPD: Stable On-Policy Distillation for Autoregressive Spiking Language Models

ArXiv ID: 2608.27857

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Enqiao Lu, Xingrui Yu, Yiwei Fu, Zhenglin Wan, Pengfei Zhou, Wangbo Zhao, Muqing Jian, Xueyi Zhang, Yang You, Ivor Tsang

Abstract: Spiking neural networks (SNNs) offer a path to energy-efficient language modeling through sparse encoding and event-driven computation, but training capable spiking language models from scratch remains difficult. A practical alternative is ANN-to-SNN migration through knowledge distillation (KD), where a pretrained artificial neural network (ANN) teacher supervises an SNN student. Existing migration approaches distill on fixed corpus prefixes, whereas autoregressive inference conditions on self-generated prefixes, creating prefix-source mismatch. It manifests as output-policy mismatch with the ANN teacher and internal spiking-dynamics drift between self-generated and matched corpus prefixes. On-policy distillation (OPD) offers a natural way to mitigate both manifestations by continuing teacher supervision on self-generated prefixes. We evaluate a teacher-only full-KL variant, Vanilla OPD, via a controlled stress test and observe it may suffer from delayed rollout-feedback collapse. This result shows that on-policy coverage alone does not ensure stable adaptation. Motivated by these findings, we propose SpikeOPD, a stable on-policy distillation framework for autoregressive SNNs that learns from self-generated prefixes while maintaining rollout stability. It applies full-KL teacher correction to reduce output-policy mismatch, while matched-prefix policy anchoring constrains policy departure from the frozen reference SNN on the same prefixes. Layerwise spike regularization further limits firing-rate deviations during on-policy adaptation. Across three model scales, SpikeOPD improves average accuracy over the corresponding KD SNNs by 0.8, 1.7, and 2.9 points at 0.125B, 0.35B, and 1.3B, respectively, while preserving their sparse-compute profiles.

Comment: Matched-prefix anchoring and spike-rate regularization stabilize autoregressive spiking-model distillation against rollout collapse.

Topic Match: Spiking-dynamics drift and rollout collapse provide a concrete architecture-specific training contribution; sparsity-preserving ANN-to-SNN conversion adds an efficiency connection.

Relevance: 7 Novelty: 6


9. Equivariant Sheaf Neural Networks: Learning Geometric Transport on Graphs

ArXiv ID: 2608.28853

Primary Topic: Architecture and Training Dynamics

Authors: Alessio Borgi, Mario Severino, Fabrizio Silvestri, Pietro Liò

Abstract: Equivariant graph neural networks provide a principled way to model geometric systems, but efficient first-order architectures remain limited in how vector information can be transformed as it moves across a graph. We introduce \textsc{ESNN}, an Equivariant Sheaf Neural Network that enriches this interaction by learning directed, matrix-valued transport between neighboring vector features while preserving exact Euclidean equivariance. Rather than increasing the order of the representation, ESNN keeps scalar and vector features first-order and places the additional geometric flexibility in the edge transport itself. We characterize this transport theoretically, showing that when relative displacement is the only covariant geometric input, every linear $O(n)$-equivariant map decomposes into independent radial and tangential components, while learned covariant features enable richer feature-conditioned transformations. We also introduce controlled symmetry relaxation for systems with a preferred ambient direction, which may be prescribed or inferred from data while recovering full $E(n)$-equivariance when the directional pathway is inactive. Across particle dynamics, mesh-based simulation, point-cloud classification, and molecular property prediction, ESNN improves dynamics prediction, recovers the gravity axis when symmetry is broken, yields substantial gains on selected mesh tasks and long-horizon rollouts, and remains robust to unseen rotations. These results show that learning how geometric information is transported across edges offers a complementary route to expressive equivariant message passing without requiring higher-order representations.

Comment: Learned matrix-valued edge transport increases equivariant message-passing expressivity without increasing representation order.

Topic Match: Introduces a theoretically analyzed graph-layer mechanism, although its connection to large-model training is indirect.

Relevance: 6 Novelty: 7


10. Transformer-Encoder Trees for Efficient Multilingual Machine Translation and Speech Translation

ArXiv ID: 2509.17930

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Yiwen Guan, Jacob Whitehill

Abstract: Multilingual translation suffers from computational redundancy, especially when translating into multiple languages simultaneously. In addition, translation quality can suffer for low-resource languages. To address this, we introduce Transformer Encoder Tree (TET), a hierarchical, non-autoregressive encoder-only architecture trained with Connectionist Temporal Classification (CTC) for multilingual translation. TET shares intermediate representations among linguistically similar target languages, improving accuracy on low-resource languages while reducing computational redundancy and enabling the generation of all target languages in a single forward pass. TET eliminates the sequential bottleneck of autoregressive models and supports fully parallel decoding of all tokens across all target languages. Compared to a naive one-to-many multilingual design, TET reduces the total parameter count by 66% and lowers inference computation by 60%. In speech translation, combining TET with a non-autoregressive speech recognition backbone (Wav2Vec2) shows competitive translation quality compared to autoregressive systems while speeding up inference by approximately 7-14 times.

Comment: Hierarchical encoder sharing reduces duplicated computation while CTC enables parallel generation across target languages.

Topic Match: Shared encoder branches introduce a computational architecture with parameter and inference savings, although the design and evidence remain specific to multilingual translation.

Relevance: 6 Novelty: 6


11. PolicyLong: Towards On-Policy Context Extension

ArXiv ID: 2604.07809

Primary Topic: Architecture and Training Dynamics

Authors: Junlong Jia, Jiang Zhou, Ziyang Chen, Xing Wu, Chaochen Gao, TingHao Yu, Feng Zhang, Songlin Hu

Abstract: Extending LLM context windows is hindered by scarce high-quality long-context data. Recent methods synthesize data with genuine long-range dependencies via information-theoretic verification, selecting contexts that reduce a base model's predictive entropy. However, their single-pass offline construction with a fixed model creates a fundamental off-policy gap: the static screening landscape misaligns with the model's evolving capabilities, causing the training distribution to drift. We propose PolicyLong, shifting data construction towards a dynamic on-policy paradigm. By iteratively re-executing data screening (entropy computation, retrieval, and verification) using the current model, PolicyLong ensures the training distribution tracks evolving capabilities, yielding an emergent self-curriculum. Crucially, both positive and hard negative contexts derive from the current model's entropy landscape, co-evolving what the model learns to exploit and resist. Experiments on RULER, HELMET, and LongBench-v2 (Qwen2.5-3B) show PolicyLong consistently outperforms EntropyLong and NExtLong, with gains growing at longer contexts (e.g., +2.54 at 128K on RULER), confirming the value of on-policy data evolution.

Comment: Refreshes entropy-based data selection using the current model to keep long-context curricula aligned with learning progress.

Topic Match: Curriculum adaptation is adjacent to training dynamics; the core contribution is model-dependent data construction for context extension.

Relevance: 6 Novelty: 6


Efficiency, Compression, and Large-Scale Training (18)

1. A Probabilistic Interpretation of KV Cache Eviction

ArXiv ID: 2608.28293

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Renato Geh, Alex Chen, Daniel Israel, Aditya Grover, Guy Van den Broeck

Abstract: The premise and promise of KV (cache) eviction is simple: higher throughput can be achieved by evicting some entries from the KV cache, at a negligible cost to quality. This holds empirically for many existing methods, though most rely on creative heuristics for selecting which entries to drop. Despite recent advances, the problem of KV eviction has remained informal in the literature. This paper aims to properly formalize this problem through the lens of probabilistic reasoning and reveal what can be learned from this perspective. Concretely, we (1) formalize the problem of KV eviction and, unfortunately, prove that it is computationally hard, (2) show that by framing it probabilistically, KV eviction reduces to the problem of expectation estimation, which can be approximated through sampling, (3) show that through this probabilistic interpretation, correcting for evicted entries during decoding---a previously ignored problem---becomes feasible, and (4) reveal that existing methods in the literature are zero-variance biased estimators that can be easily adapted in order to enable decode time correction. In practice, we show that this probabilistic version of KV eviction coupled with decode time correction is more robust to different tasks compared to existing eviction methods and achieves competitive performance at the same compression budget.

Comment: Formulates KV eviction as expectation estimation, enabling sampling-based approximation and correction during decoding.

Topic Match: The central contribution is a principled cache-compression mechanism with decoding corrections under limited memory.

Relevance: 9 Novelty: 8


2. MultiHashFormer: Hash-based Generative Language Models

ArXiv ID: 2606.28057

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Huiyin Xue, Atsuki Yamaguchi, Nikolaos Aletras

Abstract: Language models (LMs) represent tokens using embedding matrices that scale linearly with the vocabulary size. To constrain the parameter footprint, prior work proposes hashing many tokens into a single vector within encoder-only models. While this offers parameter efficiency, many-to-one collisions prevent its use in causal LMs. In this paper, we propose MultiHashFormer, a new framework that allows hash-based autoregression. Each token is represented as a unique hash signature, a short sequence of discrete hash IDs, generated by multiple independent hash functions. A Hash Encoder compresses this signature into a single latent vector for processing by a Transformer decoder. Then, a Hash Decoder generates the hash signature of the next token, which is then mapped back to text. We evaluate our approach at the 100M, 1B and 3B parameter scales, demonstrating that MultiHashFormer consistently outperforms standard Transformer LMs across multiple benchmarks. Furthermore, we show that our model handles multilingual vocabulary expansion with a constant parameter footprint without any modifications.

Comment: Multi-hash token signatures enable autoregressive generation and vocabulary expansion without enlarging the model's parameter footprint.

Topic Match: The primary contribution changes vocabulary-dependent parameter scaling through a new hash-signature encoder and decoder, with validation up to 3B parameters.

Relevance: 9 Novelty: 8


3. HyQuant: Hybrid-Precision Quantization for LLM Attention

ArXiv ID: 2608.27875

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Jiatong Ding, Bingxin Xing, Yu Zhang, Dian Ding, Xiaodong Yi, Xianbin Ouyang, Feihu Zhou, Kun Zhang, Zhenyu Guo, Hao Pan, Guangtao Xue, Yiming Zhang

Abstract: Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the \emph{attention} module often introduces large errors at very low bit-widths, causing performance degradation. Existing methods mainly rely on smoothing techniques to handle outliers, while we propose a hybrid quantization design to better balance accuracy and efficiency. Specifically, we propose \textbf{HyQuant}, an efficient hybrid quantization framework for LLM attention. HyQuant quantizes most attention states into low-bit formats while retaining a small set of vertical-line tokens and local-window states in high precision. These accuracy-critical regions are selected using lightweight vertical-line-aware attention-pattern signals, reducing quantization error with limited overhead. In the Prefill stage, HyQuant uses a hybrid-precision quantized attention operator that preserves vertical-line tokens and a local sliding window in full precision while quantizing the remaining context. In the Decode stage, HyQuant applies the same principle to KV-cache compression and fuses KV dequantization with attention computation to improve memory and hardware efficiency. Across diverse tasks, models, and datasets, HyQuant maintains nearly lossless accuracy with an extremely simple design, demonstrating the efficiency and practical feasibility of hybrid quantization for LLM attention. Code is available at: https://github.com/jerrysfls/HyQuant .

Comment: Preserving attention-critical tokens and local windows in high precision enables selective low-bit attention and KV-cache compression.

Topic Match: Hybrid attention quantization and fused KV-cache dequantization directly target large-model memory and computation costs.

Relevance: 9 Novelty: 7


4. SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM Inference

ArXiv ID: 2608.28911

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Daeha Lee, Do-Hyung Kim, Jae-Hong Kim

Abstract: The key-value (KV) cache is the dominant memory bottleneck of long-context large language model (LLM) inference, growing linearly with context length. We show that uniform KV quantization on a fractional-bit grid does not degrade gracefully: under a prespecified multi-seed statistical protocol, Llama-3.1-8B-Instruct with an affine quantizer is statistically indistinguishable from FP16 KV down to 2.322 code bits/value and collapses at 2.0 bits - a quality cliff in (2.0, 2.322] that reappears in generation-time quantization and multi-turn dialogue and transfers to Mistral-7B. The cliff reframes importance-aware mixed precision: above it, eight model-internal importance indicators are statistically interchangeable, so the benefit of mixing is grid interpolation, reaching average precisions uniform quantization cannot realize. SemKV preserves every token, ranks tokens by a model-internal score, and assigns two adjacent above-cliff precisions, achieving a measured 6.0x storage reduction with no statistically detectable quality difference from full KV (n=900, three seeds), and outperforming FP16 token pruning granted a 1.5x larger memory budget. Replacing the affine base with a distortion-optimized quantizer (TurboQuant-MSE) lowers the cliff in every protocol tested, raising the no-detectable-loss operating point to 7.9x. The recipe: measure the cliff for the target deployment setting, then interpolate above it.

Comment: Cliff-calibrated mixed-precision KV quantization interpolates between safe precisions while retaining every token.

Topic Match: KV-cache compression is directly in scope; the main contribution is an operating-point selection principle supported by measured storage savings and statistical quality comparisons.

Relevance: 9 Novelty: 6


5. Trust the Mass: Forced Weights in KV-Cache Eviction

ArXiv ID: 2608.25230

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Jack Shi, Jerry Gu

Abstract: Every deployed sparse-attention or KV-cache-eviction rule keeps a subset of the keys, discards the rest, and renormalizes the attention weights over the kept set. Enumerating the exact best subset under that constraint on $168{,}192$ attention rows from five models shows that keeping the largest weights is already near-optimal, since the best subset closes only a median $2$ to $5\%$ of the remaining gap to full attention. If selection closes this little, published margins between eviction methods must come from elsewhere, so we measure the bytes each method holds. In the shared evaluation pipeline, the strongest query-agnostic methods hold the full cache because their per-head selections are stored as masks, and only ragged per-head storage frees that memory. Enforcing a nominal budget on one fixed selection costs $14$ to $62$ benchmark points. We trace an $87.6$-point retrieval margin to rankings computed while the question is visible. ContourKV, a training-free allocator built from the dropped-mass statistic, wins $93$ of $160$ paired comparisons against that state of the art and loses $22$ at the byte count of the budget-enforcing baselines, and it ties the strongest of them.

Comment: Introduces dropped-mass KV-cache allocation evaluated under actual byte budgets.

Topic Match: The cache allocator and analysis of physical memory savings provide a concrete compression contribution alongside the evaluation audit.

Relevance: 8 Novelty: 7


6. SABER: Stability-Aware Early Exit for LLM Reasoning via Adversarial Branch Probing

ArXiv ID: 2608.27963

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Wanli Cheng, Haiya Xiang, Juntao Li, Hongling Wang, Wenliang Chen

Abstract: Large Reasoning Models (LRMs) achieve strong reasoning capabilities, yet long-chain reasoning becomes inefficient once the intermediate answer stabilizes across reasoning steps: additional reasoning yields little marginal benefit while incurring substantial inference cost. Existing early-exit methods based on confidence or entropy poorly capture reasoning stability, while consistency-based approaches rely on multi-step trajectory agreement, requiring sequential evaluations that delay exit. To better balance efficiency and reliability, we propose SABER, a training-free framework for stability-aware early exit via adversarial branch probing. SABER constructs simple yet effective semantic perturbations around intermediate reasoning states to form adversarial branches, and applies lightweight probing to estimate their likely final outcomes without full trajectory rollouts. When the probed outcomes remain consistent across branches, SABER exits early; otherwise, it continues reasoning. Experiments across multiple reasoning benchmarks and model architectures show that SABER reduces reasoning token consumption by 30.2\%--39.8\% on average while maintaining competitive accuracy with full-length reasoning.

Comment: Uses adversarial branch probes to decide when reasoning can terminate early.

Topic Match: Its central mechanism adaptively reduces inference computation through a stability-based early-exit criterion.

Relevance: 8 Novelty: 7


7. Learning What Matters: Supervising Global Context Pruning with Causal Evidence Sets

ArXiv ID: 2607.21692

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: James E. Allchin

Abstract: Pruning a long context means committing to the blocks a model will keep, and the usual selector is distilled from a dense teacher's attention. That assumes attention shows which context the answer depends on. We test the assumption on retrieval tasks where the evidence is known exactly, by masking context and measuring whether the answer changes. Attention and causal dependence disagree. Teachers attend to outdated facts that the answer does not depend on, and they attend differently across training runs that use the same evidence. Selectors trained on that attention copy both failures. On a multi-hop retrieval task, a selector distilled from attention routes at 36% to 98% depending on the training run. The same selector trained on causal evidence sets reaches 99% or better on every run. Dense accuracy does not tell the teachers apart. Masking the frozen teacher recovers the causal sets of these tasks without annotations. Frozen pretrained models show the same conflict, and selectors supervised with known evidence labels beat attention-based eviction through 32B when context must be pruned before the question arrives.

Comment: Trains context-pruning selectors using causal evidence sets recovered by masking a frozen teacher.

Topic Match: The core contribution changes the supervision used to learn context pruning, directly supporting memory-efficient long-context computation.

Relevance: 8 Novelty: 7


8. BRACE: Taming Sharp Irregularities via Barycentric Rational Forecasting for Fast Diffusion Transformers Inference

ArXiv ID: 2608.07572

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Jinlong Yang, Jinke Wu, Lizilin, Yao Zhou

Abstract: Diffusion Transformers (DiTs) have demonstrated exceptional performance in high-fidelity image and video generation. To alleviate their massive computational overhead, temporal feature caching has been proposed to bypass redundant computations. However, existing cache-then-forecast methods driven by derivative-based polynomials often cause severe quality degradation under high acceleration due to unstable long-step predictions. To address this bottleneck, we propose Barycentric Rational Forecasting with Chebyshev Enhancement (BRACE). Motivated by the observation that DiT feature trajectories are globally smooth yet frequently exhibit sharp irregularities and local non-smoothness, BRACE shifts the paradigm from derivative-driven polynomial extrapolation to feature-driven rational forecasting. Specifically, it maintains a local sliding window to cache sparse historical features and leverages adapted Chebyshev weights to formulate a barycentric rational function, directly aggregating these raw features to ensure numerical stability. Extensive experiments demonstrate that BRACE achieves state-of-the-art quality-efficiency trade-offs across various DiT architectures with negligible computational overhead.

Comment: Forecasts cached DiT features with barycentric rational functions to bypass repeated computation.

Topic Match: The core contribution is a feature-cache forecasting mechanism that reduces diffusion-transformer inference computation while controlling prediction instability.

Relevance: 8 Novelty: 7


9. The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning

ArXiv ID: 2608.28859

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Dylan Jayabahu, Tinuade Adeleke

Abstract: Reasoning models do not stop when they know the answer. On DeepSeek-R1-Distill-Qwen-7B the chain of thought runs about twice as long as the model's own answer probability takes to settle, and how much of that excess is removable varies from problem to problem, so a global length penalty cannot take it out. We take it out by internalizing a causal interpretability finding into the weights. The mechanism is a halt vector: a difference-of-means direction at layer 18 of this model whose steering strength controls how long it thinks, while a replicated value axis does nothing. Installing that intervention in the weights is harder than it looks. Maximizing the scalar projection onto the direction corrupts the off-axis dimensions a frozen downstream reader depends on, and generation gets longer instead of shorter; what works is reconstructing the whole steered activation with those dimensions pinned to their natural values. Fit from 24 problems and no reinforcement learning, the halt removes about a quarter of the thinking at held accuracy across five unseen benchmarks, and the cut tracks each problem's own removable slack at 0.70. It also closes a non-termination pathology that grows with difficulty and that a decoding-time confidence hook makes worse. We do not claim to beat a well-tuned length penalty or decoding-time early exit on the raw trade-off; the contribution is how the halt is obtained.

Comment: Internalizing a halt intervention through activation reconstruction reduces reasoning tokens by about 25% at held accuracy.

Topic Match: Learned stopping directly reduces inference computation and introduces a mechanism for problem-dependent computation length.

Relevance: 8 Novelty: 7


10. Parser States Already Know: Structure-Conditioned KV Persistence for Structured Generation

ArXiv ID: 2608.28276

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Linze Wu, Xinrui Chen

Abstract: Structured generation underpins large language model (LLM) agents that produce JSON, SQL, and function calls, where a single wrong field can cause the downstream action to fail. Constrained decoding already tracks parser transitions to enforce formal validity, and these transitions expose how generated tokens participate in schema-critical decisions such as required fields, arguments, and structural boundaries under the active grammar. Existing KV compression largely leaves this task-relevant structural signal unused. We introduce PASK (Parser-Aware Structural KV Persistence), which turns parser-derived structure into layer-group-specific KV persistence decisions. PASK addresses the mismatch between model-side KV sensitivity and task-level structured risk by using task-error sensitivity to set minimum protection floors and attention-output distortion to allocate residual KV capacity. An offline calibration stage compiles these signals into a persistence policy, leaving only lightweight structure-conditioned lookup online. At a targe total KV budget of 0.33, PASK outperforms the strongest compressed baseline by 17.39 percentage points on average across eight BFCL non-live and Live subcategories on Qwen3-4B. In end-to-end serving, PASK achieves up to 2.2x higher throughput and 3.3x lower TPOT, while using 0.53x the peak GPU memory of Full KV.

Comment: Parser-conditioned KV retention allocates layer-group cache budgets using task-error sensitivity and attention distortion.

Topic Match: A new KV compression policy is the core contribution, with structured generation providing the retention signal.

Relevance: 8 Novelty: 7


11. DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory Architectures

ArXiv ID: 2511.15503

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Peiming Yang, Sankeerth Durvasula, Ivan Fernandez, Mohammad Sadrosadati, Onur Mutlu, Gennady Pekhimenko, Christina Giannoula

Abstract: High-performance Host processors can integrate Processing-In-Memory (PIM) devices, which can accelerate memory-intensive kernels of Machine Learning (ML) models, including Large Language Models (LLMs), by leveraging the large memory bandwidth available at PIM cores. However, Host processor needs consecutive elements distributed across DRAM banks, while PIM cores need consecutive elements within their local banks. This necessitates data rearrangements in ML kernel execution that pose significant performance and programmability challenges, further exacerbated by the need to support diverse PIM devices. Current compilation approaches lack systematic optimization for diverse ML kernels and multiple PIM devices, and may largely ignore data rearrangement costs during the compute code optimization step. We show that data rearrangements and compute code optimization are interdependent, and need to be jointly optimized during the tuning process. Therefore, we design DCC, the first data-centric ML compiler for PIM systems that jointly co-optimizes data rearrangements and compute code in a unified tuning process. DCC integrates a multi-layer PIM abstraction to support multiple PIM backends. DCC enables effective co-optimization of data partitioning strategies with compute loop partitioning schemes. DCC applies PIM-specific code optimizations, and leverages a fast and accurate performance prediction model to select the bestperforming code schedule for a given kernel on a target PIM architecture. Our evaluations in various individual ML kernels show that DCC achieves up to 7.68x speedup (2.21x average) on HBM-PIM, and up to 13.17x speedup (3.92x average) on AttAcc PIM, over GPU-only execution. In end-to-end LLM inference, DCC on AttAcc accelerates GPT-3 and LLaMA-2 by 4.52x average (up to 7.71x in LLaMA-2) over GPU. DCC is open-sourced at https://github.com/SPIN-Research-Group/DCC.

Comment: Jointly optimizing PIM data rearrangements and compute schedules substantially reduces LLM inference cost.

Topic Match: The compiler introduces a concrete mechanism for reducing memory-intensive execution costs, with end-to-end LLM results; its specialized PIM focus makes it somewhat peripheral to training.

Relevance: 8 Novelty: 7


12. Speculative Probing: LLM Monitoring at Speculative-Decoding Cost

ArXiv ID: 2608.28099

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Collin Zhang, Tingwei Zhang, Vitaly Shmatikov

Abstract: Real-time classification during language model inference is valuable for safety filtering, behavioral analysis, and model monitoring, but current approaches force a trade-off between accuracy and efficiency. Hidden-state probes are fast but limited: they are either not context-aware: operating on a single vector and cannot model interactions across positions; or they are very costly: having dedicated classifier models (Llama Guard, Qwen Guard, LLM-as-judge) or performing computation on hidden states for all tokens and then pooling the results (MultiMax). This shows an intrinsic trade-off between efficiency and accuracy. However, we find that the speculative-decoding module in recent LLMs can be repurposed for efficient high-quality classification. By appending a trained soft prompt at the end of the target sequence, we can repurpose the speculative-decoding module into a sequence classifier. At inference time in a speculative-decoding pipeline, the KV cache is already in GPU memory, so classification adds negligible overhead. We evaluate on four classification tasks across four models (Qwen3.5-4B, 9B, 27B, MiniCPM4.1-8B). Our small probes consistently outperform zero-shot GPT-5.4-mini and, on multilingual prompt safety, match or beat specialized 8B safety classifiers (Qwen3Guard-Gen-8B, Llama-Guard-3-8B) without running a full LLM.

Comment: Reuses speculative-decoding modules and resident KV caches for contextual classification with negligible additional inference computation.

Topic Match: The core mechanism avoids running a separate large classifier by reusing existing decoder computation and cached context, making inference efficiency the strongest fit.

Relevance: 7 Novelty: 7


13. The information geometry of product-reference discrete diffusion: Interaction growth complexity and optimal scheduling

ArXiv ID: 2608.28949

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Martin J. Wainwright

Abstract: We study a class of product-reference diffusion algorithms for sampling from a discrete distribution. We show that their sampling performance can be characterized using a path-based measure of data geometry that we call the interaction growth complexity (IGC). We show that a bivariate IGC kernel gives an exact representation of both the KL discretization error and a simple one-step upper bound. The simpler univariate IGC density can be used to study the effect of stepsize choices on the iteration complexity required to obtain $ε$-accurate samples in KL divergence. Samplers that traverse the path with equi-spaced steps in log-squared-reliability-odds have performance that depends on the aggregate IGC mass, whereas refined choices of stepsizes have a lower complexity depending on a square-root functional. In the fine-grid limit, both of these characterizations become sharp. We also allow general product reference distributions and show that the reference law can substantially reshape the IGC profile and the resulting sampling complexity; in particular, references far from both the uniform and the data marginals can yield dimension-dependent improvements. Finally, the aggregate IGC mass admits bounds in terms of total correlation and dual total correlation, thereby connecting the pathwise geometry to classical measures of multivariate dependence.

Comment: Derives geometry-adaptive timestep schedules that reduce discrete-diffusion sampling complexity.

Topic Match: Sampling-complexity and schedule analysis are closest to efficiency_scaling; the abstract does not establish a connection to large-model training or execution costs.

Relevance: 6 Novelty: 8


14. AERA: Adaptive Evidence Residual Allocation for Efficient Test-Time Reasoning

ArXiv ID: 2608.27964

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Ziming Wang, Ivor Tsang, Hangwei Qian

Abstract: Test-time scaling improves language-model reasoning by generating additional candidate solutions, but allocating the same inference budget to every problem is computationally wasteful. Existing adaptive stopping methods commonly rely on confidence, agreement, or answer stability, implicitly assuming that stronger current evidence indicates that further computation is unnecessary. We show that this assumption can fail: checkpoint-level correctness evolves non-monotonically, and observable evidence may strengthen before an answer collapses or weaken before it recovers. Motivated by this mismatch, we introduce Adaptive Evidence Residual Allocation (AERA), a sequential controller that learns whether additional computation is likely to recover a better answer from checkpoint-observable evidence. AERA characterizes cumulative response prefixes using answer-distribution, temporal, re-solving, semantic, and compute features, and repeatedly decides whether to stop or allocate the next response block. Future checkpoint correctness is used only to construct offline supervision and is never available to the controller at inference time. Across GSM8K and GPQA Diamond, AERA identifies question-specific residual opportunities while substantially reducing inference computation. In a frozen-threshold incremental-generation evaluation on 300 untouched GSM8K questions, AERA achieves 92.61% accuracy versus 93.01% with 128 responses while reducing completion tokens by 95.99%. These results suggest that adaptive reasoning should estimate the future value of computation rather than equating present confidence with correctness.

Comment: A learned stop-or-continue controller estimates whether additional sampling can recover a better answer and allocates inference compute accordingly.

Topic Match: Adaptive sampling-budget allocation directly changes inference cost through a computational policy, providing a substantive but peripheral match to this training-centered feed.

Relevance: 7 Novelty: 6


15. Pro-Router: Token-Aware Progressive Model Routing with Adaptive Edge-Cloud Collaboration for Efficient Multimodal LLM Inference

ArXiv ID: 2608.28726

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Xinyuan Gui, Shaowen Wang, Sheng Sun, Zijian Wang, Zishu Yu, Zheming Yang

Abstract: The remarkable performance of multimodal large language models (MLLMs) comes at the cost of substantial computational overhead, posing significant challenges to real-time deployment and cost effectiveness. Existing model routing approaches either decide from coarse request-level features alone or spend one or several extra language model passes to inspect the generated response, leaving the token-level uncertainty signals that emerge during generation unused. To address these limitations, we propose Pro-Router, a token-aware progressive model routing method with adaptive edge-cloud collaboration for efficient multimodal LLM inference. Pro-Router employs a two-stage progressive decision mechanism. First, a lightweight prompt pre-scorer module performs rapid pre-screening before token generation begins, guiding apparently simple requests to small models. Second, a token-aware verifier reads the sampling probability distribution of each token the small model generates, estimating the model's confidence in its own output to determine, per request, whether the answer ships or escalates to the cloud-based high-precision model. Furthermore, we design an adaptive edge-cloud serving pipeline that sizes every dispatch to each device's measured service rate, so both the edge and the cloud tiers stay fully utilized without manual parameter tuning and are not impacted by the network latency. Extensive experiments on multiple multimodal benchmark datasets and models demonstrate the effectiveness of Pro-Router. Compared to other methods, it achieves the highest routing accuracy and improves routing speed by more than 10x. Its serving pipeline also reaches more than 75% higher end-to-end throughput than the existing model routing pipeline. Our code is available at https://github.com/xinyuangui2/pro-router.

Comment: Progressive prompt screening and token-confidence verification route requests between small and large models with reduced inference overhead.

Topic Match: The token-aware model cascade provides a concrete inference-efficiency mechanism, with relevance concentrated in serving costs.

Relevance: 7 Novelty: 6


16. Nested Byte-Level Vocabularies Are Cheap to Deploy and Expensive to Share: A Pre-Registered Negative Result

ArXiv ID: 2608.28151

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Christos Koutsiaris

Abstract: A byte-level BPE tokenizer is an ordered list of merge rules, so applying only a prefix yields a vocabulary whose token identifiers are the first rows of the full vocabulary. This prefix nesting allows one language model to operate at several vocabulary sizes, use a control token to indicate the active size, and be deployed at any trained size by slicing its embedding and output head. We pre-registered five claims, including margins, seeds, contrasts, and a stop rule, and trained 30 models with 3.1M- and 10.6M-parameter bodies on 200M tokens each. Slicing is numerically exact: across 76 checks, a sliced model reproduces the restricted full model's logits bit for bit and removes 66% of deployed weights without changing latency. However, the shared model trails a fixed-cap specialist by 3.64% bits per byte at 32k against a 1% margin, and by 2.96% at 8k against a 2% margin. A 2x2 ablation separating the control token from output restriction finds that the token changes performance by +0.07% to +0.13%, with all intervals crossing zero, while output restriction costs +0.47% to +1.19%; the factors are substitutes rather than complements. Multi-cap training nevertheless improves robustness: under typographical noise, the same checkpoint degrades 12.5--15.4 points less in its fine mode and outperforms each fixed-cap specialist at that specialist's vocabulary size. A control with neither cap token nor output restriction is equally robust, attributing this benefit to multi-granularity training rather than conditioning. The per-cap penalty tracks each cap's share of training rows, yielding a falsifiable prediction for future work.

Comment: Nested vocabularies enable exact embedding and output-head slicing while exposing the accuracy cost of shared multi-vocabulary training.

Topic Match: Vocabulary-based parameter reduction is central, with controlled training ablations explaining accuracy and robustness trade-offs on small models.

Relevance: 7 Novelty: 6


17. CoRe-MoE: Compact Reusable MoE for Continual Multimodal Instruction Tuning

ArXiv ID: 2608.27867

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Runze Liu, Naibin Gu, Mingxu Ai, Yuqing Li, Peng Fu, Zheng Lin, Weiping Wang

Abstract: Continual multimodal instruction tuning requires multimodal large language models to acquire new task abilities sequentially while preserving previously learned knowledge. LoRA-MoE provides a promising solution by introducing expert-based capacity, but repeatedly learning and maintaining full LoRA experts leads to substantial parameter overhead. This raises a natural question: is full expert expansion necessary for every new task? To answer it, we analyze the SVD of task-specific LoRA updates and observe substantial overlap in their input- and output-side LoRA direction subspaces, with task-specific adaptation largely captured by lightweight coordinates over these subspaces. Motivated by this observation, we propose CoRe-MoE, a Compact Reusable MoE framework for parameter-efficient continual multimodal instruction tuning. CoRe-MoE extracts reusable input- and output-side direction bases from an initial expert bank, and for subsequent tasks trains only compact coordinate experts together with task-specific low-rank routers. Experiments on two representative MLLMs show that CoRe-MoE improves final average performance over the strongest competing baseline by up to 5.90 points, while using less than 1% of the trainable parameters required by sequential LoRA for later tasks. The code is publicly available at https://github.com/runzezz/CoRe-MoE.

Comment: Reusable input and output LoRA bases reduce later-task training to compact coordinate experts.

Topic Match: The substantive contribution is a compact adaptation parameterization that reduces trainable parameters in continual instruction tuning.

Relevance: 7 Novelty: 6


18. Sliding-window beats linear attention

ArXiv ID: 2608.28444

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Alexia Jolicoeur-Martineau, Rhea Sanjay Sukthanker, Pashmina Cameron, Emy Gervais

Abstract: Due to the nature of quadratic attention, Large Language Models (LLMs) consume a lot of memory and energy. Every new token costs more than the previous one. For each additional token, the keys and values must be stored in memory indefinitely, which is unsustainable. Several alternatives have been proposed to fix the quadratic scaling problem, one of which is retrofitting LLMs to use Linear Attention. This idea has attracted a lot of attention, given its promise to solve the quadratic scaling problem with state-of-the-art performance at low cost. However, this line of research has not been properly compared to simpler baselines. In this work, we show that Sliding Window Attention (SWA) with sinks performs as well or better than post-trained Linear Attention models. We observe this across multiple LLMs on various downstream tasks. For long-context reasoning tasks (Needle-in-a-Haystack and BABILong), SWA achieves massively higher performance (2 to 10 times higher than linear attention). SWA requires no post-training, is extremely fast, and requires low memory; therefore, making it an extremely cheap and reliable solution. To reduce inference memory cost, we strongly recommend switching to SWA instead of post-training linear models. Linear attention models may have shown some promise, but they likely require to be trained from scratch or extensive post-training in order to even match SWA.

Comment: Shows that sliding-window attention with sinks limits KV-cache memory while matching or exceeding retrofitted linear attention.

Topic Match: KV-cache efficiency is directly relevant, but the contribution evaluates existing inference mechanisms and falls under the evaluation-only exclusion.

Relevance: 6 Novelty: 6


Paper Selection Prompt

System Prompt

You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.

User Prompt

Instructions

Respond in JSONL. Output exactly one JSON object per paper, one per line:

{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}

Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only: daily_hot, new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return []. - daily_hot means the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. - new_frontier means the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.

Scoring Criteria

Relevance and Novelty are independent axes. Score both from 1 to 10.

Relevance Scoring

  • 9-10: directly centered on the target foundational topics; highest when the core contribution is clearly within them.
  • 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
  • 5-6: touches the target topics, but the main contribution is elsewhere.
  • 3-4: largely outside the target topics, often application-focused or domain-specific.
  • 1-2: unrelated.

Important: Broad frontier relevance, major launch status, or daily buzz is not enough for a high Relevance score here. Those cases belong in the hotspot digest unless the paper strongly matches the specialized paper topics.

Novelty Scoring

  • 9-10: new paradigm, theory, or major methodological breakthrough.
  • 7-8: substantial methodological advance or strong new insight.
  • 5-6: meaningful but incremental extension or refinement.
  • 3-4: minor, narrow, or mostly engineering or domain-specific improvement.
  • 1-2: little originality; mainly standard application of existing methods.

Topic Registry

Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - moe_training: MoE Training - Mixture-of-Experts training: routing, load balancing, expert granularity and shared experts, training stability, MoE parallelism and communication, and MoE kernels. - training_systems: Large-Scale Training Systems and Efficiency - Large-scale training systems and efficiency: distributed training algorithms, optimizers, and communication. - architecture_training: Architecture and Training Dynamics - Core architectural or computational mechanisms, dynamic computation, and training-stability dynamics. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, memory or cache efficiency, and large-scale training systems that materially change cost or behavior.

Papers

[PAPER LIST HERE]

Relevant Topics

This is a training-side feed. The centre of gravity is Mixture-of-Experts training and the large-scale training systems it sits on; architecture and efficiency work is kept when it bears on how a large model is trained or what it costs to train.

Keep a paper when its CORE CONTRIBUTION falls in one of the four topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.

  1. MoE Training (primary) - Keep: expert routing and gating mechanisms; load balancing, with or without auxiliary losses; expert granularity, fine-grained and shared experts; MoE training stability (loss spikes, router collapse, z-loss and related fixes); expert and hybrid parallelism; all-to-all dispatch/combine and the communication schedules MoE needs; MoE kernels and grouped-GEMM style implementations; expert capacity, token dropping, and upcycling dense checkpoints into MoE. - Filter: papers that merely train on top of a MoE model without a new routing, balancing, stability, parallelism, or kernel contribution; "mixture of experts" in the classical ensemble or recommender sense.

  2. Large-Scale Training Systems - Keep: distributed training algorithms (data, tensor, pipeline, sequence, and expert parallelism, FSDP and ZeRO-style sharding); optimizers and preconditioners for large-scale pretraining; communication algorithms, overlap, and collective design; throughput and MFU work that materially changes what a training run costs or whether it converges; scaling-law work that informs how such runs are configured. - Filter: routine infrastructure tuning with no new algorithmic or systems idea.

  3. Architecture and Training Dynamics - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, dynamic or modular computation); optimisation and training-dynamics analysis that explains why large models train the way they do. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight.

  4. Efficiency and Compression - Keep: quantization, sparsity, pruning, low-rank adaptation, KV-cache and memory-efficient designs, and other work that materially changes the cost of running or training a large model, especially when the mechanism is new rather than a tuned variant. - Filter: straightforward application of standard efficiency methods, and deployment or serving engineering with no new idea.

Leave these to the hotspot digest unless the core contribution clearly falls in one of the four topics above: - major model or product releases - agent, tooling, and RAG launches - benchmark, leaderboard, or evaluation-only papers - representation-learning theory, interpretability, and feature-formation analysis - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning) - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains