This is a remedial run for missed papers from 07/06/2026 to 07/06/2026.
Results generated on 09/12/2026.
Personalized Daily ArXiv Papers 2026-07-07
| Model | Metric | Usage | Papers | ||||
|---|---|---|---|---|---|---|---|
| Prompt | Completion | Total | Total arXiv | Scanned | Relevant | ||
local agent CLI |
Tokens | not reported | not reported | not reported | 361 | 361 | 17 |
| Cost | not reported | not reported | not reported | ||||
Token counts are not reported for this run. 4 of 4 model calls succeeded, 1,530s of model wall clock.
Topic Coverage:
| Topic | Papers |
|---|---|
| MoE Training | 1 |
| Large-Scale Training Systems and Efficiency | 2 |
| Architecture and Training Dynamics | 9 |
| Efficiency, Compression, and Large-Scale Training | 5 |
Table of contents by topic:
MoE Training (1)
- Localized LoRA-MoE: Block-wise Low-Rank Experts With Adaptive Routing Authors: Babak Barazandeh, Subhabrata Majumdar, Vinay Prithyani, George Michailidis
Large-Scale Training Systems and Efficiency (2)
-
Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training Authors: Jingwei Zuo, Cong Zeng, Ilyas Chahed, Maksim Velikanov, Dhia Eddine Rhaiem, Pasquale Balsebre, Abhay Kumar, Younes Belkada, Hakim Hacid
-
FAST: A Holistic Framework for Optimizing Memory-I/O, Computation, and Sampling in Temporal GNN Training Authors: Yushu Cai, Qingrui Zhu, Lei Liu, Kai Sheng, Hao Chen, Xin He
Architecture and Training Dynamics (9)
-
ResonatorLM: Causal Resonant Field Mixing for Efficient Long-Context Language Modeling Authors: Archie Chaudhury
-
BCMT: Blockwise Causal Memory Transformer Authors: Rachid Arezki
-
Advances in Neural Controlled Differential Equations Authors: Benjamin Walker
-
The Active Ingredient in Muon's Grokking Authors: Yufeng Wang
-
$\mathbfλ$-VAE: Variance Equalization for Posterior Collapse Authors: Girum Demisse
-
Hyperparameter Transfer in Graph Neural Networks Authors: Gage DeZoort, Boris Hanin
-
Non-asymptotic Convergence of Stochastic Gradient Descent in Score-based Generative Models Authors: Stanislas Strasman, Sobihan Surendran, Sylvain Le Corff
-
To Retain or to Adapt? Generalizing Continual Learning Authors: Giulia Lanzillotta, Mandana Samiei, Doina Precup, Razvan Pascanu, Claire Vernade
-
Minimum Block Width for Universal Approximation by Residual Neural Networks with Inner Width One Authors: Qi Zhou, Xuan Zhou, Xiao-Song Yang
Efficiency, Compression, and Large-Scale Training (5)
-
KVpop -- Key-Value Cache Compression with Predictive Online Pruning Authors: Lukas Hauzenberger, Niklas Schmidinger, Anamaria-Roberta Hartl, David Stap, Thomas Schmied, Sebastian Böck, Günter Klambauer, Sepp Hochreiter
-
TORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language Models Authors: Riccardo Renzulli, Gabriele Spadaro, Shruthi Gowda, Alaa Eddine Mazouz, Van-Tam Nguyen
-
Elastic Gang: Per-Token Membership Change for a Hard-Barriered LLM Inference Gang Co-Scheduled with OS Processes Authors: Daeyeon Son
-
Layer-Parallel Inference Reduces Encrypted Nonlinear Depth in Transformers Authors: Ligong Han, Kai Xu, Hao Wang, Ruijiang Gao, Han Gao, Akash Srivastava
-
Learning Only What Valid Adapters Can Express: Subspace-Constrained Adaptation Against Fine-Tuning Poisoning Authors: Fabien Polly
MoE Training (1)
1. Localized LoRA-MoE: Block-wise Low-Rank Experts With Adaptive Routing
ArXiv ID: 2607.05114
Primary Topic: MoE Training
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Babak Barazandeh, Subhabrata Majumdar, Vinay Prithyani, George Michailidis
Abstract: Large Language Models (LLMs) and high-dimensional perception networks increasingly rely on parameter-efficient fine-tuning (PEFT) to adapt to diverse operational contexts. However, standard methods like LoRA are structurally limited by a monolithic bottleneck, making them highly susceptible to gradient warfare. Interleaved multi-task streams may trigger destructive optimization feedback, collapsing adapter weights into unspecialized averages. While recent spatial partitioning methods have introduced block-wise isolation, they remain trapped in static topologies, unable to adapt to dynamic task-switching or environmental sensor failure. In this work, we introduce Localized LoRA-MoE, a unified framework that fuses localized spatial blocking with dynamic, context-conditioned routing. We propose and evaluate two novel architectural paradigms: Block-Wise LoRA-MoE (Centralized Macro-Routing), which modulates the entire structural grid via a monolithic context signal, and Cell-Wise LoRA-MoE (Decentralized Micro-Routing), which empowers every coordinate cell in the matrix grid with autonomous, localized expert gating. Through a comprehensive suite of benchmarks, ranging from high-dimensional SVD matrix simulations and real-world tabular transformations to spatial vision perception under sensor degradation, we demonstrate that both architectures resolve optimization deadlocks inherent in static baselines. Our empirical results establish that decentralized cell-level gating achieves complete statistical parity with an omniscient global coordinator, providing a robust "gradient firewall" that protects surviving pathways from fault-propagated corruption. Our proposals consistently outperform static baselines, offering a scalable and parameter-efficient solution for dynamic model adaptation across granular coordinate fields and shifting operational regimes.
Comment: Context-conditioned routing selects low-rank experts at block and individual matrix-cell granularity.
Topic Match: Expert gating and granularity are central contributions, implemented through parameter-efficient adapters; evidence emphasizes adaptation rather than large-scale pretraining.
Relevance: 8 Novelty: 7
Large-Scale Training Systems and Efficiency (2)
1. Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training
ArXiv ID: 2607.04969
Primary Topic: Large-Scale Training Systems and Efficiency
Also Matches: Architecture and Training Dynamics
Authors: Jingwei Zuo, Cong Zeng, Ilyas Chahed, Maksim Velikanov, Dhia Eddine Rhaiem, Pasquale Balsebre, Abhay Kumar, Younes Belkada, Hakim Hacid
Abstract: The training paradigm of large language models has shifted from traditional one-pass training to multi-epoch training, as reasonable reuse of limited high-quality data can improve both model performance and sample efficiency. Meanwhile, excessive repetition introduces the risk of overfitting and diminishing returns. Determining when and how to reuse data effectively thus emerges as a natural but under-explored question. Through a novel observation of model's "Memorization Window" signals derived from loss retention dynamics and downstream evaluation scores, we propose "Memorization-guided Data Reuse", a training paradigm that adaptively determines when and how data should be reused, enabling principled decisions on the number of training epochs and the scheduling of data replays. Our preliminary experiments reveal a consistent memorization-driven regime: performance continues to improve with repetition far beyond current practice (e.g., the commonly cited four-epoch limit). While a full scheduler remains future work, these insights provide a foundation for memorization-aware training schedules, helping to determine reuse budgets and move toward training LLMs smarter rather than longer with limited high-quality data.
Comment: Loss-retention signals inform multi-epoch pretraining data-reuse budgets.
Topic Match: Pretraining epoch allocation directly informs run configuration, supported by memorization dynamics; a complete adaptive scheduler remains future work.
Relevance: 8 Novelty: 6
2. FAST: A Holistic Framework for Optimizing Memory-I/O, Computation, and Sampling in Temporal GNN Training
ArXiv ID: 2607.05095
Primary Topic: Large-Scale Training Systems and Efficiency
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Yushu Cai, Qingrui Zhu, Lei Liu, Kai Sheng, Hao Chen, Xin He
Abstract: Temporal Graph Neural Networks (TGNNs) are widely used for learning from dynamic graphs in applications such as recommendation, social network analysis, and traffic forecasting. However, scaling TGNN training to large dynamic graphs remains challenging due to three intertwined bottlenecks: memory I/O, irregular computation, and temporal neighbor sampling. Existing systems often optimize these stages in isolation, leaving substantial performance headroom on the table. We present FAST, a holistic framework that accelerates end-to-end TGNN training by jointly optimizing sampling, memory I/O, and computation. FAST introduces SlimCache, which exploits within-batch compression and cross-batch caching to reduce host-device data movement under limited GPU memory budgets. It further designs thread-efficient graph operators tailored to sparse temporal subgraphs, improving GPU cache locality and reducing the latency of aggregation and edge softmax. In addition, FAST employs a topology-aware sampling strategy that improves CPU cache locality and accelerates temporal neighbor sampling. Extensive experiments on real-world large dynamic graphs show that FAST achieves an average of 2.1x (up to 4.7x) speedup over state-of-the-art systems without sacrificing model accuracy.
Comment: Within-batch compression and cross-batch caching reduce host-device traffic during temporal GNN training.
Topic Match: The contribution changes training data movement, sampling, and GPU operators; its scope is specialized temporal GNN systems.
Relevance: 7 Novelty: 6
Architecture and Training Dynamics (9)
1. ResonatorLM: Causal Resonant Field Mixing for Efficient Long-Context Language Modeling
ArXiv ID: 2607.05583
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Archie Chaudhury
Abstract: Contemporary language models are dominated by the transformer architecture, which leverages self-attention mechanisms to enable more efficient, parallelized training across a wide set of documents and corpora. This has allowed transformers to effectively model data across a wide range of modalities and contexts. However, transformers, along with their conventional counterparts such as recurrent neural networks (RNNs) and convolutional neural networks (CNNs), often struggle to maintain efficiency when processing long contexts. We introduce ResonatorLM, a new mechanism that replaces attention with a physics-derived alternative. ResonatorLM treats token sequences as a single, driven one-dimensional latent field and replaces attention dot products with causal functions of damped resonators. We implement ResonatorLM on a traditional network architecture and test it on standard long-context modeling tasks. We find that in a small, 6M matched setting, training and prefill speedups increase with sequence length, decode speed reaches 6.47x compared to that of a standard, optimized transformer at 32K tokens, and accuracy reaches 61.31 percent (compared to 55.32 percent) on WikiText.
Comment: Replaces dot-product attention with causal damped-resonator field mixing for long sequences.
Topic Match: The attention-replacement mechanism directly fits sequence architecture, with training and inference efficiency evidence limited to a 6M-parameter model.
Relevance: 9 Novelty: 7
2. BCMT: Blockwise Causal Memory Transformer
ArXiv ID: 2608.13578
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Rachid Arezki
Abstract: Transformer architectures rely on dense self-attention to model long-range dependencies, but this mechanism exhibits quadratic complexity with respect to sequence length. We introduce BCMT (Blockwise Causal Memory Transformer), an architecture for long-context language modeling that decouples local token interactions from global context propagation. Dense causal self-attention is applied independently within local blocks, while each block produces an adaptive summary aggregated through an exponential causal memory. This memory is subsequently injected back into the token representations, enabling efficient propagation of long-range contextual information without relying on explicit global attention. Unlike standard Transformers and recurrent memory architectures, BCMT maintains neither dense interactions between distant tokens nor learned memory states. Its memory mechanism is fully parallelizable and remains compatible with standard implementations of dense self-attention. Experiments on language modeling with context lengths of up to 1024 tokens show that BCMT achieves validation performance comparable to that of Dense Transformers while significantly improving training throughput and reducing memory consumption. An ablation study further confirms that these improvements arise from the proposed memory mechanism. These results demonstrate that an exponential causal memory constructed from block summaries provides an effective alternative to dense global attention mechanisms for long-context language modeling.
Comment: Block-local attention and parallel exponential aggregation replace dense global token interactions.
Topic Match: It changes the sequence-model computation and training cost; reported experiments extend only to 1,024-token contexts.
Relevance: 9 Novelty: 6
3. Advances in Neural Controlled Differential Equations
ArXiv ID: 2607.05280
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Benjamin Walker
Abstract: Many real-world systems evolve continuously, yet most machine learning models interpret time series as discrete sequences. Continuous-time approaches instead treat time series as samples from an underlying input path, a formulation that naturally accommodates irregularly sampled or oversampled data. Among these, Neural Controlled Differential Equations (NCDEs) are a maximally expressive class of models that parametrise a vector field using a neural network and evolve their hidden state by solving a dynamical system driven by the input path. NCDEs typically use a non-linear vector field, so their expressive power and continuous-time flexibility come at the cost of a forward pass that is both computationally expensive and inherently sequential, limiting their scalability and practical applicability. This thesis advances the training and scalability of NCDEs through three complementary contributions. First, building on neural rough differential equations, Log-NCDEs apply the Log-ODE method to efficiently approximate an NCDE's solution during training, improving both computational speed and empirical performance. Second, Linear NCDEs replace the non-linear vector field with a linear one, enabling closed-form solutions and parallel-in-time computation without sacrificing theoretical expressivity. Third, Structured Linear NCDEs use structured linear vector fields to further enhance efficiency while maintaining theoretical expressiveness and empirical performance. Collectively, these methods reduce the time per training step for an NCDE by up to three orders of magnitude while achieving state-of-the-art performance across diverse time series benchmarks.
Comment: Linear and structured NCDE vector fields enable parallel-in-time computation while preserving theoretical expressivity.
Topic Match: The core contribution changes recurrent state evolution to remove sequential computation bottlenecks and substantially accelerate training.
Relevance: 8 Novelty: 7
4. The Active Ingredient in Muon's Grokking
ArXiv ID: 2607.20512
Primary Topic: Architecture and Training Dynamics
Also Matches: Large-Scale Training Systems and Efficiency
Authors: Yufeng Wang
Abstract: The Muon optimizer reaches the grokking threshold on modular arithmetic faster than AdamW. Prior work attributes this to "spectral-norm constraints plus orthogonalized momentum" but does not isolate which mechanism matters. To better understand Moun's behavior, we run multi-seed and multi-learning-rate sweeps to decompose and stress-test the effect. First, an ablation shows the speedup comes from orthogonalization (the Newton-Schulz iteration): orthogonalize-only matches full Muon, whereas spectral-only is no faster than AdamW and is unreliable, and this verdict holds across learning rates. Second, a mechanistic analysis finds that orthogonalizing optimizers reach generalization at roughly 3x lower spectral norm and, controlling for how much the embedding actually moves, settle into a lower-norm solution rather than simply perturbing the embedding less. Third, reducing the Newton-Schulz iteration count from five to one accelerates reaching the threshold but makes the grokked solution fragile, prone to transient collapse, with fragility that grows with learning rate; a single iteration is fast and stable only at small learning rate, while the canonical five iterations are the learning-rate-robust choice. We also show spectral scaling can be dropped at no measured cost. A methodological thread runs throughout: under a stability-aware metric, "faster" claims about grokking optimizers can invert, so we report both first-crossing and sustained-grok times. To support reproducibility, we release our full training and analysis code at https://github.com/louiswang524/muon-grokking-frontier
Comment: Isolates Newton-Schulz orthogonalization as Muon's grokking driver and links iteration count to optimization stability.
Topic Match: The core is optimizer-mechanism and stability analysis relevant to Muon configuration, with evidence restricted to modular-arithmetic grokking.
Relevance: 8 Novelty: 6
5. $\mathbfλ$-VAE: Variance Equalization for Posterior Collapse
ArXiv ID: 2607.05531
Primary Topic: Architecture and Training Dynamics
Authors: Girum Demisse
Abstract: Variational Autoencoders (VAEs) frequently suffer from posterior collapse, a failure mode in which the approximate posterior converges to the prior, rendering the latent code uninformative. Despite extensive research, a unified account of why collapse occurs has remained an open question. We identify and formalize two logically independent but coupled causes. \emph{Gradient imbalance} occurs when the decoder's reconstruction signal vanishes faster than the $\mathbb{KL}$ regularization pressure as the posterior widens. \emph{Information gap} occurs when the stochastic sampling step discards a substantial fraction of the encoder's computed representation, attenuating decoder sensitivity and making collapse inexpensive. Both causes share the same collapse trajectory, and we show that the information gap is algebraically equivalent to mismatch between the aggregate posterior and the prior, unifying two pathologies. Subsequently, we introduce $λ$-VAE, which resolves both causes through a single modification to the reparameterization step: the sampling noise is scaled by per-dimension exponent, while the $\mathbb{KL}$ penalty retains the original posterior variance. This asymmetry shifts the stable training attractor away from the degenerate collapsed state, driving all latent dimensions toward the same equilibrium -- a mechanism we term \emph{variance equalization}. A closed-form optimal exponent per dimension follows from a net information gain objective, with a single hyperparameter controlling the reconstruction-generation tradeoff. We validate on standard benchmarks (Binary MNIST, Binary Omniglot, CIFAR-10, CelebA-64), showing consistent reductions in collapsed dimensions, information capacity gains of up to $2.8\times$ nats, and reconstruction quality improvements of up to $+0.33$ BPD.
Comment: Variance-dependent reparameterization noise shifts VAE training away from posterior-collapse attractors.
Topic Match: A mechanistic account of training collapse leads directly to a reparameterization fix, although its demonstrated scope is VAEs.
Relevance: 7 Novelty: 7
6. Hyperparameter Transfer in Graph Neural Networks
ArXiv ID: 2607.05017
Primary Topic: Architecture and Training Dynamics
Also Matches: Large-Scale Training Systems and Efficiency
Authors: Gage DeZoort, Boris Hanin
Abstract: The performance of deep learning models crucially depends on the settings of hyperparameters like learning rate, initialization scale, and weight decay. Hyperparameter transfer aims to make near-optimal hyperparameter settings consistent across model scale, so that large models can be optimized by proxy tuning their smaller, cheaper-to-optimize counterparts. While transfer principles are well-studied in the context of dense neural networks in language and vision tasks, they remain comparatively under-explored for graph neural networks (GNNs). We develop and validate a transfer parameterization for GNNs trained with SGD, Adam, and AdamW. Through theoretical scaling analyses and controlled experiments, we show that the proposed parameterization yields stable feature updates, learning rate transfer, and improved performance as width and depth increase. For SGD, we identify graph-dependent first-layer correction factors and show that their use can accelerate early training in graphs with sparse bag-of-words inputs. For Adam, we explore how different message passing normalizations affect early- and late-training transfer behavior, illustrating the importance of message passing normalization and advocating for an associated hyperparameter. For AdamW, we adapt a parameterization that allows for the joint transfer of weight decay and learning rate. Together, these results provide a practical recipe for scaling GNNs across a variety of learning tasks and training scenarios.
Comment: Graph-aware parameterization transfers learning rates and weight decay across model width and depth.
Topic Match: Parameterization and normalization control scale-dependent training dynamics, while transferable optimizer settings reduce tuning costs.
Relevance: 7 Novelty: 7
7. Non-asymptotic Convergence of Stochastic Gradient Descent in Score-based Generative Models
ArXiv ID: 2607.04775
Primary Topic: Architecture and Training Dynamics
Authors: Stanislas Strasman, Sobihan Surendran, Sylvain Le Corff
Abstract: Score-based Generative Models (SGMs) have achieved impressive performance in data generation across a wide range of applications. While the statistical properties of their sampling procedures are increasingly well understood, the optimization dynamics underlying their training remain less explored. SGMs are typically trained by minimizing a weighted denoising scorematching objective, yet optimization guarantees with stochastic gradients remain limited. In this work, we study Stochastic Gradient Descent (SGD) for SGMs, contributing results in two complementary regimes. First, for general score parameterizations, we establish a non-convex convergence rate for SGD on the weighted denoising score-matching objective, with explicit dependence on the schedule-dependent weighting factors. Second, for overparameterized two-layer ReLU networks, we develop a Neural Tangent Kernel analysis tailored to diffusion training with stochastic gradients, yielding score-approximation error bounds along the SGD trajectory. Finally, our analysis quantifies the role of the reweighting factor in the score approximation error, providing theoretical guidance for weighting choices used in practice.
Comment: SGD convergence bounds expose how diffusion loss weighting affects score-approximation error.
Topic Match: It directly analyzes stochastic training dynamics and objective weighting, with scope focused on diffusion score learning.
Relevance: 7 Novelty: 7
8. To Retain or to Adapt? Generalizing Continual Learning
ArXiv ID: 2607.05609
Primary Topic: Architecture and Training Dynamics
Authors: Giulia Lanzillotta, Mandana Samiei, Doina Precup, Razvan Pascanu, Claire Vernade
Abstract: The Continual Learning (CL) literature has long been driven by the goal of mitigating catastrophic forgetting. This objective rests on a pervasive, often unstated assumption: that a lifelong learner should approximate the Joint-Task Learning (JTL) solution and retain all previously acquired knowledge. We challenge this retention-centered premise, arguing that in non-stationary environments prioritizing retention can impede real-time adaptation. Shifting the focus to the Average Lifelong Error (ALE), we formalize CL as an online optimization problem governed by the interaction between environmental and learning dynamics. We introduce Transfer Efficiency as a quantitative measure of the tension between Instability, the bias inherited from conflicting past experience, and Transient Error, the optimization cost of learning new tasks from scratch. Under mild convergence conditions, holding across linear and neural network models, this decomposition yields a Critical Task Duration: a closed-form threshold beyond which historical knowledge transitions from a warm-start advantage to an optimization liability whenever retention induces a positive stationary bias. We validate these theoretical predictions on continual image classification and reinforcement learning benchmarks. Finally, by connecting continual learning to the online learning framework of predictable sequences, we show that JTL is only one instance of a broader family of objectives, and we propose a new general class of continual learning algorithms, which we call Predictive Continual Learning. Predictive CL algorithms optimize expected future performance under an explicit, dynamically updated model of future tasks. As a proof of concept, we analyze a Window algorithm that interpolates between JTL and Independent-Task Learning (ITL), outperforming both under controlled distributional drift.
Comment: Derives a critical task duration at which retention-induced bias outweighs the optimization benefit of warm starts.
Topic Match: Retention-versus-adaptation dynamics are the theoretical core; applicability to large-model pretraining is indirect.
Relevance: 7 Novelty: 7
9. Minimum Block Width for Universal Approximation by Residual Neural Networks with Inner Width One
ArXiv ID: 2607.04597
Primary Topic: Architecture and Training Dynamics
Authors: Qi Zhou, Xuan Zhou, Xiao-Song Yang
Abstract: In this paper, we study the universal approximation property of residual neural networks. For input and output dimensions $d_x$ and $d_y$, and LeakyReLU, ReLU, ReLU-like activation functions, the upper and lower bounds of the minimum block width are established. To achieve $L^p$ approximation $(1\leq p <+\infty)$ on any compact set, we show that the exact minimum block width is $\max{d_x,d_y}$ when each residual branch has inner width 1. Furthermore, we show that residual neural networks with block width $\min{d_x+d_y, \max{2d_x+1,d_y}}$ can achieve uniform approximation on any compact set under the constraint that each residual branch has inner width 1. Besides, for any activation function family, we prove that there exist functions that cannot be approximated by residual neural networks with block width less than $\max{d_x, d_y}$, both in the $L^p$ sense and the uniform sense, regardless of inner width. Consequently, for LeakyReLU, ReLU, ReLU-like activation functions and $d_y\geq 2d_x+1$, the exact minimum block width for uniform approximation is $d_y$ when each residual branch has inner width 1.
Comment: Derives minimum residual-stream widths for universal approximation with one-neuron residual branches.
Topic Match: Residual-network expressivity is the architectural connection; implications for training behavior and practical training cost remain indirect.
Relevance: 6 Novelty: 6
Efficiency, Compression, and Large-Scale Training (5)
1. KVpop -- Key-Value Cache Compression with Predictive Online Pruning
ArXiv ID: 2607.05061
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Lukas Hauzenberger, Niklas Schmidinger, Anamaria-Roberta Hartl, David Stap, Thomas Schmied, Sebastian Böck, Günter Klambauer, Sepp Hochreiter
Abstract: Key-value (KV) cache growth is a major bottleneck in autoregressive decoding, as memory and bandwidth scale linearly with context length. Existing KV eviction methods often rely on static heuristics or proxy scores, which poorly track future token utility and cause brittle eviction as relevance shifts. To address this, we introduce KVpop, which learns a fixed-budget KV eviction policy by directly supervising the keep-or-drop decision. The scorer is trained against a novel future-attention target, computed efficiently without materializing dense attention maps. We further introduce a delayed memory-based scorer that, uniquely among learned eviction methods, defers scoring for a fixed number of steps to exploit near-future context. On AIME and HMMT mathematical reasoning, KVpop retains 98% of full-attention performance on Qwen3-4B at 75% KV cache compression and 97% at 88% compression, consistently outperforming established eviction baselines. Qwen3-8B shows even stronger results, reaching near-full teacher performance. These results show that supervising eviction with future-attention signals cuts memory costs while maintaining quality.
Comment: Future-attention supervision and delayed scoring learn fixed-budget KV-cache eviction decisions.
Topic Match: The core contribution is a learned cache-pruning mechanism with substantial memory reductions demonstrated on 4B-8B models.
Relevance: 9 Novelty: 7
2. TORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language Models
ArXiv ID: 2607.04593
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Riccardo Renzulli, Gabriele Spadaro, Shruthi Gowda, Alaa Eddine Mazouz, Van-Tam Nguyen
Abstract: Vision-Language Models (VLMs) have demonstrated impressive capabilities across different tasks, but their computational cost is dominated by the large number of visual tokens fed to the language model. Existing token reduction methods rely on attention-based scores or pairwise similarity, without an explicit semantic representation of each token. We introduce TORINO (TOken Reduction via Interpretable coNcept Overlap), a plug-and-play framework for adaptive visual token reduction in VLMs that requires no fine-tuning of the underlying model. TORINO leverages Sparse Autoencoders (SAEs) to project visual tokens into an interpretable latent space where token relationships can be analyzed through shared concept activations. Specifically, we define concept overlap as the degree of agreement between active SAE latents and use it to group tokens that share semantic content. Reduction within each group is then performed by either pruning or merging, providing a unified framework that preserves semantically important visual information while removing redundancy. Unlike fixed-budget approaches, TORINO dynamically adapts the reduction rate to input complexity, allowing different images to retain different numbers of tokens. Experiments across multiple vision-language benchmarks show that TORINO achieves favorable efficiency-accuracy trade-offs, reducing the number of visual tokens with minimal performance loss.
Comment: SAE concept overlap drives input-adaptive visual-token pruning and merging.
Topic Match: The central mechanism compresses visual tokens to reduce computation in large VLM decoders.
Relevance: 8 Novelty: 6
3. Elastic Gang: Per-Token Membership Change for a Hard-Barriered LLM Inference Gang Co-Scheduled with OS Processes
ArXiv ID: 2607.04668
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Daeyeon Son
Abstract: On-device LLM decoding is a hard-barriered CPU-SIMD computation that wants every core for milliseconds per token, while the rest of the OS wants those same cores continuously. A barriered gang cannot simply be dropped into a preemptive scheduler: an unannounced departure deadlocks a barrier, and an unannounced arrival silently corrupts logits. I present the elastic gang of Anima OS, a bare-metal x86-64 Rust kernel in which the inference gang is a first-class schedulable entity whose core membership may change between any two tokens. The core mechanism is an ACK-latched epoch protocol that never waits on a named core: a seqlock-style generation-tagged latch composed with RCU/epoch-style membership consent, so each token's participant set is the intersection of the cores the gang requested and the cores that acked the current epoch. An un-acked core is outside this token and joins at most one token later. Displaced general processes migrate and keep running; cores return to them the moment a generation ends. On a real AMD Zen 5 machine (8C/16T), inference output is bit-exact under verified per-token membership change on both a 135M and a 7B model, the property that makes elasticity safe in a kernel whose safety gate reads logits. Against fair static core partitions, elastic membership Pareto-dominates: at intermediate inference duty cycles it delivers 1.75x (25%), 1.52x (50%), and 1.28x (75%) the general throughput of a static 8-core split at equal or better inference throughput, recovers all eight stranded cores when inference is idle, and converges to the split at saturation. Returning a lent core costs 0.22 us (p50); acquiring a busy, tenant-occupied core costs one scheduling quantum (~16 ms): a running tenant is never preempted mid-slice. Decode throughput saturates at gang width 8, so ceding cores past the knee is nearly free: elasticity auto-sizes the gang online.
Comment: An acknowledgment-latched epoch protocol safely changes CPU gang membership between decoded tokens.
Topic Match: The new synchronization mechanism improves inference resource sharing, with scope limited to CPU and OS scheduling.
Relevance: 7 Novelty: 7
4. Layer-Parallel Inference Reduces Encrypted Nonlinear Depth in Transformers
ArXiv ID: 2607.04819
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Ligong Han, Kai Xu, Hao Wang, Ruijiang Gao, Han Gao, Akash Srivastava
Abstract: Fully homomorphic encryption (FHE) enables computation on encrypted data, but practical encrypted Transformer inference is bottlenecked by the sequential composition of many nonlinear blocks. We study whether Structured Newton Layer Parallelism (SNLP) can make this inter-layer composition more FHE-friendly: each Transformer block still requires polynomial approximations for operations such as softmax and RMSNorm, but SNLP reduces the layerwise sequential nonlinear depth from L stages to a small number of solver iterations plus linear structured corrections. Using a simulation framework based on Chebyshev polynomial approximations, we measure error accumulation under sequential versus SNLP inference across 8 models and 4 architecture families. On a 0.5B IDN-trained model, SNLP reduces symbolic bootstraps from 53 to 20 (2.65x) with only +1.2% perplexity degradation, while lowering error amplification (1.36x vs. 1.42x). Across all tested models, SNLP has lower amplification than sequential inference. Ablations show that softmax approximation dominates the error budget and CKKS arithmetic noise is negligible in our setting, suggesting that SNLP is complementary to block-level FHE-friendly operator design rather than a replacement for it.
Comment: Layer-parallel Newton iterations reduce sequential nonlinear depth in encrypted Transformer inference.
Topic Match: The core targets encrypted inference cost, with evidence limited to approximation simulations and symbolic bootstrap counts.
Relevance: 7 Novelty: 6
5. Learning Only What Valid Adapters Can Express: Subspace-Constrained Adaptation Against Fine-Tuning Poisoning
ArXiv ID: 2607.05300
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Fabien Polly
Abstract: Parameter-efficient fine-tuning still leaves a broad space of behavior-changing updates reachable, so a poisoned objective can be represented and optimized. We study an alternative: adaptation constrained to the subspace estimated from a trusted pool of existing task adapters. On flan-t5-large with 196 public LoRA adapters, we show that (1) the functionally relevant content of an adapter lies in a low-dimensional shared subspace, 30 to 38 percent of its weight norm being redundant under the evaluated task distributions; (2) gradient adaptation restricted to 128 coordinates on this subspace matches full LoRA fine-tuning on clean classification data, while under targeted label inversion LoRA collapses to 3-26 percent exact match and the constrained learner keeps 62-96 percent on the tasks the pool covers; (3) the constrained learner cannot fit corrupted data, its adaptation loss separating clean from garbage by two orders of magnitude (120x), an out-of-distribution signal without an extra detector; and (4) against an adaptive backdoor attacker who optimizes within the subspace, the attack is blocked (8 percent success versus 100 for LoRA) on the task where its target behavior is unlike anything in the pool, and only partially blocked (85 percent) when the target coincides with a common pool behavior. On these two tasks the outcome is consistent with how close the target is to the pool's directions, which suggests but does not establish a pool-relative boundary. The mechanism trades peak plasticity for these properties: on tasks the pool covers poorly, unconstrained fine-tuning wins, and the protection assumes the pool itself is trusted. Code and data are public.
Comment: Restricts adapter updates to 128 coordinates in a shared subspace estimated from trusted task adapters.
Topic Match: The shared subspace touches parameter-efficient adaptation; the central evaluated benefit is poisoning resistance, with computational savings unestablished.
Relevance: 6 Novelty: 6
Paper Selection Prompt
System Prompt
You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.
User Prompt
Instructions
Respond in JSONL. Output exactly one JSON object per paper, one per line:
{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}
Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only:
daily_hot,new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return[]. -daily_hotmeans the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. -new_frontiermeans the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.Scoring Criteria
Relevance and Novelty are independent axes. Score both from 1 to 10.
Relevance Scoring
- 9-10: directly centered on the target foundational topics; highest when the core contribution is clearly within them.
- 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
- 5-6: touches the target topics, but the main contribution is elsewhere.
- 3-4: largely outside the target topics, often application-focused or domain-specific.
- 1-2: unrelated.
Important: Broad frontier relevance, major launch status, or daily buzz is not enough for a high Relevance score here. Those cases belong in the hotspot digest unless the paper strongly matches the specialized paper topics.
Novelty Scoring
- 9-10: new paradigm, theory, or major methodological breakthrough.
- 7-8: substantial methodological advance or strong new insight.
- 5-6: meaningful but incremental extension or refinement.
- 3-4: minor, narrow, or mostly engineering or domain-specific improvement.
- 1-2: little originality; mainly standard application of existing methods.
Topic Registry
Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - moe_training: MoE Training - Mixture-of-Experts training: routing, load balancing, expert granularity and shared experts, training stability, MoE parallelism and communication, and MoE kernels. - training_systems: Large-Scale Training Systems and Efficiency - Large-scale training systems and efficiency: distributed training algorithms, optimizers, and communication. - architecture_training: Architecture and Training Dynamics - Core architectural or computational mechanisms, dynamic computation, and training-stability dynamics. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, memory or cache efficiency, and large-scale training systems that materially change cost or behavior.
Papers
[PAPER LIST HERE]
Relevant Topics
This is a training-side feed. The centre of gravity is Mixture-of-Experts training and the large-scale training systems it sits on; architecture and efficiency work is kept when it bears on how a large model is trained or what it costs to train.
Keep a paper when its CORE CONTRIBUTION falls in one of the four topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.
MoE Training (primary) - Keep: expert routing and gating mechanisms; load balancing, with or without auxiliary losses; expert granularity, fine-grained and shared experts; MoE training stability (loss spikes, router collapse, z-loss and related fixes); expert and hybrid parallelism; all-to-all dispatch/combine and the communication schedules MoE needs; MoE kernels and grouped-GEMM style implementations; expert capacity, token dropping, and upcycling dense checkpoints into MoE. - Filter: papers that merely train on top of a MoE model without a new routing, balancing, stability, parallelism, or kernel contribution; "mixture of experts" in the classical ensemble or recommender sense.
Large-Scale Training Systems - Keep: distributed training algorithms (data, tensor, pipeline, sequence, and expert parallelism, FSDP and ZeRO-style sharding); optimizers and preconditioners for large-scale pretraining; communication algorithms, overlap, and collective design; throughput and MFU work that materially changes what a training run costs or whether it converges; scaling-law work that informs how such runs are configured. - Filter: routine infrastructure tuning with no new algorithmic or systems idea.
Architecture and Training Dynamics - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, dynamic or modular computation); optimisation and training-dynamics analysis that explains why large models train the way they do. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight.
Efficiency and Compression - Keep: quantization, sparsity, pruning, low-rank adaptation, KV-cache and memory-efficient designs, and other work that materially changes the cost of running or training a large model, especially when the mechanism is new rather than a tuned variant. - Filter: straightforward application of standard efficiency methods, and deployment or serving engineering with no new idea.
Leave these to the hotspot digest unless the core contribution clearly falls in one of the four topics above: - major model or product releases - agent, tooling, and RAG launches - benchmark, leaderboard, or evaluation-only papers - representation-learning theory, interpretability, and feature-formation analysis - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning) - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains