Previous Day 2026-07-10
Monthly Overview 2026-07
Next Day 2026-07-14

This is a remedial run for missed papers from 07/10/2026 to 07/12/2026.

Results generated on 09/12/2026.

Personalized Daily ArXiv Papers 2026-07-13

Model Metric Usage Papers
Prompt Completion Total Total arXiv Scanned Relevant
local agent CLI Tokens not reported not reported not reported 595 595 40
Cost not reported not reported not reported

Token counts are not reported for this run. 8 of 16 model calls succeeded, 7,204s of model wall clock.

Topic Coverage:

TopicPapers
Large-Scale Training Systems and Efficiency3
Architecture and Training Dynamics22
Efficiency, Compression, and Large-Scale Training15

Table of contents by topic:

Large-Scale Training Systems and Efficiency (3)

  1. WSqD: A Horizon-Free Learning Rate Schedule for Large Model Training Authors: Jianhao Ma, Yuxin Chen

  2. Understanding Schedule-Free Methods in Nonconvex Optimization: Rate Guarantees and Escaping Saddles Authors: Jiseok Chae, Donghwan Kim

  3. LionVote: Per-Layer Learning Rate Adaptation for Lion Authors: Kris Atallah

Architecture and Training Dynamics (22)

  1. LeRoPE: Learnable RoPE Frequencies Improve Language Modeling Authors: Petros Karypis, Sean O'Brien, Shreyas Kadekodi, Rui Zhu, Julian McAuley

  2. LayerNorm as Implicit Gain Control in Looped Transformers Authors: Matthias M. M. Buehlmaier

  3. Next-Dense-Stride Prediction for Multimodal Autoregressive Visual Modeling Authors: Chicago Y. Park, Jialin Mao, Xiaojian Xu, Taha Kass-Hout, Ulugbek S. Kamilov, Cao Xiao

  4. From Atoms to Entropy: Optimal Noise Allocation for Diffusion Training in the Convex Regime Authors: Luca Ambrogioni, Giulio Franzese, Alberto Foresti, Gabriel Raya, Bac Nguyen, Georgios Batzolis, Yuhta Takida, Naoki Murata, Chieh-Hsin Lai, Yuki Mitsufuji

  5. Foveation-Guided Dynamic Token Selection for Robust and Efficient Vision Transformers Authors: Ibrahim Batuhan Akkaya, Kishaan Jeeveswaran, Bahram Zonooz, Elahe Arani

  6. Adaptive Compute in Latent World Models: When Depth Helps, Hurts, or Doesn't Matter Authors: Achyuthan Sivasankar

  7. CHASE: Cache-Hole-Adapted Skip Exit for Looped State-Space Language Models Authors: Zhenxuan Yu, Takeshi Kojima, Yutaka Matsuo, Yusuke Iwasawa

  8. Learning in Curved Weight Space:Exponential-Linear Weight Reparameterization for Improved Optimization Authors: Ethan Smith

  9. MLPs are Hebbians: Constructing Efficient Fact-Storing MLPs for Transformers Authors: Roberto Garcia, Jerry Liu, Ronny Junkins, Sabri Eyuboglu, Atri Rudra, Christopher Ré

  10. When Data Imbalance Helps: Robust Generalization Through Shortcut Saturation Authors: Cheng-Ting Chou, Duc Binh Hoang

  11. Robustly Invertible Nonlinear Dynamics and the BiLipREN: From Inversion-Based Control to Generative Trajectory Modelling Authors: Yurui Zhang, Ruigang Wang, Ian R. Manchester

  12. Conservation Laws for Diffusion Models Authors: Ziv Aharoni, Henry D. Pfister

  13. Singular perturbations and hierarchical learning in two-layer neural networks Authors: Cédric Gerbelot, Jean-Christophe Mourrat

  14. From Self-Attention to Connection Laplacian: A Unified Operator View of Transformers Authors: Binbin Lin, Wei Chen, Yalun Li, Wenxiao Wang, Jieping Ye, Xiaofei He

  15. AutoNorm: Understanding Adaptive Normalization in Transformers through Differentiable Gating Authors: Piyush Kaushik Bhattacharyya, Divyanshu Rai, Swastik Singh, Kumar Aakash, Ayush Ranjan, Krutika Verma

  16. Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation Authors: Keqin Peng, Chen Li, Yuanxin Ouyang, Yancheng Yuan, Liang Ding

  17. Interference and Retention in Continual Learning Authors: Julius Störk

  18. Source-Lifted Flow Matching for Intervenable Multimodal Imitation Authors: He Zhang, Ying Sun, Pengteng Li, Ziyang Chen, Yiren Zhao, Ziyang Rao, Weiyu Guo, Yandong Guo, Hui Xiong

  19. Exploring More to Solve More: Boosting Diversity in Text Diffusion Models via Entropy-Based Guidance Authors: Jingwei Zhang, Haoyu Lei, Zijin Feng, Jiacheng Sun, Farzan Farnia

  20. Energy-guided Recursive Model Authors: Yifei Zhao, Ying Tang

  21. Fully Trainable Deep Differentiable Logic Gate Networks and Lookup Table Networks Authors: Wout Mommen, Lars Keuninckx, Matthias Hartmann, Werner Van Leekwijck, Piet Wambacq

  22. On Locality and Length Generalization in Visual Reasoning Authors: Pulkit Madan, Sanjay Haresh, Reza Ebrahimi, Sunny Panchal, Apratim Bhattacharyya, Roland Memisevic

Efficiency, Compression, and Large-Scale Training (15)

  1. CARE-LoRA: Compressed Activation REconstruction for Memory-Efficient LoRA Authors: Gengyu Zhang, Haiyin Ran, Zhengbao He, Yuhang Liu, Hanling Tian, Zhehao Huang, Xiaolin Huang

  2. Super-Tuning: From Activation-Aware Pruning to Sparse Fine-Tuning Authors: Ivan Ilin, Philip Zmushko, Peter Richtárik

  3. Codec-Gauge: Learning Compression-Friendly Gauges for Transformer KV Caches Authors: Yitao Jiang, Yaoqing Yang, Luyang Zhao, Muhao Chen, Devin Balkcom

  4. MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference Authors: Venkatesha Matam, Keon Kim

  5. STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Authors: Victor J. B. Jung, Gagandeep Singh, Joseph Melber, Kristof Denolf, Francesco Conti, Luca Benini

  6. KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling Authors: Peng Kuang, Haibo Jin, Xiaoyu Han, Yanli Wang, Xiaopeng Yuan, Ye Yu, Kaidi Xu, Haohan Wang

  7. Adaptive Model Compression (AMC): Saliency-Driven Resource Allocation for Ultra-Low-Power Transformer Inference Authors: Jiayin Hu, Kai Yuan, Vanessa Hu, Xuetao Yin, Jianhua Li, Sean Suchter

  8. Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM Inference Authors: Junfei Zhan, Haoxun Shen, Mingang Guo, Zixuan Huang, Tengjiao He

  9. IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation Authors: Yiting Wang, Jingyi Zhang, Wenhu Zhang, Ke Chao, Yves Liang, Kun Cheng, Kang Zhao

  10. Unlocking Parallelism in Autoregressive Language Models via Speculative Decoding with Progressive Tree Drafting Authors: Zipeng Gao, Zhi Zheng, Qingrong Xia, Junda Lin, Ziwei Zhao, Tong Xu, Zhefeng Wang, Enhong Chen

  11. Reliability Scaling Laws for Quantized Large Language Models Authors: Sirine Ayadi, Sándor Daróczi, Stephan Günnemann, Bertrand Charpentier

  12. IterCOMP: Reasoning-aware Adaptive Prompt Compression for Multi-hop Question Answering Authors: JungMin Yun, YoungBin Kim

  13. Self-Guided Test-Time Training for Long-Context LLMs Authors: Xinyu Zhu, Zhe Xu, Xiaohan Wei, Yunchen Pu, Fei Tian, Chonglin Sun, Kaushik Rangadurai, Hua Zhi, Frank Shyu, Sandeep Pandey, Luke Simon, Yu Meng, Xi Liu

  14. Dependency-Aware Chain-of-Thought Compression for Financial Reasoning Authors: Wenjun Wu, Lei Fu, Kejian Tong, Tao Ning, Sichen Zhao

  15. Structured Thoughts For Improved Reasoning And Context Pruning Authors: Zain Sarwar, Supriyo Chakraborty, Berkcan Kapusuzoglu, Chia-Hsuan Lee, Anirban Das, Stephen Rawls, Kartik Balasubramaniam, Sambit Sahu


Large-Scale Training Systems and Efficiency (3)

1. WSqD: A Horizon-Free Learning Rate Schedule for Large Model Training

ArXiv ID: 2607.10959

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Jianhao Ma, Yuxin Chen

Abstract: Standard learning rate schedules such as cosine annealing are tied to a fixed training horizon, limiting their ability to accommodate post hoc horizon extension. Warmup-stable-decay (WSD) partially addresses this issue by maintaining a long constant-rate phase before a short linear cooldown, allowing training to resume from a pre-decay checkpoint. However, its peak learning rate is still tuned based on the original training horizon and can become suboptimal when training is extended. Motivated by stochastic convex optimization, we propose WSqD (Warmup with Square-root base and linear Decay), a learning rate schedule that replaces WSD's constant stable phase with a shifted inverse-square-root base while retaining the final linear cooldown. In the stochastic convex setting, WSqD provably attains the minimax-optimal $O(1/\sqrt{T})$ last-iterate convergence rate. Importantly, its base learning rate schedule is horizon-independent, and the training horizon is needed only to determine when to begin the final cooldown. Empirically, on language-model pretraining using the SlimPajama corpus, WSqD matches or outperforms carefully tuned WSD and other baselines across multiple training horizons while reusing a single peak learning rate.

Comment: Introduces a horizon-independent inverse-square-root learning-rate base with optimal last-iterate convergence theory.

Topic Match: The core contribution is a pretraining schedule that permits horizon extension without retuning the peak learning rate.

Relevance: 9 Novelty: 7


2. Understanding Schedule-Free Methods in Nonconvex Optimization: Rate Guarantees and Escaping Saddles

ArXiv ID: 2607.09167

Primary Topic: Large-Scale Training Systems and Efficiency

Also Matches: Architecture and Training Dynamics

Authors: Jiseok Chae, Donghwan Kim

Abstract: Schedule-Free methods have attracted growing interest for alleviating the burden of designing and tuning a learning rate scheduler, while matching and sometimes even outperforming optimizers with tuned schedulers. Despite their strong empirical results, their convergence theory in nonconvex optimization, where modern machine learning objectives typically arise, has remained largely unexplored. In this paper, we provide worst-case analyses of Schedule-Free gradient descent and Schedule-Free stochastic gradient descent, in their standard form and without auxiliary modifications or restrictive conditions, for smooth but possibly nonconvex objectives. Based on a Lyapunov analysis derived from the continuous-time limiting ordinary differential equation associated with these methods, we show that Schedule-Free gradient descent and Schedule-Free stochastic gradient descent achieve the optimal worst-case convergence rates attainable among first-order methods. We further formulate Schedule-Free gradient descent as a nonautonomous dynamical system and prove strict-saddle avoidance under an arbitrarily small one-time perturbation. These theoretical results provide a better understanding of the strong performance that Schedule-Free methods demonstrate.

Comment: Establishes nonconvex convergence rates and saddle avoidance for schedule-free optimization.

Topic Match: The paper directly explains the convergence dynamics of a large-model optimizer family.

Relevance: 8 Novelty: 7


3. LionVote: Per-Layer Learning Rate Adaptation for Lion

ArXiv ID: 2607.09266

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Kris Atallah

Abstract: Per-layer diagnostics reveal that, at the prescribed learning rate, Lion's effective scale is 2.6-2.8x too high for attention and MLP parameters and ~2x too high for normalization layers on ViT-Tiny/CIFAR-100; this 32% cross-layer-type disparity cannot be reproduced by a single global rate. The measurement comes from LionVote, a per-layer learning rate mechanism in which each parameter tensor maintains a compound level, a persistent integer updated every c epochs by two diagnostics (gradient direction stability and momentum health) resolved by a validation loss tiebreaker. Voting thresholds derive from geometric identities, the EMA time constant, and a noise-floor estimate; cadence is bounded structurally and selected by ablation. On ViT-Tiny/CIFAR-100, LionVote achieves 69.7% top-1 accuracy vs. Lion's 69.0% (p < 0.02, Welch's t-test) and AdamW's 68.8%. Per-layer adaptation value depends on both architectural heterogeneity and task; on uniform CNN architectures tuned SGD with cosine annealing remains dominant, and on ViT architectures gains are task-dependent.

Comment: Adapts Lion learning rates per parameter tensor using gradient and momentum diagnostics.

Topic Match: It contributes a new optimizer-control mechanism for heterogeneous model layers.

Relevance: 7 Novelty: 6


Architecture and Training Dynamics (22)

1. LeRoPE: Learnable RoPE Frequencies Improve Language Modeling

ArXiv ID: 2607.10134

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Petros Karypis, Sean O'Brien, Shreyas Kadekodi, Rui Zhu, Julian McAuley

Abstract: Rotary Positional Encodings (RoPE) are currently the most popular positional encodings used in modern language models. RoPE rotates two-dimensional chunks of query and key vectors, operating as a function of their relative positional offset. The position-wise rates of rotation in RoPE typically follow a geometric sequence specified by a fixed base-frequency hyperparameter. Prior work has improved performance by either increasing this parameter to slow rotation or by applying RoPE to only a subset of QK dimensions. In this work we modify RoPE by learning a scalar per frequency, treating frequencies as learnable parameters rather than hyperparameters. We validate Learned RoPE by training a ladder of language models from scratch, ranging from 52M to 2.5B parameters. We observe and analyze the emergence of a high-norm, positional LeRoPE band. LeRoPE consistently outperforms RoPE and partial RoPE across all scales, with RoPE requiring 3.4% more compute (FLOPs) to match LeRoPE at the largest scale.

Comment: Makes individual RoPE frequencies learnable and shows consistent from-scratch gains through 2.5B parameters.

Topic Match: Learnable positional frequencies are the core architectural change, with reduced compute-to-quality as a secondary efficiency result.

Relevance: 9 Novelty: 7


2. LayerNorm as Implicit Gain Control in Looped Transformers

ArXiv ID: 2607.10681

Primary Topic: Architecture and Training Dynamics

Authors: Matthias M. M. Buehlmaier

Abstract: In pre-LayerNorm looped transformers, LayerNorm inside the recurrent block acts as an implicit gain controller: by coupling the block's local Lipschitz constant inversely to the activation scale, it renders the recurrence Jacobian non-normal -- asymptotically contractive at every verified fixed point even where its operator norm exceeds 1 -- so the true stability budget is the spectral margin, not an operator-norm bound. That margin depletes as the carry $ρ\to 1$, and a minority of initializations never converge to a fixed point at all, so the diagonal carry constraint $ρ(\bar{A}) < 1$ is necessary but not sufficient for convergence of the full recurrence. Training experiments across six tasks, including a controlled ablation, reveal that the linear carry is not the depth-memory mechanism: gradient descent routes memory through the block's more expressive nonlinear recurrence and leaves the stability-constrained carry at rest -- the carry's role is stabilization, not memory. We characterize the boundary of this claim: on tasks with axis-aligned per-channel structure, gradient descent does recruit the carry. All results are derived analytically and verified in a from-scratch, CPU-scale implementation; verification at larger scale is needed.

Comment: Shows that LayerNorm induces gain-controlled non-normal stability in looped Transformer recurrences.

Topic Match: It directly analyzes normalization, recurrence stability, and learned memory pathways.

Relevance: 8 Novelty: 8


3. Next-Dense-Stride Prediction for Multimodal Autoregressive Visual Modeling

ArXiv ID: 2607.09892

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Chicago Y. Park, Jialin Mao, Xiaojian Xu, Taha Kass-Hout, Ulugbek S. Kamilov, Cao Xiao

Abstract: We introduce DenseAR, a new generative paradigm that reformulates autoregressive image generation as coarse-to-fine next-dense-stride prediction using a compact single-scale tokenizer. Our key insight is that traversing a single-scale latent grid with progressively denser strides naturally captures the transition from global structure to fine detail. This addresses two limitations of existing autoregressive models at once: the slow inference of raster-order autoregression, which DenseAR avoids by predicting multiple tokens in parallel, and the heavy cost of multi-scale approaches, which need long, multi-resolution token sequences to achieve coarse-to-fine prediction. Building on our efficient framework and the flexibility of autoregressive modeling, we further extend DenseAR to a unified model that handles multiple modalities and imaging tasks within a single backbone. We validate DenseAR on both medical and natural images. On multi-contrast brain MRI, a single DenseAR model unifies cross-modal translation, modality-conditioned generation, and tumor segmentation, while remaining competitive with task-specific methods. On ImageNet, DenseAR improves class-conditional generation quality (FID and IS) over both a single-grid baseline without stride ordering and a multi-scale tokenizer-based baseline.

Comment: Replaces raster autoregression with parallel coarse-to-fine prediction over progressively denser strides.

Topic Match: The main contribution is a new autoregressive computation that shortens generation sequences.

Relevance: 8 Novelty: 8


4. From Atoms to Entropy: Optimal Noise Allocation for Diffusion Training in the Convex Regime

ArXiv ID: 2607.20540

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Luca Ambrogioni, Giulio Franzese, Alberto Foresti, Gabriel Raya, Bac Nguyen, Georgios Batzolis, Yuhta Takida, Naoki Murata, Chieh-Hsin Lai, Yuki Mitsufuji

Abstract: How should a diffusion model decide which noise levels to train on, and how much? Despite the importance of this choice, current noise schedules are based largely on heuristics or empirical tuning. Here, we develop a general statistical framework for studying asymptotically optimal noise-level allocation in diffusion training. Our first main result concerns the fully coupled regime, where information can spread between different time points. Under convexity or Polyak-Lojasiewicz-type assumptions, we show that the optimized training schedule admits an atomic minimizer, concentrated on finitely many noise levels. Our second main result specializes this framework to an idealized independent-learner regime, intended to model temporal specialization in neural networks. Under an additional feature-noise decoupling condition, a random-matrix analysis leads to an information-theoretic proxy: the decoupled sampling density is proportional to the square root of the generative entropy rate, the rate at which conditional entropy grows along the forward process. We test these predictions in controlled settings where the coupled objective can be optimized directly, including Dirac mixtures, low-dimensional manifolds, and MNIST. In these settings, the optimized schedules are consistently finite-support, while the smooth entropic proxy closely tracks the atomic optimum in neural-network models and breaks down mainly in the fully coupled parametric case, as the theory suggests. We then evaluate the entropic schedule in larger-scale experiments, where full schedule optimization is currently intractable. The results indicate that square-root entropy scheduling can substantially improve training efficiency on discrete domains and remains competitive with standard EDM-style heuristics on continuous images.

Comment: Derives theoretically optimal diffusion noise-level allocations and an entropy-based training schedule.

Topic Match: Training dynamics is primary because the theory characterizes where optimization effort should be allocated across diffusion time.

Relevance: 8 Novelty: 8


5. Foveation-Guided Dynamic Token Selection for Robust and Efficient Vision Transformers

ArXiv ID: 2607.09480

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Ibrahim Batuhan Akkaya, Kishaan Jeeveswaran, Bahram Zonooz, Elahe Arani

Abstract: The human visual system (HVS) employs foveated sampling and eye movements to achieve efficient perception, conserving both metabolic energy and computational resources. Drawing inspiration from this robustness and adaptability, we introduce the Foveated Dynamic Transformer (FDT), a foveation-guided dynamic token-selection architecture that integrates these mechanisms into a vision transformer framework. The FDT exhibits strong resilience to various types of noise and adversarial attacks, despite not being explicitly trained for such challenges. This inherent robustness is achieved through the use of fixation and foveation modules: the fixation module identifies fixation points to filter out irrelevant information, while the foveation module generates foveated embeddings with multi-scale information. At the 50% fixation-budget setting, FDT achieves higher accuracy than DeiT-S (81.9% vs. 80.9%) while reducing multiply-accumulate operations by 34.57%, highlighting one operating point on its accuracy-efficiency trade-off. These attributes position FDT as an HVS-inspired step toward artificial neural networks that combine adaptive computation with improved resilience.

Comment: Selects vision tokens dynamically through learned fixation and multiscale foveation.

Topic Match: Dynamic token computation is the architectural mechanism producing the compute reduction.

Relevance: 8 Novelty: 7


6. Adaptive Compute in Latent World Models: When Depth Helps, Hurts, or Doesn't Matter

ArXiv ID: 2607.10203

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Achyuthan Sivasankar

Abstract: Adaptive compute for world models -- early-exit or mixture-of-depths predictors that spend variable depth per rollout step -- presumes that extra depth buys better predictions. In autoregressive rollouts, where planning actually happens, that premise requires depth's per-step precision to survive composition. We test it directly with one pre-registered instrument, the shallow penalty rho = err(shallowest-exit rollout)/err(full-depth rollout), on nine DeepMind Control tasks under matched single-step (K=1) and multi-step (K=4) training, eight seeds each. Three regimes emerge: depth helps (intrinsic, 6/9 tasks, rho up to 8x), depth actively hurts (inversion, 2/9, rho down to 0.87x), or depth barely matters (flat). The inversion is created by training, not the dynamics: supervising early exits only at the first rollout step erases it (Delta=+0.28, n=8, non-overlapping distributions) -- a routability catch-22: the per-step deep supervision that makes exits routable also trains them to out-roll the full stack. The regime is predictable: a frozen dimensionality-only classifier, committed before training, labels held-out tasks correctly out-of-sample, including an extreme extrapolation. The inversion reproduces under a transformer predictor, yet its manifestation is configuration-dependent, shifting with metric space, horizon, encoder, backbone, and -- most strongly -- training data: on the two tasks we retrained, competent-policy data removes both the inversion and the intrinsic tradeoff, loss unchanged. In a CEM planner, rho predicts whether planning benefits from depth. Every threshold and gate was committed before the corresponding compute, including a pre-registered negative for the motivating hypothesis. Whether more compute helps a world model is not a task property; it is a property of the operating configuration, with a stable, predictable, mechanism-backed core.

Comment: Identifies training-dependent regimes in which mixture-of-depths exits help, hurt, or become irrelevant during rollouts.

Topic Match: The central result explains the training dynamics of adaptive-depth computation, with compute savings as a secondary consequence.

Relevance: 8 Novelty: 7


7. CHASE: Cache-Hole-Adapted Skip Exit for Looped State-Space Language Models

ArXiv ID: 2607.10110

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Zhenxuan Yu, Takeshi Kojima, Yutaka Matsuo, Yusuke Iwasawa

Abstract: Recent work on looped language models suggests that many reasoning problems benefit from greater computational depth rather than from additional independent parameters. Existing studies, however, focus almost exclusively on Transformer backbones, leaving open whether this principle also applies to state-space language models. We investigate Looped Mamba and Looped Hybrid Mamba-Transformer architectures, which repeatedly apply a shared Mamba (or hybrid) block to introduce explicit finite-depth recurrent computation. On two controlled reasoning tasks-Mano (modular-arithmetic manipulation) and p-hop induction-Looped Mamba consistently outperforms parameter-matched non-looped baselines and, in several settings, matches or exceeds non-looped models of equal effective depth. We then extend the study to language model pre-training under matched iso-parameter and iso-FLOPs protocols, which jointly disentangle the effects of parameter sharing and effective depth: looped models remain competitive on downstream benchmarks with substantially fewer distinct parameters, although deeper non-looped models retain an advantage in validation perplexity under strict iso-FLOPs comparisons. Finally, we adapt Ouro's two-stage exit gate to Looped Mamba for threshold-controlled selection among recurrent-step outputs. Executing such exits on a state-space backbone, however, leaves the recurrent state without its deeper updates, and validation perplexity then degrades severely. We therefore introduce a cache-hole adaptation that aligns continued training with skipped-state inference. At the scales studied, the adapted model keeps perplexity close to full computation and matches or exceeds full-compute exit-state selection on downstream benchmarks while executing roughly half of the recurrent steps, which translates into measured inference speedups once the prefill is compute-bound.

Comment: Combines looped Mamba depth with cache-hole-aware exit training to make recurrent-step skipping viable.

Topic Match: The primary advance is an adaptive recurrent architecture and training procedure, with reduced inference computation as a secondary benefit.

Relevance: 8 Novelty: 7


8. Learning in Curved Weight Space:Exponential-Linear Weight Reparameterization for Improved Optimization

ArXiv ID: 2607.09967

Primary Topic: Architecture and Training Dynamics

Also Matches: Large-Scale Training Systems and Efficiency

Authors: Ethan Smith

Abstract: Many neural networks operations have a multiplicative nature rather than additive: halving or doubling a norm are analogous relatively but require unequal optimization distances when taking linear steps. Adaptive optimizers such as Adam normalize updates per coordinate, but update steps remain additive; weights with very different magnitudes receive similarly sized absolute changes, producing very different relative perturbations. We introduce \textbf{\method} (\textbf{\methodshort}), a weight reparameterization for neural networks that combines a sign-aware symmetric-exponential pathway with an identity-like linear pathway. The symmetric-exponential pathway is near-linear for small raw weights but increasingly curved at larger magnitudes. Additive updates in logarithmic space map to magnitude-proportional changes in effective weight space. The linear pathway provides a direct route through the transform that we hypothesize stabilizes optimization, while learnable scale, curvature, and offset parameters control balance between pathways and the curvature of the exponential pathway. These components create a curved parameter-space geometry that empirically improves speed of loss descent over standard linear parameterization. We also identify a useful \emph{mismatched initialization}: raw weights are chosen so a symmetric version of the transform matches Xavier statistics, but training uses an asymmetric forward transform that leaves positive weights at full strength while making negative weights smaller in magnitude; in small-model ablations, this improves early optimization and may act as a form of symmetry breaking. We train transformers on OpenWebText over nine width$\times$depth configurations, \methodshort reaches matched validation loss in 1.32--1.49$\times$ fewer training steps, with the largest widths seeing the biggest gains.

Comment: Reparameterizes weights with exponential and linear pathways to reduce transformer pretraining steps.

Topic Match: The central mechanism changes parameter-space geometry and optimization dynamics, while its step reduction also affects training-system cost.

Relevance: 8 Novelty: 7


9. MLPs are Hebbians: Constructing Efficient Fact-Storing MLPs for Transformers

ArXiv ID: 2607.10034

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Roberto Garcia, Jerry Liu, Ronny Junkins, Sabri Eyuboglu, Atri Rudra, Christopher Ré

Abstract: Large language models (LLMs) store factual knowledge in their parameters. While recent work has shown that this knowledge resides in MLP layers, existing constructive and mechanistic interpretability models of fact-storage in LLMs fail to explain the surprising empirical phenomenon that they store facts at an information-theoretically optimal rate. In this work, we develop a theoretical account of this phenomenon. We develop the first Transformer-compatible fact-storing MLP closed-form construction that satisfies the following three properties empirically observed in LLMs: it (i) attains optimal fact storage scaling, (ii) handles arbitrary input/output geometries, and (iii) works inside Transformers. Key to our work is to analyze the decoding margin of MLPs, whereas prior work only studies MLP fact storage. Under isotropic embeddings, our construction achieves information-theoretically optimal storage capacity scaling and requires $10$-$104\times$ fewer parameters at matched fact count than prior constructions. For arbitrary key and value embeddings, we show that our construction attains the same storage capacity scaling, up to penalization factors depending on the embedding geometries. Moreover, we demonstrate that our constructed MLPs can be used within Transformer blocks for factual recall tasks at optimal capacity scaling, requiring $15$-$63\times$ fewer parameters at matched fact count than prior constructions. Finally, as a proof-of-concept, we show that fact-storing MLPs enable modular fact editing by swapping a Transformer's MLP with a new one.

Comment: Constructs Transformer-compatible MLPs with information-theoretically optimal fact-storage scaling.

Topic Match: Its main contribution is a new constructive account of the Transformer MLP's computational role.

Relevance: 7 Novelty: 8


10. When Data Imbalance Helps: Robust Generalization Through Shortcut Saturation

ArXiv ID: 2607.10116

Primary Topic: Architecture and Training Dynamics

Authors: Cheng-Ting Chou, Duc Binh Hoang

Abstract: We study robust generalization under spurious correlations: tasks where a shortcut feature is correlated with the true label in training but anti-correlated in an adversarial held-out split. Varying the spurious ratio $r$ (the fraction of training examples where shortcut = true label) and model capacity, we find a counterintuitive result: data imbalance promotes generalization in sufficiently capable models. On a synthetic task where the true label is sum parity of an integer sequence and the shortcut is the parity of the maximum-valued element, a 2-layer, 2-head transformer generalized (reached $100\%$ adversarial accuracy) in 0% of seeds at $r{=}0.50$ but 77% of seeds at $r{=}0.90$. The effect is absent in 1-layer models, where imbalance instead traps the model on the shortcut. Through mechanistic analysis -- gradient conflict dynamics, circuit evolution, and QK/OV circuit ablations -- we characterize a mechanistic pathway consistent with imbalance promoting generalization.

Comment: Explains how shortcut saturation can redirect Transformer optimization toward robust circuits.

Topic Match: Its core contribution is a mechanistic account of feature-learning and gradient-conflict dynamics.

Relevance: 7 Novelty: 8


11. Robustly Invertible Nonlinear Dynamics and the BiLipREN: From Inversion-Based Control to Generative Trajectory Modelling

ArXiv ID: 2607.10026

Primary Topic: Architecture and Training Dynamics

Authors: Yurui Zhang, Ruigang Wang, Ian R. Manchester

Abstract: This paper proposes a new notion of robust invertibility for nonlinear dynamical systems, and introduces constructive parameterizations of recurrent neural network which are robustly invertible by design. We define robust invertibility as the existence of a causal inverse system such that both the forward and inverse systems are contracting and have bounded incremental input-output gains (the system is bi-Lipschitz), implying that both forward prediction and input reconstruction are robust to signal perturbations and initial-state mismatch. We construct robustly invertible recurrent models via series composition of static orthogonal layers and dynamic layers satisfying a strong input-output monotonicity property, and provide a differentiable neural network parameterizations in the form of the bi-Lipschitz recurrent equilibrium network (BiLipREN). Additionally, composition with dynamic orthogonal layers yields a nonlinear minimum-phase/all-pass (a.k.a. inner--outer) factorization. We illustrate the utility of the framework through a series of application examples in data-driven internal model control, dynamic surrogate loss learning, and signal-space normalizing flows, illustrating its utility for robust control, trajectory optimization, and generative modeling of complex trajectory distributions.

Comment: Introduces contractive, bi-Lipschitz recurrent parameterizations with robust forward and inverse dynamics.

Topic Match: The core contribution is a new recurrent architecture with provable stability and invertibility properties.

Relevance: 7 Novelty: 8


12. Conservation Laws for Diffusion Models

ArXiv ID: 2607.10067

Primary Topic: Architecture and Training Dynamics

Authors: Ziv Aharoni, Henry D. Pfister

Abstract: While autoregressive models optimize the exact data likelihood via the chain rule, diffusion models are typically trained with denoising objectives. We develop conservation laws based on generalized extrinsic information transfer (GEXIT) functions for a broad class of memoryless noise processes, showing that the data--model cross-entropy (CE) can be characterized exactly as an integral of local information-theoretic derivatives along the noise path. This yields a unified characterization of the likelihood for discrete and continuous diffusion, with the Gaussian case reducing to the well-known mutual information--minimum mean-square error (I-MMSE) relationship. An immediate implication is a locality property: one can compute the information-theoretic derivatives using only the marginal posteriors along the noise path. As a result, training reduces to learning the marginal posteriors by minimizing the negative log-likelihood. While the conservation law implies that the entropy does not depend on the noise path, finite-capacity denoisers approximate the posteriors with varying accuracy across noise types, leading to differences in performance. We validate these predictions on synthetic Markov sources and standard benchmarks, including text8 and CIFAR-10.

Comment: Derives exact information-theoretic conservation laws connecting diffusion denoising along a noise path to cross-entropy.

Topic Match: The paper provides foundational insight into diffusion training objectives and how finite-capacity denoisers interact with noise paths.

Relevance: 7 Novelty: 8


13. Singular perturbations and hierarchical learning in two-layer neural networks

ArXiv ID: 2607.10869

Primary Topic: Architecture and Training Dynamics

Authors: Cédric Gerbelot, Jean-Christophe Mourrat

Abstract: We study the population gradient flow of an infinitely wide two-layer neural network learning a misspecified single-index model in high dimension. The two layers are optimized jointly, with a perturbative parameter tuning the relative training speed between the first and second layer. This setting was considered by Berthier, Montanari and Zhou in \cite{berthier2024learning}, who conjectured a hierarchical learning scenario with explicit timescales as the second layer is trained faster than the first. In this paper, we prove that the constant and linear components of the hidden link function are indeed recovered within the predicted timescales, at sharp explicit thresholds. We then analyze the onset of learning of the quadratic component and show that the components learned at earlier stages continue to influence the dynamics in an essential way. Our proof is based on quantitative approximation results for singularly perturbed flows evolving near a manifold defined by integral constraints. At a phenomenological level, we also show that the empirical measure of the weights displays singular behaviour when reaching the quadratic component of the hidden link, with a small fraction of neurons growing significantly while the remaining ones rearrange to preserve the components already learned.

Comment: Proves explicit timescales for hierarchical feature learning in two-layer population gradient flow.

Topic Match: The paper directly advances mechanistic understanding of optimization stages and feature-learning dynamics.

Relevance: 7 Novelty: 8


14. From Self-Attention to Connection Laplacian: A Unified Operator View of Transformers

ArXiv ID: 2607.10677

Primary Topic: Architecture and Training Dynamics

Authors: Binbin Lin, Wei Chen, Yalun Li, Wenxiao Wang, Jieping Ye, Xiaofei He

Abstract: Self-attention is a ubiquitous primitive in modern sequence models, yet its operator-level geometry is only partially understood. We view a token sequence as a vector field over the token-position graph and identify attention as a connection walk: messages are aggregated by a nonnegative walk matrix while being transported along each edge by a learned linear map. Within this framework, we prove that single-head attention (SHA) is exactly a connection propagation step with constant transport, and that multi-head attention (MHA) is exactly a single edge-dependent connection walk whose effective transport is an attention-gated mixture of headwise transports. We further clarify the conditions under which the corresponding generator reduces to a random-walk connection Laplacian, highlighting the roles of stochasticity, reversibility, and metric-compatible transports. Empirically, we find that trained Transformers across scales (from 124M to 8B) and structures (encoder/decoder) exhibit geometric structure consistent with our theory: effective attention graphs converge to stable geometric operators in deeper layers, learned transports self-organize into approximate scaled isometries, and both phenomena strengthen consistently with scale. Overall, the paper provides a precise connection-walk formalism that links self-attention to classical geometric operators, along with a set of operator-level tools for analyzing transformer models from a geometric perspective.

Comment: Recasts multi-head attention exactly as an edge-dependent connection walk and studies its scale-dependent geometry.

Topic Match: The operator formalism directly analyzes the computational mechanism and emergent geometry of transformer attention.

Relevance: 7 Novelty: 8


15. AutoNorm: Understanding Adaptive Normalization in Transformers through Differentiable Gating

ArXiv ID: 2607.10593

Primary Topic: Architecture and Training Dynamics

Authors: Piyush Kaushik Bhattacharyya, Divyanshu Rai, Swastik Singh, Kumar Aakash, Ayush Ranjan, Krutika Verma

Abstract: Normalization is a critical component for stabilizing Transformer training, yet the choice between static strategies such as Layer Normalization (LN) and adaptive alternatives remains largely task-dependent. In this paper, we investigate a key optimization challenge in differentiable normalization gating. Our experiments show that, on relatively stationary vision tasks, the high gradient variance introduced by Gumbel-Softmax gating can hinder convergence of the routing mechanism, causing learned gates to underperform simple random selection. In contrast, on non-stationary language modeling and classification tasks, sustained gating diversity enables the model to learn more effective layer-wise normalization policies. Motivated by these observations, we propose AutoNorm-S (Stabilized), a training strategy that mitigates optimization instability through a gate-freezing schedule. AutoNorm-S achieves competitive or improved performance across multiple benchmarks, outperforming adaptive normalization baselines on NLP datasets, including PTB and SST-2, while remaining competitive on standard vision benchmarks. These results suggest that decoupling normalization selection from optimization noise provides a practical and principled approach for adaptive normalization in Transformer architectures.

Comment: Stabilizes differentiable normalization selection by freezing gates to control Gumbel-Softmax gradient variance.

Topic Match: Adaptive normalization and its gating-instability mechanism directly concern Transformer architecture and training stability.

Relevance: 8 Novelty: 6


16. Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation

ArXiv ID: 2607.10805

Primary Topic: Architecture and Training Dynamics

Authors: Keqin Peng, Chen Li, Yuanxin Ouyang, Yancheng Yuan, Liang Ding

Abstract: On-Policy Self-Distillation (OPSD) has emerged as a crucial paradigm for enhancing and aligning Large Language Models (LLMs). However, in complex reasoning tasks, OPSD paradoxically degrades downstream performance. In this paper, we systematically investigate this pathology and identify a severe optimization trap we define as \textbf{Thinking Collapse} -- a sharp decline in the model's native intermediate reasoning behavior, measured by epistemic-token density (ET per 1k). Through entropy-based gradient masking and token-level target analysis, we show that this collapse is triggered by aggressive teacher gradients at high-student-entropy decision forks, where student epistemic tokens are frequently suppressed into teacher non-epistemic targets and are highly concentrated in high pointwise student-teacher divergence regions. To resolve this optimization pathology, we propose \textbf{Adaptive Dual-Perspective OPSD (AD-OPSD)}, a robust control framework that dynamically moderates the self-distillation objective. AD-OPSD selectively anchors high-suppression-risk sandboxed tokens to a reference prior derived from the frozen base model via an asymmetrical pointwise divergence gate, preserving native thinking capacity while retaining OPSD's error-correcting power. Extensive experiments across competitive mathematical benchmarks show that AD-OPSD improves over standard OPSD by up to \textbf{+4.1\%} absolute average accuracy across diverse model scales and datasets. Further analysis demonstrates that AD-OPSD mitigates thinking collapse and generalizes robustly to different post-training paradigms.

Comment: Diagnoses a token-level optimization trap in on-policy self-distillation and introduces adaptive gradient control.

Topic Match: Its strongest fit is mechanistic training-dynamics analysis of collapse and a targeted stabilization method.

Relevance: 7 Novelty: 7


17. Interference and Retention in Continual Learning

ArXiv ID: 2607.09202

Primary Topic: Architecture and Training Dynamics

Authors: Julius Störk

Abstract: Continual learning commonly relies on post-hoc mechanisms such as replay, elastic regularization, or distillation. This work argues that forgetting should instead be modeled directly as interference between tasks. In the frozen-feature regime, forgetting from learning a new task is exactly the interference energy induced on the old task. In deep networks, the same quantity is recovered through path-averaged curvature with minimal additional forward passes. When task supports are disjoint, forgetting can be eliminated structurally and when task supports overlap in conflicting directions, a non-zero distortion floor is unavoidable. The same geometry optimally merges models through task-aware orthogonalization. From this analysis we derive Interference-Gated Functional Allocation (IGFA), a replay-free, Fisher-free method that shares directions when tasks align and protects them when they conflict. Across benchmarks, IGFA achieves lossless retention when tasks are structurally separable and moves unavoidable cost from irreversible forgetting into deferred but recoverable plasticity when they are not. It matches the strongest replay-free structural baselines on dissimilar-task streams and improves on unconditional projection when similarity makes transfer worth preserving.

Comment: Models forgetting as task interference and derives a replay-free gating and allocation rule.

Topic Match: Its core is a mechanistic account of optimization interference and retention dynamics.

Relevance: 6 Novelty: 8


18. Source-Lifted Flow Matching for Intervenable Multimodal Imitation

ArXiv ID: 2607.10206

Primary Topic: Architecture and Training Dynamics

Authors: He Zhang, Ying Sun, Pengteng Li, Ziyang Chen, Yiren Zhao, Ziyang Rao, Weiyu Guo, Yandong Guo, Hui Xiong

Abstract: Flow-matching policies are promising for imitation learning because they model complex multimodal action distributions. However, their stochasticity is largely passive: repeated sampling may yield diverse behaviors, but users cannot directly choose among valid continuations from the same state. We propose Source-Lifted Flow Matching (SL-FM), a source-intervenable flow-matching policy that exposes such a handle while keeping the velocity field shared and latent-free. The handle selects only the source endpoint of the conditional flow, not a mode-specific field, preserving the standard formulation while avoiding decomposition into separate mode-conditioned dynamics. The core mechanism is \textbf{Orthogonal Source Lifting}, designed to prevent path-crossing ambiguity. Instead of partitioning target actions by mode, SL-FM lifts handle-specific sources into auxiliary orthogonal coordinates and keeps targets in the original action subspace. This preserves the demonstrated action distribution while allowing one shared field to carry different branches without merging at crossings. To keep handles usable across states, we learn a state-dependent source mixture end to end and use a responsibility floor, giving each handle weak supervision and mitigating dead modes. Experiments on crossing-flow diagnostics and robot-control benchmarks show that SL-FM converts passive source randomness into an actionable intervention variable. It removes crossing-induced composite trajectories, changes future routes in 91.1\% of matched-prefix interventions, and achieves strong free-deployment performance, with improvements in several benchmark settings. Overall, source geometry provides actionable multimodal control without conditioning the velocity field on the selected mode.

Comment: Uses orthogonal source lifting to make branches of a shared flow field independently selectable.

Topic Match: Its core is a new source-intervenable flow-matching computation.

Relevance: 6 Novelty: 8


19. Exploring More to Solve More: Boosting Diversity in Text Diffusion Models via Entropy-Based Guidance

ArXiv ID: 2608.00024

Primary Topic: Architecture and Training Dynamics

Authors: Jingwei Zhang, Haoyu Lei, Zijin Feng, Jiacheng Sun, Farzan Farnia

Abstract: Although diffusion models have revolutionized continuous domains like image synthesis through high quality generations and controllable guidance mechanisms, bringing this controllability to the discrete, sequential nature of text remains an open challenge. Meanwhile, current sampling strategies and guidance methods adjust token likelihoods without capturing the broader semantic landscape, leading to a suboptimal balance between fidelity and diversity. In this work, we introduce a novel training-free Semantic-Aware Kernel Entropy (SAKE) guidance method. Our method computes the order-2 Rényi entropy over a kernel Gram matrix that captures both cross-token semantic interactions and relative token positions. By linearizing this objective in the embedding space, we derive a tractable guidance signal that dynamically adjusts the sampling distribution, flattening it to encourage exploration during redundancy and sharpening it for fidelity when diverse. Empirical experiments demonstrate that our approach achieves a superior Pareto frontier between fidelity and diversity, and improves multi-sample performance on reasoning-intensive tasks, such as code and mathematics generation, compared to temperature scaling and discrete guidance baselines.

Comment: Guides text-diffusion sampling with a semantic kernel-entropy objective.

Topic Match: It introduces a new computational mechanism for controlling discrete diffusion generation.

Relevance: 6 Novelty: 7


20. Energy-guided Recursive Model

ArXiv ID: 2607.10128

Primary Topic: Architecture and Training Dynamics

Authors: Yifei Zhao, Ying Tang

Abstract: Recursive reasoning models address structured problems by repeatedly updating latent states of small neural networks. However, their test-time scaling lacks a principled inference mechanism: increasing depth or stochastic breadth generates more trajectories without a clear criterion for selection, and existing methods predominantly rely on additional q-heads or heuristic voting. Here, we develop the Energy-guided Recursive Model (ERM), which introduces an intrinsic selection principle based on explicit Hopfield energies. ERM leverages Hopfield-type memories of valid local or global structures to define the selector over candidate trajectories. The resulting energy seamlessly integrates with energy-based techniques such as parallel tempering to enhance sampling efficiency and ranking. With $D=64$ recurrent steps and $K=128$ candidates, ERM reaches optimal solutions on Sudoku ($98.97\%$), Pencil Puzzle Bench (PPBench, $88.04\%$) and Maze ($99.30\%$), improving upon recent Probabilistic Tiny Recursive Model and Equilibrium Reasoners. These results suggest that incorporating explicit energy functions into recursive reasoning offers a principled path toward more effective inference.

Comment: Selects recursive reasoning trajectories using explicit Hopfield energies.

Topic Match: It contributes a new selection mechanism for recurrent dynamic computation.

Relevance: 6 Novelty: 7


21. Fully Trainable Deep Differentiable Logic Gate Networks and Lookup Table Networks

ArXiv ID: 2607.09399

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Wout Mommen, Lars Keuninckx, Matthias Hartmann, Werner Van Leekwijck, Piet Wambacq

Abstract: We introduce a novel method for both partial and full optimization of the connections in deep differentiable logic gate networks (LGNs) and lookup table networks (LUTNs). Our training method utilizes a probability distribution over a set of connections per gate/lookup table (LUT) input pin, selecting the connection with highest merit, all whilst the optimal gate types or LUT-entries are learned in parallel. We show that the connection-optimized LGNs outperform standard fixed-connection LGNs on the Yin-Yang, MNIST Handwritten Digits and Fashion-MNIST benchmarks, while requiring only a fraction of the number of logic gates. We achieve 98.92% on the MNIST dataset with two layers of 8000 gates. With only one layer of 8000 gates, we obtain 98.45%, showing that our method requires almost 50 times fewer gates compared to fixed-connection LGNs. Training stability up to ten layers has been ensured by employing a high learning rate, straight-through estimators and trimming constant-output gate types. Additionally, we present a LUT neuron description that enables stable training with backpropagation, tested up to 6-layer deep networks. The model requires four times fewer trainable parameters and still achieves a higher accuracy compared to the fixed-connection LGN training algorithm. Our connection-training algorithm also works well for the LUTNs, achieving an accuracy of 98.88% for two layers of 2000 6-input LUTs.

Comment: Jointly learns gate or LUT connectivity and functions, reducing required logic-gate counts by nearly 50x.

Topic Match: Trainable discrete connectivity is the primary architectural contribution, with network-size reduction providing an efficiency match.

Relevance: 6 Novelty: 7


22. On Locality and Length Generalization in Visual Reasoning

ArXiv ID: 2607.09061

Primary Topic: Architecture and Training Dynamics

Authors: Pulkit Madan, Sanjay Haresh, Reza Ebrahimi, Sunny Panchal, Apratim Bhattacharyya, Roland Memisevic

Abstract: A striking feature of the human visual system is that it ingests visual information through a series of local foveated glimpses, rather than a single global computation. This makes human vision distinctly different from most popular computer vision models in use today, which input images globally and in a single shot. A natural question therefore is whether local, sequential vision models may provide any fundamental computational benefits in addition to being biologically more plausible than global models. In this work, we investigate this question from the perspective of visual state tracking and length generalization. Inspired by recent studies of length generalization in language models, we study the behavior of vision models trained on simple vision tasks that require the aggregation of local information across an image. Our experiments reveal that, similar to language models, vision models can learn to exploit global shortcuts and thereby fail to generalize over task length or complexity. We also show that recurrent vision policies based on strictly local perception can mitigate these failures, thereby allowing models to generalize on these tasks. Our results show that local attention may be an essential overlooked requirement for robust compositional generalization.

Comment: Shows that strictly local recurrent visual computation avoids global shortcuts and improves length generalization.

Topic Match: The central result is a mechanistic comparison of local recurrent and global visual architectures.

Relevance: 6 Novelty: 6


Efficiency, Compression, and Large-Scale Training (15)

1. CARE-LoRA: Compressed Activation REconstruction for Memory-Efficient LoRA

ArXiv ID: 2607.11940

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Gengyu Zhang, Haiyin Ran, Zhengbao He, Yuhang Liu, Hanling Tian, Zhehao Huang, Xiaolin Huang

Abstract: As the scale of large pre-trained models continues to grow, fine-tuning them under limited memory budgets has become increasingly challenging. Low-Rank Adaptation (LoRA), currently one of the most widely adopted parameter-efficient fine-tuning (PEFT) methods, mitigates this challenge by optimizing only low-rank adaptation matrices, thereby greatly reducing the number of trainable parameters. With the parameter overhead substantially reduced, the activations retained for backpropagation have emerged as the primary remaining memory bottleneck during LoRA fine-tuning. To address this, we propose CARE-LoRA, a data-aware Compressed Activation REconstruction framework. By exploiting the inherent projection structure of LoRA, CARE-LoRA replaces the full input activation with the low-rank compressed activation naturally produced by the LoRA branch. It further computes a lightweight reconstruction matrix during the forward pass with negligible additional computation cost, which is used during backpropagation to reconstruct the gradient signal, thereby keeping LoRA matrices fully trainable. Extensive experiments across diverse models and downstream tasks demonstrate that, while substantially reducing the overall memory footprint, CARE-LoRA achieves competitive or even superior performance compared with standard LoRA and representative LoRA variants. Our code is publicly available at https://github.com/fishandyu/CARE-LoRA .

Comment: Reconstructs LoRA gradient signals from compressed branch activations to reduce fine-tuning memory.

Topic Match: It directly targets the activation-memory bottleneck in parameter-efficient training.

Relevance: 9 Novelty: 7


2. Super-Tuning: From Activation-Aware Pruning to Sparse Fine-Tuning

ArXiv ID: 2607.09287

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Ivan Ilin, Philip Zmushko, Peter Richtárik

Abstract: Large language models (LLMs) remain expensive to fine-tune because full-parameter updates require substantial memory, compute, and per-task storage. We study whether saliency signals originally developed for pruning can be reused to choose where a model should adapt. We propose Super, a sparse parameter-efficient fine-tuning (PEFT) method that fixes a small trainable support using a Wanda-style activation-weighted magnitude score [Sun et al., 2023] computed from a calibration pass. We then introduce Supra, a hybrid adapter that combines this sparse update with LoRA while preserving a matched trainable-parameter budget through a simple budget-splitting rule. In single-seed Math17K arithmetic experiments on Llama-3.2-1B and Meta-Llama-3-8B, the best Super/Supra variants achieve the highest average accuracy among the tested schedule-selected adapter configurations. We also include a PaFi-style magnitude-only support as a closest training-free sparse baseline and find that low-score supports under both magnitude and Wanda-style orderings can be effective. These results suggest that simple pruning-inspired orderings can provide useful fixed sparse supports for PEFT, especially when combined with low-rank adapters.

Comment: Uses activation-aware pruning scores to choose sparse trainable weights and combine them with LoRA.

Topic Match: It directly connects structured sparsity, pruning saliency, and parameter-efficient fine-tuning.

Relevance: 9 Novelty: 7


3. Codec-Gauge: Learning Compression-Friendly Gauges for Transformer KV Caches

ArXiv ID: 2607.20538

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Yitao Jiang, Yaoqing Yang, Luyang Zhao, Muhao Chen, Devin Balkcom

Abstract: Long-context Transformer inference increasingly relies on KV-cache compression or quantization. Prior rotation and transform-coding results suggest that the channel basis of each key/value vector affects how faithfully a fixed backend preserves model behavior. We introduce Codec-Gauge, a post-training cache-coordinate layer that learns small orthogonal channel transforms around existing compression and quantization backends. Its frequency-distribution objective combines a token-channel DCT spectral-centroid loss with a smooth rate proxy to concentrate KV energy in low-frequency codec-facing layouts. We evaluate actual compression and decompression using measured bytes and rolling compressed-history scoring. Across six models at $3$, $4$, and $6$ bits/value, learned gauges reduce zfp KL divergence by $44.0\%$ on average relative to raw coordinates and outperform random, Hadamard, DCT, and PCA/KLT controls. The same gauges improve quality preservation for block-uniform and KIVI-style quantization. Experiments on a 27B model and long-context task prompts reproduce the quality trend, while serial storage and timing measurements validate the implemented compressed-cache paths. These results establish cache-coordinate geometry as a practical post-training variable for improving compression fidelity without changing model weights, attention semantics, or backend coding rules.

Comment: Learns orthogonal KV-channel transforms that make existing cache codecs and quantizers preserve model behavior better.

Topic Match: It directly improves KV-cache compression without changing model weights or attention semantics.

Relevance: 9 Novelty: 7


4. MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference

ArXiv ID: 2607.10582

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Venkatesha Matam, Keon Kim

Abstract: Large language model (LLM) agents accumulate heterogeneous context, including system instructions, plans, user turns, retrieved documents, tool outputs, and intermediate reasoning, whose key-value (KV) cache can become a major memory bottleneck. Existing eviction policies generally apply the same attention- or recency-based rule to every token, ignoring semantic structure already available to the agent orchestrator. We introduce MemDecay, a training-free, region-aware KV-cache eviction policy. MemDecay assigns tokens region-specific base priorities and decay rates, refreshes retention scores when tokens receive attention, and evicts the lowest-scoring pages under a fixed cache budget while allowing critical regions to be pinned. We also provide a procedure for calibrating decay rates from measured attention lifetimes. We evaluate MemDecay at approximately 450 and 1,700 token contexts using Qwen2.5-1.5B and 3B. Across all settings, attention lifetimes differ by an order of magnitude across regions: system-token half-lives range from 148 to 189 decoding steps, compared with 14 to 16 for scratchpad tokens. Pinning preserves system-region facts at full-cache accuracy in every setting, while no baseline preserves more than 13 of 24. Region-aware retention remains effective as context grows, whereas recency-based retention collapses. Accumulated-attention retention performs better on unpinned content, however, and ablations identify attention-score normalization as the main limitation of the current formulation. These results establish semantic prompt structure as a robust signal for KV-cache management while clarifying how it should be combined with attention-based importance.

Comment: Evicts KV-cache pages using region-specific priorities, decay rates, and attention refreshes.

Topic Match: It introduces a new semantic mechanism for controlling KV-cache memory cost.

Relevance: 8 Novelty: 7


5. STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU

ArXiv ID: 2607.09385

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Victor J. B. Jung, Gagandeep Singh, Joseph Melber, Kristof Denolf, Francesco Conti, Luca Benini

Abstract: The growing adoption of large language model-based agents within operating system workflows has increased the importance of energy-efficient inference on laptop-class systems-on-chip (SoCs). While cloud offloading remains common, it introduces reliability and privacy concerns that are particularly problematic for agentic workloads. Recent laptop SoCs, therefore, incorporate neural processing engines (NPUs) optimized for energy efficiency; however, effectively mapping attention mechanisms onto NPUs remains challenging due to architectural diversity and explicit data-movement programming models. In this work, we present STEEL, the first open-source implementation of FlashAttention targeting XDNA-like NPUs. STEEL introduces a dataflow formulation of prefill attention, enabling efficient exploitation of spatial parallelism and on-chip memory. Furthermore, STEEL addresses the load imbalance induced by the causal mask by leveraging a sparsity-aware pipeline placement onto the NPU array, reducing synchronization overhead and improving utilization. We evaluate STEEL on the AMD Ryzen AI 9 HX 370 SoC and compare its performance against optimized CPU and GPU implementations. Experimental results show that STEEL reduces energy consumption by an average of 9.17x and 1.75x relative to CPU and GPU baselines, respectively. On XDNA 1, STEEL achieves an average 9.6x latency reduction over the prior state of the art, and delivers a 22.8x speedup on average compared to a layer-by-layer attention implementation on XDNA 2.

Comment: Sparsity-aware fused-attention dataflow cuts long-context NPU inference energy and latency.

Topic Match: The new attention kernel and dataflow materially reduce LLM inference cost.

Relevance: 8 Novelty: 7


6. KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling

ArXiv ID: 2607.09153

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Peng Kuang, Haibo Jin, Xiaoyu Han, Yanli Wang, Xiaopeng Yuan, Ye Yu, Kaidi Xu, Haohan Wang

Abstract: Process Reward Models (PRMs) have been proven to be highly effective in guiding test-time scaling (TTS) methods, which significantly boost the capabilities of LLM-based multi-agent systems. However, existing PRMs are text-based: they re-encode the entire trajectory text from scratch. In long multi-agent rollouts, the scoring cost, growing quadratically with respect to sequence length L, creates a severe computational bottleneck, severely limiting PRMs' application in long-context scenarios. To resolve this, we introduce KV-PRM, a highly efficient process reward model that eliminates the heavy text re-encoding by directly reading the KV cache produced naturally during the LLM's generation phase. By processing a single "verify token" against the pre-existing KV cache, KV-PRM reduces the scoring cost from O(L^2) to O(L). We formally prove that the KV cache contains strictly greater information capacity than text, and is more efficient for downstream reward modeling. Empirically, across the MATH, GSM8K, and AIME benchmarks, KV-PRM matches or strictly outperforms text-PRMs under various TTS methods such as Beam Search, MCTS, and Weighted Voting, with up to a 5,000x reduction in scoring FLOPs, a 37x reduction in latency, and a 34x reduction in per-sequence memory footprint compared to text-based PRMs.

Comment: Reuses generation KV caches so process-reward scoring falls from quadratic to linear sequence cost.

Topic Match: KV-cache transfer is the mechanism responsible for its large inference-cost reduction.

Relevance: 7 Novelty: 8


7. Adaptive Model Compression (AMC): Saliency-Driven Resource Allocation for Ultra-Low-Power Transformer Inference

ArXiv ID: 2607.10109

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Jiayin Hu, Kai Yuan, Vanessa Hu, Xuetao Yin, Jianhua Li, Sean Suchter

Abstract: Deploying large-scale transformer models on resource-constrained edge devices remains a challenge due to the high energy and memory overhead inherent in static inference, which processes simple and complex tokens with uniform intensity. To address this, we propose Adaptive Model Compression (AMC), a saliency-driven framework that dynamically allocates hardware resources based on token importance. By implementing a multi-tier architecture, our system identifies critical high-saliency information for full-precision processing while aggressively reducing the rank and bit-width of less significant data. Experimental results demonstrate that AMC achieves a 59.2% reduction in system energy and a 2.24x increase in throughput on 45nm CMOS hardware. This approach effectively extends the battery life of mobile devices by utilizing high-definition compute only where necessary, maintaining robust performance with a marginal 3.6% accuracy trade-off.

Comment: Dynamically assigns token-level rank and numerical precision from saliency.

Topic Match: The core contribution is adaptive compression that changes inference energy and throughput.

Relevance: 8 Novelty: 6


8. Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM Inference

ArXiv ID: 2607.09520

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Junfei Zhan, Haoxun Shen, Mingang Guo, Zixuan Huang, Tengjiao He

Abstract: Vision-Language Models (VLMs) are the perceptual backbone of embodied AI, but their energy footprint on edge hardware remains poorly understood. Existing efficiency efforts focus predominantly on reducing visual tokens, implicitly treating visual processing as the dominant energy cost. We overturn this implicit assumption through the first systematic energy profiling of on-device VLM inference, spanning five models across three architecture families, four input resolutions, and two hardware platforms (NVIDIA RTX 3070 and Jetson Orin NX). Our analysis yields three findings. First, average inference power is a model-intrinsic constant, invariant to input resolution, image complexity, and prompt type, with less than 5% variation across all conditions. This means that all energy variation across inputs must arise from variation in inference time, not from variation in power draw. Second, each output token costs 11 to 39x more wall-clock time than each input token due to the compute-bound and memory-bound asymmetry between prefill and decode, making output token count the dominant driver of both latency and energy. Third, image complexity, measured by the number of objects in an image, induces up to 4.1x energy differences at identical resolution. This variation arises not from increased visual processing cost, but from differences in output length. These findings expose a fundamental limitation of visual token pruning: even removing all visual tokens saves at most 10% of total energy for fixed-token models. Across models spanning 1 billion to 8 billion parameters, controlling output length saves up to 97% of total energy, with the energy dominance of decoding growing stronger at larger model scale. In short, the true energy bottleneck in edge VLM inference is not what the model sees, but how much it says. Code is available at https://github.com/Junfei-Z/seeing-is-free.

Comment: Shows through energy profiling that autoregressive decoding length dominates edge-VLM energy consumption.

Topic Match: It identifies the dominant systems-level cost and the limits of visual-token pruning.

Relevance: 7 Novelty: 7


9. IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation

ArXiv ID: 2607.09133

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Yiting Wang, Jingyi Zhang, Wenhu Zhang, Ke Chao, Yves Liang, Kun Cheng, Kang Zhao

Abstract: While large-scale text-to-image generative models have achieved unprecedented visual performance, their inherent reliance on multi-step iterative solvers incurs severe inference latency. Few-step distillation targeting the Classifier-Free Guidance (CFG) trajectory has emerged as the prevalent dual-dimensional compression paradigm. However, existing frameworks remain subjugated by a coarse-grained blind injection paradigm that perpetually enforces a globally static guidance strength while indiscriminately sampling the supervisor timestep. This state-agnostic design completely disregards the intrinsic nature of image generation as a dynamic evolutionary process characterized by progressive entropy reduction, which not only restricts the performance boundary of few-step compression but also precipitates severe CFG over-conditioning artifacts. To transcend these limitations, we re-examine the distillation procedure through the theoretical lens of Information Theory, formally modeling it as a dynamic mutual information game constrained by the Information Bottleneck (IB) principle. Specifically, we dismantle traditional blind assumptions via a dual-track adaptive framework. To determine the injection target, we propose an instance-aware selection mechanism that transmutes the intractable KL divergence constraint into a zero-overhead closed-form solution predicated on the local vector field norm. To regulate the injection strength, we introduce an entropy-aware schedule that dynamically decays alongside the SNR, applying maximal thrust for initial structural anchoring before smoothly reverting to the natural manifold to refine micro-details. Extensive empirical evaluations corroborate that our framework fundamentally eradicates over-conditioning artifacts, shattering the performance ceiling to achieve SOTA generative fidelity under extremely stringent 2-step configurations.

Comment: Uses information-bottleneck-guided adaptive CFG distillation to enable high-quality two-step diffusion sampling.

Topic Match: Few-step distillation directly compresses the inference computation of a large generative model.

Relevance: 7 Novelty: 7


10. Unlocking Parallelism in Autoregressive Language Models via Speculative Decoding with Progressive Tree Drafting

ArXiv ID: 2607.10661

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Zipeng Gao, Zhi Zheng, Qingrong Xia, Junda Lin, Ziwei Zhao, Tong Xu, Zhefeng Wang, Enhong Chen

Abstract: Speculative decoding has significantly accelerated Large Language Model (LLM) inference by alleviating memory-bound bottlenecks. However, traditional speculative decoding typically relies on auxiliary draft modules, incurring significant training and communication overhead. Although recent methods attempt to generate drafts within the target model itself, they often fail to fully exploit its latent parallel capacity due to a lack of structural coordination. In this paper, we propose \textbf{Progressive Tree Drafting (PTD)}, which employs a structured, guided parallel drafting strategy to harness the model's parallel potential. By coupling a progressive tree structure with a stepwise pruning mechanism, PTD actively guides the LLM to explore multiple semantic paths in a single forward pass, ensuring both draft diversity and coherence. Experiments demonstrate that PTD achieves up to $2\times$ decoding speedup across various benchmarks while remaining training-free and model-agnostic. Our code is available at: https://github.com/MINE-USTC/PTD.

Comment: Introduces training-free progressive tree drafting for parallel speculative decoding.

Topic Match: The method directly reduces autoregressive inference cost through a new structured drafting and pruning mechanism.

Relevance: 7 Novelty: 7


11. Reliability Scaling Laws for Quantized Large Language Models

ArXiv ID: 2607.10855

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Sirine Ayadi, Sándor Daróczi, Stephan Günnemann, Bertrand Charpentier

Abstract: Quantization is a powerful strategy to build capable and resource-efficient large language models (LLMs) by reducing the bitwidth of the parameters. While quantized LLMs achieve state-of-the-art performance on unperturbed inputs using standard predictive metrics, their performance on perturbed inputs, measured using reliability metrics, remains underexplored, despite its importance for reliable deployment. To address this gap, we first conduct a comprehensive reliability evaluation of quantized LLMs consisting of three key components: (1) Uncertainty: We assess the trustworthiness of LLMs quantized to 2, 3, 4, and 8 bits using six different quantization methods, employing established uncertainty metrics. (2) Calibration: We assess how well-calibrated the uncertainty estimates of quantized models are across model scales and bit precisions. (3) Robustness: We design character-level and word-level input perturbations to evaluate the reliability of quantized models under semantically-preserving variations in the inputs that arise in real-world applications. Second, we characterize how reliability scales with the total number of model bits. Our study reveals that while the performance scales monotonically with the total number of bits, the reliability scalings are nonlinear. A reliability peak occurs for 4-bit quantized models, indicating that quantizing moderately sized models offers the best reliability-efficiency trade-off. Additionally, our empirical findings reveal that quantization enhances the robustness of LLMs to natural input perturbations.

Comment: Derives empirical reliability scaling against total model bits and identifies a 4-bit reliability-efficiency optimum.

Topic Match: Quantized-model behavior and bitwidth trade-offs make compression the central fit.

Relevance: 7 Novelty: 6


12. IterCOMP: Reasoning-aware Adaptive Prompt Compression for Multi-hop Question Answering

ArXiv ID: 2608.13588

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: JungMin Yun, YoungBin Kim

Abstract: Multi-hop question answering requires complex reasoning across multiple evidence segments, which often overwhelms retrieval-augmented generation systems with lengthy and noisy contexts, thereby undermining both efficiency and accuracy. While existing prompt compression methods attempt to address this issue, they are typically designed for single-turn queries and fail to capture interdependent reasoning steps. We propose IterCOMP, a unified, training-free prompt compression framework that incorporates multi-hop reasoning within an iterative compression loop. IterCOMP decomposes documents into evidence segments, evaluates question answerability, and generates targeted follow-up questions to iteratively integrate essential evidence, producing a compact, reasoning-oriented prompt. Experiments on MusiQue, 2WikiMultiHopQA, and HotpotQA demonstrate that IterCOMP achieves substantial improvements in Exact Match and F1 scores while reducing the token budget, outperforming existing baselines and exhibiting robustness as reasoning complexity increases.

Comment: Compresses multi-hop contexts through iterative evidence selection and targeted follow-up questions.

Topic Match: The core mechanism lowers inference context cost through adaptive prompt compression.

Relevance: 6 Novelty: 6


13. Self-Guided Test-Time Training for Long-Context LLMs

ArXiv ID: 2607.09415

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Xinyu Zhu, Zhe Xu, Xiaohan Wei, Yunchen Pu, Fei Tian, Chonglin Sun, Kaushik Rangadurai, Hua Zhi, Frank Shyu, Sandeep Pandey, Luke Simon, Yu Meng, Xi Liu

Abstract: Long-context processing has become increasingly important for large language models (LLMs), but simply extending the context window does not guarantee effective utilization of long inputs. As input length grows, accuracy often degrades, indicating that models still struggle to identify and use the evidence most relevant to a question. A promising way to improve long-context utilization is test-time training (TTT), which treats the test context as a training example for instance-specific parameter adaptation. However, applying TTT to the entire long context is prohibitively expensive, while adapting on randomly sampled spans introduces severe noise. Because most spans in a long context are irrelevant to the specific question, training on them may even degrade the base model's performance. Our preliminary study shows that TTT is highly sensitive to training-span quality: on LongBench-v2, TTT on randomly sampled spans hurts performance, whereas TTT on oracle spans substantially improves it. Motivated by this, we propose a simple method, Self-Guided TTT (S-TTT): before adaptation, the model identifies the evidence spans it should learn from, and the standard language-modeling training objective is applied only to those selected spans. On two challenging long-context reasoning benchmarks, LongBench-v2 and LongBench-Pro, S-TTT improves accuracy for both Qwen3-4B-Thinking-2507 and Llama-3.1-8B-Instruct, achieving up to a 15% relative improvement.

Comment: Selects evidence spans before instance-specific adaptation, avoiding the cost and noise of full-context test-time training.

Topic Match: Selective adaptation materially reduces the computation required by long-context test-time training.

Relevance: 6 Novelty: 6


14. Dependency-Aware Chain-of-Thought Compression for Financial Reasoning

ArXiv ID: 2609.00413

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Wenjun Wu, Lei Fu, Kejian Tong, Tao Ning, Sichen Zhao

Abstract: Chain of thought prompting improves complex reasoning, but its long intermediate traces create substantial inference cost and hinder practical deployment in financial settings. We present a Hierarchical Semantic Distillation Network, HSDN, for compressing reasoning chains while preserving answer accuracy and logical coherence. The framework combines semantic segmentation, dependency graph construction, dual encoder importance scoring, constrained segment selection, and local boundary rewriting. A frozen Qwen3 4B model is used only for feature extraction and final answer generation, while the compression process remains structured and interpretable. On the AFAC2025 benchmark, HSDN achieves 91.0% accuracy with 68.4% compression, outperforming strong compression baselines in overall score and reasoning coherence. The results show that graph guided compression is effective for high stakes financial reasoning tasks.

Comment: Compresses reasoning traces through dependency-aware segment selection to reduce generated-token cost.

Topic Match: Inference efficiency is primary because the method materially shortens reasoning context while attempting to preserve correctness.

Relevance: 6 Novelty: 6


15. Structured Thoughts For Improved Reasoning And Context Pruning

ArXiv ID: 2607.10386

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Zain Sarwar, Supriyo Chakraborty, Berkcan Kapusuzoglu, Chia-Hsuan Lee, Anirban Das, Stephen Rawls, Kartik Balasubramaniam, Sambit Sahu

Abstract: Large language models (LLMs) excel at generating long chains of thought, but long reasoning traces are often verbose and memory-inefficient. In this work, we introduce Structured Thoughts, a framework that organizes reasoning into alternating and blocks: captures exploratory scratch work, while contains the distilled conclusion of that step. We construct a dataset of structured thoughts by segmenting reasoning traces into blocks and prompting an LLM to summarize each step into its corresponding . Fine-tuning pretrained foundation models on this reformatted data produces models that adopt the structured reasoning style, leading to performance gains of up to 8.08\% on reasoning benchmarks compared to standard SFT. The explicit structure also enables context pruning: after each / pair, the can be pruned, allowing the model to retain conclusions without keeping the full scratch work in the context. A proof-of-concept pruning implementation achieves an average of 85\% memory / context savings with an 8.67\% performance drop across mathematical tasks.

Comment: Structures reasoning into prunable scratch work and retained conclusions, reducing live context consumption.

Topic Match: The strongest foundational connection is explicit context-memory reduction through selective pruning.

Relevance: 6 Novelty: 6


Paper Selection Prompt

System Prompt

You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.

User Prompt

Instructions

Respond in JSONL. Output exactly one JSON object per paper, one per line:

{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}

Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only: daily_hot, new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return []. - daily_hot means the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. - new_frontier means the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.

Scoring Criteria

Relevance and Novelty are independent axes. Score both from 1 to 10.

Relevance Scoring

  • 9-10: directly centered on the target foundational topics; highest when the core contribution is clearly within them.
  • 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
  • 5-6: touches the target topics, but the main contribution is elsewhere.
  • 3-4: largely outside the target topics, often application-focused or domain-specific.
  • 1-2: unrelated.

Important: Broad frontier relevance, major launch status, or daily buzz is not enough for a high Relevance score here. Those cases belong in the hotspot digest unless the paper strongly matches the specialized paper topics.

Novelty Scoring

  • 9-10: new paradigm, theory, or major methodological breakthrough.
  • 7-8: substantial methodological advance or strong new insight.
  • 5-6: meaningful but incremental extension or refinement.
  • 3-4: minor, narrow, or mostly engineering or domain-specific improvement.
  • 1-2: little originality; mainly standard application of existing methods.

Topic Registry

Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - moe_training: MoE Training - Mixture-of-Experts training: routing, load balancing, expert granularity and shared experts, training stability, MoE parallelism and communication, and MoE kernels. - training_systems: Large-Scale Training Systems and Efficiency - Large-scale training systems and efficiency: distributed training algorithms, optimizers, and communication. - architecture_training: Architecture and Training Dynamics - Core architectural or computational mechanisms, dynamic computation, and training-stability dynamics. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, memory or cache efficiency, and large-scale training systems that materially change cost or behavior.

Papers

[PAPER LIST HERE]

Relevant Topics

This is a training-side feed. The centre of gravity is Mixture-of-Experts training and the large-scale training systems it sits on; architecture and efficiency work is kept when it bears on how a large model is trained or what it costs to train.

Keep a paper when its CORE CONTRIBUTION falls in one of the four topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.

  1. MoE Training (primary) - Keep: expert routing and gating mechanisms; load balancing, with or without auxiliary losses; expert granularity, fine-grained and shared experts; MoE training stability (loss spikes, router collapse, z-loss and related fixes); expert and hybrid parallelism; all-to-all dispatch/combine and the communication schedules MoE needs; MoE kernels and grouped-GEMM style implementations; expert capacity, token dropping, and upcycling dense checkpoints into MoE. - Filter: papers that merely train on top of a MoE model without a new routing, balancing, stability, parallelism, or kernel contribution; "mixture of experts" in the classical ensemble or recommender sense.

  2. Large-Scale Training Systems - Keep: distributed training algorithms (data, tensor, pipeline, sequence, and expert parallelism, FSDP and ZeRO-style sharding); optimizers and preconditioners for large-scale pretraining; communication algorithms, overlap, and collective design; throughput and MFU work that materially changes what a training run costs or whether it converges; scaling-law work that informs how such runs are configured. - Filter: routine infrastructure tuning with no new algorithmic or systems idea.

  3. Architecture and Training Dynamics - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, dynamic or modular computation); optimisation and training-dynamics analysis that explains why large models train the way they do. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight.

  4. Efficiency and Compression - Keep: quantization, sparsity, pruning, low-rank adaptation, KV-cache and memory-efficient designs, and other work that materially changes the cost of running or training a large model, especially when the mechanism is new rather than a tuned variant. - Filter: straightforward application of standard efficiency methods, and deployment or serving engineering with no new idea.

Leave these to the hotspot digest unless the core contribution clearly falls in one of the four topics above: - major model or product releases - agent, tooling, and RAG launches - benchmark, leaderboard, or evaluation-only papers - representation-learning theory, interpretability, and feature-formation analysis - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning) - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains