Previous Day 2026-06-29
Monthly Overview 2026-06
Next Day 2026-07-01

This is a remedial run for missed papers from 06/29/2026 to 06/29/2026.

Results generated on 09/12/2026.

Personalized Daily ArXiv Papers 2026-06-30

Model Metric Usage Papers
Prompt Completion Total Total arXiv Scanned Relevant
local agent CLI Tokens not reported not reported not reported 418 418 21
Cost not reported not reported not reported

Token counts are not reported for this run. 6 of 6 model calls succeeded, 1,526s of model wall clock.

Topic Coverage:

TopicPapers
MoE Training1
Large-Scale Training Systems and Efficiency2
Architecture and Training Dynamics13
Efficiency, Compression, and Large-Scale Training5

Table of contents by topic:

MoE Training (1)

  1. Residual-Guided Expert Specialization for Incomplete Multimodal Learning Authors: Seunghun Baek, Jihwan Park, Jaeyoon Sim, Minjae Jeong, Hoseok Lee, Won Hwa Kim

Large-Scale Training Systems and Efficiency (2)

  1. CuraWeb: Joint Optimization of Quality, Redundancy, and Diversity for Web-Scale Pretraining Data Authors: Peiguang Li, Yongwei Zhou, Juncheng Diao, Yuchun Fan, Jian Yang, Jianxiao Yang, Zhongda Su, Shuguang Jiao, Xiao Wei, Zhiye Zou, Gan Dong, Zhizhao Zeng, Rongxiang Weng, Jingang Wang, Xunliang Cai

  2. Extrapolating from Regularised Solutions for Solving Ill-Conditioned Linear Systems in Machine Learning Authors: Disha Hegde, Jon Cockayne, Chris. J. Oates

Architecture and Training Dynamics (13)

  1. MATCH: Modulating Attention via In-Context Retrieval for Long-Context Transformers Authors: Linrui Ma, Chun Hei Lo, Xinyu Wang, Peng Lu, Xihao Yuan, Hanting Chen, Kai Han, Xinghao Chen, Chengjun Zhan, Hanlin Xu, Yichun Yin, Lifeng Shang, Feng Wen, Boxing Chen, Yufei Cui

  2. Stabilizing Extrapolation in Looped Transformers via Learned Stochastic Stopping Authors: Hsun-Yu Kuo, El Mahdi Chayti, Patrik Reizinger, Wieland Brendel, Martin Jaggi

  3. A Stochastic--Geometric Theory of Scaling Laws in Grokking Authors: Róisín Luo, Christian Gagné, Jonas Ngnawé, Ihsan Ullah, Karyn Morrissey

  4. SGD at the Edge of Stability: Stochastic Stabilization with Large Learning Rates Authors: Konstantinos Emmanouilidis, Lachlan MacDonald, Salma Tarmoun, Rene Vidal

  5. Why Do Few-Step Text Latents Fail When Image Latents Work? Non-Commitment at Sharp Categorical Readouts Authors: Zhongyao Wang

  6. Predictable GRPO: A Closed-Form Model of Training Dynamics Authors: Rajat Ghosh, Datta Nimmaturi, Aryan Singhal, Vaishnavi Bhargava, Henry Wong, Johnu George, Debojyoti Dutta

  7. Beyond Block Boundaries: Multi-Block Editing for Diffusion Large Language Models Authors: Xingyu Mou, Zijin Huang, Tianze Zhang, Yuxin Ma, Lanning Wei, Zengfeng Huang, Da Zheng, Lun Du

  8. Convergence of Continual Learning in Homogeneous Deep Networks Authors: Matan Schliserman, Gon Buzaglo, Itay Evron, Daniel Soudry

  9. Simplifying Flow Matching Transformations with Low-Rank Mixture Models Authors: Liam A. Kruse, Houjun Liu, Alexandros E. Tzikas, Mansur M. Arief, Mykel J. Kochenderfer

  10. Quality-Aware Modulation for Diffusion Transformers Authors: Luke Budny, Yuhong Guo, Kevin Cheung

  11. Relevance Is Not Permission: Localizing and Controlling Metric-Facing Attention Contributions Authors: Minwoo Yu, Young-guk Ha

  12. Characterizing Optimizer-Dependent Training Dynamics Through Hessian Eigenvector Displacement and Localization Authors: Marcelina Marjankowska, Valerio Modugno, Paolo Barucca

  13. Curvature-Weighted Gradient Diversity: A Noise Measure for Geometry-Adaptive SGD Schedules Authors: Muhammad Hamza, Ayush Goel

Efficiency, Compression, and Large-Scale Training (5)

  1. Hierarchical Global Attention (HGA) Authors: Woernle Frank, Fedosov Vladimir, Grinenko Artemiy

  2. Beyond Single-Dimensional Compression: The Compound Sparsity Frontier of Large Language Models Authors: Chao Han, Haozhe Hu, Xiaoyu Shen

  3. When Is a Draft Accepted? A Theory of Acceptance in Speculative Decoding Authors: Aaryam Sharma

  4. PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF Authors: Doo Hwan Hwang, Kee-Eung Kim

  5. Diffusion Fine-tuning with Rewarded Moment Matching Distillation Authors: Alexis Jacq, Guillaume Couairon, Valentin De Bortoli, Quentin Berthet, Arnaud Doucet, Romuald Elie


MoE Training (1)

1. Residual-Guided Expert Specialization for Incomplete Multimodal Learning

ArXiv ID: 2606.30355

Primary Topic: MoE Training

Authors: Seunghun Baek, Jihwan Park, Jaeyoon Sim, Minjae Jeong, Hoseok Lee, Won Hwa Kim

Abstract: As real-world prediction systems often face missing modalities at inference, incomplete multimodal learning (IML) remains a practical challenge. While prior methods aim to learn representations robust to missing inputs, representations from incomplete modalities inevitably deviate from their full-modality counterparts due to missing evidence. To explicitly leverage these deviations, we propose MARS (Missingness-Aware Residual-guided Specialization), a mixture-of-experts framework that guides expert specialization based on how representations are reshaped by missingness. By contrasting task representations derived from incomplete inputs with their complete counterparts during training, we derive a privileged residual signal that captures this representational gap. The residual signal guides a residual router to assign samples to experts specialized for the corresponding deviation patterns. In parallel, a feature router learns to imitate this routing behavior using only incomplete inputs, enabling deployment without access to full modalities. To mitigate this train-test router gap, we develop a discrepancy-aware noise regularization that adaptively perturbs the residual router's decisions when the feature router deviates, enhancing expert robustness under imperfect imitation. Experiments on multimodal classification (CASIA-SURF, CREMA-D, UPMC Food-101) and segmentation (MCubeS) under missing scenarios show that MARS consistently surpasses baselines while remaining efficient and extensible to diverse backbones and tasks.

Comment: Privileged residuals supervise expert routing, with discrepancy-aware noise improving robustness to the deployed router.

Topic Match: New expert-assignment supervision and router-training regularization constitute a substantive MoE contribution, although validation is confined to incomplete multimodal tasks.

Relevance: 7 Novelty: 7


Large-Scale Training Systems and Efficiency (2)

1. CuraWeb: Joint Optimization of Quality, Redundancy, and Diversity for Web-Scale Pretraining Data

ArXiv ID: 2607.22662

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Peiguang Li, Yongwei Zhou, Juncheng Diao, Yuchun Fan, Jian Yang, Jianxiao Yang, Zhongda Su, Shuguang Jiao, Xiao Wei, Zhiye Zou, Gan Dong, Zhizhao Zeng, Rongxiang Weng, Jingang Wang, Xunliang Cai

Abstract: Open-web corpora curated via highly selective filters, such as FineWeb-Edu and DCLM, constitute the core of LLM pretraining data and have significantly advanced LLM performance. However, these pipelines typically rely on singular optimization objectives, which inevitably narrows distributional diversity and marginalizes long-tail knowledge, thereby restricting data coverage and underutilizing the vast potential of the open web. To address this limitation, we propose a novel curation paradigm that shifts from linear pruning to the joint optimization of quality, redundancy, and diversity. This framework synergizes dual-track cleaning (rule-based and model-driven) with hybrid deduplication (n-gram and semantic), while employing a multi-objective sampler to balance informational quality with distributional breadth. Applying this framework to Common Crawl, we construct CuraWeb, a 2T-token English corpus. Unlike existing resources, CuraWeb establishes an industrial-grade standard for data curation by recovering a more holistic data distribution with enhanced diversity and minimal redundancy, achieving broader coverage of long-tail knowledge across diverse domains. Experimental evaluations at the 3B scale demonstrate that CuraWeb significantly outperforms state-of-the-art baselines, yielding an average performance gain of 1.8\% across a wide range of benchmarks, particularly in knowledge-intensive and reasoning tasks.

Comment: A multi-objective pretraining-data sampler balances quality, redundancy and distributional coverage.

Topic Match: Pretraining inputs are closely adjacent to training_systems, but the core contribution is corpus curation rather than optimization, parallelism or training execution.

Relevance: 6 Novelty: 6


2. Extrapolating from Regularised Solutions for Solving Ill-Conditioned Linear Systems in Machine Learning

ArXiv ID: 2606.30328

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Disha Hegde, Jon Cockayne, Chris. J. Oates

Abstract: Rapid prototyping of algorithms is a critical step in modern machine learning. Most algorithms exploit linear algebra, creating a need for lightweight numerical routines which -- while potentially sub-optimal for the task at hand -- can be rapidly implemented. For the numerical solution of ill-conditioned linear systems of equations, the standard solution for prototyping is Tikhonov-regularised inversion using a nugget. However, selection of the size of nugget is often difficult, and the use of data-adaptive procedures precludes automatic differentiation, introducing instabilities into end-to-end training. Further, while data-adaptive procedures perform multiple linear solves to select the size of nugget, only the result of one such solve is returned, which we argue is wasteful. This paper aims to circumvent the above difficulties, presenting autonugget; a Python package for automatic and stable numerical solution of linear systems suitable for rapid prototyping, and fully compatible with automatic differentiation using JAX. autonugget combines multiple linear solves using Richardson extrapolation to determine the solution of the ill-conditioned system, improving in accuracy over approximations based on a single nugget.

Comment: Richardson extrapolation combines regularized linear solves into a more accurate autodifferentiable solver.

Topic Match: Stable differentiable linear algebra can support training systems, but the contribution is a prototyping solver without demonstrated large-scale pretraining benefits.

Relevance: 6 Novelty: 6


Architecture and Training Dynamics (13)

1. MATCH: Modulating Attention via In-Context Retrieval for Long-Context Transformers

ArXiv ID: 2606.29844

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Linrui Ma, Chun Hei Lo, Xinyu Wang, Peng Lu, Xihao Yuan, Hanting Chen, Kai Han, Xinghao Chen, Chengjun Zhan, Hanlin Xu, Yichun Yin, Lifeng Shang, Feng Wen, Boxing Chen, Yufei Cui

Abstract: The quadratic computational cost of traditional attention mechanisms poses a major bottleneck to the scalability and practical deployment of large language models (LLMs), particularly in long-context scenarios. To improve efficiency, existing approaches often enforce rigid structural constraints such as local attention windows. However, these strategies typically lead to substantial performance degradation on tasks requiring precise long-range recall. In this work, we propose MATCH, a scalable and efficient framework that augments sparsified attention mechanisms with dynamically integrated in-context information through an efficient retrieval system. Empirical results show that MATCH significantly improves the performance of sparse-attention models on both synthetic and real-world natural-language tasks. These findings highlight the versatility of MATCH as a general approach for enhancing in-context retrieval capabilities while maintaining the efficiency benefits of sparse attention architectures.

Comment: Augments sparse attention with dynamic in-context retrieval to preserve long-range recall.

Topic Match: The contribution modifies attention computation while retaining the efficiency of sparse attention.

Relevance: 9 Novelty: 7


2. Stabilizing Extrapolation in Looped Transformers via Learned Stochastic Stopping

ArXiv ID: 2606.29983

Primary Topic: Architecture and Training Dynamics

Authors: Hsun-Yu Kuo, El Mahdi Chayti, Patrik Reizinger, Wieland Brendel, Martin Jaggi

Abstract: Looped Transformers, which repeatedly apply a shared transformer block, are an architecturally natural fit for variable-length algorithmic tasks. Although they can exhibit strong length generalization beyond the length of training sequences, this behavior is brittle, yielding high out-of-distribution (OOD) variance, even across well-performing in-distribution solutions. We trace this variance to the spurious correlation in simple algorithmic tasks between sequence length and number of loops. Introducing stochasticity into the number of loops during training sharply reduces OOD variance and stabilizes predictions across inference-time loop counts. To improve upon heuristic randomization schemes, we further analyze RL-Halting as a learned stochastic schedule and find that it generally improves the accuracy-stability trade-off. Across binary addition, Dyck-1, Unique Set, and Copy, learned stochastic stopping often improves this trade-off but can also stabilize a suboptimal computation. Our work suggests that "when to stop" should be treated as a training-time design choice, not merely an inference-time computation-allocation rule.

Comment: Randomizes and learns stopping during training to stabilize looped-transformer extrapolation.

Topic Match: Loop-count scheduling directly changes training of shared-block recurrent computation and explains its extrapolation instability.

Relevance: 9 Novelty: 7


3. A Stochastic--Geometric Theory of Scaling Laws in Grokking

ArXiv ID: 2606.30388

Primary Topic: Architecture and Training Dynamics

Authors: Róisín Luo, Christian Gagné, Jonas Ngnawé, Ihsan Ullah, Karyn Morrissey

Abstract: Delayed generalization (\ie~grokking) refers to the phenomenon in which a neural network fits its training data early in training but only begins to generalize after a prolonged delay, often through an abrupt transition. Despite extensive empirical study, its underlying mechanism remains poorly understood. In this work, we first theoretically characterize a shell--core topological configuration of the reachable solution space induced by Adam's optimization dynamics with weight-shrinkage regularization, supported by empirical evidence. This optimization-induced topological configuration gives rise to grokking. In model's parameter space, random initialization solutions concentrate on a thin outer spherical shell, enclosing another spherical shell of memorization solutions, which in turn contains a core corresponding to the generalization solutions. Leveraging stopping-time theory, we then analyze the geometry of this topological configuration and the solution transition time at which optimization trajectories escape the memorization manifold and first reach the boundary of the generalization manifold. Our theoretical analysis derives grokking scaling laws for the learning rate, batch size, and $\ell_2$ regularization coefficient, which are further validated through experiments and shown to recover results from prior literature.

Comment: Derives grokking transition-time scaling in learning rate, batch size and regularization from Adam-induced geometry.

Topic Match: The central contribution explains optimization-driven delayed generalization and quantifies how training hyperparameters affect the transition.

Relevance: 8 Novelty: 8


4. SGD at the Edge of Stability: Stochastic Stabilization with Large Learning Rates

ArXiv ID: 2606.30930

Primary Topic: Architecture and Training Dynamics

Authors: Konstantinos Emmanouilidis, Lachlan MacDonald, Salma Tarmoun, Rene Vidal

Abstract: Modern deep learning has been shown to operate at the edge of stability, routinely using learning rates far larger than those justified by classical optimization theory. Most prior analyses of the edge of stability phenomenon focus on deterministic gradient descent, leaving the stochastic setting largely unexplored. In this work, we provide sharp convergence guarantees for Stochastic Gradient Descent (SGD) applied to the multiclass cross-entropy loss, for both linear classifiers and two-layer neural networks. We show that the stochasticity of SGD may cause the dynamics to alternate between an edge-of-stability regime that is dominated by curvature-driven oscillations, and a stable regime in which the expected loss decreases at a controlled rate. Despite that, we prove that SGD self-stabilizes the dynamics, ensuring that the iterates return to stability in a fixed number of iterations and allowing convergence in the best-iterate sense even with large learning rates. Experiments validate our theoretical findings and illustrate the benefits of SGD in the large-stepsize regime.

Comment: Proves stochastic self-stabilization and best-iterate convergence beyond classical learning-rate stability limits.

Topic Match: Directly explains large-step optimization dynamics, with theoretical guarantees currently restricted to linear classifiers and two-layer networks.

Relevance: 8 Novelty: 8


5. Why Do Few-Step Text Latents Fail When Image Latents Work? Non-Commitment at Sharp Categorical Readouts

ArXiv ID: 2606.30705

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Zhongyao Wang

Abstract: Deterministic few-step generation succeeds on continuous image latents but collapses to incoherent text on continuous text latents, and we show the cause is geometric rather than a training or scaling deficiency: a smooth, regularity-limited deterministic map cannot resolve a discrete branch choice before a sharp categorical readout, so few-step failure is governed by decoder sharpness, not transport accuracy. In the overlapping regime of real text autoencoders, we prove (Theorem 3) that the posterior-mean terminal step flips tokens at the rate of the latent mass in an $O(s(t))$ tube around decision boundaries. Two diagnostics, DABI (readout sharpness) and CCI (categorical commitment), measured on published checkpoints show that four independently built continuous-text decoders amplify a boundary-aligned perturbation far beyond a norm-matched isotropic one (DABI from $5\times10^{2}$ to $>10^{5}$), while image decoders have DABI $\approx 1$. Two mechanisms escape the continuous bound: categorical commitment (autoregressive decoders succeed despite sharper readouts) and stochastic re-injection (deterministic ODE at $K=4$ gives PPL 294 versus SDE 50 on the same model). In the idealized separated regime we prove matching sharp transport laws, including a dimension phase diagram: the deterministic stiffness needed to separate $M$ modes grows as $Θ(\sqrt{\log M})$ once the latent dimension is $Ω(\log M)$ (and as $M^{1/n}$ in fixed dimension), with a depth-$B$ hierarchy giving a $\sqrt{B}$-smaller per-step peak (Theorems 5-7); a coarea identity links these to the overlapping tube (Theorem 17). The result is an accuracy-depth-stiffness tradeoff: within the deterministic-continuous class the cost is irreducible, and both escapes step outside it.

Comment: Sharp categorical readouts impose an accuracy-depth trade-off on deterministic continuous text generation.

Topic Match: Derives a computational limit on few-step text-generation architectures and identifies categorical commitment and stochasticity as mechanisms that escape it.

Relevance: 8 Novelty: 8


6. Predictable GRPO: A Closed-Form Model of Training Dynamics

ArXiv ID: 2606.30789

Primary Topic: Architecture and Training Dynamics

Authors: Rajat Ghosh, Datta Nimmaturi, Aryan Singhal, Vaishnavi Bhargava, Henry Wong, Johnu George, Debojyoti Dutta

Abstract: We develop a first-principles reduced-order model of these dynamics. Under a single mean-field assumption that summarizes the policy by its expected reward, we reduce the GRPO update to a stochastically-forced damped oscillator whose mass, damping, and stiffness are fixed in closed form by the optimizer hyperparameters together with a single measured curvature scale -- momentum supplies the inertia, off-policy lag erodes the damping, and the group size enters, to leading order, as a noise temperature. The reduction has three consequences. First, it subsumes the empirical single-exponential saturation law as its overdamped limit, recasting the fitted plateau, timescale, and size exponent as the fixed point, inverse stiffness, and curvature-scaling exponent of the underlying potential, and adding, through the retained inertial term, the slow-start phase the single exponential cannot represent. Second, it yields predictions tied to independently measurable quantities rather than fitted ones: group-size invariance of the deterministic trajectory with a $1/G$ stationary fluctuation, a sharp stability threshold in the refresh interval, and an overdamped-to-oscillatory transition. Third, it furnishes diagnostics that separate failure modes a reward curve alone conflates -- reward hacking, advantage degeneracy, policy concentration, and dynamical instability. Across three models and two group sizes, the closed-form trajectory fits training reward to $R^2 \geq 0.91$ and the mean trajectory is group-size invariant to leading order -- on both the reward curve and out-of-distribution transfer to eight math benchmarks -- while the within-group reward spread retains a residual $G$-dependence that the leading-order temperature picture does not capture.

Comment: Derives optimizer- and off-policy-lag-dependent stability predictions for GRPO training.

Topic Match: Its central contribution is mechanistic training-stability analysis through a reduced dynamical model, making it a focused post-training exception.

Relevance: 8 Novelty: 7


7. Beyond Block Boundaries: Multi-Block Editing for Diffusion Large Language Models

ArXiv ID: 2607.22663

Primary Topic: Architecture and Training Dynamics

Authors: Xingyu Mou, Zijin Huang, Tianze Zhang, Yuxin Ma, Lanning Wei, Zengfeng Huang, Da Zheng, Lun Du

Abstract: Block diffusion is the dominant approach for scaling discrete diffusion language models (dLLMs), as fixed-size blocks preserve parallel decoding while keeping quadratic attention costs tractable. Yet blockwise generation creates a structural weakness: tokens near a block boundary lack future cross-block context, and errors in finalized blocks become irreversible context for later generation. We call this the block boundary problem. Measuring predictions with and without later-block context shows that boundary sensitivity rises sharply: on AIME 2025, mean self-containedness divergence (SCD) in the last quarter of a block is 61.3 times that in the first quarter. We propose Multi-Block Editing (MBE), which revises decoded tokens using cross-block context. Training-Free MBE reopens a full-attention window over selected blocks without parameter updates. To address the mismatch between block-diffusion training and MBE inference, Multi-Block Edit SFT introduces bidirectional attention masks and progressively enlarges the editing span. We also extend SGLang with a multi-shape CUDA Graph pool and fine-grained KV-cache control for efficient variable-length editing. Experiments on LLaDA2.1-Mini across 12 benchmarks show broad, consistent gains. Training-Free MBE improves or matches standard decoding on every benchmark. Full MBE raises the 12-benchmark average from 61.45 to 64.24, with gains of up to 13.33 points on AIME 2025, while retaining 87.3--96.7% of standard-decoding end-to-end throughput across four datasets.

Comment: Introduces cross-block token revision with bidirectional training masks for block-diffusion language models.

Topic Match: Cross-block computation and its training masks address a structural limitation of block-diffusion architectures.

Relevance: 8 Novelty: 7


8. Convergence of Continual Learning in Homogeneous Deep Networks

ArXiv ID: 2606.30559

Primary Topic: Architecture and Training Dynamics

Authors: Matan Schliserman, Gon Buzaglo, Itay Evron, Daniel Soudry

Abstract: We characterize weakly regularized continual classification in homogeneous models as sequential projections onto task margin sets. This result generalizes prior analyses restricted to either stationary (single-task) deep models or continual linear models. We show that global convergence generally fails, even for simple models linear in data but nonlinear in parameters. Nevertheless, by leveraging results from nonconvex projection theory, we identify regularity properties of homogeneous deep networks that guarantee local linear convergence under random and cyclic task sequences. Finally, we extend our analysis to continual regression, unifying the framework for homogeneous models.

Comment: Characterizes continual optimization as sequential projections and establishes conditions for local linear convergence.

Topic Match: Convergence dynamics of homogeneous deep networks are central, although applicability to large-model training remains indirect.

Relevance: 7 Novelty: 7


9. Simplifying Flow Matching Transformations with Low-Rank Mixture Models

ArXiv ID: 2606.29724

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Liam A. Kruse, Houjun Liu, Alexandros E. Tzikas, Mansur M. Arief, Mykel J. Kochenderfer

Abstract: Normalizing flows are powerful generative models that learn an invertible mapping between complex data distributions and simple latent distributions, typically a standard normal density. However, this choice of latent density can impose unnecessary complexity on the learned flow transformation due to the topological mismatch between the latent and data densities, leading to slower training and suboptimal performance. In this work, we propose using mixtures of probabilistic principal component analyzers (MPPCA) as the latent density for normalizing flows. We simplify the learned flow transformation by learning a latent distribution that more closely aligns with the data distribution in terms of KL divergence, thus enabling faster convergence and improved generative performance. Critically, MPPCA models can be fit quickly and cheaply using the expectation-maximization algorithm, making them a practical choice for initializing latent distributions even in high-dimensional generative tasks. We validate our method on both tabular and image datasets, demonstrating consistent gains in training efficiency and generation quality compared to baselines.

Comment: A cheaply fitted low-rank mixture base distribution simplifies learned flow transformations and accelerates convergence.

Topic Match: Changes the generative model's latent distribution to reduce optimization difficulty and training cost, with evidence currently confined to tabular and image flows.

Relevance: 7 Novelty: 6


10. Quality-Aware Modulation for Diffusion Transformers

ArXiv ID: 2606.30934

Primary Topic: Architecture and Training Dynamics

Authors: Luke Budny, Yuhong Guo, Kevin Cheung

Abstract: Modern text-to-image diffusion models, such as diffusion transformers (DiT), rely on timestep or prompt embeddings to modulate the strength of the denoising process in each timestep. While this modulation communicates the current noise level, it does not provide any quality-aware information, which can lead to generated images that are unaligned, visually inconsistent, and lacking in fidelity. In this paper, we propose the Quality Representation Module (QRM), a lightweight transformer module that learns a quality-aware representation based on existing model inputs, and produces a set of vectors $M_{qrm}$. These vectors adjust the adaptive LayerNorm modulation within the DiT transformer blocks, thereby injecting a quality-sensitive signal into the denoising parameters. The QRM introduces no significant changes to the sampling schedule or diffusion backbone. Experiments include ablations on QRM training losses and architectures, as well as empirical results demonstrating consistent image quality improvements over baseline DiT-based models.

Comment: Learns quality-dependent conditioning vectors for adaptive LayerNorm in diffusion transformers.

Topic Match: The contribution changes the conditioning mechanism inside transformer normalization, with a narrower focus on diffusion quality.

Relevance: 7 Novelty: 6


11. Relevance Is Not Permission: Localizing and Controlling Metric-Facing Attention Contributions

ArXiv ID: 2606.30139

Primary Topic: Architecture and Training Dynamics

Authors: Minwoo Yu, Young-guk Ha

Abstract: Attention identifies items relevant to a current query, but does not separately determine whether their value contributions support the prediction. We propose Warrant, a unified method for locating and controlling metric-facing attention contributions. Warrant identifies and exposes the item-wise contribution path that reaches the reported metric, then applies current-query-conditioned permission on that same path. Full Warrant improves the primary metric in 27 of 32 model-dataset comparisons across CTDG, MTPP, RAG, STPP, and TKG. Exact item-removal analysis in five representative settings finds near-zero correlation between attention and marginal prediction utility; even the highest-attention item reduces target utility in 43.5-54.4% of examples. Decomposition over the complete benchmark shows that the contributions of path exposure and learned permission vary by task. In a five-seed HotpotQA analysis, the opened path assigns more attention mass to distractors than to gold evidence, whereas learned permission preserves gold contributions, suppresses distractor contributions, and recovers evidence ranking in four of five seeds. These results show why attention-selected contributions must be localized and authorized again on the metric-facing path.

Comment: Adds query-conditioned gates to attention contributions on the path to model predictions.

Topic Match: The transferable mechanism controls attention value contributions, although validation emphasizes downstream architectures and tasks.

Relevance: 7 Novelty: 6


12. Characterizing Optimizer-Dependent Training Dynamics Through Hessian Eigenvector Displacement and Localization

ArXiv ID: 2606.30226

Primary Topic: Architecture and Training Dynamics

Authors: Marcelina Marjankowska, Valerio Modugno, Paolo Barucca

Abstract: Hessian spectral properties are a standard tool in analysing neural-network training, with eigenvalues linked to sharpness, generalization, and optimization dynamics. Eigenvalues quantify curvature magnitude, while eigenvectors identify which parameters generate that curvature. In this work, we study how the leading Hessian eigenvectors evolve during training and how they affect the learning trajectories. We track the training dynamics of multilayer perceptrons on a classification problem and measure eigenvector dynamics through two complementary statistics: (i) displacement over time, inspired by analyses of glassy systems, and (ii) localization via the inverse participation ratio. The metrics are compared against a random null model of the Hessian induced by the architecture. Our results reveal clear optimizer-dependent behaviour. SGD leads to progressively more stable leading curvature directions, while Adam exhibits substantially stronger reorganization of eigenvectors throughout training. We also observe a localization phenomenon under Adam, where a small subset of parameters contributes disproportionately to the leading curvature directions. These results suggest that Hessian eigenvector dynamics capture key differences in optimizer behaviour and the resulting training trajectories.

Comment: Tracks optimizer-dependent Hessian eigenvector reorganization and localization during training.

Topic Match: The core is SGD-versus-Adam optimization geometry, with evidence limited to small multilayer perceptrons.

Relevance: 7 Novelty: 6


13. Curvature-Weighted Gradient Diversity: A Noise Measure for Geometry-Adaptive SGD Schedules

ArXiv ID: 2606.30455

Primary Topic: Architecture and Training Dynamics

Authors: Muhammad Hamza, Ayush Goel

Abstract: The standard convergence analysis of mini-batch stochastic gradient descent (SGD) models gradient noise using a single variance term that treats all parameter directions equally, ignoring the fact that noise in high-curvature directions has less impact because learning rates are already constrained there. We introduce Curvature-Weighted Gradient Diversity (CWGD), a geometry-aware measure that weights per-sample gradient diversity by the inverse square root of the Hessian, providing a tighter proxy for the effective optimization noise. For strongly convex quadratic objectives with diagonal Hessians and isotropic noise, we prove that a CWGD-modulated cosine learning-rate schedule can reduce the asymptotic optimization error floor by up to a factor of two compared with standard cosine annealing. We implement this idea as CWGD-Cosine using a Hutchinson-based diagonal Hessian estimator that is exact for quadratic objectives. Across a range of condition numbers, batch sizes, and noise structures, CWGD-Cosine consistently achieves approximately 20% lower final optimization error than standard cosine annealing while incurring negligible overhead in the quadratic setting. We also identify and correct a degenerate curvature estimator, analyze the robustness of the proposed estimator, and explicitly discuss the limitations of the method, including Hessian staleness in non-convex optimization. These results establish CWGD as a principled geometry-aware measure of optimization noise and motivate future extensions to more general learning problems.

Comment: Introduces a curvature-weighted gradient-noise measure to adapt SGD learning-rate schedules.

Topic Match: The core analyzes how curvature and gradient noise affect optimization dynamics, though its guarantees and experiments are limited to quadratic objectives.

Relevance: 6 Novelty: 6


Efficiency, Compression, and Large-Scale Training (5)

1. Hierarchical Global Attention (HGA)

ArXiv ID: 2606.30709

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Woernle Frank, Fedosov Vladimir, Grinenko Artemiy

Abstract: Hierarchical Global Attention (HGA) is a drop-in replacement for dense causal attention in pretrained long-context transformers. HGA preserves the original checkpoint parameters: the pretrained $W_Q$, $W_K$, $W_V$, and $W_O$ projections remain unchanged, no calibration parameters are introduced, and no retraining is required. Applied to Qwen3-30B-A3B-Instruct-2507-FP8 on a single RTX~5090 (32GB), the patched model runs out of the box at a 64K-token context, where token-level K/V storage is not feasible on this hardware. Unlike previous sparse-attention methods, HGA performs hierarchical two-level routing. It first retrieves relevant chunks using compact RoPE-aware summaries and then refines the selection by routing only the most relevant groups before performing exact token-level attention. This hierarchical retrieval significantly reduces the number of fetched tokens while preserving exact attention over the retrieved token set, making RAM- and NVMe-backed storage practical. The full historical token K/V resides in host RAM or NVMe storage, while only a small routed working set is transferred to GPU memory during attention. Consequently, GPU memory consumption depends primarily on model weights and the routed working set rather than on the total context length. Across all tested context lengths (4K - 64K tokens), routed attention remains within approximately $0.01$--$0.02$ nats of dense attention while the sparsity used is just about 3%. These results suggest that the approximation introduced by hierarchical routing is small, and that the remaining quality gap is likely dominated by long-context positional encoding rather than by the routing algorithm itself.

Comment: Two-level attention retrieval keeps a small GPU working set while offloading historical KV to RAM or NVMe.

Topic Match: The central contribution reduces GPU KV residency through a new hierarchical sparse-attention mechanism that operates on pretrained checkpoints.

Relevance: 9 Novelty: 7


2. Beyond Single-Dimensional Compression: The Compound Sparsity Frontier of Large Language Models

ArXiv ID: 2607.18280

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Chao Han, Haozhe Hu, Xiaoyu Shen

Abstract: Large language models (LLMs) are often compressed through static parameter pruning or dynamic token-level computation, yet aggressive sparsification can trigger rapid performance degradation beyond an essential sparsity boundary. This work asks \emph{whether combining these two mechanisms can delay such degradation by distributing the compression burden}. We study a minimalist compound sparsity framework that first applies low-rank approximation and channel pruning to obtain a statically compressed backbone, and then introduces lightweight routers for per-token dynamic layer skipping. This design enables independent control of parameter sparsity and token-level computation sparsity. Experiments across language understanding and modeling benchmarks show that compound sparsity consistently outperforms single-mechanism compression under the same total sparsity, delaying the decay point on understanding tasks and preserving stronger modeling performance. Further analysis reveals cross-dimensional interference between parameter pruning and token skipping, and shows that near-balanced allocation is most effective under a fixed sparsity budget. These results demonstrate that compound compression provides a practical way to improve LLM compression, while revealing a broader cross-dimensional sparsity boundary that ultimately limits further compression. Code will be available at https://github.com/EIT-NLP/LLM-Pruning.

Comment: Analyzes how parameter pruning and token-dependent layer skipping interact under a fixed compression budget.

Topic Match: The core contribution concerns LLM compression limits and allocation between static sparsity and dynamic computation.

Relevance: 9 Novelty: 6


3. When Is a Draft Accepted? A Theory of Acceptance in Speculative Decoding

ArXiv ID: 2606.30265

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Aaryam Sharma

Abstract: Speculative decoding accelerates language model inference by using a fast drafter to propose candidate tokens that are then verified by a larger target model. Existing theory largely studies the stochastic, distribution-preserving setting, where the goal is to exactly sample from the target distribution. In contrast, many practical systems use greedy decoding, relaxed acceptance rules, or tree-based candidate sets, where success is governed by local ranking and threshold events rather than exact distributional equality. We develop a theory for these regimes. We identify that many common acceptance criteria have rejection regions that can be characterized as lower level sets of the target distribution. For these, we characterize the exact KL divergence required for rejection yielding exact certificates and sharp margin-based bounds for strict greedy decoding, additive and multiplicative relaxed acceptance, top-(m) relaxed criteria, and entropy-thresholded acceptance. We then extend the framework to greedy tree decoding, deriving exact and margin-only certificates for when the target greedy token remains covered by the drafter's top-(m) candidates. Finally, we evaluate the resulting certificates on Qwen3 models, showing that relaxed and tree-based criteria substantially enlarge the region of certified acceptance, especially on decoding steps with low target model distribution margin. These results complement existing distribution-preserving analyses of speculative decoding by characterizing the deterministic local acceptance events common in practical inference systems.

Comment: Sharp KL and margin certificates characterize token acceptance in greedy and tree speculative decoding.

Topic Match: The theory addresses acceptance rules that govern draft-token reuse and speculative decoding efficiency.

Relevance: 8 Novelty: 7


4. PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF

ArXiv ID: 2606.29758

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Doo Hwan Hwang, Kee-Eung Kim

Abstract: Reinforcement Learning from Human Feedback (RLHF) for Large Language Models increasingly relies on critic-free methods as a practical alternative to actor--critic training. Despite their simplicity, existing critic-free approaches propagate a trajectory-level learning signal uniformly across all tokens in a trajectory. This requires full-trajectory policy updates for every rollout, leading to substantial optimization cost for long reasoning traces, even though intermediate prefixes often contain enough information to largely determine the final outcome. We propose Prefix-Sampling Proximal Policy Optimization (PS-PPO), a compute-efficient critic-free method for RLHF that exploits this temporal redundancy. PS-PPO introduces a prompt-conditioned cutoff distribution and samples a cutoff timestep for each trajectory. During the update pass, PS-PPO backpropagates only through the sampled prefix of each trajectory and applies an importance-weighting correction so that the resulting truncated gradient estimator remains unbiased with respect to the full-trajectory objective. Experiments on mathematical reasoning and RLHF benchmarks show that PS-PPO achieves large reductions in training compute and peak GPU memory, while maintaining accuracy comparable to strong critic-free baselines.

Comment: Uses importance-weighted prefix backpropagation to reduce training compute and peak GPU memory.

Topic Match: The unbiased truncated-gradient estimator is a concrete LLM training-efficiency mechanism, with RLHF as its application setting.

Relevance: 8 Novelty: 7


5. Diffusion Fine-tuning with Rewarded Moment Matching Distillation

ArXiv ID: 2606.30414

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Alexis Jacq, Guillaume Couairon, Valentin De Bortoli, Quentin Berthet, Arnaud Doucet, Romuald Elie

Abstract: Distillation and Reinforcement Learning (RL) fine-tuning are the primary pillars of diffusion post-training. While traditionally studied in isolation, the interaction between these phases remains poorly understood, and in particular how fine-tuning impacts the generative quality of distilled models. We introduce Rewarded Moment Matching Distillation (RMMD), a novel framework that simultaneously distills diffusion models and maximizes a reward function. RMMD preserves the high-fidelity ``naturalness'' characteristic of advanced distillation (such as 8-step Moment Matching) by adapting the sampling loop for on-policy training and repurposing the distillation loss as a proxy for integral KL regularization. By evaluating the FID-Reward Pareto fronts on ImageNet, we demonstrate that RMMD achieves superior trade-offs compared to single-step baselines (DI++) and multi-step competitors (DRaFT, HyperNoise). Finally, we apply RMMD to GenCast, a state-of-the-art weather forecasting model, to distill it while optimizing the Continuous Ranked Probability Score (CRPS) metric. The resulting distilled model achieves a 7.5x speedup while outperforming the teacher model on 93% of target weather variables, and being better calibrated. This proves that RMMD scales to complex, high-dimensional scientific domains.

Comment: Develops reward-aware moment-matching distillation for faster diffusion sampling.

Topic Match: The new distillation objective directly targets reduced sampling cost while preserving quality, giving it a substantive efficiency contribution.

Relevance: 7 Novelty: 7


Paper Selection Prompt

System Prompt

You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.

User Prompt

Instructions

Respond in JSONL. Output exactly one JSON object per paper, one per line:

{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}

Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only: daily_hot, new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return []. - daily_hot means the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. - new_frontier means the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.

Scoring Criteria

Relevance and Novelty are independent axes. Score both from 1 to 10.

Relevance Scoring

  • 9-10: directly centered on the target foundational topics; highest when the core contribution is clearly within them.
  • 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
  • 5-6: touches the target topics, but the main contribution is elsewhere.
  • 3-4: largely outside the target topics, often application-focused or domain-specific.
  • 1-2: unrelated.

Important: Broad frontier relevance, major launch status, or daily buzz is not enough for a high Relevance score here. Those cases belong in the hotspot digest unless the paper strongly matches the specialized paper topics.

Novelty Scoring

  • 9-10: new paradigm, theory, or major methodological breakthrough.
  • 7-8: substantial methodological advance or strong new insight.
  • 5-6: meaningful but incremental extension or refinement.
  • 3-4: minor, narrow, or mostly engineering or domain-specific improvement.
  • 1-2: little originality; mainly standard application of existing methods.

Topic Registry

Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - moe_training: MoE Training - Mixture-of-Experts training: routing, load balancing, expert granularity and shared experts, training stability, MoE parallelism and communication, and MoE kernels. - training_systems: Large-Scale Training Systems and Efficiency - Large-scale training systems and efficiency: distributed training algorithms, optimizers, and communication. - architecture_training: Architecture and Training Dynamics - Core architectural or computational mechanisms, dynamic computation, and training-stability dynamics. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, memory or cache efficiency, and large-scale training systems that materially change cost or behavior.

Papers

[PAPER LIST HERE]

Relevant Topics

This is a training-side feed. The centre of gravity is Mixture-of-Experts training and the large-scale training systems it sits on; architecture and efficiency work is kept when it bears on how a large model is trained or what it costs to train.

Keep a paper when its CORE CONTRIBUTION falls in one of the four topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.

  1. MoE Training (primary) - Keep: expert routing and gating mechanisms; load balancing, with or without auxiliary losses; expert granularity, fine-grained and shared experts; MoE training stability (loss spikes, router collapse, z-loss and related fixes); expert and hybrid parallelism; all-to-all dispatch/combine and the communication schedules MoE needs; MoE kernels and grouped-GEMM style implementations; expert capacity, token dropping, and upcycling dense checkpoints into MoE. - Filter: papers that merely train on top of a MoE model without a new routing, balancing, stability, parallelism, or kernel contribution; "mixture of experts" in the classical ensemble or recommender sense.

  2. Large-Scale Training Systems - Keep: distributed training algorithms (data, tensor, pipeline, sequence, and expert parallelism, FSDP and ZeRO-style sharding); optimizers and preconditioners for large-scale pretraining; communication algorithms, overlap, and collective design; throughput and MFU work that materially changes what a training run costs or whether it converges; scaling-law work that informs how such runs are configured. - Filter: routine infrastructure tuning with no new algorithmic or systems idea.

  3. Architecture and Training Dynamics - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, dynamic or modular computation); optimisation and training-dynamics analysis that explains why large models train the way they do. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight.

  4. Efficiency and Compression - Keep: quantization, sparsity, pruning, low-rank adaptation, KV-cache and memory-efficient designs, and other work that materially changes the cost of running or training a large model, especially when the mechanism is new rather than a tuned variant. - Filter: straightforward application of standard efficiency methods, and deployment or serving engineering with no new idea.

Leave these to the hotspot digest unless the core contribution clearly falls in one of the four topics above: - major model or product releases - agent, tooling, and RAG launches - benchmark, leaderboard, or evaluation-only papers - representation-learning theory, interpretability, and feature-formation analysis - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning) - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains