This is a remedial run for missed papers from 05/28/2026 to 05/28/2026.
Results generated on 09/11/2026.
Personalized Daily ArXiv Papers 2026-05-29
| Model | Metric | Usage | Papers | ||||
|---|---|---|---|---|---|---|---|
| Prompt | Completion | Total | Total arXiv | Scanned | Relevant | ||
local agent CLI |
Tokens | not reported | not reported | not reported | 651 | 651 | 25 |
| Cost | not reported | not reported | not reported | ||||
Token counts are not reported for this run. 6 of 7 model calls succeeded, 2,808s of model wall clock.
Topic Coverage:
| Topic | Papers |
|---|---|
| MoE Training | 1 |
| Large-Scale Training Systems and Efficiency | 5 |
| Architecture and Training Dynamics | 9 |
| Efficiency, Compression, and Large-Scale Training | 10 |
Table of contents by topic:
MoE Training (1)
- ConMoE: Expert-Pool Consolidation via Prototype Reassignment for MoE Compression Authors: Yilun Yao, Jiaming Pan, Elsie Dai, Peizhuang Cong, Yaoming Li, Tong Yang
Large-Scale Training Systems and Efficiency (5)
-
AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training Authors: Ling Chen, Houming Wu, Wenjie Yu
-
On the Optimizer Dependence of Neural Scaling Laws Authors: Vansh Ramani, Shourya Vir Jain
-
Singularity-aware Optimization via Randomized Geometric Probing: Towards Stable Non-smooth Optimization Authors: Ruoran Xu, Borong She, Xiaobo Jin, Qiufeng Wang
-
FP8 is All You Need (Part 2): Full-FP64 3-D FFT on FP8-Generation Tensor CoresThe Integer-Epilogue Wall and the Minimal Hardware That Would Remove It Authors: Satoshi Matsuoka
-
Convergence of Steepest Descent and Adam under Non-Uniform Smoothness Authors: Sharan Vaswani, Yifan Sun, Reza Babanezhad
Architecture and Training Dynamics (9)
-
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't Authors: Anej Svete, William Merrill, Ryan Cotterell, Ashish Sabharwal
-
Attention as In-Context Empirical Bayes: A Two-Stage View via Particle Dynamics Authors: Matthew Smart, Soumya Ganguly, Nilava Metya, Alexandre V. Morozov, Anirvan M. Sengupta
-
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models Authors: Rohan Shravan
-
Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders Authors: Pierre-Antoine Lequeu, Camille Barboule, Benjamin Piwowarski
-
DisjunctiveNet: Neural Symbolic Learning via Differentiable Convexified Optimization Layers Authors: Shraman Pal, Can Li
-
Anti Mode-Collapse in Mean-Field Transformer via Auxiliary Variables Authors: Masaaki Imaizumi, Masanori Koyama, Noboru Isobe, Kohei Hayashi
-
On the Construction and Implications of Low-Loss Valleys in LoRA-based Bayesian Inference Authors: Daniel Dold, Emanuel Sommer, Julius Kobialka, Oliver Dürr, David Rügamer
-
The Long-Term Effects of Data Selection in LLM Fine-Tuning Authors: Yuxin Yang, Aoxiong Zeng, Xiangquan Yang
-
Faithful Embeddings of Irregular and Asynchronous Data for Online Log-NCDEs Authors: Benjamin Walker, Alexandre Bloch, Lingyi Yang, Sam Morley, Terry Lyons
Efficiency, Compression, and Large-Scale Training (10)
-
HARP: Hadamard-Preconditioned Adaptive Rotation Processor for Extreme LLM Quantization Authors: Artur Zagitov, Gleb Molodtsov, Aleksandr Beznosikov
-
OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning Authors: Geng Li, Guohao Chen, Ting Chen, Shilin Shan, Kuangji Zuo, Bofan Lyu, Tuo An, Gen Li, Jianfei Yang
-
Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware Initialization Authors: Junlin He, Yihong Tang, Tong Nie, Guilong Li, Binyu Yang, Jinxiao Du, Lijun Sun, Wei Ma
-
VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion Authors: Hidir Yesiltepe, Jiazhen Hu, Tuna Han Salih Meral, Adil Kaan Akan, Kaan Oktay, Hoda Eldardiry, Pinar Yanardag
-
Accelerating Constrained Decoding with Token Space Compression Authors: Michael Sullivan, Alexander Koller
-
Midpoint Generative Models Authors: Daniil Shlenskii, Nikita Gushchin, Lev Novitskiy, Dmitry V. Dylov, Alexander Korotin
-
Efficient Test-Time Finetuning of LLMs via Convex Reconstruction and Gradient Caching Authors: Alaa Khamis, Alaa Maalouf
-
MarginGate: Sparse Margin-Triggered Verification for Batch-Invariant LLM Inference Authors: Kexin Chu, Yang Zhou, Wei Zhang
-
How LoRA Remembers? A Parametric Memory Law for LLM Finetuning Authors: Ziwen Xu, Haiwen Hong, Linsong Yu, Benglei Cui, Longtao Huang, Hui Xue, Ningyu Zhang
-
Bastion: Budget-Aware Speculative Decoding with Tree-structured Block Diffusion Drafting Authors: Soowon Oh, Nam Cao, Yujin Kim, Hojung Jung, Huzama Ahmad, Sangmin Bae, Se-Young Yun
MoE Training (1)
1. ConMoE: Expert-Pool Consolidation via Prototype Reassignment for MoE Compression
ArXiv ID: 2605.29350
Primary Topic: MoE Training
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Yilun Yao, Jiaming Pan, Elsie Dai, Peizhuang Cong, Yaoming Li, Tong Yang
Abstract: Mixture-of-Experts (MoE) language models reduce per-token computation but still require storing and serving all experts, making deployment memory-intensive. Existing post-training compression methods mainly shrink this cost by pruning experts or merging their weights. We formulate post-training MoE compression as expert-pool consolidation: retaining a smaller set of pretrained experts as reusable prototypes and deterministically remapping each original expert reference to one selected prototype. This view separates the reduced expert pool from the reuse structure that represents the original expert slots, and allows prototype sharing within local layer scopes while preserving the original router interface. We propose ConMoE, a train-free prototype remapping framework that selects retained experts using calibration-based contribution and replaceability signals, then redirects original expert calls to the selected prototypes without weight updates or post-compression fine-tuning. Experiments on three pretrained MoE language models show that ConMoE matches or outperforms strong pruning and merging baselines in several settings, achieving the best average score on deepseek-moe-16b-base at both 25% and 50% routed-expert reduction, while remaining competitive on Qwen3-30B-A3B and OLMoE-1B-7B-0125. Ablations indicate that deterministic reassignment is the most stable component, whereas broader cross-layer sharing and post-hoc weight fusion are model-dependent.
Comment: Formulates post-training MoE compression as expert-pool consolidation: retain prototype experts and deterministically remap expert references while preserving the original router interface.
Topic Match: The mechanism is specific to MoE expert pools and the router interface, separating expert storage from the reuse structure.
Relevance: 8 Novelty: 6
Large-Scale Training Systems and Efficiency (5)
1. AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training
ArXiv ID: 2605.29664
Primary Topic: Large-Scale Training Systems and Efficiency
Authors: Ling Chen, Houming Wu, Wenjie Yu
Abstract: Pipeline parallelism is essential for large-scale model training, but existing asynchronous approaches often degrade convergence due to parameter mismatch between forward and backward passes. We propose Asynchronous Multi-Directional Pipeline parallelism (AMDP) to mitigate this issue while sustaining high utilization. AMDP limits the first stage of each pipeline to process at most two minibatches before backpropagation, bounding the number of parameter updates between forward and backward passes. To alleviate the resulting pipeline bubbles, AMDP launches multiple concurrent pipelines and adapts their number according to pipeline depth. In addition, AMDP accumulates gradients across minibatches and applies them in a single update, ensuring that only a bounded number of minibatches experience parameter mismatch, limited to within one optimization step. Experiments on GPT- and BERT-style models demonstrate that AMDP significantly accelerates training while preserving convergence.
Comment: Asynchronous pipeline parallelism that bounds forward/backward parameter mismatch to within one optimization step and fills bubbles with multiple concurrent pipelines.
Topic Match: Directly a distributed-training algorithm: pipeline parallelism with a convergence-preserving staleness bound.
Relevance: 9 Novelty: 6
2. On the Optimizer Dependence of Neural Scaling Laws
ArXiv ID: 2605.29387
Primary Topic: Large-Scale Training Systems and Efficiency
Authors: Vansh Ramani, Shourya Vir Jain
Abstract: The scaling exponent $α$ in neural scaling laws $L(N) \propto N^{-α}$ is commonly treated as a fixed constant set by architecture and data. We present evidence that $α$ depends systematically on the optimizer. In controlled random-feature regression experiments -- the canonical theoretical framework for neural scaling -- we measure $α$ across five optimizer variants and six spectral conditions. Preconditioned optimizers consistently yield steeper scaling (larger $α$), with the $α$-shift increasing across most of the tested spectral range, peaking near $s = 1.5$, and remaining large at $s = 2.0$. At $s \approx 1.0$ (characteristic of natural language), the full natural gradient achieves $α\approx 0.31$ versus $α\approx 0.12$ for gradient descent -- a $2.6\times$ larger fitted exponent that, within the random-feature model, compounds with each model-size doubling. Whether and how this exponent shift transfers to large-scale LLM training -- where recent evidence suggests the advantage may attenuate with scale -- remains an important open question. Our results imply that scaling-law forecasts should account for optimizer choice, and we provide a spectral diagnostic predicting when advanced optimizers will pay off.
Comment: Evidence that the neural scaling exponent is not architecture-fixed but shifts systematically with the optimizer, with preconditioned methods more than doubling the fitted exponent in random-feature regression.
Topic Match: Scaling-law work that directly informs how large runs are configured, plus a spectral diagnostic for when advanced optimizers pay off.
Relevance: 8 Novelty: 6
3. Singularity-aware Optimization via Randomized Geometric Probing: Towards Stable Non-smooth Optimization
ArXiv ID: 2605.29547
Primary Topic: Large-Scale Training Systems and Efficiency
Authors: Ruoran Xu, Borong She, Xiaobo Jin, Qiufeng Wang
Abstract: Deep learning optimization relies heavily on the assumption of smooth loss landscapes, a condition systematically violated by modern architectures due to non-smooth components such as ReLU activations and quantization operators. In such non-smooth regimes, adaptive optimizers such as Adam suffer from gradient chattering, violent oscillations caused by conflicting signals within the Clarke subdifferential, leading to poor convergence and suboptimal generalization. To address this, we introduce Singularity-aware Adam (S-Adam), a novel optimizer that stabilizes training by dynamically modulating step sizes based on local geometric instability. Our key contribution is the Local Geometric Instability (LGI) metric, a computationally efficient estimator of the Clarke subdifferential diameter derived from the variance of randomized directional derivatives. S-Adam incorporates an adaptive damping mechanism exp(-$λ$$ρ$) that decelerates updates in high-instability regions while preserving fast convergence in smooth basins. We provide a rigorous convergence analysis using differential inclusions, proving that S-Adam converges almost surely to ($δ$,$ε$)-Clarke stationary points at the optimal O(1/$\sqrt(T)$) rate. Empirical evaluations on Quantization-Aware Training (QAT) and high-noise small-batch learning demonstrate that S-Adam consistently outperforms AdamW and Prox-SGD, achieving accuracy gains of up to +4.54% on CIFAR-100 and +4.27% on TinyImageNet while effectively mitigating gradient oscillations.
Comment: Randomized estimates of nonsmooth gradient instability damp Adam updates to stabilize quantization-aware training.
Topic Match: The core contribution is a new optimizer with geometry-dependent damping; validation is limited to smaller vision and quantization-aware training experiments.
Relevance: 7 Novelty: 7
4. FP8 is All You Need (Part 2): Full-FP64 3-D FFT on FP8-Generation Tensor CoresThe Integer-Epilogue Wall and the Minimal Hardware That Would Remove It
ArXiv ID: 2606.23698
Primary Topic: Large-Scale Training Systems and Efficiency
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Satoshi Matsuoka
Abstract: The NVIDIA Blackwell Ultra (B300) GPU cuts FP64 vector throughput $\sim 30\times$ while multiplying FP8 tensor throughput. After the recovery of FP64 GEMM via Ozaki Scheme II on FP8 tensor cores and the Tensor-Memory Equilibrium model of the companions ("FP8 is All You Need, Part 1" and "Ozaki 2.5") we ask whether the fifth canonical HPC primitive, the full-FP64 $1024^3$ 3-D FFT, can be carried by the same substrate, and answer with a design and its limit. It is a Bailey six-step transform with no FP64 arithmetic: FP8-tensor DFT GEMMs with fused twiddles, residue-domain Karatsuba combines and exact CRT reconstruction whose bulk is a small GEMM on the FP16 tensor path and whose remainder is a Kulisch fixed-point accumulation with a two-sided modulo-$M$ lift, so the only rounding is the final conversion; constants are machine-generated and verified bit-exactly. The central finding: the binding resource is not floating point but a per-output integer epilogue with floor $(c_{\rm epi}/8),B_{\rm mem}$, $c_{\rm epi} \approx 203$-$281$ instructions per output: on B300 it holds the transform at 63-87 ms against a 12.9 ms roof ($4.9$-$6.7\times$ short); at most $1.3$-$1.9\times$ faster than the collapsed native path, possibly no faster at realised issue rates; no software route reaches the roof; on the NVIDIA Rubin GPU emulation loses $8$-$11\times$. An FP32 variant meets the same wall: the cause is per-scalar reconstruction, not FP64. Each floor term names its remedy: the NVIDIA B200 GPU's INT8 tensor core restored with a position-weighted cross-column accumulation primitive, a load-path deconstruction datapath shared with the companions, two ISA idioms and modular reduction at the MMA output give 16.0-23.5 ms with minor hardware and 12.9-15.0 ms with one moderate ask. All figures are projected floors, not measurements, with sensitivities and the FP8 layout condition given.
Comment: Shows the binding resource for FP64-accurate 3-D FFT on FP8 tensor cores is a per-output integer epilogue, and names the minimal hardware primitives that would remove the wall.
Topic Match: Low-precision tensor-core substrate and hardware co-design analysis that bears on what accelerator arithmetic can carry; figures are projected floors, not measurements.
Relevance: 6 Novelty: 7
5. Convergence of Steepest Descent and Adam under Non-Uniform Smoothness
ArXiv ID: 2605.30648
Primary Topic: Large-Scale Training Systems and Efficiency
Also Matches: Architecture and Training Dynamics
Authors: Sharan Vaswani, Yifan Sun, Reza Babanezhad
Abstract: Recent work has analyzed the convergence of first-order methods under non-uniform smoothness assumptions that better model the loss landscape in machine learning tasks. We generalize this assumption to objectives whose curvature is an affine function of the objective value. This property is satisfied by a broad class of problems, including logistic regression, generalized linear models with a logistic link function, softmax policy gradient in reinforcement learning, and a class of neural networks. Under this assumption and gradient domination conditions, we establish a general convergence rate for the steepest descent method, and deterministic, diagonal variants of RMSProp and Adam. Our results imply that for logistic regression on separable data and the softmax policy gradient objective, sign GD converges linearly and is provably faster than GD. Furthermore, we show that for a class of two-layer neural networks on separable data, RMSProp and Adam can converge at a linear rate with a constant step-size and momentum parameter. Finally, we present a lower bound demonstrating that, under our assumption, RMSProp and Adam are provably faster than AdaGrad, AMSGrad, gradient descent, and heavy-ball momentum.
Comment: Convergence rates for steepest descent, RMSProp, and Adam under curvature that is affine in the objective value, with a lower bound separating them from AdaGrad and heavy-ball.
Topic Match: Optimizer theory explaining when adaptive methods beat GD, relevant to large-scale pretraining optimizer choice though analysed in deterministic simplified settings.
Relevance: 6 Novelty: 6
Architecture and Training Dynamics (9)
1. Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
ArXiv ID: 2605.30523
Primary Topic: Architecture and Training Dynamics
Authors: Anej Svete, William Merrill, Ryan Cotterell, Ashish Sabharwal
Abstract: Recent work describes what transformers can and cannot compute through connections to boolean circuits, but existing results lack exact characterizations and are sensitive to modeling choices. Padded transformers -- to whose input filler symbols such as ``...'' are appended -- emerge as a useful gadget for establishing equivalences to circuit classes by providing polynomial space for adaptive parallel computation. However, only a limited set of padded transformer idealizations has been studied, leaving open how robustly these equivalences hold under changes to attention type, model width, and uniformity. We find that, under practical assumptions, padded transformers are surprisingly robust to all of these, and identify numeric precision and model depth as the main factors affecting expressivity. Concretely, we prove that polynomially padded $\text{L-uniform}$ constant-precision transformers are equivalent to $\text{L-uniform AC}^0$, while growing-precision ones achieve $\text{L-uniform TC}^0$ regardless of width. Furthermore, looping enables sequential processing analogous to circuits: $\log^d N$-looped constant-precision transformers reach $\text{FO-uniform AC}^d$, and growing-precision ones reach $\text{FO-uniform TC}^d$. Interestingly, growing width or precision beyond logarithmic does not increase expressivity, and all our results hold for both softmax and average hard attention transformers.
Comment: Identifies numerical precision and looping depth as decisive architectural controls on padded-transformer expressivity.
Topic Match: Exact expressivity results isolate the computational roles of attention, width, precision, and recurrent depth, although training behavior is not studied.
Relevance: 8 Novelty: 8
2. Attention as In-Context Empirical Bayes: A Two-Stage View via Particle Dynamics
ArXiv ID: 2605.29351
Primary Topic: Architecture and Training Dynamics
Authors: Matthew Smart, Soumya Ganguly, Nilava Metya, Alexandre V. Morozov, Anirvan M. Sengupta
Abstract: We study minimal attention-only transformers under all-token corruption and show they admit a two-stage empirical Bayes interpretation. A single attention step computes a kernel-weighted posterior mean with respect to the empirical distribution defined by the context. Depth refines this distribution through particle dynamics (Stage 1), while a long-range skip-connection carries the noisy input as a query for posterior inference (Stage 2), revealing distinct statistical roles for depth and attention residuals. The framework isolates a minimal setting in which the context itself induces a depth-dependent energy landscape governing in-context inference. We show that effective denoising can emerge without an explicit noise schedule: a fixed kernel bandwidth and finite integration horizon suffice, yielding a principled depth-noise relationship. We further establish a posterior-mean recovery guarantee for a class of well-behaved priors, where the empirical estimator converges to the Bayes-optimal predictor under asymptotic conditions. Connecting these dynamics to reverse-diffusion limits, our results provide a statistical interpretation of attention as in-context inference via sample-based posterior estimation, without explicit density modeling.
Comment: Derives distinct computational roles for attention depth and long-range residuals through empirical-Bayes particle dynamics.
Topic Match: The analysis directly decomposes attention and residual mechanisms, although its guarantees concern a minimal denoising setting.
Relevance: 7 Novelty: 7
3. Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models
ArXiv ID: 2605.29459
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Rohan Shravan
Abstract: Large language models route every input through a learned embedding table of shape |V| x d_model, consuming hundreds of millions to billions of trainable parameters at frontier scale. We introduce Kronecker Embeddings, a deterministic byte-level character-position factorization that replaces this table with a fixed encoder and a single learned projection, compatible with standard BPE tokenizers, eliminating 91--94% of input-side trainable parameters at frontier scale. We provide five contributions. First, a cross-model probe across six LMs (135M-671B parameters) shows trained input embeddings cluster typographic variants of the probe word far more than morphological relatives; Kronecker escapes this clustering at the embedding layer. Second, a controlled three-seed comparison on nanoGPT GPT-2 124M over 2.5B tokens of FineWeb-Edu shows Kronecker reaching 2.5 +- 0.2% lower validation loss than the BPE-tied baseline (gap 0.083 +- 0.007 nats, ~9% lower perplexity), needing ~1.43x fewer steps to reach BPE's converged loss. Third, a spelling-robustness probe over 110 clean/typo pairs shows Kronecker preserves the top-1 prediction on 55.5% of pairs vs. 47.3% for BPE (+8.2 pp) and lowers KL by 7.6%, winning or tying in 10 of 11 categories; a generation probe shows Kronecker echoes byte-novel strings and typos through generation where BPE forgets them. Fourth, BPE embedding norm drifts during training while Kronecker projection norm stays near 1.0, consistent with a stable representational target. Fifth, an on-the-fly runtime variant reconstructs embeddings from a 4.5 MB byte buffer rather than a 2.15 GB table at vocabulary 131,072, with 0.01--0.24% step-time overhead. Byte-level locality has a tradeoff: byte-similar but semantically distant pairs (compute/commute, nation/notion) cluster together, shifting disambiguation to early attention layers.
Comment: Replaces the |V| x d embedding table with a deterministic byte-level character-position factorization plus one learned projection, removing 91-94% of input-side trainable parameters while improving validation loss.
Topic Match: A structural change to the input layer with measured effects on convergence speed and embedding-norm stability during training; evidence is GPT-2 124M scale.
Relevance: 7 Novelty: 6
4. Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders
ArXiv ID: 2605.30022
Primary Topic: Architecture and Training Dynamics
Authors: Pierre-Antoine Lequeu, Camille Barboule, Benjamin Piwowarski
Abstract: Positional encoding (PE) underpins how permutation-invariant Transformers represent sequence order, yet how positional information is processed and stored remains poorly understood. Modern PE methods such as RoPE still struggle on tasks such as long-context understanding or retrieval \cite{chen-etal-2025-hope}. Hence, a better understanding of the internal positional mechanism could help design better PE. Building on evidence that positional and semantic signals occupy nearly orthogonal subspaces in trained Transformers, we modify an encoder Transformer to process three explicitly disentangled streams: semantic, absolute positional (AP) and relative positional (RP), and confine the masked-language-modeling (MLM) objective to the semantic stream. This decoupling enables a clean mechanistic study and yields three take-aways. (1) The isolated AP subspace spontaneously collapses into a low-frequency two-dimensional manifold that captures the structure of the document; (2) Attention heads specialize into structure and semantic-oriented groups, with RP exclusively supporting the latter; (3) Standard positional encodings do not robustly retain macroscopic structure: RoPE and RP only weakly encode it, and entangled AP loses it in the final layers under MLM pressure. The disentangled approach preserves positional encoding, which improves linguistic representation on 49 of the 65 linguistic phenomena of the Flash-Holmes probing benchmark.
Comment: Splits an encoder into separate semantic, absolute-positional and relative-positional streams, revealing that the isolated AP subspace collapses to a low-frequency 2D document manifold and that RoPE fails to retain that macrostructure.
Topic Match: A mechanistic study of the positional-encoding mechanism itself, with head specialization findings that bear on how future PE should be designed.
Relevance: 7 Novelty: 6
5. DisjunctiveNet: Neural Symbolic Learning via Differentiable Convexified Optimization Layers
ArXiv ID: 2605.30456
Primary Topic: Architecture and Training Dynamics
Authors: Shraman Pal, Can Li
Abstract: Many learning tasks in science and engineering are characterized by sparse datasets, which limits the effectiveness of purely data-driven approaches. At the same time, these problems are often accompanied by rich domain knowledge derived from physical laws, operational requirements, and expert heuristics. Such knowledge is frequently expressed as rules involving logical propositions and linear inequalities. Existing neuro-symbolic methods typically enforce these rules approximately through soft penalties, assume input-independent rules when designing specialized architectures, or rely on non-differentiable post-processing at inference time to achieve hard constraint satisfaction. While recent advances in differentiable optimization layers enable end-to-end feasibility enforcement within neural networks, extending these approaches to logical or mixed-integer rules remains challenging due to inherent nonconvexity. In this work, we propose a unified end-to-end framework for enforcing hard, input-dependent mixed integer linear constraints within neural networks. Our approach represents rules as disjunctive constraints and applies hierarchical convex relaxations to obtain convex hull formulations. These relaxations yield tractable linear constraints that can be embedded as differentiable optimization layers while enabling exact rule satisfaction. We demonstrate the effectiveness of the proposed framework on real-world datasets, achieving perfect rule satisfaction and strong predictive performance.
Comment: Embeds convexified logical and linear constraints as differentiable optimization layers.
Topic Match: Differentiable constraint layers are architectural mechanisms, but the paper targets rule-constrained predictors without establishing relevance to large-model training.
Relevance: 6 Novelty: 7
6. Anti Mode-Collapse in Mean-Field Transformer via Auxiliary Variables
ArXiv ID: 2605.30229
Primary Topic: Architecture and Training Dynamics
Authors: Masaaki Imaizumi, Masanori Koyama, Noboru Isobe, Kohei Hayashi
Abstract: We use a mean-field-based transformer model to theoretically investigate how auxiliary variables, such as positional encoding, prevent mode collapse of self-attention mechanisms. The use of mean-field transformers to analyze the properties of self-attention mechanisms has garnered significant attention in recent years due to their ability to comprehensively analyze token interactions. However, analysis of this simple model suggests that mode collapse, where token distributions degenerate to a single point, occurs during long inferences (i.e., many layers), indicating a discrepancy with reality. This study investigates this mean-field transformer model and demonstrates that the introduction of auxiliary variables, such as positional encoding, acts as a counterforce against theoretical mode collapse. Specifically, we show that in the theoretical scheme, the energy-maximizing distribution does not degenerate to a single point; instead, it is characterized by a pushforward of the auxiliary variable distribution, thereby avoiding concentration in the Dirac measure. Our main examples are the positional encoding and the fixed prompt insertion treated as a parallel auxiliary-variable mechanism. Furthermore, we demonstrate that positional encoding and prompt insertion possess universality of representation in the limit, meaning that the limit distribution of inference can exactly represent a wide class of distributions. We also analyze several key properties of positional encoding and metastability, and validate our theoretical results through mathematical experiments.
Comment: Shows in a mean-field transformer that auxiliary variables such as positional encoding prevent depth-wise token-distribution collapse, with the limit being a pushforward rather than a Dirac measure.
Topic Match: Mechanistic analysis of self-attention depth dynamics and what an architectural component actually contributes.
Relevance: 6 Novelty: 6
7. On the Construction and Implications of Low-Loss Valleys in LoRA-based Bayesian Inference
ArXiv ID: 2605.29580
Primary Topic: Architecture and Training Dynamics
Authors: Daniel Dold, Emanuel Sommer, Julius Kobialka, Oliver Dürr, David Rügamer
Abstract: While parameter-efficient fine-tuning methods like low-rank adaptation (LoRA) are standard for large language models, principled estimation of epistemic uncertainty remains challenging. Recent results in the LoRA regime suggest that discrete multi-mode approaches such as deep ensembles offer little benefit over single-mode methods. This contradicts broader observations in deep learning, where ensembling independent optima typically improves generalization, and linking these modes through continuous low-loss valleys further enhances Bayesian model averaging (BMA). Whether such structure exists in the LoRA space and whether it yields functional diversity missed by local or discrete methods has not been studied. We introduce LoRA-Curve, a segmented Bézier curve parameterization in the LoRA space, with two variants: a free configuration that jointly optimizes all control points, and an anchored configuration that connects independently fine-tuned LoRA optima. We prove pathwise continuity and Lipschitz regularity of the loss along the curve and empirically show, across reasoning and classification benchmarks with Qwen2.5 7B, that linear interpolation encounters loss barriers, while our anchored multi-segment curves connect independent optima through continuous low-loss valleys. Combined with flat-minima perturbations and a Jensen-Shannon divergence regularizer, LoRA-Curve yields measurably higher mutual information of the predictive distribution without sacrificing performance, and links continuous parameter-space traversal to functional diversity.
Comment: Segmented Bézier curves connect independently fine-tuned LoRA optima through continuous low-loss valleys.
Topic Match: Loss-landscape geometry is the closest connection, but the central objective is Bayesian uncertainty estimation through model averaging.
Relevance: 6 Novelty: 6
8. The Long-Term Effects of Data Selection in LLM Fine-Tuning
ArXiv ID: 2605.30537
Primary Topic: Architecture and Training Dynamics
Authors: Yuxin Yang, Aoxiong Zeng, Xiangquan Yang
Abstract: Data selection is increasingly used to reduce the cost of large language model (LLM) fine-tuning, with recent methods prioritizing samples by current utility, diversity, quality, or influence. This paper studies a different question: when fine-tuning occurs over multiple stages, can selection strategies that look optimal now make the model less adaptable later? We introduce a long-horizon view of LLM data selection in which a selector is evaluated not only by immediate task performance, but also by future adaptation speed, forgetting, capability imbalance, and out-of-distribution robustness. We compare representative random, loss-based, gradient-based, diversity-based, quality-based, and utility-diversity selection families under a unified multi-stage protocol. Through controlled experiments designed to instantiate this protocol, we show how short-term selectors can exhibit rank reversal: they improve the current stage while slowing subsequent learning and increasing forgetting. We formalize this behavior as \emph{myopic selection}, provide a simple local analysis of why it can occur, and propose a diagnostic Long-Horizon Aware Selection (LHAS) objective that augments immediate utility with coverage, future-proxy transfer, and anti-concentration terms. The study argues that data selection should be evaluated as a training intervention that shapes the model's learning trajectory, rather than only as a local data-efficiency mechanism.
Comment: Links fine-tuning data selection to later adaptation speed, forgetting, and reversals in selector rankings.
Topic Match: Training trajectories are the closest match, but the contribution centers on evaluating multi-stage fine-tuning selection and proposing a diagnostic objective.
Relevance: 6 Novelty: 6
9. Faithful Embeddings of Irregular and Asynchronous Data for Online Log-NCDEs
ArXiv ID: 2605.30213
Primary Topic: Architecture and Training Dynamics
Authors: Benjamin Walker, Alexandre Bloch, Lingyi Yang, Sam Morley, Terry Lyons
Abstract: Continuous-time models are a natural choice for irregular and asynchronous data. A central design choice is how to embed discrete observations into continuous time. Interpolation- and imputation-based embeddings reconstruct a continuous observation path, making the model sensitive to the choice of reconstruction. We show that this reconstruction step is unnecessary; under mild conditions, compact-set universality on the model input space transfers to the data space whenever the embedding from data to input is continuous and injective. Guided by this result, and building on the rectilinear control path for Neural Controlled Differential Equations (NCDEs), we introduce a continuous and injective embedding for Log-NCDEs, a universal class of continuous-time models. Our approach records observations as increments and composes them over arbitrary query intervals to directly form log-signatures. This provides interval-level summaries without first interpolating the observed variables, while supporting online computation. Experiments on synthetic controlled dynamics and real-world time-series datasets show that the representation is accurate, efficient, and robust to irregular, asynchronous, and sparse observations.
Comment: Injective observation embeddings enable interpolation-free online computation for Log-NCDE sequence models.
Topic Match: Sequence-input construction is the closest architectural connection; the contribution focuses on irregular time-series modeling without demonstrated large-model training implications.
Relevance: 6 Novelty: 6
Efficiency, Compression, and Large-Scale Training (10)
1. HARP: Hadamard-Preconditioned Adaptive Rotation Processor for Extreme LLM Quantization
ArXiv ID: 2605.29843
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Artur Zagitov, Gleb Molodtsov, Aleksandr Beznosikov
Abstract: Post-training quantization (PTQ) is essential for deploying LLMs under memory and bandwidth constraints. However, extreme low-bit quantization remains highly sensitive to activation outliers and anisotropic weight curvature. Existing incoherence-based PTQ methods mitigate this issue with fixed randomized Hadamard transforms (RHTs), which improve quantization robustness but cannot adapt the rotated basis to the layer, calibration distribution, or quantizer. We introduce HARP (Hadamard-preconditioned Adaptive Rotation Processor), a learnable structured two-sided orthogonal processor that replaces fixed Hadamard mixing while preserving exact full-precision equivalence. HARP represents each rotation as a product of sparse butterfly-like block-orthogonal stages, supports non-power-of-two dimensions through Mixed-Radix schedules, and initializes to the RHT processor up to a fixed permutation. Fitted only on calibration data, HARP adapts the quantization basis to each layer and backend. Across 2--4-bit settings on Llama models from 1B to 70B, HARP consistently improves perplexity and yields its clearest zero-shot gains at 2 bits; a 2-bit Qwen3-8B experiment shows the same transfer beyond the Llama family. HARP also preserves deployment efficiency: on Llama 2 7B at 2 bits, it reaches 128 tok/s, retaining 90% of RHT throughput (142 tok/s) and running approximately $2.1\times$ faster than FP16 (61 tok/s).
Comment: Replaces fixed randomized Hadamard transforms in PTQ with a learnable two-sided butterfly rotation that adapts the quantization basis per layer while keeping exact FP equivalence and 90% of RHT throughput.
Topic Match: A new quantization mechanism at 2-4 bits across Llama 1B-70B, exactly the quantization strand of the efficiency topic rather than a tuned variant.
Relevance: 9 Novelty: 7
2. OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning
ArXiv ID: 2605.29657
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Geng Li, Guohao Chen, Ting Chen, Shilin Shan, Kuangji Zuo, Bofan Lyu, Tuo An, Gen Li, Jianfei Yang
Abstract: Vision-language models (VLMs) rely on long visual token sequences for visual understanding, making the prefill stage expensive in both computation and memory. Most existing pruning methods follow an absolute-ranking paradigm, assigning importance scores to visual tokens and retaining a fixed top-K subset. In this work, we argue that this paradigm is fundamentally brittle: attention sinks distort token importance rankings, while image redundancy and query-dependent visual evidence make fixed token budgets unreliable across inputs. We propose OccamToken, a training-free framework that replaces absolute token ranking with register-anchored relative evidence testing. Instead of asking which tokens are globally important, OccamToken evaluates whether a visual token provides information beyond a register-based reference. Our key insight is that register tokens naturally absorb low-information attention patterns, making them a stable reference for identifying genuinely informative visual evidence. Based on this principle, OccamToken performs both image-adaptive redundancy pruning and query-adaptive relevance pruning through dynamic thresholds derived from register attention. Across LLaVA-NeXT, LLaVA-v1.5, and Qwen3-VL, OccamToken consistently improves the accuracy-efficiency trade-off without additional training. Notably, on LLaVA-NeXT, it reduces 2,880 visual tokens to approximately 40 while preserving over 93% of the original accuracy, enabling stable visual token compression even in the extreme 1.4% retention regime.
Comment: Register-referenced evidence tests enable adaptive visual-token pruning without additional training.
Topic Match: The central contribution is a token-compression criterion that directly reduces VLM prefill computation and memory.
Relevance: 9 Novelty: 7
3. Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware Initialization
ArXiv ID: 2605.29327
Primary Topic: Efficiency, Compression, and Large-Scale Training
Also Matches: Architecture and Training Dynamics
Authors: Junlin He, Yihong Tang, Tong Nie, Guilong Li, Binyu Yang, Jinxiao Du, Lijun Sun, Wei Ma
Abstract: Efficient Distillation (EDistill) compresses large language models (LLMs) by structured pruning parameters and tuning lightweight modules with high training efficiency. Although these EDistilled LLMs achieve state-of-the-art (SOTA) performance on general ability benchmarks relative to similarly sized LLMs, we identify a severe degradation in their multi-step reasoning ability, which we term reasoning collapse. We systematically analyze the geometric origins of reasoning collapse and show that the SOTA EDistill method based on width-reducing projection matrices suffers from eRank collapse, in which the effective rank (eRank) of hidden representations drops. We theoretically explain how singular values of randomly initialized projection matrices become unevenly distributed, leading to eRank collapse and thus token indistinguishability. To address this issue, we propose RED (Reasoning-preserved Efficient Distillation) for LLMs, which introduces activation-aware initialization to initialize projection matrices as channel-selection matrices, thus theoretically mitigating eRank collapse. Experiments on Llama and Qwen series demonstrate that RED substantially recovers reasoning while maintaining high training efficiency and SOTA general ability.
Comment: Activation-aware channel-selection initialization mitigates effective-rank collapse during width-reducing LLM distillation.
Topic Match: The method preserves reasoning during efficient model compression, supported by a mechanistic explanation of projection-induced rank collapse.
Relevance: 9 Novelty: 7
4. VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
ArXiv ID: 2605.30351
Primary Topic: Efficiency, Compression, and Large-Scale Training
Also Matches: Architecture and Training Dynamics
Authors: Hidir Yesiltepe, Jiazhen Hu, Tuna Han Salih Meral, Adil Kaan Akan, Kaan Oktay, Hoda Eldardiry, Pinar Yanardag
Abstract: Long-rollout causal video diffusion has converged on a fixed-size sliding-window KV cache, with recent progress innovating within this layout by changing which tokens occupy the window or how their positions are encoded. The per-head KV layout itself, a dominant contributor to streaming memory and latency, has been mostly left unchanged. In this paper, we present the first study of Multi-Head Latent Attention (MLA) in video diffusion. VideoMLA replaces per-head keys and values with a shared low-rank content latent and a shared decoupled 3D-RoPE positional key, reducing per-token KV memory by 92.7% at every cached layer. We further investigate why MLA succeeds in video diffusion even though the spectral assumption often used to motivate it in language models does not hold: pretrained video attention is not low-rank, with 99%-energy effective rank far above any practical latent dimension. VideoMLA retains quality at compression ratios where direct spectral approximation would predict large reconstruction error. We show that the MLA bottleneck, rather than the pretrained spectrum, determines the effective rank: both spectral and random initialization occupy nearly the full rank budget from initialization, and training preserves this budget while adapting within it. On VBench, VideoMLA matches short-horizon streaming video diffusion baselines, achieves the best overall score at long horizons among evaluated methods, and improves throughput by 1.23x on a single B200.
Comment: Shared low-rank content latents and decoupled positional keys substantially reduce video-attention KV storage.
Topic Match: Cache compression is the core mechanism, and the rank-budget experiments explain how training preserves quality despite a restrictive latent bottleneck.
Relevance: 9 Novelty: 7
5. Accelerating Constrained Decoding with Token Space Compression
ArXiv ID: 2605.29986
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Michael Sullivan, Alexander Koller
Abstract: To guarantee that an LLM's outputs conform to a specified structure, context-free grammar (CFG) decoding engines force the selection of next tokens that produce strings that conform to a given CFG. While current CFG-constrained decoding engines are highly optimized, the inherent costs arising from the massive per-step search space -- i.e. the entire token vocabulary -- result in intractably high overhead for more complex CFGs: precisely the situation where CFG engines are most useful. In this paper, we introduce CFGzip, an offline technique for compressing the token search space, which massively reduces CFG engine overhead. In experiments, we report latency reduction of up to two orders of magnitude when CFGzip is used with a SoTA grammar engine, yielding an up to 7.5x speedup in total constrained generation time: with CFGzip, constrained decoding is now feasible at scale for complex CFGs.
Comment: Offline token-space compression reduces grammar-checking overhead during constrained LLM decoding.
Topic Match: A new search-space compression algorithm materially reduces generation latency, with a focused application to grammar-constrained inference.
Relevance: 8 Novelty: 7
6. Midpoint Generative Models
ArXiv ID: 2605.29920
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Daniil Shlenskii, Nikita Gushchin, Lev Novitskiy, Dmitry V. Dylov, Alexander Korotin
Abstract: We introduce Midpoint Generative Models (MGM), a principled framework for training one-step generative models. MGM is based on a simple symmetry of Flow Matching with linear interpolation: when the two endpoint distributions coincide, the corresponding drift field vanishes at the midpoint time, $t=1/2$. We show that the norm of this field defines a valid discrepancy between distributions, which we call the Midpoint Divergence. We extend this discrepancy beyond the midpoint by introducing randomly flipped interpolations and further generalize it by replacing deterministic linear Flow Matching interpolations with symmetric stochastic interpolants, yielding a generalized Midpoint Divergence. Finally, we derive a variational formulation of our generalized divergence, yielding a tractable objective for training a one-step generator. The resulting MGM algorithm offers an effective and theoretically grounded approach to generative modeling, achieving competitive performance against existing one-step generative modeling methods.
Comment: A midpoint-divergence objective directly trains one-step generators, targeting the cost of iterative generation.
Topic Match: The main connection is generation efficiency through a new training objective; large-scale training savings are not established.
Relevance: 7 Novelty: 8
7. Efficient Test-Time Finetuning of LLMs via Convex Reconstruction and Gradient Caching
ArXiv ID: 2605.30337
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Alaa Khamis, Alaa Maalouf
Abstract: Test-time finetuning (TTFT) is a rapidly evolving paradigm that adapts a language model to each prompt by retrieving related sequences, updating the model on them, and then evaluating the prompt. However, TTFT is only practical if it is fast: selection and finetuning both happen per query, making each a direct bottleneck. Existing methods trade speed for quality: fast retrieval is often redundant, while stronger diversity-aware selection adds prohibitive per-query cost. We introduce HullFT, a geometric approach to TTFT that addresses both bottlenecks. Given a query, HullFT first represents the query embedding as a sparse convex combination of few training sequences, using efficient projection-free Frank-Wolfe optimization. This yields a support set that is inherently relevant and diverse. We then convert the fractional convex weights into an exact integer multiset for finetuning through a geometric integerization procedure. The resulting multiplicities naturally create repeated examples, which we exploit with Gradient Reuse to amortize forward-backward computation across repeated finetuning steps. Our experiments show that HullFT improves the quality-efficiency tradeoff over current state-of-the-art TTFT methods, achieving lower bits-per-byte at substantially lower total runtime.
Comment: Gradient reuse amortizes forward-backward computation across repeated examples during test-time fine-tuning.
Topic Match: The method explicitly reduces adaptation cost through sparse convex sample selection and gradient reuse, with a per-query fine-tuning scope.
Relevance: 7 Novelty: 7
8. MarginGate: Sparse Margin-Triggered Verification for Batch-Invariant LLM Inference
ArXiv ID: 2605.30218
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Kexin Chu, Yang Zhou, Wei Zhang
Abstract: Temperature-zero BF16 LLM inference is often treated as reproducible, yet the same request can emit different tokens when decoded alone or inside a larger batch. Existing fixes use batch-invariant operators or LLM-42's per-token verification, incurring cost even when most steps are stable. We ask whether verification can be applied exclusively to flipped tokens. Across five models, batch-induced token flips are sparse on the flip-rate benchmarks: on MATH500, Llama-3.1-8B flips on $0.48\%$ of synchronous decode steps, and all tested models stay within the 0.3-1.3% range on MATH500, GSM8K, and HumanEval. K/V perturbations remain flat before flips, while low top-1/top-2 logit margins expose much of the flip risk. MarginGate turns these observations into a verifier policy: it keeps BF16 decoding on high-margin steps, verifies only low-margin steps, and repairs confirmed mismatches by replacing the current K/V column. We evaluate on four datasets, calibrating on MATH500 and transferring to GSM8K, SharedGPT, and HumanEval. MarginGate restores 100% sequence-level deterministic decoding on Llama-3.1-8B and Qwen2.5-14B with 18.56%/15.05% verifier trigger rates, reducing LLM-42's latency increment by 2.23x/1.99x relative to always-on verification. On DSR1-Distill-Qwen-7B, the same policy reaches determinism in a harder regime at 49.50% triggers.
Comment: Margin-triggered verification reduces the computation needed to make BF16 decoding deterministic.
Topic Match: Selective verification and local KV repair introduce an inference-efficiency mechanism, with benefits specific to deterministic decoding.
Relevance: 7 Novelty: 6
9. How LoRA Remembers? A Parametric Memory Law for LLM Finetuning
ArXiv ID: 2605.30260
Primary Topic: Efficiency, Compression, and Large-Scale Training
Also Matches: Large-Scale Training Systems and Efficiency
Authors: Ziwen Xu, Haiwen Hong, Linsong Yu, Benglei Cui, Longtao Huang, Hui Xue, Ningyu Zhang
Abstract: Large Language Models (LLMs) must continuously learn and update knowledge to remain effective in dynamic real-world environments. While Low-Rank Adaptation (LoRA) is widely used for such memory updates, existing studies mainly rely on qualitative downstream evaluations, leaving the quantitative capacity limits and underlying dynamics of exact parametric memory largely unexplored. To bridge this gap, we employ LoRA as a controlled memory capacity probe within the latent space to systematically quantify exact parametric memory. We introduce the Parametric Memory Law, a robust power law linking loss reduction Delta L to effective parameters and sequence length. At the token level, fine-grained analysis reveals a deterministic phase transition, demonstrating that a prediction probability of p > 0.5 constitutes a sufficient condition for verbatim recall under greedy decoding. Driven by these insights, we introduce MemFT, a threshold-guided optimization strategy that dynamically redistributes the training budget toward sub-threshold tokens. Empirical evaluations demonstrate that MemFT can enhance memory fidelity and efficiency. Code will be released at https://github.com/zjunlp/ParametricMemoryLaw.
Comment: A power law relating loss reduction to effective LoRA parameters and sequence length, plus a token-level p>0.5 phase transition for verbatim recall that drives a budget-redistributing finetuning rule.
Topic Match: Quantifies the memory capacity of low-rank adaptation as a scaling law, which is squarely the low-rank-adaptation strand of the efficiency topic.
Relevance: 6 Novelty: 6
10. Bastion: Budget-Aware Speculative Decoding with Tree-structured Block Diffusion Drafting
ArXiv ID: 2605.29727
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Soowon Oh, Nam Cao, Yujin Kim, Hojung Jung, Huzama Ahmad, Sangmin Bae, Se-Young Yun
Abstract: Block-diffusion drafters have recently emerged as a powerful alternative for speculative decoding by predicting multiple future-token distributions in a single parallel step. However, since these parallel predictions are sampled from position-wise marginals rather than fully conditioned sequences, committing to a single greedy path often fails to capture the target model's preferred trajectory. To address this, we propose BASTION, a budget-aware speculative decoding framework with tree-based diffusion drafting. Unlike existing methods that rely on static tree topologies, BASTION dynamically constructs query-dependent trees by balancing draft quality against hardware constraints. Our framework integrates three synergistic components: (1) an acceptance surrogate that estimates expected accepted length via path confidence, (2) an online latency estimator that calibrates a hardware-aware roofline model, and (3) an adaptive best-first expansion that grows the tree until marginal gains no longer justify incremental verification costs. BASTION is training-free, preserves the target model's distribution, and requires no per-setting tuning. Across diverse benchmarks and GPU architectures, BASTION achieves up to a 6.61x speedup over standard autoregressive decoding, outperforming state-of-the-art block-diffusion baselines by 39%.
Comment: Builds query-dependent speculative draft trees from a block-diffusion drafter, expanding until a roofline latency estimator says marginal accepted length no longer pays for verification.
Topic Match: The acceptance surrogate plus hardware-calibrated budget is a genuinely new decoding mechanism, though it changes inference rather than training cost.
Relevance: 6 Novelty: 6
Paper Selection Prompt
System Prompt
You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.
User Prompt
Instructions
Respond in JSONL. Output exactly one JSON object per paper, one per line:
{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}
Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only:
daily_hot,new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return[]. -daily_hotmeans the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. -new_frontiermeans the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.Scoring Criteria
Relevance and Novelty are independent axes. Score both from 1 to 10.
Relevance Scoring
- 9-10: directly centered on the target foundational topics; highest when the core contribution is clearly within them.
- 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
- 5-6: touches the target topics, but the main contribution is elsewhere.
- 3-4: largely outside the target topics, often application-focused or domain-specific.
- 1-2: unrelated.
Important: Broad frontier relevance, major launch status, or daily buzz is not enough for a high Relevance score here. Those cases belong in the hotspot digest unless the paper strongly matches the specialized paper topics.
Novelty Scoring
- 9-10: new paradigm, theory, or major methodological breakthrough.
- 7-8: substantial methodological advance or strong new insight.
- 5-6: meaningful but incremental extension or refinement.
- 3-4: minor, narrow, or mostly engineering or domain-specific improvement.
- 1-2: little originality; mainly standard application of existing methods.
Topic Registry
Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - moe_training: MoE Training - Mixture-of-Experts training: routing, load balancing, expert granularity and shared experts, training stability, MoE parallelism and communication, and MoE kernels. - training_systems: Large-Scale Training Systems and Efficiency - Large-scale training systems and efficiency: distributed training algorithms, optimizers, and communication. - architecture_training: Architecture and Training Dynamics - Core architectural or computational mechanisms, dynamic computation, and training-stability dynamics. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, memory or cache efficiency, and large-scale training systems that materially change cost or behavior.
Papers
[PAPER LIST HERE]
Relevant Topics
This is a training-side feed. The centre of gravity is Mixture-of-Experts training and the large-scale training systems it sits on; architecture and efficiency work is kept when it bears on how a large model is trained or what it costs to train.
Keep a paper when its CORE CONTRIBUTION falls in one of the four topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.
MoE Training (primary) - Keep: expert routing and gating mechanisms; load balancing, with or without auxiliary losses; expert granularity, fine-grained and shared experts; MoE training stability (loss spikes, router collapse, z-loss and related fixes); expert and hybrid parallelism; all-to-all dispatch/combine and the communication schedules MoE needs; MoE kernels and grouped-GEMM style implementations; expert capacity, token dropping, and upcycling dense checkpoints into MoE. - Filter: papers that merely train on top of a MoE model without a new routing, balancing, stability, parallelism, or kernel contribution; "mixture of experts" in the classical ensemble or recommender sense.
Large-Scale Training Systems - Keep: distributed training algorithms (data, tensor, pipeline, sequence, and expert parallelism, FSDP and ZeRO-style sharding); optimizers and preconditioners for large-scale pretraining; communication algorithms, overlap, and collective design; throughput and MFU work that materially changes what a training run costs or whether it converges; scaling-law work that informs how such runs are configured. - Filter: routine infrastructure tuning with no new algorithmic or systems idea.
Architecture and Training Dynamics - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, dynamic or modular computation); optimisation and training-dynamics analysis that explains why large models train the way they do. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight.
Efficiency and Compression - Keep: quantization, sparsity, pruning, low-rank adaptation, KV-cache and memory-efficient designs, and other work that materially changes the cost of running or training a large model, especially when the mechanism is new rather than a tuned variant. - Filter: straightforward application of standard efficiency methods, and deployment or serving engineering with no new idea.
Leave these to the hotspot digest unless the core contribution clearly falls in one of the four topics above: - major model or product releases - agent, tooling, and RAG launches - benchmark, leaderboard, or evaluation-only papers - representation-learning theory, interpretability, and feature-formation analysis - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning) - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains