Previous Day 2026-08-10
Monthly Overview 2026-08
Next Day 2026-08-12

This is a remedial run for missed papers from 08/10/2026 to 08/10/2026.

Results generated on 09/13/2026.

Personalized Daily ArXiv Papers 2026-08-11

Model Metric Usage Papers
Prompt Completion Total Total arXiv Scanned Relevant
local agent CLI Tokens not reported not reported not reported 546 546 39
Cost not reported not reported not reported

Token counts are not reported for this run. 7 of 10 model calls succeeded, 4,110s of model wall clock.

Topic Coverage:

TopicPapers
Large-Scale Training Systems and Efficiency5
Architecture and Training Dynamics18
Efficiency, Compression, and Large-Scale Training16

Table of contents by topic:

Large-Scale Training Systems and Efficiency (5)

  1. SwiftQK: Fast and Communication-Efficient Tensor Parallelism for Query-Key Normalization Authors: Gyudong Kim, Wonjun Han, Young Geun Kim

  2. SeFoRA: Sketch-Aggregated Federated Low-Rank Adaptation with Heterogeneous Client Ranks Authors: Yue Xia, Tayyebeh Jahani-Nezhad, Mayank Bakshi, Rawad Bitar

  3. FedA2L: Adaptive layer-wise learning rate adjustment in decentralized federated learning Authors: Van Truong Vo, Khoa Nguyen, Taehong Kim

  4. Distributed Optimization with Streaming Data: A Temporal Weighting Perspective Authors: Muhammad Faraz Ul Abrar, Nicolò Michelusi, Erik G. Larsson

  5. FEAST: Federated Shared-Space Training for Resource-Heterogeneous Clients Authors: Bostan Khan, Masoud Daneshtalab

Architecture and Training Dynamics (18)

  1. Generalized Convexity and Smoothness via Conjugate Duality: Optimization Theory for Deep Neural Networks Authors: Binchuan Qi

  2. Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference Authors: Burc Gokden

  3. MixFormer: Linear Transformer with Mixture of Memory Experts Authors: Yu Guo, Lei Duan

  4. The Matching Principle: When Does a Training Penalty Cover Deployment Shift? Authors: Vishal Rajput

  5. ZetaGPT: A Reference Implementation of Positional--Encoding--Free State--Space--Attention Language Models Authors: Róisín Luo

  6. GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning Authors: Zhaoxin Yu, Qi Shen, Hengli Li, Zhaowei Zhang, Song-Chun Zhu, Chi Zhang, Zilong Zheng

  7. BDH-CQ: In-Context Learning with Recurrent Latent Reasoning Authors: Björn Engdahl, Adrian Kosowski, Jan Chorowski, Zuzanna Stamirowska, Przemysław Uznański, Junlin Jiang, Rohan Phadke, Remigiusz Kinas, Richard Zhong

  8. Dynamic gain neuromodulation attenuates the stability gap under joint training Authors: Alejandro Rodriguez-Garcia, Anindya Ghosh, Srikanth Ramaswamy

  9. On the global convergence of gradient flow for wide shallow models beyond homogeneous nonlinearities Authors: Romain Petit, Clarice Poon, Gabriel Peyré

  10. A Tight Lower Bound for Smooth Nonconvex Stochastic Optimization with Bounded Gradient Noise Authors: Jikai Jin

  11. Dual Space Preconditioning for Gradient Descent in the Overparameterized Regime Authors: Reza Ghane, Danil Akhtiamov, Babak Hassibi

  12. MoNo: Multiscale Optimal Transport Neural Operator for Solving PDEs on General Geometries Authors: Zijiang Yang, Xiaomeng Wu, Dongmei Fu

  13. The Kuramoto Neural Operator: Learning to Solve PDEs via Coupled Oscillator Dynamics Authors: Petr Badolia, Leonid Obukhov, Dmitry Bylinkin, Aleksandr Beznosikov

  14. Deciding When to Switch: E-Processes for Adaptive Minimax Training for Generative Adversarial Nets Authors: Hyunjoo Kim, Sicheng Wu, Agastya Venkatraman, Guang Lin, Sehwan Kim

  15. Minimum Block Width for Universal Approximation by Residual Neural Networks with Inner Width One Authors: Qi Zhou, Xuan Zhou, Xiao-Song Yang

  16. Deep Residual Echo State Networks: exploring residual orthogonal connections in untrained Recurrent Neural Networks Authors: Matteo Pinna, Andrea Ceni, Claudio Gallicchio

  17. Recurrent Neural Networks Beyond Time: Learning from Multiple Ordered Projections Authors: Vagan Terziyan, Artur Terziian, Oleksandra Vitko

  18. Closing the loop in learning with missing data Authors: Dimitrios Pylorof, Humberto E. Garcia

Efficiency, Compression, and Large-Scale Training (16)

  1. RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation Authors: Yuhan Li, Fangao Zeng, Sicong Kang, Mengfei Xu, Hao Zhou, Wei Li, Pipei Huang, Bingbing Ni

  2. SemPIC: Learning Semantic Position-Independent KV Caches Authors: Hui Xie, Peng Xiao, Yutong Deng, Shuoran Dou, Jian Yang, Jinyang Guo

  3. From Sweep to Seam: Interleaved Cross-Block Post-Training Quantization Authors: Achille Jacquemond, Yuma Ichikawa, Akira Sakai

  4. Goose: Anisotropic Speculation Trees for Training-Free Speculative Decoding Authors: Tao Jin, Phuong Minh Nguyen, Naoya Inoue

  5. NeuroAda: Activating Each Neuron's Potential for Parameter-Efficient Fine-Tuning Authors: Zhi Zhang, Yixian Shen, Congfeng Cao, Ekaterina Shutova

  6. ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention Authors: Xinyan Wang, Xiaogeng Liu, Ming Pei, Chaowei Xiao

  7. Not All Visual Tokens Are Equally Safe to Remove:Consequence-Sensitive Visual Token Compression Authors: Jingbo Wen, Liang He, Mingyu Cao, Haoyu Wang, Minxuan Hu, Kangning Cui, Xilu Wang

  8. Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching Authors: Kristian Schwethelm, Daniel Rueckert, Georgios Kaissis

  9. MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs Authors: Xinfeng Xia, Xiaofeng Hou, Jiacheng Liu, Wenfeng Wang, Mingxuan Zhang, Peng Tang, Chao Li, Minyi Guo

  10. Omni2LoRA: Coherence-Preserving Parametric Memory for Efficient Omni Language Models Authors: Puneet Mathur, Manan Suri, Dinesh Manocha

  11. LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Authors: Zeyu Xu, Xingzhong Hou, Pengkai Guo, Siling Lin, Xiao Xu, Menghua Zhai, Haoyu Chen, Yunke Zhang, Fei Huang

  12. UniDFKD: A Unified Semantic Prior Framework for Architecture-Agnostic Data-Free Knowledge Distillation Authors: Xuewan He, Tong Chu, Zihan Cheng, Yuchen Su, Qianxin Xia, Guoming Lu, Jielei Wang, Wen Li

  13. Rethinking Factor Sharing in Federated LoRA: A Rank-Aware Adaptive Approach Authors: Xinyi Xu, Bingnan Xiao, Shuang Qin, Gang Feng, Tony Q. S. Quek

  14. DUET: A Diversity-Quality Duet of Distillation Experts for Two-Step Video Generation Authors: Zian Li, Litong Gong, Borui Liao, Pengfei Liu, Xinyu Wang, Xinyuan Wei, Yifan Gao, Tiezheng Ge, Muhan Zhang

  15. On the Effect of Sampling Diversity in Scaling LLM Inference Authors: Tianchun Wang, Zichuan Liu, Yuanzhou Chen, Jonathan Light, Weiyang Liu, Haifeng Chen, Xiang Zhang, Wei Cheng

  16. A Tale of Two Temperatures: Simple, Efficient, and Diverse Sampling from Diffusion Language Models Authors: Theo X. Olausson, Metod Jazbec, Xi Wang, Armando Solar-Lezama, Christian A. Naesseth, Stephan Mandt, Eric Nalisnick


Large-Scale Training Systems and Efficiency (5)

1. SwiftQK: Fast and Communication-Efficient Tensor Parallelism for Query-Key Normalization

ArXiv ID: 2608.09160

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Gyudong Kim, Wonjun Han, Young Geun Kim

Abstract: Query-Key Normalization (QK-Norm) improves the training stability and quality of modern Large Language Models (LLMs). However, under Tensor Parallelism (TP), layerwise QK-Norm introduces additional cross-GPU communication because the normalization factor depends on the full hidden vector. We present SwiftQK, a multi-GPU RMSNorm kernel that exchanges only scalar normalization statistics and overlaps the remaining Peer-to-Peer reduction with independent element-wise computation in a deadlock-safe persistent kernel. Evaluations on recent LLMs show that SwiftQK reduces QK-Norm latency by 81.4--93.9% relative to the standard TP QK-Norm using full-vector All-Gather. In end-to-end serving, SwiftQK reduces TPOT on average by 29.5% over the All-Gather-based baseline and by 14.3% over an optimized scalar-aggregation implementation.

Comment: Replaces full-vector QK-Norm all-gather with scalar-statistic exchange and overlapped persistent execution.

Topic Match: The paper directly contributes tensor-parallel communication and kernel design for a stability-critical LLM operation.

Relevance: 8 Novelty: 7


2. SeFoRA: Sketch-Aggregated Federated Low-Rank Adaptation with Heterogeneous Client Ranks

ArXiv ID: 2608.10144

Primary Topic: Large-Scale Training Systems and Efficiency

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Yue Xia, Tayyebeh Jahani-Nezhad, Mayank Bakshi, Rawad Bitar

Abstract: We consider federated parameter efficient fine-tuning of large neural networks with low-rank adaptation (LoRA,~Hu et al.\ 2022). Combining LoRA with federated PEFT introduces challenges absent from either setting alone: clients may use different LoRA ranks, making their factor matrices dimension-incompatible, and factor-wise averaging suffers from a bilinear mismatch. We propose SeFoRA, a sketch-aggregated federated LoRA algorithm in which each client transmits a linear sketch of its local updates, enabling direct aggregation at the federator. As a result, SeFoRA alleviates the bilinear mismatch, and allows for aggregation in a small subspace of the full model. We introduce a rank-homogeneous version called SeFoRA-Ho which allows for direct adapter aggregation in this setting. We prove convergence to a neighborhood of the first-order stationary point at rate $\cO(1/T)$ for the rank-homogeneous setting. Numerical experiments on fine-tuning RoBERTa-Large on GLUE datasets show how our algorithms outperform the state-of-the-art.

Comment: Linear sketches aggregate heterogeneous-rank LoRA updates while alleviating the bilinear mismatch of factor-wise averaging.

Topic Match: The central advance is distributed update aggregation; communication through sketches and low-rank adaptation provide an additional efficiency match.

Relevance: 7 Novelty: 7


3. FedA2L: Adaptive layer-wise learning rate adjustment in decentralized federated learning

ArXiv ID: 2608.09208

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Van Truong Vo, Khoa Nguyen, Taehong Kim

Abstract: Decentralized intelligence systems with heterogeneous devices and limited coordination increasingly rely on decentralized federated learning (DFL). However, DFL suffers from convergence inefficiency under data heterogeneity due to the use of a uniform learning rate (LR) that ignores layer-specific optimization needs. Foundational layers are responsible for maintaining network consensus, while specialized layers adapt to local data characteristics, leading to conflicting gradients and degraded performance under non-IID conditions. To address this fundamental tension, this work introduces FedA2L, a method that dynamically adjusts layer-wise LRs based on model divergence signals. By leveraging local update intensity and network consensus constraints, FedA2L seamlessly integrates into existing DFL protocols without additional communication or coordination. Extensive evaluations across DFL algorithms, various model architectures, and datasets demonstrate that FedA2L achieves up to 4.94 times faster convergence than vanilla DFL and reduces communication rounds by up to 59% compared to scheduler-based baselines. Furthermore, FedA2L exhibits resilience to severe data heterogeneity, larger network sizes, and sparse topologies, reducing communication overhead and establishing it as a versatile optimization tool for resource-constrained or large-scale distributed learning in edge and IoT deployments. The code is released at https://github.com/nclabteam/FedA2L.

Comment: Adapts layer-wise learning rates from local divergence and consensus signals in decentralized training.

Topic Match: It proposes a distributed optimization mechanism that reduces convergence rounds under heterogeneous data.

Relevance: 7 Novelty: 6


4. Distributed Optimization with Streaming Data: A Temporal Weighting Perspective

ArXiv ID: 2608.09565

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Muhammad Faraz Ul Abrar, Nicolò Michelusi, Erik G. Larsson

Abstract: Optimization theory is a widely used tool for intelligent decision-making. While classical optimization deals with fixed, time-invariant objective functions, many modern applications operate in dynamic environments where data arrive sequentially, and the learning objective evolves over time, often under decentralized data and communication constraints. Motivated by these trends, we study decentralized optimization from streaming data through a structured time-varying formulation in which the global objective is a temporally weighted average of losses observed across the network. We analyze multi-iteration decentralized first-order methods, including decentralized gradient descent. For strongly convex and smooth losses, we develop guarantees for the Euclidean-norm \emph{tracking error} through a contraction-mapping viewpoint. The resulting bounds decompose the tracking error into a fixed-point tracking component and a bias term induced by decentralization and data heterogeneity. We specialize our analysis to uniform and exponentially discounted weights, as well as their finite-memory \emph{windowed} counterparts. The bounds explicitly characterize the roles of the temporal weighting rule, per-step iteration budget, step size, and network connectivity. Uniform weighting yields a vanishing fixed-point tracking contribution of order $\mathcal O(1/t)$, whereas discounted and windowed strategies generally induce non-vanishing tracking floors governed by the discount factor and effective memory, respectively. In all cases, decentralization induces an additional non-zero bias floor under a constant step size. Numerical experiments illustrate the predicted trends.

Comment: Analyzes decentralized first-order optimization when streaming observations receive explicit temporal weights.

Topic Match: Its core results concern convergence and bias in distributed optimization under evolving data.

Relevance: 7 Novelty: 6


5. FEAST: Federated Shared-Space Training for Resource-Heterogeneous Clients

ArXiv ID: 2608.09250

Primary Topic: Large-Scale Training Systems and Efficiency

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Bostan Khan, Masoud Daneshtalab

Abstract: Federated learning (FL) must serve devices with varying computational capabilities. A fixed model cannot suit all devices, while training one model per deployment limit is costly. Federated supernet training instead learns one elastic model with differently sized subnetworks, then deploys a suitable one to each device. When client inference budgets differ, however, parameters exclusive to high-cost subnetworks are reachable by fewer clients. We propose FEAST, a federated shared-space training framework that counters this imbalance by jointly training multiple subnetworks within each client's limit. Budget-tailored sub-supernet routing sends only the relevant supernet portion, and sparse aggregation merges the returned parameter slices. The trained supernet directly serves the subnetworks used during federation and supports post-hoc extraction of additional subnetworks without federated retraining. We further show that independently assigning clients' training-data volumes and inference budgets can distort accuracy--inference-cost comparisons in heterogeneous FL simulations, and introduce a one-parameter $γ$-allocation protocol to control this coupling. In our experimental setup, the SuperFedNAS and DeepFedNAS supernet training procedures remain near chance at 25M and reach at most $17.09\%$ at $596$M inference MACs; FEAST reaches $71.06\%$ at $596$M, $2.4$ points above the strongest model-heterogeneous weight-sharing baseline at its largest tier. Across CIFAR-100, CINIC-10, and TinyImageNet-200, FEAST achieves the highest population-averaged accuracy among the evaluated weight-sharing methods when each client receives its largest affordable subnetwork. Sub-supernet routing reduces aggregate model-parameter traffic by $6.8\times$ relative to full-supernet transmission.

Comment: Trains elastic federated supernets through budget-specific routing and sparse slice aggregation.

Topic Match: The core is a distributed training protocol that also reduces communication and model-family training costs.

Relevance: 6 Novelty: 6


Architecture and Training Dynamics (18)

1. Generalized Convexity and Smoothness via Conjugate Duality: Optimization Theory for Deep Neural Networks

ArXiv ID: 2608.09523

Primary Topic: Architecture and Training Dynamics

Authors: Binchuan Qi

Abstract: Deep neural network (DNN) training with stochastic gradient descent (SGD) and its variants achieves strong empirical performance, yet classical optimization theory does not fully explain this success. This limitation arises because conventional analyses rely on assumptions such as differentiability, convexity, or smoothness, which are often violated by DNN objectives. In this paper, we establish a unified optimization framework for DNN training by generalizing classical convexity and smoothness through Legendre functions and convex conjugation. Specifically, we introduce $\mathcal{H}(ψ)$-convexity and $\mathcal{H}(Ψ)$-smoothness, which unify convex and non-convex as well as smooth and non-smooth objectives within a single formalism and reveal a natural duality between generalized smoothness and convexity. Building on these generalized properties, we introduce generalized gradient descent (GD) and generalized SGD through convex conjugation. We theoretically prove that generalized GD admits an optimal learning rate of exactly $1$, and derive rigorous gradient-energy-based convergence rates for both proposed optimizers. We further reformulate DNN training as a composite optimization problem, demonstrating that its convergence relies on jointly reducing the gradient energy and controlling the induced norm of the network Jacobian. To characterize the practical influences of network architectures and training configurations, we introduce the gradient correlation factor and model capacity risk, and quantitatively analyze how architectural designs, batch size, and model capacity shape training convergence. Extensive experiments across diverse network architectures, datasets, optimizers, and loss functions validate our theoretical bounds and demonstrate precise alignment between our theoretical predictions and empirical training dynamics.

Comment: Develops generalized convexity and smoothness tools for analyzing DNN optimization.

Topic Match: The work directly targets optimization theory and training-dynamics explanations for deep networks.

Relevance: 8 Novelty: 8


2. Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference

ArXiv ID: 2608.10288

Primary Topic: Architecture and Training Dynamics

Authors: Burc Gokden

Abstract: The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attention (PLGA), replace the fixed bilinear form of scaled dot-product attention (SDPA) with a learned, input-generated bilinear operator $G_{LM}$, built from a positive tensor $A_{LM}$ by elementwise power laws. The architecture is fully specified, verified against pinned reference releases; claims are labeled theorem, conditional theorem, measurement, or conjecture. Unconditionally: PLGA contains SDPA exactly at $G_{LM}=I$; $A_{LM}$ and $A_P$ are strictly entrywise positive, with Perron-Frobenius structure on $A_{LM}$; the DAG regularizer has the NOTEARS walk-counting form and positivity obstructs exact acyclicity; and, under nonresonance (satisfied by standard rotary frequencies), a commutant criterion identifies which operators preserve relative-position dependence. An inference-collapse theorem: exact input invariance of deductive outputs collapses inference to generalized SDPA with a constant operator. Measured invariance: relative fluctuations of $10^{-6}$ and below; perturbation bounds quantify but do not certify cached inference; the assembled proxy misses the decoding margin. A conditional three-stage mechanism (rotary twirl, concentration, row-map contraction) is measured on a released checkpoint. Blockwise training and scoring under the global Gram are stated with explicit target exposure; on tested samples, block and sequential scoring select identical answers and agree on the published TruthfulQA probability-mass metric within $5\times 10^{-5}$ per item. Self-organized criticality enters as a phenomenological framework with an intrinsic order parameter; open claims become falsifiable conjectures. Selected proof cores are machine-checked in Lean 4.

Comment: Analyzes a learned bilinear attention operator and proves conditions under which it collapses to generalized SDPA.

Topic Match: The core is a mechanistic and theoretical analysis of an attention variant.

Relevance: 8 Novelty: 7


3. MixFormer: Linear Transformer with Mixture of Memory Experts

ArXiv ID: 2608.09468

Primary Topic: Architecture and Training Dynamics

Authors: Yu Guo, Lei Duan

Abstract: State Space Models (SSMs), as a mainstream research direction of linear Transformers, aim to achieve higher efficiency than standard Transformers in long-context modeling. However, existing SSMs suffer from limited input adaptivity and constrained memory capacity, leading to information loss when modeling ultra-long sequences. To address these limitations, we propose MixFormer, a novel linear Transformer that integrates a Mixture-of-Memory-Experts (MoE) mechanism. Specifically, the model maintains differentiated memory states through multiple collaborating memory experts and employs a novel Time-Aware Linear Attention (TALA) mechanism, which leverages learnable exponential decay functions and positional biases to dynamically update memory. This design enables the model to selectively reinforce important historical information while effectively mitigating memory dilution, substantially improving long-range dependency modeling. Experiments on long-sequence text and image generation tasks demonstrate that MixFormer not only achieves significant performance gains but also provides a more sustainable computational backbone for the next generation of web infrastructure.

Comment: Introduces multiple learned memory states and time-aware linear attention for adaptive long-context computation.

Topic Match: Its main contribution is a new linear sequence-model architecture rather than conventional routed feed-forward MoE training.

Relevance: 8 Novelty: 7


4. The Matching Principle: When Does a Training Penalty Cover Deployment Shift?

ArXiv ID: 2605.22800

Primary Topic: Architecture and Training Dynamics

Authors: Vishal Rajput

Abstract: Ordinary training optimises the task loss and then stops. It never pays for internal representation energy: Jacobians can stay large in directions that never helped the label, so even small label-preserving noise throws the model off---a design gap that classical noise-injection theory fixes at second order, but only when applied as default regularisation, which current practice does not do. We make that precise with a Matching Principle: name deployment directions (Sigma_task) and the training penalty Sigma', and ask whether the second covers the first. The no-thinking default is even-spread / isotropic penalty (Sigma' proportional to I)---classical Gaussian / Tikhonov at second order: no axis estimate, no architecture change, and---in a simple linear ridge model---strictly less deployment drift than task-only training, with no coverage miss by construction. When axes are known, matching is sharper; when they are missed, a residual floor remains. Across seven domains a named second-moment penalty beats unregularised training; a controlled illustration recovers match > even-spread > wrong-axis when axes are forced. The ridge theorems are proved; deep nets remain experiments under a specified perturbation. Design rule: fix internal energy by default (even-spread); match when axes are known; treat losses that control representation sensitivity as first-class design.

Comment: Relates deployment robustness to whether training regularization covers the relevant perturbation directions.

Topic Match: It provides a mechanistic account of how representation-sensitive penalties shape training and robustness.

Relevance: 8 Novelty: 7


5. ZetaGPT: A Reference Implementation of Positional--Encoding--Free State--Space--Attention Language Models

ArXiv ID: 2608.09432

Primary Topic: Architecture and Training Dynamics

Authors: Róisín Luo

Abstract: Transformer-based language models rely on self-attention, whose computation is permutation-equivariant and therefore lacks an intrinsic mechanism for representing token order. Existing architectures address this limitation by explicitly incorporating positional information through learned positional embeddings or hand-crafted positional encodings, such as rotary positional encoding (RoPE), treating positional information as an architecturally acquired capability rather than an inherent property of the model. Motivated by the pursuit of positional-encoding-free architectures, this work explores a language model architecture that integrates causal state-space equations to implicitly encode positional information before attention computation. Specifically, each model block applies a causal state-space equation before self-attention, allowing recurrent state dynamics to encode sequential information into token representations. Consequently, subsequent attention layers operate on position-aware representations without requiring explicit positional encodings while retaining the expressive modeling capacity of self-attention. We present \textsc{ZetaGPT}, a compact hybrid language model designed for research, rapid prototyping, algorithm verification, and educational applications. In addition to the proposed architecture, \textsc{ZetaGPT} provides a fully open-source, end-to-end training pipeline encompassing dataset construction, tokenizer training, pretraining, supervised fine-tuning, reinforcement learning from human feedback (RLHF), and chain-of-thought (CoT) reasoning via pure reinforcement learning. To the best of our knowledge, \textsc{ZetaGPT} is the first open-source small language model without explicit positional encoding and establishes a compact, reproducible reference implementation for the development and empirical study of positional-encoding-free language models.

Comment: Precedes attention with causal state-space dynamics to eliminate explicit positional encoding.

Topic Match: The paper centers on a hybrid state-space-attention architectural mechanism.

Relevance: 8 Novelty: 6


6. GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning

ArXiv ID: 2608.02585

Primary Topic: Architecture and Training Dynamics

Authors: Zhaoxin Yu, Qi Shen, Hengli Li, Zhaowei Zhang, Song-Chun Zhu, Chi Zhang, Zilong Zheng

Abstract: Optimization-based latent reasoning improves large language model outputs by optimizing instance-specific continuous states at test time while keeping model parameters frozen. Existing methods, however, typically connect these states to the reasoning trajectory through decoded tokens, making sequence-level credit assignment indirect and obscuring how latent updates shape subsequent reasoning. We introduce GradCuit (gradient through circuit), which inserts optimizable latent states at a selected Transformer layer between the hidden representations of the prompt and the generated continuation. Causal self-attention provides every continuation-token log-probability with a differentiable path to every preceding latent state through the remaining Transformer blocks, enabling reward-weighted gradients from the entire continuation to be assigned directly to the latents. Across five instruction-tuned backbones, three reasoning benchmarks, and two answer formats, GradCuit achieves an average accuracy of 64.5%, outperforming chain-of-thought prompting by 6.6 percentage points and the strongest competing method by 2.4 points. GradCuit also demonstrates greater robustness: across seven learning-rate settings, it consistently outperforms LatentSeek while reducing the standard deviation of accuracy from 1.53 to 0.82, and even its random-walk variant remains competitive with LatentSeek. For interpretability, token-level gradient attribution reveals that latent influence concentrates on reasoning-connector tokens, while layer analysis identifies early-to-middle Transformer layers as the most effective optimization space. By directly optimizing internal reasoning from outcome feedback, GradCuit opens a new axis of robust and interpretable test-time scaling, where LLMs adapt how they reason rather than merely regenerate, sample, or rerank outputs.

Comment: Adds optimizable intermediate latents with direct sequence-level gradient paths for test-time computation.

Topic Match: The central contribution is a new dynamic-computation mechanism inside Transformer layers.

Relevance: 7 Novelty: 7


7. BDH-CQ: In-Context Learning with Recurrent Latent Reasoning

ArXiv ID: 2608.09888

Primary Topic: Architecture and Training Dynamics

Authors: Björn Engdahl, Adrian Kosowski, Jan Chorowski, Zuzanna Stamirowska, Przemysław Uznański, Junlin Jiang, Rohan Phadke, Remigiusz Kinas, Richard Zhong

Abstract: We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning. Inputs presented at inference time continuously update the model's recurrent memory; the model then solves a query through iterative computation in a high-dimensional latent space, without verbalizing its intermediate reasoning. We evaluate the model on the public ARC-AGI-1 evaluation set and use controlled ARC-like interventions to study what it learns from demonstrations, how consistently it applies an inferred transformation, and which concepts remain difficult. A 150M-parameter configuration reaches 29.5% pass@2 at a computed inference cost of \$0.0007 per task. This operating point breaks through the previously reported ARC-AGI-1 cost-accuracy Pareto frontier, establishing a new state of the art in benchmark cost efficiency.

Comment: Combines recurrent memory with iterative latent computation in a compact reasoning architecture.

Topic Match: Recurrent state and latent iterative computation are the paper's central architectural mechanisms.

Relevance: 7 Novelty: 7


8. Dynamic gain neuromodulation attenuates the stability gap under joint training

ArXiv ID: 2507.14056

Primary Topic: Architecture and Training Dynamics

Authors: Alejandro Rodriguez-Garcia, Anindya Ghosh, Srikanth Ramaswamy

Abstract: Recent work in continual learning has highlighted the stability gap -- a temporary performance drop on previously learned tasks when new ones are introduced. This phenomenon reflects a mismatch between rapid adaptation and strong retention at task boundaries, underscoring the need for optimization mechanisms that balance plasticity and stability over abrupt distribution changes. While optimizers such as momentum-SGD and Adam introduce implicit multi-timescale behavior, they still exhibit pronounced stability gaps. Importantly, these gaps persist even under ideal joint training, making it crucial to study them in this setting to isolate their causes from other sources of forgetting. Motivated by how noradrenergic (neuromodulatory) bursts transiently increase neuronal gain under uncertainty, we introduce a dynamic gain scaling mechanism as a two-timescale optimization technique that balances adaptation and retention by transiently increasing the effective update magnitude while dynamically reparameterizing the weights governing the forward pass, thereby empirically mitigating transition-induced curvature amplification. Across domain- and class-incremental MNIST, CIFAR, and mini-ImageNet benchmarks under task-agnostic joint training, dynamic gain scaling effectively attenuates stability gaps while maintaining competitive accuracy, improving robustness at task transitions.

Comment: Introduces dynamic gain scaling to reduce transition-induced training instability.

Topic Match: The paper centers on an optimization mechanism and analysis of stability-gap dynamics.

Relevance: 7 Novelty: 7


9. On the global convergence of gradient flow for wide shallow models beyond homogeneous nonlinearities

ArXiv ID: 2605.10775

Primary Topic: Architecture and Training Dynamics

Authors: Romain Petit, Clarice Poon, Gabriel Peyré

Abstract: A surprising phenomenon in the training of neural networks is the ability of gradient descent to find global minimizers of the training loss despite its non-convexity. Following earlier work, we investigate this behavior for wide shallow models. Existing global convergence results primarily concern models with positively one-homogeneous nonlinearities, such as ReLU activations, and models with scalar output weights and bounded nonlinearities, such as sigmoid activations. We study a broader class of models, including multi-head attention layers and two-layer networks with bounded or asymptotically positively one-homogeneous activations and vector output weights. Building upon [Chizat and Bach, 2018], we prove that, in the limit of many hidden neurons or attention heads, non-global minimizers of the training loss are unstable under mean-field gradient flow dynamics by constructing "escape regions" in the parameter space. Our global convergence statements are conditional in the following sense: if the mean-field gradient flow converges in W2, then its limit must be a global minimizer. We revisit the bounded nonlinearity, scalar-output setting of [CB18], giving an escape region construction adapted to unbounded nonlinear parameter domains. We also propose new constructions for nonlinearities with at most linear growth under a non-degeneracy assumption and for asymptotically positively one-homogeneous nonlinearities. Finally, we show the well-posedness and stability estimates for the mean-field training dynamics under sub-Gaussian initializations.

Comment: Extends conditional mean-field global-convergence guarantees to multi-head attention and broader nonlinearities with vector output weights.

Topic Match: Directly analyzes optimization dynamics, although the guarantees concern idealized wide shallow models and require convergence in W2.

Relevance: 7 Novelty: 7


10. A Tight Lower Bound for Smooth Nonconvex Stochastic Optimization with Bounded Gradient Noise

ArXiv ID: 2608.09004

Primary Topic: Architecture and Training Dynamics

Authors: Jikai Jin

Abstract: We prove a sharp lower bound for smooth nonconvex stochastic optimization with uniformly bounded gradient noise. In the (K=1) fresh-sample model, every randomized adaptive algorithm requires $$Ω\left( \frac{ΔL}{ε^2} + \frac{ΔLσ^2}{ε^4} \right)$$ queries to find a point with expected gradient norm at most (ε). This matches the standard upper bound and, to the best of our knowledge, resolves the question raised by [Arjevani et al. 2023] of whether almost-surely bounded oracle error permits a better rate than bounded variance. The proof was independently generated with GPT-5.6 Sol in Codex's Ultra mode during a two-hour session. The human author supplied the prompt and was responsible only forchecking the proof and revising and polishing the manuscript.

Comment: Proves a tight query lower bound for stochastic nonconvex optimization with bounded gradient noise.

Topic Match: The strongest fit is foundational optimization theory relevant to training dynamics.

Relevance: 6 Novelty: 8


11. Dual Space Preconditioning for Gradient Descent in the Overparameterized Regime

ArXiv ID: 2603.10485

Primary Topic: Architecture and Training Dynamics

Authors: Reza Ghane, Danil Akhtiamov, Babak Hassibi

Abstract: In this work, we study the convergence properties of the Dual Space Preconditioned Gradient Descent, encompassing optimizers such as Normalized Gradient Descent and Gradient Clipping. We consider preconditioners of the form $\nabla K$, where $K: \mathbb{R}^{d \times k} \to \mathbb{R}$ is convex and apply $\nabla K(\cdot)$ to train an over-parameterized linear model with a convex loss of the form $\ell(X W - Y)$, for weights $W \in \mathbb{R}^{d \times k}$, labels $Y \in \mathbb{R}^{n \times k}$ and data $X \in \mathbb{R}^{n \times d}$. Under the aforementioned assumptions, we prove that the iterates of the full-batch preconditioned gradient descent converge at an exponential rate to a point $W_{\infty} \in \mathbb{R}^{d \times k}$ satisfying $XW_{\infty} = Y$. We also study the implicit bias of Dual Space Preconditioned Gradient Descent. First, we demonstrate analytically and empirically that, for general $K(\cdot)$, $W_\infty$ depends on the chosen constant step size, hindering a precise characterization of the implicit bias. We also provide an approximate implicit bias property for general preconditioners, namely, $|W_0 - W_{\infty}|F \le c |W_0 - W|}, \inftyF$ for a constant $c>0$ and $W$ denoting the convergence point of GD initialized at $W_0$. Furthermore, for preconditioners of the form $K(G) = h(|G|}, \inftyF)$, known as {\it isotropic preconditioners}, and for the stochastic variation of the algorithm with arbitrary batch-size, we prove linear convergence to $W$. Finally, in the experiments, we demonstrate faster convergence on a nonlinear model obtained using the smoothed matrix elastic-net as a preconditioner.}, \infty

Comment: Characterizes convergence and implicit bias for dual-space preconditioned gradient methods.

Topic Match: The main contribution is optimization-dynamics theory for preconditioned learning.

Relevance: 6 Novelty: 7


12. MoNo: Multiscale Optimal Transport Neural Operator for Solving PDEs on General Geometries

ArXiv ID: 2608.09764

Primary Topic: Architecture and Training Dynamics

Authors: Zijiang Yang, Xiaomeng Wu, Dongmei Fu

Abstract: Transformer-based neural operators have achieved substantial progress in solving Partial Differential Equations (PDEs) by projecting spatial observations into compact latent tokens and learning physical interactions in latent spaces. However, we reveal that existing learnable projection mechanisms cannot ensure stable and balanced assignments from observation points to latent tokens, causing some latent tokens to be over-assigned while others remain underutilized. This limitation further restricts the design of hierarchical architectures, as assignment imbalance is continuously inherited and amplified across latent spaces, eventually causing severe token collapse in deeper spaces. To address these issues, we propose MoNo (Multiscale Optimal Transport Neural Operator), a progressive multiscale neural operator that efficiently solves PDEs on general geometries through stable latent-space construction. At its core is CoTAP (Cross-scale Optimal Transport Assignment and Projection), a novel latent-space construction method that formulates cross-space assignment between adjacent spaces as an entropy-regularized optimal transport problem, thereby constructing balanced bidirectional projections and stable latent spaces. CoTAP also ensures stable information transfer across multiple latent spaces, further enabling multiscale architectures on general geometries, which in turn support more efficient learning of long-range physical interactions. Extensive experiments demonstrate that MoNo outperforms existing state-of-the-art neural operators in both prediction performance and computational efficiency. Code is available at https://github.com/ZijiangY1116/MoNo.

Comment: Uses entropy-regularized optimal transport to prevent latent-token assignment collapse across scales.

Topic Match: The balanced multiscale latent projection is a new architectural mechanism, despite its PDE focus.

Relevance: 6 Novelty: 7


13. The Kuramoto Neural Operator: Learning to Solve PDEs via Coupled Oscillator Dynamics

ArXiv ID: 2608.10234

Primary Topic: Architecture and Training Dynamics

Authors: Petr Badolia, Leonid Obukhov, Dmitry Bylinkin, Aleksandr Beznosikov

Abstract: Operator learning is a rapidly advancing area of computational science. It is particularly well suited to problems where a partial differential equation (PDE) must be solved repeatedly under varying physical configurations. Most existing architectures represent the solution operator in a fixed basis. While this assumption is well aligned with global structures, it is less suitable for phenomena governed by local interactions in physical space. We explore an alternative perspective motivated by the observation that the continuum limit of coupled oscillator systems can describe a broad class of PDEs. Building on this idea, we introduce the Kuramoto Neural Operator (KNO), which represents the solution through the evolution of a latent field of interacting oscillators. Across a diverse collection of PDE benchmarks, KNO achieves strong predictive performance, with improvements over competing approaches. Our experimental evaluation also includes an extensive ablation study that quantifies the contribution of each architectural component incorporated into KNO. Furthermore, we show that the model's prediction error is closely linked to the collective dynamics of the latent oscillators. It varies systematically with their degree of synchronization, providing insights into the underlying mechanisms.

Comment: Represents neural-operator computation through interacting latent oscillators with synchronization-linked error.

Topic Match: The latent oscillator dynamics constitute a new analyzed architectural mechanism.

Relevance: 6 Novelty: 7


14. Deciding When to Switch: E-Processes for Adaptive Minimax Training for Generative Adversarial Nets

ArXiv ID: 2608.10096

Primary Topic: Architecture and Training Dynamics

Authors: Hyunjoo Kim, Sicheng Wu, Agastya Venkatraman, Guang Lin, Sehwan Kim

Abstract: Modern data science increasingly gives rise to hypothesis-testing problems that are not naturally formulated in terms of parameters within prespecified statistical models. One important example is the dynamic evaluation of optimization algorithms, where decisions must be made during training about whether further updates remain beneficial or the algorithm should switch to a different phase. This issue is particularly relevant in stochastic min-max optimization. Generative adversarial networks (GANs) provide a canonical example, as their training requires repeated decisions about when to switch between discriminator and generator updates, yet existing methods typically rely on fixed update ratios or heuristic criteria. We formulate this switching problem as sequential hypothesis testing and develop an e-process-based adaptive training procedure. During discriminator updates, one e-process tests the null that the discriminator-induced separation between the empirical data distribution and the generator law remains below a target level. During generator updates, with the discriminator fixed, a second e-process tests the reverse null that this separation remains above a refresh level. Conditional on the observed training sample, we prove that fresh empirical indices and latent draws yield conditional e-values that can be accumulated into e-processes, providing anytime-valid Type I error control under adaptive model updates and data-dependent switching. Across multimodal synthetic distributions and image benchmark datasets, the proposed method matches or outperforms the best fixed-ratio baselines under several widely used GAN objectives.

Comment: Uses anytime-valid e-processes to adaptively switch GAN generator and discriminator updates.

Topic Match: The strongest fit is a new optimization-control mechanism for unstable minimax training dynamics.

Relevance: 6 Novelty: 7


15. Minimum Block Width for Universal Approximation by Residual Neural Networks with Inner Width One

ArXiv ID: 2607.04597

Primary Topic: Architecture and Training Dynamics

Authors: Qi Zhou, Xuan Zhou, Xiao-Song Yang

Abstract: In this paper, we study the universal approximation property of residual neural networks. For input and output dimensions $d_x$ and $d_y$, and LeakyReLU, ReLU, ReLU-like activation functions, the upper and lower bounds of the minimum block width are established. To achieve $L^p$ approximation $(1\leq p <+\infty)$ on any compact set, we show that the exact minimum block width is $\max{d_x,d_y}$ when each residual branch has inner width 1. Furthermore, we show that residual neural networks with block width $\min{d_x+d_y, \max{2d_x+1,d_y}}$ can achieve uniform approximation on any compact set under the constraint that each residual branch has inner width 1. Besides, for any activation function family, we prove that there exist functions that cannot be approximated by residual neural networks with block width less than $\max{d_x, d_y}$, both in the $L^p$ sense and the uniform sense, regardless of inner width. Consequently, for LeakyReLU, ReLU, ReLU-like activation functions and $d_y\geq 2d_x+1$, the exact minimum block width for uniform approximation is $d_y$ when each residual branch has inner width 1.

Comment: Establishes exact or bounded minimum widths for universal approximation by narrow residual blocks.

Topic Match: The work directly analyzes how residual architectural width controls representational capacity.

Relevance: 6 Novelty: 7


16. Deep Residual Echo State Networks: exploring residual orthogonal connections in untrained Recurrent Neural Networks

ArXiv ID: 2508.21172

Primary Topic: Architecture and Training Dynamics

Authors: Matteo Pinna, Andrea Ceni, Claudio Gallicchio

Abstract: Echo State Networks (ESNs) are a particular type of untrained Recurrent Neural Networks (RNNs) within the Reservoir Computing (RC) framework, popular for their fast and efficient learning. However, traditional ESNs often struggle with long-term information processing. In this paper, we introduce a novel class of deep untrained RNNs based on temporal residual connections, called Deep Residual Echo State Networks (DeepResESNs). We show that leveraging a hierarchy of untrained residual recurrent layers significantly boosts memory capacity and long-term temporal modeling. For the temporal residual connections, we consider different orthogonal configurations, including randomly generated and fixed-structure, and study their effect on network dynamics. A thorough mathematical analysis outlines necessary and sufficient conditions to ensure stable dynamics within DeepResESN. Empirically, the proposed approach consistently outperforms traditional shallow and deep RC on a variety of time series tasks. Overall, DeepResESN offers a promising approach for designing hierarchical ESNs with better prediction accuracy on long sequences, without sacrificing the computational advantages that make RC attractive.

Comment: Introduces temporally residual, orthogonally connected reservoirs with analyzed stability conditions.

Topic Match: The core contribution is a recurrent architectural design coupled with stability analysis.

Relevance: 6 Novelty: 6


17. Recurrent Neural Networks Beyond Time: Learning from Multiple Ordered Projections

ArXiv ID: 2608.09690

Primary Topic: Architecture and Training Dynamics

Authors: Vagan Terziyan, Artur Terziian, Oleksandra Vitko

Abstract: Recurrent neural networks (RNNs) are widely used for sequence learning, yet their application is commonly associated with temporal data, although recurrent computation fundamentally operates on ordered sequences rather than on time itself. Building on this observation, we introduce the Ordered Structural Dependency Hypothesis (OSDH), which proposes that multiple admissible orderings of the same observations may reveal complementary structural dependencies inaccessible through a single sequential organization. To operationalize this hypothesis, we propose the Independent Structural Expert Principle (ISEP), whereby projection-specific sequence models are trained independently before their learned representations are integrated through a dedicated fusion model. As a concrete realization, we present Structural Evolution RNNs (SE-RNNs), which employ conventional RNNs as projection-specific structural experts while preserving the underlying recurrent computation unchanged. Proof-of-concept experiments on three synthetic datasets with substantially different levels of structural complexity demonstrate that the proposed architecture consistently benefits from multiple ordered projections when hidden structural dependencies are present, while remaining competitive on simpler datasets. Since OSDH is independent of the underlying sequence-processing model, the proposed framework naturally extends beyond recurrent networks and may be instantiated using alternative architectures. The results suggest a general computational perspective for exploiting complementary ordered representations across diverse structured learning problems.

Comment: Combines independently trained recurrent experts over multiple orderings of the same observations.

Topic Match: The contribution is a modular recurrent architecture for exploiting complementary sequence orderings.

Relevance: 6 Novelty: 6


18. Closing the loop in learning with missing data

ArXiv ID: 2608.09030

Primary Topic: Architecture and Training Dynamics

Authors: Dimitrios Pylorof, Humberto E. Garcia

Abstract: What should a machine learning model learn when data is missing during training? We look at the learning process from a dynamical systems perspective, cast data missingness as a structured loss of actuation that limits controllability of the parameter error dynamics, and ultimately derive adaptation mechanisms with Lyapunov stability characteristics that throttle model updates in ways that preserve learning coherence under partial, intermittent observability. Under recurrent excitation, our analysis provides ISS-type residual-to-state bounds with respect to a bounded closed-loop mismatch between the loss residual and the preconditioned update geometry. We evaluate the efficacy of our directional observability-aware adaptive learning approach on multimodal contexts, reinforcing its premise in promoting learning coherence and stability even in pathologically sparse domains and problems.

Comment: Derives observability-aware update rules with Lyapunov and ISS-style stability bounds for training under intermittent missing data.

Topic Match: Optimization stability is the core contribution, although its connection to large-model pretraining remains indirect.

Relevance: 6 Novelty: 6


Efficiency, Compression, and Large-Scale Training (16)

1. RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation

ArXiv ID: 2608.09226

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Yuhan Li, Fangao Zeng, Sicong Kang, Mengfei Xu, Hao Zhou, Wei Li, Pipei Huang, Bingbing Ni

Abstract: Efficient text-to-image generation requires both reinforcement-learning (RL)-based reward alignment and few-step distillation, yet these procedures are typically performed sequentially, increasing training cost and risking the loss of reward gains during compression. We instead take an RL-native perspective: diffusion RL already generates reward-scored finite-step trajectories, whose intermediate states provide a natural source of distillation supervision rather than a disposable byproduct of sampling. Based on this insight, we propose REST (Reward-Enhanced Scored-Trajectory Distillation), a single-stage RL-distillation co-training framework that attaches a decoupled student to an arbitrary RL teacher. The student learns segment-wise from the teacher's evolving rollout trajectories while leaving the original teacher optimization unchanged. To prevent uniform imitation from preserving undesirable low-reward behaviors, we further introduce Advantage-Modulated Distillation (AMD), which transforms rollout advantages into signed weights over a base distillation loss. AMD strengthens supervision from preferred trajectories and mildly repels the student from low-reward ones. The resulting framework is lightweight and plug-and-play, requires no extra image rollouts, no separate distillation dataset, and no adversarial training. Experiments on compositional generation, visual text rendering, and human-preference alignment show that REST enables few-step CFG-free inference that matches or surpasses its 40-step RL teacher, with an overall additional training cost below 25% over pure RL. REST improves DrawBench PickScore over RTDMD by 0.82 while requiring only one-fifth of the training iterations.

Comment: Co-trains few-step diffusion students directly from reward-scored teacher trajectories.

Topic Match: Few-step distillation is the central mechanism and materially reduces generative-model inference cost.

Relevance: 8 Novelty: 8


2. SemPIC: Learning Semantic Position-Independent KV Caches

ArXiv ID: 2607.28069

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Hui Xie, Peng Xiao, Yutong Deng, Shuoran Dou, Jian Yang, Jinyang Guo

Abstract: Long-context retrieval and agentic workloads repeatedly reuse the same documents under changing instructions, histories, and document orders. Prefix caching cannot exploit this reuse, while position-independent caching (PIC) remains unreliable because independently compiled KV states lack the future context in which they will be consumed. Our diagnostics show that a learned boundary-conditioned baseline sharply reduces attention deviation near reusable-block boundaries but leaves interior and task-level residuals, motivating adaptation of the document representation itself. We present \emph{SemPIC}, which trains a LoRA-enabled Writer to compile native per-layer document KVs through behavioral distillation while retaining the pretrained decoder as an unchanged Reader. Adaptation is confined to offline cache construction, preserving the standard KV interface and cache-hit decoding path. We further introduce KV Gradient Checkpointing, which reduces peak training memory without severing gradients through cached KVs. Across three models and four tasks, SemPIC raises mean micro-F1 over KV Packet from 0.53 to 0.60, approaching Full Recompute at 0.62. Code: https://github.com/jn12-29/SemPIC

Comment: Trains position-independent document KV states through behavioral distillation while preserving standard cache interfaces.

Topic Match: The work introduces a reusable learned KV-cache representation and a memory-saving training procedure.

Relevance: 8 Novelty: 8


3. From Sweep to Seam: Interleaved Cross-Block Post-Training Quantization

ArXiv ID: 2608.09595

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Achille Jacquemond, Yuma Ichikawa, Akira Sakai

Abstract: Compressing large language models to two bits or fewer is increasingly feasible through block-wise post-training quantization; cross-block variants reconstruct neighboring Transformer blocks within a moving window. In the fixed two-block setting studied here, the matched sequential baseline moves this window through the network once, so errors introduced early in the sweep are not revisited. We propose Interleaved Cross-Block Quantization (ICBQ), a scheduling modification that revisits the boundary pair between consecutive chunks. Each seam pair is refined twice: first at the end of one chunk and again at the start of the next. The method retains the local two-block objective and reuses the calibration inputs of existing block-wise PTQ pipelines. Under stated local contraction and smoothness assumptions, we derive a depth-wise upper-bound comparison in which seam revisits multiply the propagated term while the residual remains bounded independently of depth. In the reported experiments, ICBQ reduces ternary-quantization perplexity relative to the matched Sequential CBQ baseline, yields finite perplexity in configurations where the baseline has severe degradation, and can also be used with 3-bit and 2-bit GPTQ.

Comment: Revisits chunk-boundary block pairs to limit error propagation in ultra-low-bit quantization.

Topic Match: The core contribution directly improves two-bit and ternary post-training quantization of LLMs.

Relevance: 9 Novelty: 6


4. Goose: Anisotropic Speculation Trees for Training-Free Speculative Decoding

ArXiv ID: 2604.02047

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Tao Jin, Phuong Minh Nguyen, Naoya Inoue

Abstract: Speculative decoding accelerates large language model inference by drafting multiple candidate tokens and verifying them in a single forward pass. Candidates are organized as a tree: deeper trees accept more tokens per step, but adding depth requires sacrificing breadth (fallback options) under a fixed verification budget. Existing training-free methods draft from a single token source and shape their trees without distinguishing candidate quality across origins. We observe that two common training-free token sources -- n-gram matches copied from the input context, and statistical predictions from prior forward passes -- differ sharply in acceptance rate (~6x median gap, range 2-18x across five models and five benchmarks). We prove that when such a quality gap exists, the optimal tree is anisotropic (asymmetric): reliable tokens should form a deep chain while unreliable tokens spread as wide branches, raising the depth ceiling of balanced trees. We realize this structure in GOOSE, a training-free framework that builds an adaptive spine tree: a deep chain of high-acceptance context-matched tokens with wide branches of low-acceptance alternatives at each node. The resulting tree provably accepts at least as many tokens per step as either source alone. On five LLMs (7B-33B) and five benchmarks, GOOSE achieves 1.9-4.3x lossless speedup, outperforming balanced-tree baselines by 12-33% under the same budget.

Comment: Anisotropic speculation trees materially reduce lossless LLM decoding cost without training.

Topic Match: The core contribution is a new inference-efficiency mechanism under a fixed verification budget.

Relevance: 8 Novelty: 7


5. NeuroAda: Activating Each Neuron's Potential for Parameter-Efficient Fine-Tuning

ArXiv ID: 2510.18940

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Zhi Zhang, Yixian Shen, Congfeng Cao, Ekaterina Shutova

Abstract: Existing parameter-efficient fine-tuning (PEFT) methods primarily fall into two categories: addition-based and selective in-situ adaptation. The former, such as LoRA, introduce additional modules to adapt the model to downstream tasks, offering strong memory efficiency. However, their representational capacity is often limited, making them less suitable for fine-grained adaptation. In contrast, the latter directly fine-tunes a carefully chosen subset of the original model parameters, allowing for more precise and effective adaptation, but at the cost of significantly increased memory consumption. To reconcile this trade-off, we propose NeuroAda, a novel PEFT method that enables fine-grained model finetuning while maintaining high memory efficiency. Our approach first identifies important parameters (i.e., connections within the network) as in selective adaptation, and then introduces bypass connections for these selected parameters. During finetuning, only the bypass connections are updated, leaving the original model parameters frozen. Empirical results on 23+ tasks spanning both natural language generation and understanding demonstrate that NeuroAda achieves state-of-the-art performance with as little as $\leq \textbf{0.02}\%$ trainable parameters, while reducing CUDA memory usage by up to 60%. We release our code here: https://github.com/FightingFighting/NeuroAda.git.

Comment: Fine-tunes selected connections through memory-efficient trainable bypass parameters.

Topic Match: The core contribution is a new parameter-efficient adaptation mechanism with very low trainable parameter and memory costs.

Relevance: 8 Novelty: 7


6. ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention

ArXiv ID: 2603.22016

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Xinyan Wang, Xiaogeng Liu, Ming Pei, Chaowei Xiao

Abstract: Large Reasoning Models (LRMs) often reach a correct solution before their long Chain-of-Thought trace ends, yet continue with redundant verification, repeated attempts, or unnecessary exploration that wastes computation and can even overturn the correct answer. We frame this behavior as a latent productive-to-redundant transition and show it is directly reflected in hidden states: around first-correct-solution (FCS) boundaries, late-layer representations separate efficient from overthinking tokens, while boundary-permutation and position controls collapse. We propose ROM, a streaming intervention framework that monitors a frozen LRM with a lightweight hidden-state detector ($\sim$0.1\% of backbone parameters) and intervenes at well-formed reasoning boundaries; Counterfactual Self-Correction (CSC) balances supervision with wrong$\rightarrow$correct trajectories, preserving useful pre-FCS self-correction. Unlike prior adaptive early-exit methods, ROM extracts no intermediate answers, launches no probe decoding, and updates no backbone weights. Across five backbones from three model families and five reasoning benchmarks, against ten recent baselines under a shared protocol, ROM$_{\text{CSC}}$ attains the highest accuracy in 19 of 25 model--benchmark settings, cuts response length by 28--77\% (mean 45\%) versus vanilla decoding, and is the only method on the accuracy--length Pareto front in every setting. The same MATH500-trained supervision transfers zero-shot across scales, families, and task domains, and end-to-end wall-clock latency drops by 46.5\% with $\sim$5\% per-token overhead. Code is available at https://github.com/SaFo-Lab/ROM.

Comment: A lightweight hidden-state detector intervenes at reasoning boundaries to stop redundant computation online.

Topic Match: Streaming control of reasoning length directly reduces large-model inference cost, supported by reported end-to-end latency improvements.

Relevance: 8 Novelty: 7


7. Not All Visual Tokens Are Equally Safe to Remove:Consequence-Sensitive Visual Token Compression

ArXiv ID: 2608.09176

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Jingbo Wen, Liang He, Mingyu Cao, Haoyu Wang, Minxuan Hu, Kangning Cui, Xilu Wang

Abstract: Visual token compression for vision--language models (VLMs) has largely relied on criteria such as attention, redundancy, and uncertainty to maximize average accuracy under a fixed compute budget, implicitly assuming that all errors carry equal cost. However, the consequence of an incorrect prediction on downstream tasks is rarely symmetric: misreading an invoice amount can be far more costly than misclassifying a background color. Motivated by this, we introduce consequence-sensitive visual token compression, which allocates visual computation across requests according to their potential error costs. Our method follows a calibrate-then-allocate procedure, estimating consequence-specific error-budget curves offline and applying the calibrated token budgets online using consequence signals available from question or task information. On a controlled within-task benchmark, high- and low-consequence questions are drawn from the same document images, so content alone cannot reveal which questions are costly to get wrong. In this setting, our method reduces high-stakes errors from 0.300 to 0.133 under the same total token budget, whereas a content-driven allocator performs no better than uniform allocation. Measuring how error rates change with token budget across different cost ratios, we derive an allocation frontier: uniform allocation is optimal when errors are equally costly, and token transfer toward high-consequence questions becomes increasingly beneficial as the cost gap grows. This allocation principle generalizes well across three dense vision-language benchmarks, two budget realization mechanisms (token deletion and resolution reallocation), two VLM architectures, and multiple token selection strategies. On a realistic mixed workload, consequence-sensitive allocation reduces cost-weighted error by 38% while achieving approximately 21% lower latency than full-resolution inference.

Comment: Allocates visual-token budgets according to the consequences of potential errors.

Topic Match: The central mechanism is adaptive token compression that reduces latency under a fixed compute budget.

Relevance: 7 Novelty: 7


8. Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching

ArXiv ID: 2608.09444

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Kristian Schwethelm, Daniel Rueckert, Georgios Kaissis

Abstract: A main promise of looped language models (LMs) is depth-adaptive inference. By iterating a block of shared layers a variable number of times, the model can use less compute for "easy" tokens and more for "hard" ones. However, this adaptivity breaks standard batching: tokens in the same batch now require a different number of loops, so there is no unified forward pass, making efficient inference difficult. Standard inference frameworks like vLLM schedule on the token level and cannot handle this because tokens need to be removed from the batch within the forward pass. Loop-level scheduling has been proposed as a solution, but never implemented end to end. The key challenge is that looped architectures also contain non-looped boundary stages (e.g., token embedding and LM head) that must be scheduled at different frequencies than the loop. We introduce continuous depth batching (CDB), which schedules at the granularity of individual loop iterations. CDB handles boundary stages and loop steps in separate priority queues, makes exit decisions one step ahead, and overlaps all scheduling work with GPU computation. On Ouro 1.4B and Huginn 3.5B, CDB can realize up to $99\%$ of the theoretical maximum speed-up from adaptive-depth, translating to $1.5$-$1.9\times$ higher offline throughput and $45$-$90\%$ lower normalized latency under dynamic serving load.

Comment: Implements continuous batching at individual loop iterations for adaptive-depth language models.

Topic Match: The central contribution realizes the computational savings of a dynamic-depth architecture through a new scheduling design.

Relevance: 7 Novelty: 7


9. MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs

ArXiv ID: 2510.19366

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: MoE Training

Authors: Xinfeng Xia, Xiaofeng Hou, Jiacheng Liu, Wenfeng Wang, Mingxuan Zhang, Peng Tang, Chao Li, Minyi Guo

Abstract: Mixture-of-Experts (MoE) scales model capacity through sparse activation, and is becoming an important architecture for large language models (LLMs). However, existing MoE serving systems typically execute all requests under a fixed routing configuration, limiting their ability to exploit heterogeneous computation requirements across requests. Routing top-$k$, which determines the number of routed experts activated per token, directly controls routed-expert computation and provides a natural mechanism for request-level compute elasticity. Realizing this capability, however, requires finer-grained routing units and efficient runtime execution for heterogeneous routing budgets. We present \textsc{MoE-Prism}, a model and system support framework for request-level compute elasticity in MoE serving. \textsc{MoE-Prism}decomposes monolithic experts into fine-grained sub-experts to expose denser routing operating points and provides a $k$-aware serving runtime that effectively serves heterogeneous routing budgets under both throughput-oriented and latency-sensitive workloads. We implement \textsc{MoE-Prism} on top of vLLM and evaluate it on three representative MoE models. \textsc{MoE-Prism} expands the number of available routing operating points by $4\times$, improves offline inference throughput by up to 33.9\%, and reduces online serving TTFT under heterogeneous workloads. These results demonstrate practical elastic MoE serving with request-level routing targets.

Comment: Decomposes monolithic experts into fine-grained sub-experts with request-adaptive top-k execution.

Topic Match: Its primary objective is elastic inference cost, while expert granularity and routing connect it directly to MoE design.

Relevance: 7 Novelty: 7


10. Omni2LoRA: Coherence-Preserving Parametric Memory for Efficient Omni Language Models

ArXiv ID: 2608.09227

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Puneet Mathur, Manan Suri, Dinesh Manocha

Abstract: Omnimodal language models (OLMs) enable unified audio-visual understanding, but processing long joint token sequences makes inference computationally prohibitive. While recent token compression methods attempt to alleviate this burden, compressing modalities in isolation often destroys the temporal cross-modal anchors necessary for coherent reasoning. We introduce Omni2LoRA, a two-stage framework for efficient parametric memory compression via coherence-preserving context distillation that bypasses the token bottleneck entirely. First, a Perceiver hypernetwork processes intermediate representations from a frozen OLM to encode the multimodal context into a full-rank Low-Rank Adaptation (LoRA) adapter in a single forward pass. To prevent the resulting parameter footprint from scaling linearly with recording length, we optimize a discrete rank allocation policy via Group Relative Policy Optimization (GRPO) that uses a modality-ablated counterfactual reward to explicitly penalize the loss of audio-visual coherence, forcing the model to allocate its fixed sub-linear rank budget to synergistic cross-modal anchors rather than isolated visual features. Across three omnimodal backbones, Omni2LoRA operating at a 30% rank budget outperforms direct full-context inference and strong token-compression baselines (OmniZip, OMAC, O-MARC) on four audio-visual question answering benchmarks, improving average accuracy by 8-12% over the strongest baseline and remaining stable under compression ratios as tight as 75%, where token-pruning methods degrade sharply. By converting multimodal memory into a fixed-budget, reusable parameter state, our method drives answer-time multimodal-token load to zero, cutting per-query Time to First Token (TTFT) by up to 12x relative to full-context inference and amortizing to under 0.5s after a handful of queries.

Comment: Distills long multimodal context into a fixed-budget, reusable LoRA parameter state.

Topic Match: Parametric context compression is the main mechanism reducing token load, latency, and memory growth.

Relevance: 7 Novelty: 7


11. LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models

ArXiv ID: 2607.16305

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Zeyu Xu, Xingzhong Hou, Pengkai Guo, Siling Lin, Xiao Xu, Menghua Zhai, Haoyu Chen, Yunke Zhang, Fei Huang

Abstract: Vision-Language Models (VLMs) have achieved strong progress in multimodal understanding. However, scaling dense or sparse Mixture-of-Experts (MoE) models to improve performance limits deployment in resource-constrained environments due to the trade-off between high memory usage from full loading and increased latency from on-demand loading. Recently, the Per-Layer Embedding (PLE) architecture addresses this by scaling models with large external embedding tables stored in read-only memory (ROM) and performing lightweight lookup to retrieve relevant embeddings to enhance token representations. Nevertheless, existing PLE-style methods are primarily designed for text embeddings due to the convenience of ID-based retrieval, limiting their effectiveness in VLMs where multimodal embeddings contain richer information for visual tasks. In this paper, we propose LookME, the first framework that enables lookup-based enhancement for multimodal embeddings in VLMs while supporting partitioned storage and on-demand loading. To efficiently lookup arbitrary continuous multimodal embeddings from large-scale embedding tables, we propose a hierarchical two-level lookup method employing a coarse-to-fine strategy that performs lookups from the scene-level to the intra-scene primitive-level. Furthermore, we integrate the lookup method with a sparse injection strategy, which adaptively prioritizes critical embeddings over voluminous multimodal embeddings within layers, and facilitates embedding table reuse across neighboring layers, improving the trade-off among efficiency, model size, and performance. Experiments on multiple visual benchmarks show that LookME outperforms text-only PLE-style methods, validating the effectiveness of lookup-based multimodal embedding enhancement.

Comment: Scales VLM capacity through hierarchical lookup and sparse injection of externally stored embeddings.

Topic Match: Its core contribution reduces loaded parameters and latency through a new memory-efficient lookup architecture.

Relevance: 7 Novelty: 7


12. UniDFKD: A Unified Semantic Prior Framework for Architecture-Agnostic Data-Free Knowledge Distillation

ArXiv ID: 2608.09287

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Xuewan He, Tong Chu, Zihan Cheng, Yuchen Su, Qianxin Xia, Guoming Lu, Jielei Wang, Wen Li

Abstract: Data-Free Knowledge Distillation (DFKD) transfers knowledge from a pretrained teacher model to a compact student model by synthesizing semantically informative data, eliminating the need for access to the original training dataset. Existing DFKD methods rely heavily on architecture-specific statistical priors (e.g., Batch Normalization statistics) to guide data synthesis, however, such architecture-dependent priors are often absent in modern architectures such as Vision Transformers (ViTs), resulting in degraded semantic quality of the synthesized data and consequently catastrophic performance degradation. In this paper, we propose \emph{UniDFKD}, a unified data-free knowledge distillation framework that replaces architecture-specific statistics with explicit, architecture-agnostic semantic priors. \emph{UniDFKD} governs the entire synthesis-distillation pipeline along three dimensions: (1) Categorical Semantic Conditioning (CSC) defines \emph{what} to synthesize by persistently modulating the generator with language-derived embeddings to capture semantic diversity; (2) Spatial Semantic Anchoring (SSA) dictates \emph{where} evidence belongs by anchoring the teacher's spatial attributions to a Gaussian prior; and (3) Spatial Semantic Distillation (SSD) controls \emph{how} knowledge is transferred by explicitly aligning teacher-student spatial evidence alongside predictions. Extensive experiments across CNNs and ViTs demonstrate that UniDFKD establishes a new state-of-the-art, outperforming existing methods by an average absolute margin of over 20\% in both homogeneous and heterogeneous settings.

Comment: Introduces architecture-agnostic data-free distillation using semantic and spatial priors.

Topic Match: Knowledge distillation into compact students is the central compression mechanism.

Relevance: 7 Novelty: 7


13. Rethinking Factor Sharing in Federated LoRA: A Rank-Aware Adaptive Approach

ArXiv ID: 2608.09742

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Large-Scale Training Systems and Efficiency

Authors: Xinyi Xu, Bingnan Xiao, Shuang Qin, Gang Feng, Tony Q. S. Quek

Abstract: Low-rank adaptation (LoRA) represents large language model (LLM) updates with two compact matrix factors, i.e., $A$ and $B$, providing an efficient way to fine-tune large models in federated learning paradigm. Inspired by the asymmetric roles of the LoRA factors, we study whether $A$ should be shared across clients while $B$ remains client-specific (Share-A/Local-B), or whether $B$ should instead be shared while $A$ remains client-specific (Share-B/Local-A). With a least-squares surrogate, we reveal that Share-A/Local-B requires the client-specific LoRA update matrices to use a common rank-$r$ input-side space, whereas Share-B/Local-A requires a common rank-$r$ output-side space. The two strategies therefore incur different projection residuals, indicating that the preferred strategy is the one with the smaller aggregate residual across clients. With this insight, we propose Federated Adaptive Factor Sharing Low-Rank Adaptation (FedAS-LoRA), which selects the sharing side before training to enhance fine-tuning performance. To enable adaptive factor selection before training, we design a Rank-Aware Shared-Subspace Sufficiency (RSS) metric, which effectively assesses whether a shared rank-$r$ input subspace is sufficient for the local data distributions using representations extracted from a frozen LLM backbone. Experiments across different tasks, data distributions, LoRA ranks, and participation settings confirm the effectiveness of RSS and the superior performance of FedAS-LoRA.

Comment: Derives a rank-aware criterion for choosing which LoRA factor is shared across federated clients.

Topic Match: Low-rank factor sharing is the central efficiency mechanism, with federated optimization as its distributed-training setting.

Relevance: 7 Novelty: 6


14. DUET: A Diversity-Quality Duet of Distillation Experts for Two-Step Video Generation

ArXiv ID: 2608.09637

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Zian Li, Litong Gong, Borui Liao, Pengfei Liu, Xinyu Wang, Xinyuan Wei, Yifan Gao, Tiezheng Ge, Muhan Zhang

Abstract: Diffusion models have enabled high-quality video generation in recent years, but the high cost of iterative sampling hinders their practical deployment. Few-step distillation alleviates this cost, yet exposes a quality--diversity trade-off between its two dominant paradigms: trajectory-level distillation (e.g., sCM) favors diversity, whereas distribution-level distillation (e.g., DMD) favors quality. Targeting extreme two-step video generation, we introduce DUET, which reconciles the two paradigms through a noise-level duet of experts: an sCM expert takes the high-noise step to lay out diverse structure, and a DMD expert takes the low-noise step to refine appearance detail. Since the two experts are trained independently with their native objectives, DUET sidesteps the optimization difficulties of loss-level combinations and delivers quality and diversity jointly rather than trading one for the other. We further identify the relay interface and the high-noise stage as the remaining bottlenecks, and address them with RL-guided expert adaptation, yielding DUET+. With the Wan2.1-T2V-1.3B backbone, DUET lifts the two-step quality of sCM close to the level of DMD while retaining nearly all of its structural diversity---about twice that of DMD---and DUET+ further improves overall quality while preserving this diversity advantage. Together, these results establish noise-level expert specialization as a simple, effective paradigm for reconciling diversity and quality in two-step video generation.

Comment: Separately distilled high- and low-noise experts combine complementary objectives within two diffusion sampling steps.

Topic Match: The core method improves extreme few-step distillation, making generation efficiency a substantive match despite its video-specific scope.

Relevance: 7 Novelty: 6


15. On the Effect of Sampling Diversity in Scaling LLM Inference

ArXiv ID: 2502.11027

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Tianchun Wang, Zichuan Liu, Yuanzhou Chen, Jonathan Light, Weiyang Liu, Haifeng Chen, Xiang Zhang, Wei Cheng

Abstract: Large language model (LLM) scaling inference is key to unlocking greater performance, and leveraging diversity has proven an effective way to enhance it. Motivated by the observed relationship between solution accuracy and meaningful response diversity, we systematically study the effect of prompt diversity in scaling inference. We theoretically explain why diversified sampling improves Best-of-$N$ scaling, showing that responses generated from diverse prompts after Best-of-$N$ selection exhibit significantly lower error rates than those produced from stationary prompts. Building on this analysis, we derive a diversity-fidelity trade-off principle, that guides the design of sampling strategies introducing diversity. From this guidance, we instantiate a family of effective perturbation styles. We theoretically and empirically characterize \textbf{when} diversified exploration remains effective, demonstrating that it works under a variety of conditions, and we further show that under majority voting, diversity may vanish. Finally, we systematically evaluate the effectiveness of sampling diversity and show that, when applied appropriately in different contexts, meaningful perturbations yield stronger, task-dependent gains as diversity increases. Overall, this work provides a systematic analysis that offers a theoretical and empirical foundation for understanding how sampling diversity affects LLM inference-time scaling.

Comment: Derives a diversity-fidelity principle for prompt perturbations in best-of-N inference scaling.

Topic Match: The paper analyzes how sampling strategy changes the returns from a fixed inference-compute budget.

Relevance: 6 Novelty: 7


16. A Tale of Two Temperatures: Simple, Efficient, and Diverse Sampling from Diffusion Language Models

ArXiv ID: 2604.09921

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Theo X. Olausson, Metod Jazbec, Xi Wang, Armando Solar-Lezama, Christian A. Naesseth, Stephan Mandt, Eric Nalisnick

Abstract: Much work has been done on designing fast and accurate sampling for diffusion language models (dLLMs). However, these efforts have largely focused on the tradeoff between speed and quality of individual samples; how to additionally ensure diversity across samples remains less well understood. In this work, we show that diversity can be increased by using softened, tempered versions of familiar confidence-based remasking heuristics, retaining their computational benefits and offering simple implementations. We motivate this approach by introducing an idealized formal model of fork tokens and studying the impact of remasking on the expected entropy at the forks. Empirically, the proposed tempered heuristics close the exploration gap (pass@k) between existing confidence-based and autoregressive sampling, hence outperforming both when controlling for cost (pass@NFE). We further study how the increase in diversity translates to downstream post-training and test-time compute scaling. Overall, our findings demonstrate that simple, efficient, and diverse sampling from dLLMs is possible.

Comment: Uses tempered remasking to improve diffusion-LM sample diversity at fixed inference cost.

Topic Match: The main fit is an inference algorithm that improves the quality-diversity trade-off per sampling step.

Relevance: 6 Novelty: 6


Paper Selection Prompt

System Prompt

You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.

User Prompt

Instructions

Respond in JSONL. Output exactly one JSON object per paper, one per line:

{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}

Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only: daily_hot, new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return []. - daily_hot means the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. - new_frontier means the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.

Scoring Criteria

Relevance and Novelty are independent axes. Score both from 1 to 10.

Relevance Scoring

  • 9-10: directly centered on the target foundational topics; highest when the core contribution is clearly within them.
  • 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
  • 5-6: touches the target topics, but the main contribution is elsewhere.
  • 3-4: largely outside the target topics, often application-focused or domain-specific.
  • 1-2: unrelated.

Important: Broad frontier relevance, major launch status, or daily buzz is not enough for a high Relevance score here. Those cases belong in the hotspot digest unless the paper strongly matches the specialized paper topics.

Novelty Scoring

  • 9-10: new paradigm, theory, or major methodological breakthrough.
  • 7-8: substantial methodological advance or strong new insight.
  • 5-6: meaningful but incremental extension or refinement.
  • 3-4: minor, narrow, or mostly engineering or domain-specific improvement.
  • 1-2: little originality; mainly standard application of existing methods.

Topic Registry

Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - moe_training: MoE Training - Mixture-of-Experts training: routing, load balancing, expert granularity and shared experts, training stability, MoE parallelism and communication, and MoE kernels. - training_systems: Large-Scale Training Systems and Efficiency - Large-scale training systems and efficiency: distributed training algorithms, optimizers, and communication. - architecture_training: Architecture and Training Dynamics - Core architectural or computational mechanisms, dynamic computation, and training-stability dynamics. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, memory or cache efficiency, and large-scale training systems that materially change cost or behavior.

Papers

[PAPER LIST HERE]

Relevant Topics

This is a training-side feed. The centre of gravity is Mixture-of-Experts training and the large-scale training systems it sits on; architecture and efficiency work is kept when it bears on how a large model is trained or what it costs to train.

Keep a paper when its CORE CONTRIBUTION falls in one of the four topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.

  1. MoE Training (primary) - Keep: expert routing and gating mechanisms; load balancing, with or without auxiliary losses; expert granularity, fine-grained and shared experts; MoE training stability (loss spikes, router collapse, z-loss and related fixes); expert and hybrid parallelism; all-to-all dispatch/combine and the communication schedules MoE needs; MoE kernels and grouped-GEMM style implementations; expert capacity, token dropping, and upcycling dense checkpoints into MoE. - Filter: papers that merely train on top of a MoE model without a new routing, balancing, stability, parallelism, or kernel contribution; "mixture of experts" in the classical ensemble or recommender sense.

  2. Large-Scale Training Systems - Keep: distributed training algorithms (data, tensor, pipeline, sequence, and expert parallelism, FSDP and ZeRO-style sharding); optimizers and preconditioners for large-scale pretraining; communication algorithms, overlap, and collective design; throughput and MFU work that materially changes what a training run costs or whether it converges; scaling-law work that informs how such runs are configured. - Filter: routine infrastructure tuning with no new algorithmic or systems idea.

  3. Architecture and Training Dynamics - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, dynamic or modular computation); optimisation and training-dynamics analysis that explains why large models train the way they do. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight.

  4. Efficiency and Compression - Keep: quantization, sparsity, pruning, low-rank adaptation, KV-cache and memory-efficient designs, and other work that materially changes the cost of running or training a large model, especially when the mechanism is new rather than a tuned variant. - Filter: straightforward application of standard efficiency methods, and deployment or serving engineering with no new idea.

Leave these to the hotspot digest unless the core contribution clearly falls in one of the four topics above: - major model or product releases - agent, tooling, and RAG launches - benchmark, leaderboard, or evaluation-only papers - representation-learning theory, interpretability, and feature-formation analysis - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning) - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains