This is a remedial run for missed papers from 07/02/2026 to 07/02/2026.
Results generated on 09/12/2026.
Personalized Daily ArXiv Papers 2026-07-03
| Model | Metric | Usage | Papers | ||||
|---|---|---|---|---|---|---|---|
| Prompt | Completion | Total | Total arXiv | Scanned | Relevant | ||
local agent CLI |
Tokens | not reported | not reported | not reported | 374 | 374 | 24 |
| Cost | not reported | not reported | not reported | ||||
Token counts are not reported for this run. 5 of 5 model calls succeeded, 1,599s of model wall clock.
Topic Coverage:
| Topic | Papers |
|---|---|
| Large-Scale Training Systems and Efficiency | 3 |
| Architecture and Training Dynamics | 10 |
| Efficiency, Compression, and Large-Scale Training | 11 |
Table of contents by topic:
Large-Scale Training Systems and Efficiency (3)
-
Revisiting Decentralized Online Convex Optimization with Compressed Communication Authors: Hao Zhou, Xiaoyu Wang, Chang Yao, Mingli Song, Yuanyu Wan
-
Decentralized Stochastic Subgradient-type Methods with Communication Compression for Nonsmooth Nonconvex Optimization Authors: Siyuan Zhang, Nachuan Xiao, Xin Liu
-
Understanding the Robustness of Distributed Self-Supervised Learning Frameworks Against Non-IID Data Authors: Xuanyu Chen, Nan Yang, Shuai Wang, Dong Yuan
Architecture and Training Dynamics (10)
-
A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets Authors: Wanyun Cui
-
Training Hybrid Block Diffusion Language Models with Partial Bidirectionality Authors: Pranshu Chaturvedi, Parth Shroff, Tarun Suresh, Hangoo Kang, Kaiyue Wen
-
The Orthogonalized Read Is a Removable Training Scaffold for Recurrent Memory Authors: Keston Aquino-Michaels
-
Benign Overfitting Does Not Occur in Diffusion Models Authors: Tyler Farghly, Benjamin Dupuis, Alain Durmus, Umut Simsekli
-
Dendritic In-Context Learning in a Single-Layer Spiking Neural Network Authors: Juwei Shen, Yujie Wu, Changwen Chen
-
SUNTA: Hierarchical Video Prediction with Surprise-based Chunking Authors: Tomoshi Iiyama, Masahiro Suzuki, Yutaka Matsuo
-
Born Discrete, Made Smooth: Variational Formulation of Shallow Neural Networks Authors: Matej Benko, Pierre Bousquet, Iwona Chlebicka, Błażej Miasojedow
-
X-LogSMask: Expand Transformer for Graph-Structured Data Authors: Leyan Li, Rennong Yang, Zhenxing Zhang, Liping Hu
-
Self-Gating Attention for Efficient Time Series Forecasting Authors: Dezheng Wang, Tong Chen, Wei Yuan, Congyan Chen, Shihua Li, Hongzhi Yin
-
LiNO: Lifting based multiresolution neural operator Authors: Himanshu Pandey, Subham Patel, Ratikanta Behera
Efficiency, Compression, and Large-Scale Training (11)
-
SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models Authors: Qi Lyu, Jiahua Dong, Baichen Liu, Xudong Wang, Mingfei Han, Yulun Zhang, Fahad Shahbaz Khan, Salman Khan, Lianqing Liu, Zhi Han
-
Lynx: Progressive Speculative Quantization for accelerating KV Transfer in Long-Context Inference Authors: Wenchen Han, Gingfung Matthew Yeung, Marco Barletta, William Toner, Amory Hoste, Adam Barker
-
WBMM: Windowed Batch Matrix Multiplication for Efficient Large Receptive Field Convolution Authors: Wan Song, Wei Zhou, Rui Wang, Jun Yu, Toru Kurihara, Jiajia Xu, Shu Zhan
-
Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning Authors: Xuehui Wang, Xuankun Yang, Wei Shen
-
Spec-AUF: Accept-Until-Fail Training under Train-Inference Misalignment for Masked Block Drafters Authors: Tianjian Yang, Meng Li
-
EPnG: Adaptive Expert Prune-and-Grow for Parameter-Efficient MoE Fine-tuning Authors: Ahin Lee, Sehyun Yun, Taesik Gong
-
LAARA: Layer-Aware Adaptive Rank Allocation for Parameter-Efficient Fine-Tuning Authors: Ashutosh Tripathi, Surya Deep Singh, Pranab Sahoo, Sriparna Saha
-
ART for Diffusion Sampling: Continuous-Time Control and Actor-Critic Learning Authors: Yilie Huang, Wenpin Tang, Xun Yu Zhou
-
PARTREP: Learning What to Repeat for Decoder-only LLMs Authors: Andikawati P Widjaja, Yongjun Kim, Hyounghun Kim, Jaeho Lee
-
Less Tokens, Better Forecasts: Sparse Residual Routing for Efficient Weather Prediction Authors: Janet Wang, Yunbei Zhang, Lin Zhao, Xi Xiao, Jihun Hamm, Xiao Wang
-
Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving Authors: Shrikara Arun, Anjaly Parayil, Srikant Bharadwaj, Renee St. Amant, Victor Rühle
Large-Scale Training Systems and Efficiency (3)
1. Revisiting Decentralized Online Convex Optimization with Compressed Communication
ArXiv ID: 2607.01665
Primary Topic: Large-Scale Training Systems and Efficiency
Authors: Hao Zhou, Xiaoyu Wang, Chang Yao, Mingli Song, Yuanyu Wan
Abstract: Decentralized online convex optimization (D-OCO) is a popular framework for distributed applications with streaming data. To tackle the communication bottleneck, previous studies have investigated D-OCO with compressed communication and proposed several algorithms that are variants of online gradient descent (OGD). However, for D-OCO with exact communication, the best existing algorithms are variants of follow-the-regularized-leader (FTRL). In this paper, for the first time, we propose two FTRL-type algorithms for D-OCO with compressed communication. Compared with OGD-type algorithms, our algorithms are more elegant in both algorithmic design and theoretical analysis. The key insight is that the dual update mechanism of FTRL allows us to make a simple application of the technique for average consensus with communication compression. More specifically, our first algorithm considers the full-information setting, and can match the existing regret bounds. Our second algorithm is designed for the bandit setting, and can significantly improve both the regret bounds and communication costs of existing algorithms.
Comment: Combines FTRL dual updates with compressed consensus to improve bandit regret and communication costs.
Topic Match: New communication-compressed distributed optimization algorithms directly fit the systems topic, with scope limited to online convex learning.
Relevance: 7 Novelty: 7
2. Decentralized Stochastic Subgradient-type Methods with Communication Compression for Nonsmooth Nonconvex Optimization
ArXiv ID: 2607.01755
Primary Topic: Large-Scale Training Systems and Efficiency
Authors: Siyuan Zhang, Nachuan Xiao, Xin Liu
Abstract: In this paper, we consider the nonsmooth nonconvex decentralized optimization problem, where inter-agent communication is compressed. We propose a general framework that unifies various decentralized stochastic subgradient-type methods with unbiased compression and contractive compression with error compensation. By relating the consensus-error iterates and the averaged iterates to the trajectories of continuous-time differential inclusions, we establish global convergence for all methods encompassed by our framework when the objective functions are nonsmooth and lack Clarke regularity. Based on our framework, we further develop several compression-based methods, including decentralized stochastic subgradient methods utilizing sign-based regularization and gradient-tracking momentum. Preliminary numerical experiments empirically support our theoretical results and highlight the communication-accuracy trade-off of the newly developed methods.
Comment: Establishes convergence for decentralized nonsmooth subgradient methods with compressed messages and error compensation.
Topic Match: Communication-efficient distributed optimization is central; evidence consists of convergence theory and preliminary numerical experiments.
Relevance: 7 Novelty: 7
3. Understanding the Robustness of Distributed Self-Supervised Learning Frameworks Against Non-IID Data
ArXiv ID: 2607.02447
Primary Topic: Large-Scale Training Systems and Efficiency
Authors: Xuanyu Chen, Nan Yang, Shuai Wang, Dong Yuan
Abstract: Recent research has introduced distributed self-supervised learning (D-SSL) approaches to leverage vast amounts of unlabeled decentralized data. However, D-SSL faces the critical challenge of data heterogeneity, and there is limited theoretical understanding of how different D-SSL frameworks respond to this challenge. To fill this gap, we present a rigorous theoretical analysis of the robustness of D-SSL frameworks under non-IID (non-independent and identically distributed) settings. Our results show that pre-training with Masked Image Modeling (MIM) is inherently more robust to heterogeneous data than Contrastive Learning (CL), and that the robustness of decentralized SSL increases with average network connectivity, implying that federated learning (FL) is no less robust than decentralized learning (DecL). These findings provide a solid theoretical foundation for guiding the design of future D-SSL algorithms. To further illustrate the practical implications of our theory, we introduce MAR loss, a refinement of the MIM objective with local-to-global alignment regularization. Extensive experiments across model architectures and distributed settings validate our theoretical insights, and additionally confirm the effectiveness of MAR loss as an application of our analysis.
Comment: Theoretical analysis links distributed pretraining robustness to network connectivity and objective choice.
Topic Match: Distributed-training behavior is central, but the setting is non-IID federated and decentralized self-supervised learning rather than large-model scaling.
Relevance: 6 Novelty: 7
Architecture and Training Dynamics (10)
1. A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets
ArXiv ID: 2607.02303
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Wanyun Cui
Abstract: Linear-attention and state-space language models compress the prefix into a fixed-size recurrent state, yielding O(1) memory at the cost of a lossy exact memory: when many key--value associations compete, earlier facts are overwritten and needle recall degrades. Inspired by Complementary Learning Systems, we give linear attention a hippocampal complement. HOLA (Hippocampal Linear Attention) keeps the usual delta-rule state as a compressive memory and adds a bounded exact KV cache, forming a semiparametric test-time memory: the state models linearly compressible structure, while the cache stores associations that should not be forced through that state. The cache writes without a learned eviction module, keeping tokens with large beta * ||e||, the prediction residual actually committed to the state; a decoupled RMSNorm-gamma cache read then turns these exact KV pairs into sharp retrieval rather than soft averaging. At 340M parameters trained on 15B SlimPajama tokens, HOLA lowers Wikitext perplexity from 27.32 to 22.92 (-16.1%), below a full-attention Transformer++ (26.88), and improves LAMBADA perplexity from 30.95 to 30.26. It also achieves the best linear in-context retrieval and remains much more robust than GDN or a matched HOLA+recency cache on RULER needle-in-a-haystack recall out to 32k tokens (16x its training length).
Comment: Residual-selected exact KV caching augments the compressive state of delta-rule linear attention.
Topic Match: The cache changes the trained attention architecture and recurrent-state read mechanism, making this a substantive architectural contribution.
Relevance: 9 Novelty: 7
2. Training Hybrid Block Diffusion Language Models with Partial Bidirectionality
ArXiv ID: 2607.02805
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Pranshu Chaturvedi, Parth Shroff, Tarun Suresh, Hangoo Kang, Kaiyue Wen
Abstract: High-throughput long-context generation is one of the central challenges for large language models. Generation is typically memory-bandwidth-bound rather than compute-bound: each decoding step must stream the accumulated key/value (KV) cache from memory, so bandwidth demand grows with context length while only one token is emitted. Two parallel approaches have therefore emerged: reducing memory access with efficient attention variants and linear-time mixers such as Mamba, or increasing parallel computation by generating blocks of tokens at once. However, technical challenges arise when combining these two ideas. Earlier hybrid diffusion models such as DiffuMamba use bidirectional Mamba mixing, including a reverse-direction scan relative to causal generation. This reverse scan needs to scan the entire sequence, so its states are not prefix-only and cannot be precisely reused as a cache even when diffusion is performed block by block. We propose a BDLM Mamba--attention hybrid that addresses this challenge by restricting the reverse Mamba scan to the active denoising block, which enables exact caching across blocks. In an 87M-parameter DCLM sweep, BDLM Mamba-H achieves the best C4-en validation perplexity compared to BDLM attention and full-sequence baselines. At 350M parameters, it remains competitive with BDLM attention. For long-context inference, BDLM Mamba-H reaches 19.7x the throughput of full-sequence DiffuMamba-H at 65K tokens and 3.7x the throughput of BDLM attention at 262K, showing that Mamba hybrids are a potential long-context diffusion architecture.
Comment: Restricting reverse Mamba scans to the active diffusion block enables exact prefix-state caching.
Topic Match: Partial bidirectionality resolves a structural incompatibility between recurrent mixing, block diffusion, and reusable caches.
Relevance: 9 Novelty: 7
3. The Orthogonalized Read Is a Removable Training Scaffold for Recurrent Memory
ArXiv ID: 2607.19390
Primary Topic: Architecture and Training Dynamics
Authors: Keston Aquino-Michaels
Abstract: A recent report finds that orthogonalizing the mLSTM memory matrix at read time (five Newton-Schulz iterations, trained through) substantially improves noisy associative recall. The effect replicates, but it is not a memory improvement. Training on this task is a long chance plateau followed by a sharp escape, and the orthogonalized read acts by re-conditioning the learning problem during the plateau. Three properties establish this. It must be self-consistent: an exact recursive least-squares read (the Mesa layer) reproduces it, while straight-through halves, delta-rule writes, frozen random keys, and plain normalization all fail. It is uniform: across a learning-rate x hardness grid it multiplies the escape hazard roughly six-fold with no detectable hardness dependence, widening the workable learning-rate corridor that narrows for the baseline. And it is removable: applied to failed models at inference it rescues none, and annealed away on an escape-triggered schedule it leaves numerically stock mLSTMs at full accuracy. Much of the published gain needs no architecture at all: solved-rate at a fixed budget measures escape hazard, which follows a heat/noise law (learning-rate elasticity +3.0, gradient-noise elasticity -1.65) under which the original vocab-96 result is a large-batch noise condition rather than a capacity one. Decoding the memory state directly shows failed models carry roughly half their associations in linearly recoverable form: the plateau is a readout failure over half-written storage. Two conclusions travel beyond the intervention: recall benchmarks used for architecture selection partly measure trainability, and the system is a fully instrumented model organism of "emergence," in which a sharp behavioral threshold demonstrably arises from a censored metric over gradually accumulating structure.
Comment: A removable orthogonalized read improves recurrent-model training by accelerating escape from optimization plateaus.
Topic Match: Controlled interventions distinguish training conditioning from memory capacity and explain apparent architectural gains through optimization dynamics.
Relevance: 8 Novelty: 8
4. Benign Overfitting Does Not Occur in Diffusion Models
ArXiv ID: 2607.02671
Primary Topic: Architecture and Training Dynamics
Authors: Tyler Farghly, Benjamin Dupuis, Alain Durmus, Umut Simsekli
Abstract: Benign overfitting and double descent have come to shape our understanding of generalization in deep learning, establishing that overfitting is not only compatible with good generalization but can actively benefit it. Diffusion models share much of the machinery of standard deep learning, so it is natural to assume that they also exhibit these properties. In this work, we show that this assumption is largely incorrect. We first establish fundamental impossibility results showing that, unless the sample size grows exponentially with the data dimension, overfitting and good generalization cannot occur simultaneously. Consequently, the population loss follows a classical U-shaped curve in model complexity rather than exhibiting double descent. Analyzing a simplified setting, we identify a key difference between regression and score matching: regression benefits from an alignment between the target and the empirical covariance; score matching admits no such alignment, leaving overfitting irreparably harmful. We further identify implicit regularization stemming from time-smoothness of the score and early stopping during training as mechanisms that prevent such overfitting and verify our findings with high-dimensional image generation experiments. Our results reveal that generalization in diffusion models is governed by mechanisms distinct from those of traditional regression, motivating the development of new theory.
Comment: Explains harmful overfitting in score matching and identifies time smoothness and early stopping as implicit regularizers in diffusion training.
Topic Match: The core contribution explains how the diffusion training objective changes overfitting and regularization, directly informing training dynamics.
Relevance: 8 Novelty: 8
5. Dendritic In-Context Learning in a Single-Layer Spiking Neural Network
ArXiv ID: 2607.02283
Primary Topic: Architecture and Training Dynamics
Authors: Juwei Shen, Yujie Wu, Changwen Chen
Abstract: In-context learning (ICL) operates via implicit gradient descent embedded in the forward pass of modern AI architectures -- Transformers, Mamba, state-space models, and MLPs. Capturing this capability in biologically plausible Spiking Neural Networks (SNNs) has remained an open challenge: existing SNNs fail the Garg-2022 benchmark at non-trivial task dimensions. We trace this failure to a structural assumption: prior SNN designs route adaptation through inference-time synaptic plasticity, viewing the dendritic compartment as a passive conduit for error or teacher signals. We challenge this assumption. The subthreshold dynamics of a single dendritic compartment already implement a complete online learning algorithm. By treating the compartment as the computational substrate rather than a passive conduit, we propose DendriCL -- a single-layer compartmental spiking architecture whose apical recurrence is structurally identical to leaky online Widrow-Hoff LMS. This dynamics-only update collapses the architectural depth required for general-purpose ICL to a single layer. DendriCL is uniquely seed-stable at super-dimensional Garg-2022 ICL -- where dense Transformers exhibit grokking-style instability and fail past moderate task dimension -- and a linear probe recovers the reference online-LMS trajectory directly from the apical membrane at R^2 = 0.93, showing the algorithm is structurally embedded in the dynamics rather than implicitly discovered during training. Taken together, ICL requires neither attention, depth, nor inference-time plasticity: a single compartment with online-LMS dynamics is sufficient.
Comment: A single dendritic recurrence structurally implements online LMS for single-layer in-context learning.
Topic Match: The contribution is an explicit recurrent learning mechanism, although evidence is limited to synthetic tasks and small spiking networks.
Relevance: 7 Novelty: 8
6. SUNTA: Hierarchical Video Prediction with Surprise-based Chunking
ArXiv ID: 2607.02087
Primary Topic: Architecture and Training Dynamics
Authors: Tomoshi Iiyama, Masahiro Suzuki, Yutaka Matsuo
Abstract: Hierarchical state-space models (HSSMs) offer a promising approach to long-horizon prediction by segmenting sequences into temporal chunks. However, their performance hinges on how chunk boundaries are determined. While prior HSSMs typically rely on fixed-length chunking or similarity-based boundary detection, these methods often misalign with the intrinsic temporal structure of the data. We argue that chunking should instead be driven by prediction errors, which more directly indicate when longer-range context becomes necessary. Nevertheless, integrating surprise-based chunking into HSSMs introduces critical challenges, including hierarchical collapse during end-to-end training and the absence of surprise signals during open-loop prediction. To address these issues, we propose Surprise-based Nested Temporal Abstraction (SUNTA), a method that employs a decoupled training strategy to preserve surprise signals and uses internal inconsistency as a top-down surprise metric to determine chunk boundaries within imagined rollouts. Experiments on video prediction tasks in 2D and 3D environments demonstrate that SUNTA outperforms baselines, uniquely maintaining accurate predictions over 250 timesteps, whereas all baselines degrade within the first 10 timesteps.
Comment: Decouples hierarchical training to preserve surprise-driven temporal chunking and prevent hierarchy collapse.
Topic Match: The core innovations concern dynamic temporal boundaries and training stability in hierarchical state-space models, demonstrated through video prediction.
Relevance: 7 Novelty: 7
7. Born Discrete, Made Smooth: Variational Formulation of Shallow Neural Networks
ArXiv ID: 2607.02003
Primary Topic: Architecture and Training Dynamics
Authors: Matej Benko, Pierre Bousquet, Iwona Chlebicka, Błażej Miasojedow
Abstract: Although neural networks are remarkably effective, their underlying optimization principles remain theoretically elusive, often characterized by non-convex landscapes and stochastic heuristics. In this work, we propose a paradigm shift by replacing the discrete training problem of shallow neural networks with a well-posed continuum variational surrogate. We identify a family of $λ$-convex functionals over parameter densities in weighted Sobolev spaces and prove that these variational problems are globally well-posed, stable, and exhibit unexpected almost $C^3$ regularity. Unlike existing Wasserstein-based or Mean-Field approaches, which often face limited regularity and discretization challenges, our formulation provides direct access to elliptic regularity and convex analysis. This allows us to prove that the optimal parameter density can be obtained by solving a single linear system, bypassing iterative optimization entirely. We establish explicit generalization error controls at a rate of $1/α$ relative to the regularization parameter, and prove that finite-width networks of size $N$ achieve the continuum optimum at an $O(1/N)$ rate. This perspective bridges the gap between the Neural Tangent Kernel (NTK) and feature-learning regimes, providing a principled framework for understanding over-parameterization through the lens of variational calculus.
Comment: A convex continuum surrogate replaces shallow-network optimization with a single linear solve.
Topic Match: The core is a new optimization formulation, although its guarantees concern shallow-network surrogates rather than observed large-model training dynamics.
Relevance: 6 Novelty: 8
8. X-LogSMask: Expand Transformer for Graph-Structured Data
ArXiv ID: 2607.01553
Primary Topic: Architecture and Training Dynamics
Authors: Leyan Li, Rennong Yang, Zhenxing Zhang, Liping Hu
Abstract: Transformers have become general-purpose architectures, but their all-to-all self-attention is poorly matched to graph data, whose interactions are sparse, structured and multi-scale. Existing Graph Transformers address this mismatch through structural encodings, hybrid message-passing modules or learned attention constraints, often introducing additional complexity and limited interpretability. Here we introduce X-LogSMask, an explainable multi-head logarithmic structural mask that injects symmetrically normalized graph topology directly into attention logits. The logarithmic transform converts structural connectivity into a topology-aware gating signal, suppressing unsupported node interactions while preserving feature-dependent attention. By assigning different powers of the normalized adjacency matrix to different attention heads, X-LogSMask gives each head a defined structural radius and supports multi-hop information propagation within a single layer. We further show that a standard Transformer encoder can be interpreted as one-step message passing on a complete graph, motivating X-LogSMask as a topology-constrained alternative to unrestricted self-attention. Across 20 node-, edge- and graph-level benchmarks, Transformers equipped with X-LogSMask achieve state-of-the-art performance on 13 datasets and remain competitive in a lightweight one-layer configuration. These results show that simple, interpretable structural masks can make self-attention an effective graph-learning operator without changing the Transformer architecture. The code is available at https://github.com/LiLeyan-0120/X-LogSMask.
Comment: Adds logarithmic topology masks to attention logits and assigns different graph-hop radii to attention heads.
Topic Match: Introduces a structural attention mechanism rather than merely applying a Transformer; its scope remains graph learning.
Relevance: 7 Novelty: 6
9. Self-Gating Attention for Efficient Time Series Forecasting
ArXiv ID: 2607.02344
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Dezheng Wang, Tong Chen, Wei Yuan, Congyan Chen, Shihua Li, Hongzhi Yin
Abstract: Transformer architectures have shown strong potential in time series forecasting, where multi-head self-attention is widely used to capture temporal dependencies across historical timestamps. However, standard self-attention has quadratic time and memory complexity with respect to the look-back length. This cost may limit its use in resource-constrained or high-throughput forecasting systems, where fast and memory-efficient inference is important. Through qualitative and quantitative analyses, we observe that self-attention maps in time series forecasting often contain redundant patterns across different timestamps. This phenomenon can be related to the repeated temporal patterns and relatively stable temporal correlations in many real-world time series. Motivated by this observation, we propose Self-Gating Attention (SGA), a plug-and-play attention mechanism that represents the attention score with a shared learnable matrix and an input-dependent residual component. The shared matrix captures common attention patterns, while the residual component captures input-dependent variations. In this way, SGA avoids the query and key projections used in standard attention score computation, leading to linear time and score-matrix memory complexity with respect to the look-back length. We integrate SGA into several forecasting backbones and compare it with standard self-attention and lightweight attention variants on nine publicly available real-world datasets covering electricity, finance, weather, medical monitoring, human activity, and climate records. The results show that SGA improves inference efficiency on public benchmarks while maintaining competitive forecasting performance against state-of-the-art attention mechanisms. These benchmark results provide deployment-oriented evidence.
Comment: Proposes linear-cost attention scoring through shared learnable patterns and input-dependent residuals that eliminate query-key projections.
Topic Match: Attention-score redesign is the central contribution, although its assumptions and empirical evidence are specific to forecasting.
Relevance: 7 Novelty: 6
10. LiNO: Lifting based multiresolution neural operator
ArXiv ID: 2607.02715
Primary Topic: Architecture and Training Dynamics
Authors: Himanshu Pandey, Subham Patel, Ratikanta Behera
Abstract: Recently, neural operators have shown promising outcomes for learning solution operators of differential equations directly from data. This framework learns a functional mapping from the parameter field to the solution field, enabling the prediction of an entire class of solutions rather than a specific instance. However, existing operators often struggle to capture both global dynamics and fine-scale structure simultaneously. To design an effective operator capable of representing multiscale features, a hierarchical multiscale decomposition framework is required. In this study, we develop the Lifting Neural Operator (LiNO), a multiresolution operator built on the second-generation wavelet lifting scheme. LiNO learns a multiresolution decomposition directly from data by parameterizing the lifting transform. This lifting transformation is adaptive to the underlying solution function and exactly invertible by construction, enabling information-preserving multiscale operator learning. In the lifted multiresolution space, the operator evolves coarse and directional detail coefficients separately, resulting in scale-aware modeling of the underlying physics. We evaluate LiNO on several benchmarks, including Darcy flow, the Poisson equation, the Allen-Cahn equation, the compressible Navier-Stokes equation, and the Gray-Scott reaction-diffusion system. Together, these benchmarks cover a wide range of physical behaviors, including multiscale phenomena, transport-dominated dynamics, and chaotic systems. LiNO demonstrates strong performance on these challenging benchmarks compared with state-of-the-art neural operators. These results suggest that adaptive multiresolution operators provide a promising direction for scientific machine learning.
Comment: Learned, exactly invertible lifting transforms provide multiresolution computation for neural operators.
Topic Match: Invertible multiscale decomposition is a concrete architectural mechanism, with relevance limited by the scientific PDE-surrogate setting.
Relevance: 6 Novelty: 6
Efficiency, Compression, and Large-Scale Training (11)
1. SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models
ArXiv ID: 2607.01876
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Qi Lyu, Jiahua Dong, Baichen Liu, Xudong Wang, Mingfei Han, Yulun Zhang, Fahad Shahbaz Khan, Salman Khan, Lianqing Liu, Zhi Han
Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable progress in multimodal understanding, yet their enormous parameter scale and cross-modal computation incur substantial memory and latency overhead, severely limiting real-world deployment on resource-constrained devices. Binarization offers an attractive solution by drastically reducing storage and computational costs. However, existing binarization methods neglect the varying importance of weights across different layers and modalities. This causes parameters irrelevant to downstream tasks to be unnecessarily retained, whereas modality-critical weights may not be adequately optimized, resulting in significant performance degradation. To address these challenges, we develop a novel \underline{S}ignificance-\underline{A}ware \underline{B}inarization for \underline{L}arge \underline{V}ision-\underline{L}anguage \underline{M}odels (SAB-LVLM). Specifically, after constructing Hessian matrices for textual and visual inputs, we propose a spatial significance map to distinguish full-precision weights activated under a single modality from those activated across modalities. We then devise a modality-guided integration strategy to obtain the significance-aware binarization map, which measures weight significance across layers and modalities. Subsequently, this binarization map is incorporated into the binarization objective as an error reweighting term, and binarization fitting is performed through an alternating significance-weighted update scheme. Extensive experiments illustrate the superiority of our SAB-LVLM over existing binary PTQ methods under an approximately 1-bit compression constraint. Our code is accessible at https://github.com/LyuQi127/SAB_LVLM.
Comment: Uses modality-specific Hessian significance to reweight one-bit quantization errors during alternating binary fitting.
Topic Match: The core contribution is a modality-aware weight-binarization mechanism for compressing large vision-language models.
Relevance: 9 Novelty: 6
2. Lynx: Progressive Speculative Quantization for accelerating KV Transfer in Long-Context Inference
ArXiv ID: 2607.01831
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Wenchen Han, Gingfung Matthew Yeung, Marco Barletta, William Toner, Amory Hoste, Adam Barker
Abstract: Long-context inference is increasingly common in large language model (LLM) serving, driven by retrieval-augmented generation and agentic systems. In disaggregated inference, these workloads require transferring large Key-Value (KV) caches across the network, where decoding cannot begin until the transfer completes. Recent KV quantization techniques reduce data volume and alleviate this bottleneck, but existing schemes fail to achieve both low network-exposed latency and high inference accuracy. We challenge the assumption that the KV cache is an indivisible unit that must be fully received before use. We leverage the observation that different bits in the KV cache contribute unequally to attention computation and inference precision: the most significant bits capture the coarse structure of attention and the least significant bits refine precision. This property enables partial use of the KV cache during decoding. We present Lynx, a system that enables progressive, split-stream KV transfer by partitioning the KV cache into a high-priority Anchor stream carrying the most significant bits and a low-priority Residual stream carrying remaining precision. Decoding begins upon receipt of the Anchor stream and proceeds speculatively while the Residual stream is transferred concurrently, followed by verification that ensures equivalence to higher-precision decoding. Across multiple models and serving workloads, Lynx achieves Time-to-First-Token (TTFT) comparable to aggressive 4-bit KV quantization, while matching the accuracy of high-precision (BF16) inference, improving TTFT over standard 8-bit KV quantization by up to $1.43\times$ and improving accuracy over state-of-the-art by up to $5.1\%$.
Comment: Progressive KV bit streams enable speculative decoding before full-precision cache transfer completes.
Topic Match: Bit-prioritized cache transfer and subsequent verification introduce a concrete mechanism for reducing exposed communication latency.
Relevance: 8 Novelty: 7
3. WBMM: Windowed Batch Matrix Multiplication for Efficient Large Receptive Field Convolution
ArXiv ID: 2607.02097
Primary Topic: Efficiency, Compression, and Large-Scale Training
Also Matches: Architecture and Training Dynamics
Authors: Wan Song, Wei Zhou, Rui Wang, Jun Yu, Toru Kurihara, Jiajia Xu, Shu Zhan
Abstract: Large kernel depthwise convolutions achieve strong performance but suffer from significant degradation as kernel size grows due to irregular memory access from gather-based computation; while Large Kernel Acceleration (LKA) helps on small feature maps, it becomes counterproductive on large feature maps, even slower than non-accelerated implementations. We propose Windowed Batch Matrix Multiplication (WBMM), which partitions input into contiguous windows and indexes a compact relative position bias table to construct weight matrices, enabling regular memory access via batched matrix multiplication. This yields a unique property: WBMM's throughput improves with larger windows, opposite to depthwise convolutions that degrade with larger kernels. Operator-level benchmarks show WBMM with 14x14 windows outperforms 5x5 depthwise convolution baselines in speed while providing a 7.8x larger per-layer receptive field. Combined with inter-block cross-window communication and hierarchical window reparameterization, WBMM achieves comparable or higher accuracy on ImageNet-1K, COCO, and ADE20K with 1.31-1.88x training speedup, and demonstrates consistent advantages across GPU, CPU, and edge devices without requiring specialized acceleration kernels. Our code is available at https://github.com/wansong-s/WBMM.
Comment: Windowed batched matrix multiplication replaces irregular large-kernel convolution access with contiguous computation.
Topic Match: Operator and memory-access redesign produce measured training speedups, making efficiency the primary contribution.
Relevance: 8 Novelty: 7
4. Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning
ArXiv ID: 2607.02484
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Xuehui Wang, Xuankun Yang, Wei Shen
Abstract: Visual token pruning is a crucial strategy for accelerating VLMs by compressing redundant image patches, yet existing methods often fail to preserve critical cues under dense instructions and fine-grained queries. In this paper, we investigate this failure and identify two underlying bottlenecks: the widespread dispersion of textual noise that corrupts dense cross-modal scoring, and the feature fragmentation inherent to standard token selection. To address these issues, we propose Entropy-Aware Dense Pruning (EADP), a framework that reformulates pruning as a structured compression problem. EADP first leverages statistical entropy to quantify and filter out textual noise, yielding a robust, fine-grained instruction relevance score. Subsequently, instead of naive Top-K selection, EADP casts token selection as a submodular maximization problem with a spatial prior, explicitly ensuring a holistic and non-redundant visual representation. Extensive experiments demonstrate that EADP improves the accuracy-efficiency trade-off of VLMs, robustly preserving fine-grained visual cues under strict token budgets while achieving SoTA performance on challenging multimodal benchmarks.
Comment: Formulates visual-token pruning as spatially informed submodular selection after entropy-based filtering of textual noise.
Topic Match: Structured token compression under tight budgets is the central method, directly targeting VLM inference efficiency.
Relevance: 8 Novelty: 6
5. Spec-AUF: Accept-Until-Fail Training under Train-Inference Misalignment for Masked Block Drafters
ArXiv ID: 2607.01893
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Tianjian Yang, Meng Li
Abstract: Speculative decoding accelerates autoregressive generation by drafting a block of tokens that the target model verifies left-to-right, committing only the longest accepted prefix. Block (DLM-style) drafters predict the whole block in parallel, which is fast but trained with a full-block cross-entropy that supervises every position against the gold continuation -- even though inference discards every token after the first rejection. Recent acceptance-aware objectives patch this by reweighting the full-block loss; we instead use teacher-forced learning as a motivation for how supervision should concentrate on the accepted prefix. A mask-only block drafter has no input-side channel for gold-prefix conditioning, so AUF approximates that prefix-sensitive supervision on the loss side by keeping the cross-entropy support only through the drafter's first predicted failure. AUF is a single, detached change to the CE support -- no auxiliary objective, no verifier rollouts, and no change to the inference pipeline or the exactness contract. Within fixed drafter backbones and serving settings on Qwen3-8B, AUF raises the DFlash drafter's average emitted length $τ$, averaged over six benchmarks, from 2.40 to 2.61, with a gain on every benchmark, and transfers to Domino's two-branch head (2.56 to 2.68). Two findings sharpen the picture: the decay-only baseline reaches higher token accuracy on the shared block mask yet decodes worse, and on DFlash, once AUF truncates the support, the standard exponential position-decay weighting becomes empirically inert.
Comment: Truncates drafter cross-entropy supervision at the first predicted failure to better match speculative acceptance.
Topic Match: The training-objective change directly improves speculative decoding efficiency by increasing accepted draft length.
Relevance: 8 Novelty: 6
6. EPnG: Adaptive Expert Prune-and-Grow for Parameter-Efficient MoE Fine-tuning
ArXiv ID: 2607.01789
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Ahin Lee, Sehyun Yun, Taesik Gong
Abstract: Mixture-of-Experts (MoE) models scale efficiently but remain costly to adapt due to redundant experts and uniform parameter allocation. Existing parameter-efficient fine-tuning (PEFT) methods such as LoRA ignore MoE routing dynamics, leading to suboptimal resource use. We propose EPnG, an adaptive prune-and-grow framework that reallocates LoRA capacity based on expert importance derived from router gate probabilities. EPnG prunes under-utilized experts and expands high-importance experts via rank growth with orthogonal initialization, while maintaining a fixed parameter budget. Across OLMoE and Qwen1.5-MoE, EPnG consistently outperforms LoRA under the same budget and achieves performance comparable to full fine-tuning while updating only 0.55%-0.72% of parameters (up to 140x-180x fewer). These results demonstrate that aligning PEFT with MoE routing yields a more effective and scalable fine-tuning strategy.
Comment: Router-informed expert pruning and LoRA rank growth reallocate adaptation capacity under a fixed parameter budget.
Topic Match: The innovation is parameter-efficient expert adaptation; existing gate probabilities provide importance signals rather than a new routing algorithm.
Relevance: 8 Novelty: 6
7. LAARA: Layer-Aware Adaptive Rank Allocation for Parameter-Efficient Fine-Tuning
ArXiv ID: 2607.19391
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Ashutosh Tripathi, Surya Deep Singh, Pranab Sahoo, Sriparna Saha
Abstract: Low-Rank Adaptation is widely used for parameter-efficient fine-tuning, yet existing methods typically assign the same adapter rank to every transformer layer despite their heterogeneous adaptation requirements. In this work, we show theoretically and empirically that uniform rank allocation is fundamentally suboptimal. Motivated by this observation, we propose LAARA (Layer Aware Adaptive Rank Allocation framework), a search-free framework that dynamically allocates ranks using lightweight diagonal Fisher estimates computed during training. LAARA combines projection-wise normalization, logarithmic compression, blended adapter importance estimation, and a vote-to-change dampening mechanism to produce stable and efficient rank adaptation. Experiments on GLUE and MathInstruct benchmark demonstrate that LAARA consistently matches or outperforms popular state of the art approaches such as LoRA, AdaLoRA, DyLoRA, and Bitfit while using significantly fewer trainable parameters. Our results show that Fisher-guided rank allocation provides a principled and effective foundation for adaptive parameter-efficient fine-tuning. The code is publicly available at: https://anonymous.4open.science/r/LAARA-D305/LAARA.py
Comment: Diagonal-Fisher importance dynamically allocates adapter ranks across transformer layers.
Topic Match: Adaptive rank allocation directly targets parameter-efficient fine-tuning, extending existing variable-rank adapter methods.
Relevance: 8 Novelty: 6
8. ART for Diffusion Sampling: Continuous-Time Control and Actor-Critic Learning
ArXiv ID: 2607.02137
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Yilie Huang, Wenpin Tang, Xun Yu Zhou
Abstract: We study timestep allocation for score-based diffusion sampling, where a learned reverse-time dynamics is discretized on a finite grid. Uniform and hand-crafted schedules are standard choices, but they rely on fixed prescriptions and can therefore be suboptimal. To address this limitation, we propose Adaptive Reparameterized Time (ART), a continuous-time control formulation that learns a time change by treating the speed of the sampling clock as the control, so that a uniform grid on the learned clock induces adaptive timesteps in the original diffusion time. Based on a leading-order Euler error surrogate, ART provides a principled objective for allocating timesteps along the sampling trajectory. To solve this deterministic control problem, we introduce ART-RL, an auxiliary randomized formulation with Gaussian policies that turns schedule learning into a continuous-time reinforcement learning problem. We prove that the randomized ART-RL formulation is equivalent to ART at the optimizer level, in the sense that its optimal Gaussian policy recovers the optimal ART time-warping rate through its mean. We further establish policy evaluation and policy improvement characterizations and derive trajectory-based moment identities that yield implementable actor--critic updates for learning the schedule. Across experiments ranging from controlled low-dimensional settings to image generation, ART-RL can be plugged into existing diffusion samplers by changing only the timestep grid, consistently improving sample quality over strong baseline schedules at matched budgets while leaving the rest of the sampling pipeline unchanged. The learned schedules also exhibit broad generalization, transferring without retraining across sampling budgets, datasets, solvers, pipelines, and representation spaces.
Comment: Continuous-time reparameterization learns how to allocate diffusion sampling steps under a fixed compute budget.
Topic Match: The core contribution improves sampling quality per computation through adaptive discretization, with scope limited to inference.
Relevance: 7 Novelty: 7
9. PARTREP: Learning What to Repeat for Decoder-only LLMs
ArXiv ID: 2607.01792
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Andikawati P Widjaja, Yongjun Kim, Hyounghun Kim, Jaeho Lee
Abstract: While decoder-only LLMs excel at a vast array of natural language tasks, it suffers from an asymmetric information flow induced by causal attention: later tokens are richer in contextual grounding than earlier ones. A simple and effective remedy is prompt repetition -- just appending a second copy of prompt before generation can redistribute grounding across positions and improve reasoning performance. However, full repetition of the original prompt doubles the KV cache footprint and quadruples attention cost during prefill, making it impractical for long-context settings. We propose PartRep, a selective augmentation method that appends only the most informative tokens -- rather than the entire prompt. We use token-wise negative log-likelihood (NLL) as a selection signal, motivated by the hypothesis that less predictable tokens are less recoverable from surrounding context and therefore benefit more from late-position repetition. To avoid the heavy cost of a full forward pass for scoring, we train a lightweight gate that predicts high-NLL tokens from early-layer hidden states, enabling token selection during mid-prefill via early exit. Across eight benchmarks (including MMLU, GSM8K, and RULER) and three model families (Qwen2.5, Llama3.2, Gemma4), PartRep retains most of the gains of full repetition while using only 59.4\% of its KV cache and 79.0\% of its prefill FLOPs.
Comment: Predicts repeat-worthy prompt tokens from early-layer states, reducing KV usage to 59.4% of full prompt repetition.
Topic Match: A learned selection gate reduces augmentation memory and prefill costs; the savings are measured against full prompt repetition.
Relevance: 7 Novelty: 6
10. Less Tokens, Better Forecasts: Sparse Residual Routing for Efficient Weather Prediction
ArXiv ID: 2607.02829
Primary Topic: Efficiency, Compression, and Large-Scale Training
Also Matches: Architecture and Training Dynamics
Authors: Janet Wang, Yunbei Zhang, Lin Zhao, Xi Xiao, Jihun Hamm, Xiao Wang
Abstract: Existing ViT-based weather forecasting models apply uniform computation across all spatial tokens, even though nearby atmospheric grid points often contain similar values and large regions evolve smoothly over time. This makes much of the intermediate per-token computation redundant. Standard token-efficiency methods, such as pruning or merging, reduce cost by removing or fusing tokens. However, weather forecasting is a spatiotemporal dense prediction problem in which a history of atmospheric states must be mapped to future values on the original latitude-longitude grid. Thus, every grid cell must retain a physically meaningful representation, especially under autoregressive rollout. We introduce Sparse-Reslim, a parameter-free plug-in routing module that makes sparse token processing compatible with this fixed-grid requirement. Sparse-Reslim routes only 25% of spatial tokens through the expensive middle transformer blocks and treats those blocks as residual updates: it computes the change produced for the routed tokens and scatters only this delta back to the full sequence. Unselected tokens keep their pre-routing representations exactly, so no grid cell is dropped or replaced by a mask token, and no fusion layer or additional parameters are introduced. Across ERA5 resolutions up to the operational 0.25\textdegree{} standard and two model families, a deterministic Transformer and a diffusion model, Sparse-Reslim improves forecast accuracy on every evaluated variable while substantially reducing cost: training is about 2.5x faster in the main settings and reaches 3.18x speedup at 0.25\textdegree{}, with over 2.2x lower peak memory. A controlled decomposition shows that the accuracy gain comes primarily from sparse routing itself, while random token selection provides an additional regularization benefit without selector overhead.
Comment: Sparse residual token routing reduces training computation while preserving the complete spatial representation.
Topic Match: Selective token execution directly changes training cost and memory use; the supporting experiments remain specific to weather models.
Relevance: 7 Novelty: 6
11. Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving
ArXiv ID: 2607.02043
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Shrikara Arun, Anjaly Parayil, Srikant Bharadwaj, Renee St. Amant, Victor Rühle
Abstract: Disaggregated LLM serving runs prefill and decode on separate GPU pools to keep the two phases from interfering. In practice, this creates a new asymmetry: under bursty, heavy-tailed workloads prefill nodes saturate while decode nodes have compute underutilized, and on a production-style A100 cluster with 2 prefill and 2 decode nodes (2P2D), we find that prefill execution accounts for only 2-23% of P95 Time-to-First-Token (TTFT). Queuing and inter-node GPU-GPU KV-cache transfer account for the rest. We present a proactive prefill-deflecting scheduler that lets decode nodes serve prefill phase of requests as chunked-prefill steps interleaved with their in-flight decode batches. For each queued request, we estimate the TTFT it would see on the prefill node, and on every decode node, search for the largest chunk schedule that keeps in-flight decodes within their Time-Between-Tokens (TBT) SLO and deflect when the decode path helps tail latency. Because the prefill phase of deflected requests runs in place on the decode node, the inter-node KV transfer is eliminated. Implemented on vLLM and evaluated on production-style traces with DeepSeek-V2-Lite, our approach reduces P95 TTFT by upto 81% and raises SLO attainment by upto 79% over state-of-the-art disaggregated schedulers, at sub-millisecond per-request routing cost.
Comment: SLO-constrained prefill deflection uses spare decode-node compute and eliminates corresponding KV transfers.
Topic Match: Adaptive request placement and chunk scheduling introduce an inference-efficiency mechanism, though the contribution is confined to serving.
Relevance: 7 Novelty: 6
Paper Selection Prompt
System Prompt
You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.
User Prompt
Instructions
Respond in JSONL. Output exactly one JSON object per paper, one per line:
{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}
Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only:
daily_hot,new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return[]. -daily_hotmeans the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. -new_frontiermeans the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.Scoring Criteria
Relevance and Novelty are independent axes. Score both from 1 to 10.
Relevance Scoring
- 9-10: directly centered on the target foundational topics; highest when the core contribution is clearly within them.
- 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
- 5-6: touches the target topics, but the main contribution is elsewhere.
- 3-4: largely outside the target topics, often application-focused or domain-specific.
- 1-2: unrelated.
Important: Broad frontier relevance, major launch status, or daily buzz is not enough for a high Relevance score here. Those cases belong in the hotspot digest unless the paper strongly matches the specialized paper topics.
Novelty Scoring
- 9-10: new paradigm, theory, or major methodological breakthrough.
- 7-8: substantial methodological advance or strong new insight.
- 5-6: meaningful but incremental extension or refinement.
- 3-4: minor, narrow, or mostly engineering or domain-specific improvement.
- 1-2: little originality; mainly standard application of existing methods.
Topic Registry
Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - moe_training: MoE Training - Mixture-of-Experts training: routing, load balancing, expert granularity and shared experts, training stability, MoE parallelism and communication, and MoE kernels. - training_systems: Large-Scale Training Systems and Efficiency - Large-scale training systems and efficiency: distributed training algorithms, optimizers, and communication. - architecture_training: Architecture and Training Dynamics - Core architectural or computational mechanisms, dynamic computation, and training-stability dynamics. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, memory or cache efficiency, and large-scale training systems that materially change cost or behavior.
Papers
[PAPER LIST HERE]
Relevant Topics
This is a training-side feed. The centre of gravity is Mixture-of-Experts training and the large-scale training systems it sits on; architecture and efficiency work is kept when it bears on how a large model is trained or what it costs to train.
Keep a paper when its CORE CONTRIBUTION falls in one of the four topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.
MoE Training (primary) - Keep: expert routing and gating mechanisms; load balancing, with or without auxiliary losses; expert granularity, fine-grained and shared experts; MoE training stability (loss spikes, router collapse, z-loss and related fixes); expert and hybrid parallelism; all-to-all dispatch/combine and the communication schedules MoE needs; MoE kernels and grouped-GEMM style implementations; expert capacity, token dropping, and upcycling dense checkpoints into MoE. - Filter: papers that merely train on top of a MoE model without a new routing, balancing, stability, parallelism, or kernel contribution; "mixture of experts" in the classical ensemble or recommender sense.
Large-Scale Training Systems - Keep: distributed training algorithms (data, tensor, pipeline, sequence, and expert parallelism, FSDP and ZeRO-style sharding); optimizers and preconditioners for large-scale pretraining; communication algorithms, overlap, and collective design; throughput and MFU work that materially changes what a training run costs or whether it converges; scaling-law work that informs how such runs are configured. - Filter: routine infrastructure tuning with no new algorithmic or systems idea.
Architecture and Training Dynamics - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, dynamic or modular computation); optimisation and training-dynamics analysis that explains why large models train the way they do. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight.
Efficiency and Compression - Keep: quantization, sparsity, pruning, low-rank adaptation, KV-cache and memory-efficient designs, and other work that materially changes the cost of running or training a large model, especially when the mechanism is new rather than a tuned variant. - Filter: straightforward application of standard efficiency methods, and deployment or serving engineering with no new idea.
Leave these to the hotspot digest unless the core contribution clearly falls in one of the four topics above: - major model or product releases - agent, tooling, and RAG launches - benchmark, leaderboard, or evaluation-only papers - representation-learning theory, interpretability, and feature-formation analysis - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning) - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains