This is a remedial run for missed papers from 08/12/2026 to 08/12/2026.
Results generated on 09/14/2026.
Personalized Daily ArXiv Papers 2026-08-13
| Model | Metric | Usage | Papers | ||||
|---|---|---|---|---|---|---|---|
| Prompt | Completion | Total | Total arXiv | Scanned | Relevant | ||
local agent CLI |
Tokens | not reported | not reported | not reported | 452 | 452 | 35 |
| Cost | not reported | not reported | not reported | ||||
Token counts are not reported for this run. 20 of 20 model calls succeeded, 3,331s of model wall clock.
Topic Coverage:
| Topic | Papers |
|---|---|
| MoE Training | 1 |
| Large-Scale Training Systems and Efficiency | 5 |
| Architecture and Training Dynamics | 16 |
| Efficiency, Compression, and Large-Scale Training | 13 |
Table of contents by topic:
MoE Training (1)
- TradingMoE: Routing the Right Experts in Evolving Markets Authors: Chang Zhou, Xingtong Yu, Minbin Huang, Zhennan Wu, Yuan Fang, Hong Cheng, Xinming Zhang
Large-Scale Training Systems and Efficiency (5)
-
Spend Less, Fit Better: Budget-Efficient Scaling Law Fitting via Active Experiment Selection Authors: Sijie Li, Shanda Li, Haowei Lin, Weiwei Sun, Ameet Talwalkar, Yiming Yang
-
Reducing Per-Sample Interference in Stochastic Optimization Authors: Apostolos Avranas
-
Epiplexity Guided Data Selection and Generation for Out-of-Distribution Generalization Authors: Ellen Su, Andres Potapczynski, Shikai Qiu, Edward Hughes, Andrew Gordon Wilson
-
MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning Authors: Shiji Zhou, Kunlin Lyu, Lei Zhang, Ruodong Wang, Yifan Sun
-
A-3PO: Accelerating Asynchronous LLM Training with Staleness-aware Proximal Policy Approximation Authors: Xiaocan Li, Shiliang Wu, Zheng Shen
Architecture and Training Dynamics (16)
-
MARCH: Scaling Recurrent Memory with Content-Routed State Anchors Authors: Ming Zhang, Kaisen Yang, Shu Yu, Ermo Hua, Ning Ding, Xia Hu, Bowen Zhou, Chaochao Lu, Youbang Sun
-
Disentangling the Expressivity of RoPE Authors: Selim Jerad, Anej Svete, Jiaoda Li, Ryan Cotterell
-
LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Authors: Qiuwu Chen, Zimo Liu, Yuchen Li, Ying Sun, Yifan Zhang, Zhijie Qiu, Zeng You, Ryan Dong, Simeng Ma, Yaofo Chen, Mingkui Tan
-
ATMA: Long-Context Language Modeling via Polar Attention and Gated-Delta Compression Memory Authors: Habibullah Akbar
-
Do Transformers Need Three Projections? Systematic Study of QKV Variants Authors: Ali Kayyam, Anusha Madan Gopal, M Anthony Lewis
-
Weak Correlations as the Underlying Principle for Linearization of Gradient-Based Learning Systems Authors: Ori Shem-Ur, Khen Cohen, Aviv Orly, Yaron Oz
-
FLARE++: Low-rank attention with dynamic attention routing Authors: Vedant Puri, Yongjie Jessica Zhang, Levent Burak Kara
-
A New First-Order Meta-Learning Algorithm with Convergence Guarantees Authors: El Mahdi Chayti, Martin Jaggi
-
Training Under Challenge: Executable Certificates and Challenge-Closed Optimality for Neural Networks Authors: Farhang Yeganegi, Arian Eamaz, Mojtaba Soltanalian
-
TESLA: Taylor Expansion of Sinusoidal Learnable Activations Authors: Daehwa Ko, Jaehyeon Kim, Seunghyun Ham, Jay Hoon Jung
-
NAE: Normalizing AutoEncoder Authors: Muhammad Abdur Rafae, Niels Landwehr
-
GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs Authors: Kai Yang, Jingwei Xu, Wanyu Wang, Kai-Yuan Guo, Zhenbo Yu, Yi Wang, Yu Qiao
-
Gauge-Fixing the Forward-Forward Objective: A Whitened Goodness Derived from a Likelihood-Ratio Account Authors: Paolo Giannitrapani
-
Chain-of-Thought Shows the Path to a Tree: Realizing Branching Complexity Authors: Debanjan Dutta, Anish Chakrabarty, Swagatam Das
-
Exploring Oversmoothing with Householder Matrices Authors: Bhaskar Karol
-
HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold Networks Authors: Zhao Su, Yuxin Xia, Haoran Li, Jun Shen, Qi Zhu, Qingguo Zhou, Binbin Yong
Efficiency, Compression, and Large-Scale Training (13)
-
CurveFP: Co-Designing Numerical Representation and Product Arithmetic for Language Models Authors: Ye Qiao
-
CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution Authors: Zihao Ye, Yingyi Huang, Hongyi Jin, Bohan Hou, Junru Shao, Zhongming Yu, Jinqi Chen, Meghan Cowan, Shiyi Cao, Shanli Xing, Hanfeng Chen, Vinod Grover, Tianqi Chen, Luis Ceze
-
Prof-K: Probabilistic One-Pass Filtering for Efficient Top-k Selection Authors: Tadeusz Dziarmaga, Witold Sikora, Åukasz Struski, Jacek Tabor, Marcin Mazur
-
WBMM: Windowed Batch Matrix Multiplication for Efficient Large Receptive Field Convolution Authors: Wan Song, Wei Zhou, Rui Wang, Jun Yu, Toru Kurihara, Jiajia Xu, Shu Zhan
-
APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference Authors: Alish Kanani, Layan Badawi, Umit Y. Ogras
-
Variational Speculative Decoding: Rethinking Draft Training from Token Likelihood to Sequence Acceptance Authors: Xiandong Zou, Jianshu Li, Jing Huang, Pan Zhou
-
LoRAQuant: Mixed-Precision Quantization of LoRA to Ultra-Low Bits Authors: Amir Reza Mirzaei, Yuqiao Wen, Yanshuai Cao, Lili Mou
-
LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration Authors: Enhuai Liu, Yunke Wang, Yutong Wang, Changming Sun, Chang Xu
-
OrderMoE: An expert similarity driven distributed edge MoE inference Authors: Xin Yuan, Ning Li, Quan Chen, Wenchao Xu, Song Guo
-
Accelerating Time Series Foundation Models with Speculative Decoding Authors: Pranav Subbaraman, Fang Sun, Jinxi Yu, Yue Yao, Huacong Tang, Xiao Luo, Yizhou Sun
-
Trie Automata for Constrained Decoding over Large Finite Sets Authors: Xingzi Xu, Karim Bouyarmane
-
QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving Authors: Yilin Liu, Rui Meng, Wangze Ni, Jianxin Yan, Heng Cao, Libin Zheng, Peng Cheng, Jinfei Liu
-
XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication Authors: Wooseong Yang, Wei-Chieh Huang, Weizhi Zhang, Yu Wang, Philip S. Yu, Junhyun Lee
MoE Training (1)
1. TradingMoE: Routing the Right Experts in Evolving Markets
ArXiv ID: 2608.11785
Primary Topic: MoE Training
Authors: Chang Zhou, Xingtong Yu, Minbin Huang, Zhennan Wu, Yuan Fang, Hong Cheng, Xinming Zhang
Abstract: Large language models (LLMs) have shown strong potential for financial analysis and trading, but direct trading remains challenging because the predictive capabilities required can vary across assets, decision fields, and market conditions. Existing LLM-based trading systems either coordinate human-defined external experts or adopt conventional internal Mixture-of-Experts (MoE) routers that do not directly evaluate how individual experts contribute to trading decisions. Moreover, these routers receive no direct signal indicating when an inactive expert has become more suitable as market conditions change. We find that native router scores poorly reflect how much individual experts improve trading decisions, frequently leaving better alternatives unselected. We further reveal that token-specific expert usefulness exhibits a compact low-dimensional structure. Based on these findings, we propose TradingMoE, a trading-oriented sparse MoE that augments a frozen dense LLM with lightweight residual experts. We introduce a Query-Key router that represents the expertise required by each token under the current market context as a low-dimensional query and matches it with learnable expert keys. We further propose a sparse expert selection update mechanism that samples a few inactive experts during training and estimates whether they should replace the weakest expert in the current Top-k route. This mechanism enables the router to update expert selection as market conditions change while preserving sparse computation. Experiments against 22 baselines on stock and cryptocurrency markets show that TradingMoE improves cumulative return over the best-performing baselines by 30.89% and 30.7%, respectively. Rolling paper-trading experiments further demonstrate that its advantage persists under forward-only deployment.
Comment: Query-key routing probes inactive experts to improve sparse top-k expert selection during training.
Topic Match: New internal routing and expert-selection updates constitute a direct MoE contribution, although evaluated on trading with a frozen backbone.
Relevance: 8 Novelty: 7
Large-Scale Training Systems and Efficiency (5)
1. Spend Less, Fit Better: Budget-Efficient Scaling Law Fitting via Active Experiment Selection
ArXiv ID: 2604.22753
Primary Topic: Large-Scale Training Systems and Efficiency
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Sijie Li, Shanda Li, Haowei Lin, Weiwei Sun, Ameet Talwalkar, Yiming Yang
Abstract: Scaling laws are used to plan multi-million-dollar training runs, but fitting those laws can itself cost millions. In modern large-scale workflows, assembling a sufficiently informative set of pilot experiments is already a major budget-allocation problem rather than a routine preprocessing step. We formulate scaling-law fitting as budget-aware sequential experimental design: given a finite pool of runnable experiments with heterogeneous costs, choose which runs to execute so as to maximize extrapolation accuracy in a high-cost target region. We propose $\mathrm{SL}^2$ (Scaling Laws, Spend Less), an uncertainty-aware method for sequentially allocating experimental budget toward the runs most useful for target-region extrapolation. Across a diverse benchmark of scaling-law tasks, $\mathrm{SL}^2$ outperforms classical design-based baselines, and often approaches the performance of fitting on the full experimental set while using only about 10\% of the total training budget. Our code is available at https://github.com/PlanarG/active-sl.
Comment: Uncertainty-aware experiment selection fits scaling laws using about 10% of the full pilot-training budget.
Topic Match: Budget-aware scaling-law fitting directly supports planning large training runs; reduced pilot-experiment cost also matches efficiency.
Relevance: 9 Novelty: 7
2. Reducing Per-Sample Interference in Stochastic Optimization
ArXiv ID: 2607.16261
Primary Topic: Large-Scale Training Systems and Efficiency
Also Matches: Architecture and Training Dynamics
Authors: Apostolos Avranas
Abstract: Modern optimizers combine gradients from the current mini-batch with historical optimization state, such as momentum or adaptive moments. While effective, this standard practice can produce parameter updates that actively increase the loss of individual samples. We term this phenomenon per-sample interference and propose redefining the parameter update as an optimization problem that explicitly minimizes it. Because the exact formulation of the problem is computationally prohibitive, we introduce a highly efficient surrogate. By reducing the problem's dimensionality to the batch size and restricting the optimization to the last linear layer, we overcome memory and speed bottlenecks. This strategy hinges on our unexpected finding that this layer alone can reliably capture core second-order statistics of the full network. The resulting surrogate problem integrates readily into standard optimizers like SGD and AdamW, and can be solved using a small number of GPU-friendly iterations. Crucially, the method exhibits favorable scaling properties, as the relative computational overhead shrinks as the model size or input grows. Experiments on image classification benchmarks confirm reduced per-sample interference and improved generalization.
Comment: A surrogate using the final linear layer reduces per-sample interference in SGD and AdamW updates.
Topic Match: The core is a scalable optimizer correction supported by analysis of training interference, with empirical evidence from image classification.
Relevance: 8 Novelty: 7
3. Epiplexity Guided Data Selection and Generation for Out-of-Distribution Generalization
ArXiv ID: 2608.11746
Primary Topic: Large-Scale Training Systems and Efficiency
Authors: Ellen Su, Andres Potapczynski, Shikai Qiu, Edward Hughes, Andrew Gordon Wilson
Abstract: Modern systems are increasingly expected to transfer across tasks not specified during training. What data facilitates generalization in these new, unanticipated settings? One hypothesis is that data with more structural information could contain shared circuits and subprograms that could be recycled in a wider array of downstream settings. Epiplexity, a recently proposed measure of the structural information a compute-bounded learner can extract from data, provides a mechanism to reason about this relationship. In this paper, we show how to operationalize epiplexity as an online training signal for data selection and synthetic data generation. For selection, we fit scaling laws to the training loss curves of natural data domains to predict the expected epiplexity gain as a function of training tokens, and use this signal to adaptively determine the sampling weights over domains during training. For synthetic data generation, we define a generator's reward as the change in learner epiplexity over a buffer of previously generated data and use REINFORCE policy gradients to guide the generator toward an epiplexity-maximizing distribution. In both cases, higher epiplexity predicts improved downstream performance on zero-shot and fine-tuning based tasks, supporting the hypothesis that data rich in structural information yield representations that transfer across domains.
Comment: Scaling-law fits to training loss curves guide adaptive sampling weights across data domains.
Topic Match: Scaling-law-guided allocation of training tokens fits training-run configuration, although the principal objective is transfer generalization rather than throughput.
Relevance: 7 Novelty: 8
4. MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning
ArXiv ID: 2608.11749
Primary Topic: Large-Scale Training Systems and Efficiency
Also Matches: Architecture and Training Dynamics
Authors: Shiji Zhou, Kunlin Lyu, Lei Zhang, Ruodong Wang, Yifan Sun
Abstract: Multi-objective optimization (MOO) has demonstrated significant success in multi-task learning by mitigating task conflicts through gradient manipulation. However, most existing methods flatten model parameters into vectors and perform gradient manipulation under Euclidean geometry, thereby overlooking the matrix structure prevalent in modern architectures such as Transformers. In this paper, we show that gradient manipulation in Euclidean space does not generally yield the steepest descent direction under matrix geometry, potentially limiting optimization efficiency. Drawing from the theory of steepest descent for matrix-valued parameters, we propose MOON (Multi-Objective OrthoNormalized Updates), which performs gradient manipulation under spectral--nuclear norm geometry and uses the orthonormalized manipulated gradient for parameter updates. Theoretically, for smooth non-convex objectives, we establish convergence of the averaged Pareto-stationarity measure at rates of $\mathcal{O}(T^{-1/2})$ in the deterministic setting and $\mathcal{O}(T^{-1/4})$ under stochastic gradients. Empirical results across various benchmarks show that MOON consistently improves both optimization efficiency and final multi-task performance. Our code is available at https://github.com/KunlinLyu/MOON.
Comment: Spectral-nuclear norm geometry produces orthonormalized multi-objective optimizer updates that respect matrix parameter structure.
Topic Match: Matrix-aware optimizer design is central, and its descent analysis informs training dynamics; evidence concerns multitask benchmarks.
Relevance: 7 Novelty: 7
5. A-3PO: Accelerating Asynchronous LLM Training with Staleness-aware Proximal Policy Approximation
ArXiv ID: 2512.06547
Primary Topic: Large-Scale Training Systems and Efficiency
Authors: Xiaocan Li, Shiliang Wu, Zheng Shen
Abstract: Decoupled PPO has been a successful reinforcement learning (RL) algorithm to deal with the high data staleness under the asynchronous RL setting. Decoupled loss used in decoupled PPO improves coupled-loss style of algorithms' (e.g., standard PPO, GRPO) learning stability by introducing a proximal policy to decouple the off-policy correction (importance weight) from the policy update constraint (trust region). However, the proximal policy requires an extra forward pass through the model at each training step, creating a computational overhead for large language models training. We observe that since the proximal policy only serves as a trust region anchor between the behavior and target policies, we can approximate it through simple interpolation without explicit computation. We call this approach A-3PO (APproximated Proximal Policy Optimization). A-3PO eliminates this overhead, accelerating training by 1.8x speedup while maintaining comparable performance. Code \& off-the-shelf example are contributed to the open-source RL training system AReaL at: https://github.com/areal-project/AReaL/blob/v1.0.0.rc1/docs/algorithms/prox_approx.md
Comment: Approximates the proximal policy to eliminate an extra forward pass in asynchronous LLM training.
Topic Match: The core algorithm reduces asynchronous training compute, making it relevant despite its narrower reinforcement-learning setting.
Relevance: 7 Novelty: 6
Architecture and Training Dynamics (16)
1. MARCH: Scaling Recurrent Memory with Content-Routed State Anchors
ArXiv ID: 2608.12435
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Ming Zhang, Kaisen Yang, Shu Yu, Ermo Hua, Ning Ding, Xia Hu, Bowen Zhou, Chaochao Lu, Youbang Sun
Abstract: Transformers owe much of their strong long-context retrieval capability to a token-level memory that grows with context length. This flexibility, however, incurs a quadratic computation complexity during training and a key--value cache that grows linearly during autoregressive inference. Recurrent alternatives offer efficient decoding by compressing the entire history into a fixed-size state, but often underperform on recall-intensive tasks since earlier associations usually get overwritten by subsequent updates, and only the most recent contextual information is retained. In this paper, we introduce Memory-Anchor Routing across Context History (MARCH), a network architecture that effectively scales state-space models beyond a fixed-size dimension, while maintaining computational efficiency over long-sequences. MARCH periodically caches cumulative recurrent-state checkpoints as state anchors and associates each anchor with a compact, content-conditioned anchor key. This lets MARCH maintain a memory bank, which can grow as context length increases, providing a controllable trade-off between historical resolution and memory cost. At each token, MARCH produces an anchor query to attend all causally available state anchors, and the output is calculated as an attention-style aggregation over all historical anchors along the current state. We show that after standard pretraining, MARCH consistently outperforms multiple linear attention variants across commonsense reasoning, LongBench, and in-context retrieval. These results demonstrate that content-routed state caching substantially strengthens recurrent long-range memory while preserving its native computation path.
Comment: Content-routed recurrent-state checkpoints expand state-space models beyond a single fixed-size state.
Topic Match: The contribution changes recurrent sequence computation through attention over state anchors, with a controllable trade-off between historical resolution and memory cost.
Relevance: 9 Novelty: 8
2. Disentangling the Expressivity of RoPE
ArXiv ID: 2608.11909
Primary Topic: Architecture and Training Dynamics
Authors: Selim Jerad, Anej Svete, Jiaoda Li, Ryan Cotterell
Abstract: Two accounts recur in explanations of the success of rotary position embeddings (RoPE). Expressivity studies associate periodic position information with modular predicates, whereas mechanistic and long-context studies emphasize positional anchors and local offsets. We formalize both accounts for fully uniform, finite-precision soft-attention transformers. We find that, if every rotary component is periodic, RoPE transformers recognize exactly the languages definable in past temporal logic with modular predicates. Conventional RoPE is different: The rotations it computes never repeat. This yields a precision-dependent bounded simulation of fixed-offset look-back operators, rather than an all-length modular characterization. Controlled experiments match this separation: Constructed periodic schedules length-generalize on modular languages, while conventional RoPE behaves more like a bounded locality bias and can impair tasks requiring position-invariant access to distant context. Altogether, our findings shed light on RoPE transformers, bringing theoretical expressivity characterizations closer to models used in practice.
Comment: Distinguishes periodic RoPE's modular expressivity from conventional RoPE's precision-bounded fixed-offset access.
Topic Match: Directly analyzes a positional-encoding mechanism and explains its consequences for locality and length generalization.
Relevance: 9 Novelty: 8
3. LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining
ArXiv ID: 2608.12419
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Qiuwu Chen, Zimo Liu, Yuchen Li, Ying Sun, Yifan Zhang, Zhijie Qiu, Zeng You, Ryan Dong, Simeng Ma, Yaofo Chen, Mingkui Tan
Abstract: Large language models (LLMs) have achieved remarkable breakthroughs across various applications. However, their architectures remain inefficient in pretraining due to two main limitations: (i) self-attention lacks an explicit inductive bias for locality, leading to redundant modeling of sequence-internal local information; (ii) mixture-of-experts (MoE) implicitly couples knowledge storage with computational pathways, hindering flexible access to sequence-external global knowledge. To overcome these limitations, we propose LoKiFormer, a novel LLM architecture that augments the standard decoder with two dedicated modules: 1) Local Fusion Attention (LFA), which incorporates a convolutional fusion to attention, explicitly capturing local patterns and allowing the attention to operate on more informative representations; 2) Knowledge Memory Module (KMM), which introduces a parametric key-value memory that explicitly stores global knowledge in addressable slots, decoupling storage from computation and enabling direct knowledge retrieval. Together, these modules enable LoKiFormer to achieve more efficient and effective integration of information at both levels. Experimental results show that LoKiFormer converges 1.33x faster in pre-training than baseline models, underscoring its superiority over existing LLM architectures.
Comment: A decoder redesign combines locality-aware attention and addressable knowledge slots to accelerate pretraining convergence.
Topic Match: The core contribution changes decoder computation during LLM pretraining; the reported 1.33x convergence improvement also supports an efficiency match.
Relevance: 9 Novelty: 7
4. ATMA: Long-Context Language Modeling via Polar Attention and Gated-Delta Compression Memory
ArXiv ID: 2606.25156
Primary Topic: Architecture and Training Dynamics
Authors: Habibullah Akbar
Abstract: Length extrapolation in language models involves competing objectives: retrieval fidelity, long-document likelihood, short-context quality, and inference cost. We present ATMA, a 378M-parameter hybrid recipe that combines Polar Attention with gated-delta recurrent memory, and study these objectives as a Pareto problem rather than claiming general architectural dominance. Polar Attention separates a normalized direction channel from a bounded participation-ratio magnitude channel. We select the recipe with a complete 120-cell, 1B-token factorial sweep, then train matched NoPE, RoPE, and Polar variants for 9.816B tokens at length 2K and evaluate them through 256K. Across the factorial, memory improves Polar's 64K retrieval score in all 20 matched cells (mean +47.8 points), whereas its effect on NoPE is small and inconsistent. At 256K, Polar retains 34.4% teacher-forced target-token accuracy and 9.0% exact five-token accuracy; exact retrieval is 18.0% on synthetic contexts but 0.0% on FinePDFs contexts. Polar also limits mean fixed-target bits-per-byte degradation to 1.26 times, at a 1.9-point mean cost on eight short-context tasks. Raven baselines lead BABILong and have length-independent decode state, illustrating a different point on the frontier. Finally, a post-hoc checkpoint audit shows that nearly identical 2K validation curves can conceal a 6.70-nat difference at 256K. Because those runs were neither seed-paired nor randomized across devices, we interpret this as checkpoint variability associated with an infrastructure transition, not a causal hardware effect. Code: https://github.com/kreasof-ai/atma
Comment: Polar Attention separates normalized direction from bounded magnitude to change long-context attention behavior.
Topic Match: The attention mechanism and controlled study of its interaction with gated-delta recurrence make architecture the strongest fit.
Relevance: 9 Novelty: 7
5. Do Transformers Need Three Projections? Systematic Study of QKV Variants
ArXiv ID: 2606.04032
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Ali Kayyam, Anusha Madan Gopal, M Anthony Lewis
Abstract: Transformers have become the standard solution for various AI tasks, with the query, key, and value (QKV) attention formulation playing a central role. However, the individual contribution of these three projections and the impact of omitting some remain poorly understood. We systematically evaluate three projection sharing constraints: a) Q-K=V (shared key-value), b) Q=K-V (shared query-key), and c) Q=K=V (single projection). The last two variants produce symmetric attention maps; to address this, we also explore asymmetric attention via 2D positional encodings. Through experiments spanning synthetic tasks, vision (MNIST, CIFAR, TinyImageNet, anomaly), and language modeling (300M and 1.2B parameter models on 10B tokens), we discovered that our transformers perform on par or occasionally better than the QKV transformer. In language modeling, Q-K=V projection sharing achieves 50% KV cache reduction with only 3.1% perplexity degradation. Crucially, projection sharing is complementary to head sharing (GQA/MQA): combining Q-K=V with GQA-4 yields 87.5% cache reduction, while Q-K=V + MQA achieves 96.9%, enabling practical on-device inference. We show that Q-K=V preserves quality because keys and values can occupy similar representational spaces and attention operates in a low-rank regime, whereas Q=K-V breaks attention directionality. Our results systematically characterize projection sharing as an underexplored instance of weight tying in attention, with direct, quantifiable inference memory benefits, particularly valuable for edge deployment. The code is publicly available at https://github.com/Brainchip-Inc/Do-Transformers-Need-3-Projections
Comment: Analyzes attention projection sharing, explaining when tied keys and values preserve quality and reduce KV-cache memory.
Topic Match: Attention projection design and its mechanistic analysis are central; measured cache reductions provide a substantial secondary efficiency contribution.
Relevance: 9 Novelty: 6
6. Weak Correlations as the Underlying Principle for Linearization of Gradient-Based Learning Systems
ArXiv ID: 2401.04013
Primary Topic: Architecture and Training Dynamics
Authors: Ori Shem-Ur, Khen Cohen, Aviv Orly, Yaron Oz
Abstract: Deep learning models, such as wide neural networks, can be conceptualized as nonlinear dynamical physical systems characterized by a multitude of interacting degrees of freedom. Such systems in the infinite limit, tend to exhibit simplified dynamics. This paper delves into gradient descent-based learning algorithms, that display a linear structure in their parameter dynamics, reminiscent of the neural tangent kernel. We establish this apparent linearity arises due to weak correlations between the first and higher-order derivatives of the hypothesis function, concerning the parameters, taken around their initial values. This insight suggests that these weak correlations could be the underlying reason for the observed linearization in such systems. As a case in point, we showcase this weak correlations structure within neural networks in the large width limit. Exploiting the relationship between linearity and weak correlations, we derive a bound on deviations from linearity observed during the training trajectory of stochastic gradient descent. To facilitate our proof, we introduce a novel method to characterise the asymptotic behavior of random tensors.
Comment: Weak cross-derivative correlations explain training linearization and bound SGD trajectories' departures from linear dynamics.
Topic Match: The core result explains parameter-training dynamics in wide networks, making it a direct dynamics match within a restricted theoretical regime.
Relevance: 8 Novelty: 7
7. FLARE++: Low-rank attention with dynamic attention routing
ArXiv ID: 2608.11519
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Vedant Puri, Yongjie Jessica Zhang, Levent Burak Kara
Abstract: Full self-attention is a strong token mixer for PDE surrogates on irregular domains, but its quadratic cost limits its use on high-resolution problems. Efficient latent-attention models such as the Fast Low-rank Attention Routing Engine (FLARE) avoid that cost by routing all N tokens through M << N learned latent queries, but those queries are parameters: once trained, the same learned query templates serve every input. We remove this restriction with FLARE++, a low-rank attention architecture with dynamic token routing. FLARE++ reuses FLARE's own encoder to build its routing queries: learned latent seeds drive one extra encode call that gathers the N input tokens into M input-conditioned queries, and those queries then determine how the same tokens are compressed and redistributed. This preserves FLARE's explicit low-rank factorization and linear O(NM) complexity, and expresses the complete routing operation with standard scaled dot-product attention (SDPA) calls alone. We also provide a multi-GPU context-parallel implementation that shards input tokens across devices without ever gathering the full token sequence on one of them. FLARE++ is competitive across a set of standard PDE surrogate benchmarks, improving on fixed-query FLARE by 24% on average, and it gains 2.3 points of average accuracy on Long Range Arena.
Comment: Constructs input-conditioned latent routing queries while preserving explicit low-rank attention and O(NM) complexity.
Topic Match: Dynamic latent-query construction changes the attention mechanism itself; its linear complexity also provides a substantive efficiency contribution.
Relevance: 8 Novelty: 7
8. A New First-Order Meta-Learning Algorithm with Convergence Guarantees
ArXiv ID: 2409.03682
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: El Mahdi Chayti, Martin Jaggi
Abstract: Learning new tasks by leveraging prior experience is a fundamental trait of intelligent systems. While Model-Agnostic Meta-Learning (MAML) is a leading approach, it suffers from significant computational and memory overhead due to the requirement of computing second-order meta-gradients. We propose \textbf{FO-B-MAML}, a novel first-order variant of MAML derived from a bi-level optimization perspective. Our framework introduces a new expression of the meta-gradient, defined as the derivative of the solution of a perturbed optimization problem. This formulation allows the meta-gradient to be estimated using various finite difference methods; in this work, we propose and analyze two simple yet effective estimators: a forward and a symmetric approximation. Unlike existing first-order methods like FO-MAML and Reptile, which suffer from irreducible bias, we prove that FO-B-MAML converges to a stationary point of the meta-objective. Notably, the symmetric estimator achieves an improved $\mathcal{O}(δ^{2/3})$ bias rate, strictly enhancing previous first-order theory. Furthermore, we demonstrate that the MAML objective violates standard smoothness assumptions; we show instead that its smoothness constant grows with the norm of the meta-gradient. This property theoretically justifies the use of normalized or clipped-gradient methods (SNGDM) over vanilla gradient descent. Our empirical results validate these advancements: FO-B-MAML achieves high accuracy, closely following second-order MAML performance. Crucially, our method bypasses the ``activation bottleneck'' of second-order approaches, maintaining a flat memory footprint even when scaling to deep, activation-heavy CNNs and Transformers.
Comment: Convergent first-order meta-gradient estimators eliminate the activation-storage bottleneck of second-order MAML.
Topic Match: Convergence and nonstandard-smoothness analysis provide substantive training-dynamics insight, with memory savings for deep models; the scope remains meta-learning.
Relevance: 7 Novelty: 8
9. Training Under Challenge: Executable Certificates and Challenge-Closed Optimality for Neural Networks
ArXiv ID: 2608.12655
Primary Topic: Architecture and Training Dynamics
Authors: Farhang Yeganegi, Arian Eamaz, Mojtaba Soltanalian
Abstract: A flat training curve does not reveal whether a neural network has reached a global optimum, is locally trapped, is representation-limited, or is mismatched to its trainer. We introduce Training Under Challenge, an executable-certificate framework in which predeclared, architecture-valid procedures construct complete alternatives in the same certified class and reevaluate the same objective. Any lower-valued candidate is a replayable witness that lower-bounds the checkpoint's empirical global-optimality gap. Passing a finite suite is only suite-relative; global-gap conclusions require a separately justified coverage mechanism. We define a resource-indexed challenge-power modulus that characterizes the largest gap compatible with passage. For squared loss, current block-decrease operators make coverage checkable and yield uniform and realized-residual bounds. We prove the converse frontier: without coverage, a first-order ReLU trainer can reach infinitely many exact conditional head optima while converging to a non-global point. On a channel-gated ResNet-18 distillation problem with known optimum, eight internal challenges cover all 240 audited output directions, and realized-residual bounds lie within factors of 1.74--3.02 of the true gap. Paired predictive certificates separate decoder under-use from representation insufficiency, while quantized-denoising studies demonstrate diagnosis, repair, and current-state recertification.
Comment: Constructive checkpoint challenges provide optimization-gap certificates when a coverage mechanism is justified.
Topic Match: Optimization guarantees and diagnosis of stalled training are central, though demonstrations concern smaller networks and do not establish large-scale pretraining applicability.
Relevance: 7 Novelty: 8
10. TESLA: Taylor Expansion of Sinusoidal Learnable Activations
ArXiv ID: 2608.11970
Primary Topic: Architecture and Training Dynamics
Authors: Daehwa Ko, Jaehyeon Kim, Seunghyun Ham, Jay Hoon Jung
Abstract: The parity problem--deciding whether the number of ones in a binary vector is odd or even--remains challenging for standard neural networks due to linear inseparability and the need for global interactions. We propose TESLA, an activation defined as a learnable combination of sine and cosine terms, enabling explicit control over polynomial degree and selective amplification of high-order components. Theoretically, we show that constraining TESLA's coefficients yields Lipschitz/Rademacher complexity bounds and shapes the training dynamics to emphasize higher-frequency structure. Empirically, on parity with input length n = 32, TESLA attains strong generalization with 100K training samples (approximately 0.002% of the 2^32 input space) and remains robust under heavy corruption, retaining high accuracy with up to 30% label noise. We also compare against periodic and frequency-based baselines (SIREN, SNAKE, and Fourier feature embeddings) on parity and Forrelation. Beyond synthetic structure, TESLA delivers comparable performance on ImageNet-100, indicating that activation-level degree control transfers to more general vision workloads. Code: https://github.com/KAU-QuantumAILab/TESLA
Comment: Learnable sinusoidal activations control polynomial components and bias training toward high-order structure.
Topic Match: A new activation mechanism and analysis of its training dynamics directly fit architecture research, though large-model training benefits remain unestablished.
Relevance: 7 Novelty: 7
11. NAE: Normalizing AutoEncoder
ArXiv ID: 2608.12084
Primary Topic: Architecture and Training Dynamics
Authors: Muhammad Abdur Rafae, Niels Landwehr
Abstract: We consider the setting of Normalizing flows with approximate inverses, an established paradigm spanning both full-dimensional ($d=D$) and bottleneck ($d<D$) settings, and group these models under the term flow autoencoders. We present a theoretical investigation into their training dynamics and prove that the proposed loss used by existing approaches is suboptimal; specifically, both encoder and decoder surrogates must be optimized in alignment with reconstruction loss. Guided by these insights, we propose Normalizing Autoencoder (NAE), which employs a novel conditional loss that aligns the surrogate loss gradient with that of reconstruction loss, directly improving upon the current standard. Extensive experiments across molecule generation, tabular data, and image benchmarks demonstrate that NAE achieves state of the art performance. Our work highlights the importance of loss alignment in flow autoencoders and establishes NAE as a powerful generative framework.
Comment: Aligns encoder and decoder surrogate gradients with reconstruction loss to improve flow-autoencoder training.
Topic Match: The contribution combines training-dynamics analysis with a corrective loss mechanism for a narrower generative architecture family.
Relevance: 7 Novelty: 7
12. GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs
ArXiv ID: 2608.11674
Primary Topic: Architecture and Training Dynamics
Authors: Kai Yang, Jingwei Xu, Wanyu Wang, Kai-Yuan Guo, Zhenbo Yu, Yi Wang, Yu Qiao
Abstract: On-policy rollout methods such as GRPO are central to post-training of large language models, yet they frequently suffer from training instabilities, cross-task capability degradation, and response-length inflation. Although prior work has characterized the subspace geometry of aggregate updates, the stepwise variation of this geometry and its relationship to model performance remain unclear. We introduce Principal-Subspace Overlap, a dimension-corrected measure of individual rollout updates relative to the dominant singular subspaces of pretrained weights. Despite low average overlap, transient spikes often precede performance degradation. To address this, we propose GCPO (Geometrically Constrained Policy Optimization), which applies hard bilateral orthogonal projections to constrain updates to the complementary subspaces, preventing such excursions by construction. Across mathematical reasoning, code generation, and tool-use tasks on Qwen3-8B and GLM4-9B, GCPO consistently outperforms GRPO and recent variants, including DAPO and GSPO, improving over the base models and the strongest baseline by up to 27.69 and 2.37 points, respectively. Furthermore, GCPO preserves general capabilities, eliminates response-length inflation, and stabilizes policy entropy. Our findings provide a new diagnostic lens and a principled design perspective for stable reinforcement learning post-training.
Comment: Hard bilateral projections constrain rollout updates to prevent excursions into dominant pretrained-weight subspaces.
Topic Match: The diagnosis and repair of update-geometry instability provide a concrete training-dynamics contribution, although evidence is confined to RL post-training.
Relevance: 7 Novelty: 7
13. Gauge-Fixing the Forward-Forward Objective: A Whitened Goodness Derived from a Likelihood-Ratio Account
ArXiv ID: 2607.12501
Primary Topic: Architecture and Training Dynamics
Authors: Paolo Giannitrapani
Abstract: The Forward-Forward algorithm trains each layer locally, so that a scalar goodness - the sum of squared activations - is high on real inputs and low on contrastive ones. Under an explicit generative model this goodness is the sufficient statistic of a likelihood-ratio test, and the pairwise form of the objective admits a gauge: a layer can lower its loss by inflating the scale of its weights rather than by separating the two populations. The analysis prescribes the repair - a whitened, scale-invariant goodness trained online within each layer - which we evaluate as a training procedure. Across three corpora, three depths and a fourfold range of layer width (13 seeds per cell), it raises linear-probe accuracy over the standard pairwise objective in every measured cell - by 4 to 7 points on eight of nine corpus-depth combinations - and closes 16-61% of the gap to end-to-end backpropagation. A control isolates the mechanism: Hinton's fixed-threshold loss also bounds the runaway, to a factor of 1.4 against 133, yet tracks the unmodified baseline - invariance to the gauge, not a bound on it, is what pays. Against the strongest published alternative - a sparse, top-k goodness - the derived objective is statistically indistinguishable on two corpora of three, yet only it removes the runaway: sparsity and gauge-invariance are independent axes, and the published variant recovers accuracy while leaving the pathology in place. We state the boundaries we measured, and every prediction was recorded before its experiment with the refutations reported.
Comment: Scale-invariant whitened goodness removes a weight-inflation shortcut in layer-local Forward-Forward optimization.
Topic Match: Diagnosing an objective's optimization pathology and deriving its repair directly fit training dynamics, with large-model applicability still untested.
Relevance: 7 Novelty: 7
14. Chain-of-Thought Shows the Path to a Tree: Realizing Branching Complexity
ArXiv ID: 2608.11716
Primary Topic: Architecture and Training Dynamics
Authors: Debanjan Dutta, Anish Chakrabarty, Swagatam Das
Abstract: Chain of Thought (CoT) lifts the expressive ceiling of bounded-depth Transformers, with characterizations tying the number of CoT steps to circuit complexity classes. What remains largely missing are concrete instantiations with explicit, depth-bounded constructions, and the traversal procedures such characterizations presuppose. We close this gap for branching complexity. We give CoT realizations of depth-first search (DFS) and of Dijkstra algorithm, the latter subsuming breadth-first search, by unique hard-attention decoders of at most two layers, and use them as a shared computational substrate: reusing the DFS decoder yields the Strahler number of an $n$-vertex tree in $2n-1$ steps with four layers, and reusing the Dijkstra decoder yields its width in $n-1$ steps with three. Since computing the Strahler number of a binary tree given as a term is \textsf{NC\textsuperscript{1}}-complete, and our constructions handle arbitrary $n$-ary trees without layer normalization or positional encodings, this is a non-trivial witness for the linear-step regime of the CoT hierarchy. Exploiting the classical bijection between ordered trees and Dyck paths, itself realized by our DFS construction, which emits the path as it traverses, we give independent constructions for both measures on the path representation.
Comment: Explicit shallow hard-attention constructions characterize the algorithmic computation enabled by iterative chain-of-thought steps.
Topic Match: Directly analyzes transformer computational mechanisms through depth-versus-step constructions, although it does not establish training or learnability results.
Relevance: 7 Novelty: 7
15. Exploring Oversmoothing with Householder Matrices
ArXiv ID: 2608.12514
Primary Topic: Architecture and Training Dynamics
Authors: Bhaskar Karol
Abstract: Deep graph neural networks(GNNs) suffer from oversmoothing- a progressive collapse of node representation towards a low information subspace as network depth increases because the normalized graph propagation operator is repeatedly applied directly to the hidden representations. In this work we study Householder Graph Neural Network (HouseGNN). Rather than updating the hidden state like standard GCN, HouseGNN uses the aggregated neighbourhood message solely to estimate a reflection direction; the node embedding is then updated by a Householder reflector followed by GroupSort, yielding a piecewise orthogonal layer that preserves Euclidean norm at every node and at every depth. We prove three core properties: (i) every internal layer preserves the node-wise Euclidean norm; (ii) the Householder reflector is scale scale and sign-invariant in the message; and (iii) pairwise distance between nodes can change through mismatch between node-wise orthogonal operators.
Comment: Message-conditioned Householder reflections and GroupSort produce layers that preserve each node's Euclidean norm.
Topic Match: The contribution is a new graph-layer mechanism with depth-related analysis; norm preservation alone does not establish prevention of oversmoothing.
Relevance: 7 Novelty: 6
16. HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold Networks
ArXiv ID: 2608.12194
Primary Topic: Architecture and Training Dynamics
Authors: Zhao Su, Yuxin Xia, Haoran Li, Jun Shen, Qi Zhu, Qingguo Zhou, Binbin Yong
Abstract: Kolmogorov-Arnold Networks (KANs) enhance nonlinear function approximation by replacing scalar weights with learnable univariate functions. However, assigning an independent function to every connection results in substantial parameter redundancy, limiting their scalability and efficiency. To reduce this redundancy, we introduce \textbf{HY}perbolic \textbf{D}ynamic \textbf{R}epresentation \textbf{A}rchitecture (HYDRA), a parameter-efficient hyperbolic extension of KAN that combines spline-based functional learning with representations in the Poincaré ball. HYDRA maps vector-valued inputs into a bounded hyperbolic latent space, performs KAN-style updates in tangent space, and employs a low-rank prototype block to share functional transformations across hidden dimensions. The resulting hyperbolic representations provide a structured radial coordinate for interpretation, while radius control improves training stability by preventing boundary saturation. Extensive experiments across eight benchmark datasets demonstrate that HYDRA consistently achieves competitive or superior predictive performance while improving parameter efficiency and representation interpretability.
Comment: Low-rank prototype sharing replaces redundant per-connection functions in a hyperbolic KAN architecture.
Topic Match: Introduces an architectural parameter-sharing mechanism with radius-based stability control, though relevance to large-model training remains un demonstrated.
Relevance: 7 Novelty: 6
Efficiency, Compression, and Large-Scale Training (13)
1. CurveFP: Co-Designing Numerical Representation and Product Arithmetic for Language Models
ArXiv ID: 2608.10010
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Ye Qiao
Abstract: Low-precision formats usually optimize scalar fidelity while inheriting conventional product arithmetic. We introduce CurveFP, a block-scaled family that distributes magnitudes across interleaved logarithmic curves. Uniform curve indices make every nonzero product an exact sign and integer-index update, while a rational radix exposes the finite phase schedule required for accumulation. We instantiate the algebra as CurveFP8 E4C3/E5C2 for training and CurveFP7 E3C3 for compact inference. On four 7B-9B models, CurveFP7 beats tensorwise FP8 perplexity with one fewer element bit and stays within 1.32% of native quality. CurveFP8 lowers error in all 36 paired training-GEMM comparisons. Across three matched 3B-token pretraining triplets, it reaches mean BF16-inference perplexity 22.5366 versus 22.5407 for FP8 and has a lower format penalty in every seed. Downstream evaluation shows transfer parity and a consistent WikiText-103 gain. In a preliminary 4x4 Nangate45 spatial accelerator tile, CurveFP8 uses one fewer product register and 4.6% less area than timing-closing FP8 at 500 MHz. These results support CurveFP as a numerical and arithmetic co-design, while leaving system-level efficiency to future study.
Comment: Logarithmic numerical formats replace nonzero floating-point products with exact sign and integer-index updates.
Topic Match: Low-precision representation and arithmetic are the core mechanisms, supported by GEMM and pretraining experiments; system-level efficiency gains remain unestablished.
Relevance: 9 Novelty: 8
2. CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution
ArXiv ID: 2608.12629
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Zihao Ye, Yingyi Huang, Hongyi Jin, Bohan Hou, Junru Shao, Zhongming Yu, Jinqi Chen, Meghan Cowan, Shiyi Cao, Shanli Xing, Hanfeng Chen, Vinod Grover, Tianqi Chen, Luis Ceze
Abstract: GPU kernel agents and GPU programming languages have advanced separately, leaving expert kernels difficult to reproduce. Agents usually treat the compiler as a fixed black box and receive only errors, correctness outcomes, and timing, while existing DSLs either hide critical scheduling decisions or expose them through difficult layout abstractions. We present CAKE, a compiler-agent co-design in which agents author CAKE IR, a typed, hardware-explicit schedule representation. CAKE exposes warp roles, memory movement, synchronization, and pipelines while supporting verification, cost modeling, and localized diagnostics. The harness itself evolves: recurring failures become verifier rules, IR primitives, model calibrations, and reusable optimization tactics. In matched implementation-hidden Flash-KMeans clean starts on B200, the best CAKE IR candidate at an 80-million-token budget runs at 1.144x the tuned FlashML baseline, compared with 0.928x for direct CUDA/PTX. Beyond this benchmark, agent-generated Kimi Delta Attention achieves a 2.05x geometric-mean speedup over official FlashKDA and passes end-to-end serving validation. Dispatcher-backed KNN and KMeans improve performance by 1.42x to 2.12x across more than 400 shapes, and four kernel changes are available as upstream PRs. CAKE targets NVIDIA GPUs from Ampere through Blackwell and separates single-shape evolution from library generalization and dispatch.
Comment: A hardware-explicit scheduling IR exposes memory movement, synchronization, and pipelines for agent-driven GPU kernel optimization.
Topic Match: The compiler representation and scheduling mechanisms directly improve kernel execution efficiency; end-to-end evidence concerns serving rather than training runs.
Relevance: 8 Novelty: 8
3. Prof-K: Probabilistic One-Pass Filtering for Efficient Top-k Selection
ArXiv ID: 2608.12573
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Tadeusz Dziarmaga, Witold Sikora, Åukasz Struski, Jacek Tabor, Marcin Mazur
Abstract: Top-k selection is a fundamental computational primitive with applications spanning databases, information retrieval, signal processing, and modern machine learning workloads, including sparse activations and attention pruning. As data sizes grow, existing approaches become inefficient: exact methods incur high memory and compute overhead, while approximate methods often rely on brittle heuristics that degrade under adversarial or heavy-tailed inputs. In this paper, we introduce Prof-K, a fast, scalable, and distribution-agnostic top-k algorithm with probabilistic correctness guarantees. Prof-K performs a single-pass filtering procedure: a small random sample estimates an adaptive threshold, the N input elements are streamed once into a compact buffer, and an exact top-k routine on this buffer recovers the true top-k elements with probability at least 1 - $ε$, where $ε$ > 0 is user specified. We derive high-probability guarantees for correctness and buffer size, together with an approximately optimal sample size that minimizes overhead as a function of N and k. Empirically, Prof-K achieves 1.5x-10x speedups over the highly optimized PyTorch topk and recent RadiK implementations, with the largest gains in the large-scale, small-to-moderate-k regime where prior methods struggle most. Unlike previous approaches, these guarantees hold independently of the input distribution, ensuring robustness to adversarial settings. By relaxing the recall target (e.g., recovering 95% of the true top-k values), Prof-K additionally provides a principled accuracy-speed trade-off. We further demonstrate its impact on training BatchTopK Sparse Autoencoders (SAEs), where top-k selection constitutes a significant portion of the training cost.
Comment: Probabilistic one-pass filtering accelerates top-k selection with explicit correctness and buffer-size guarantees.
Topic Match: The core contribution is a faster computational primitive for sparse activations and pruning, with demonstrated training savings in sparse autoencoders.
Relevance: 8 Novelty: 7
4. WBMM: Windowed Batch Matrix Multiplication for Efficient Large Receptive Field Convolution
ArXiv ID: 2607.02097
Primary Topic: Efficiency, Compression, and Large-Scale Training
Also Matches: Architecture and Training Dynamics
Authors: Wan Song, Wei Zhou, Rui Wang, Jun Yu, Toru Kurihara, Jiajia Xu, Shu Zhan
Abstract: Large kernel depthwise convolutions achieve strong performance but suffer from significant degradation as kernel size grows due to irregular memory access from gather-based computation; while Large Kernel Acceleration (LKA) helps on small feature maps, it becomes counterproductive on large feature maps, even slower than non-accelerated implementations. We propose Windowed Batch Matrix Multiplication (WBMM), which partitions input into contiguous windows and indexes a compact relative position bias table to construct weight matrices, enabling regular memory access via batched matrix multiplication. This yields a unique property: WBMM's throughput improves with larger windows, opposite to depthwise convolutions that degrade with larger kernels. Operator-level benchmarks show WBMM with 14x14 windows outperforms 5x5 depthwise convolution baselines in speed while providing a 7.8x larger per-layer receptive field. Combined with inter-block cross-window communication and hierarchical window reparameterization, WBMM achieves comparable or higher accuracy on ImageNet-1K, COCO, and ADE20K with 1.31-1.88x training speedup, and demonstrates consistent advantages across GPU, CPU, and edge devices without requiring specialized acceleration kernels. Our code is available at https://github.com/wansong-s/WBMM.
Comment: Replaces irregular convolution access with contiguous-window batched matrix multiplication, reporting 1.31-1.88x faster training.
Topic Match: The efficient convolution operator is the primary contribution; cross-window mixing and hierarchical reparameterization also alter the computational architecture.
Relevance: 8 Novelty: 7
5. APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference
ArXiv ID: 2608.11688
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Alish Kanani, Layan Badawi, Umit Y. Ogras
Abstract: Mixture-of-Experts (MoE) models are attractive for edge deployment because they provide high model capacity while activating only a small subset of parameters per token, improving compute efficiency. However, MoE inference at the edge is fundamentally limited by memory. Expert parameters are large and often reside in off-chip memory due to capacity, cost, and power constraints, putting expert loading to the critical path. We present APEX: Adaptive Expert Prefetching, a predictive resource management framework that overlaps expert loading with useful computation. APEX introduces a lightweight prefetch router that predicts candidate experts before the attention block to dynamically fetch additional experts using a learned confidence model. This adaptive strategy achieves over 99% overlap accuracy, significantly outperforming fixed top-k prefetching techniques. APEX supports two execution modes: a correctness-preserving mode that guarantees exact routing semantics, and a stall-free mode that eliminates residual stalls by operating on available experts with negligible impact on application accuracy. Across multiple MoE models, the correctness-preserving mode reduces per-token latency by up to 26% and improves energy-delay product (EDP) by up to 41% over state-of-the-art baselines, while the stall-free mode provides additional efficiency gains with negligible impact on application accuracy. These results establish adaptive, confidence-driven expert prefetching as an effective approach for efficient MoE inference on edge systems.
Comment: Confidence-driven expert prefetching overlaps off-chip expert loading with computation to reduce MoE memory stalls.
Topic Match: The central contribution is memory-efficient MoE inference through predictive expert loading, placing it in efficiency rather than MoE training.
Relevance: 8 Novelty: 7
6. Variational Speculative Decoding: Rethinking Draft Training from Token Likelihood to Sequence Acceptance
ArXiv ID: 2602.05774
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Xiandong Zou, Jianshu Li, Jing Huang, Pan Zhou
Abstract: Speculative decoding accelerates inference for (M)LLMs, yet a training-decoding discrepancy persists: while existing methods optimize single greedy trajectories, decoding involves verifying and ranking multiple sampled draft paths. We propose Variational Speculative Decoding (VSD), formulating draft training as variational inference over latent proposals (draft paths). VSD maximizes the marginal probability of target-model acceptance, yielding an ELBO that promotes high-quality latent proposals while minimizing divergence from the target distribution. To enhance quality and reduce variance, we incorporate a path-level utility and optimize via an Expectation-Maximization procedure. The E-step draws Monte Carlo samples from an oracle-filtered posterior, while the M-step maximizes weighted likelihood using Adaptive Rejection Weighting (ARW) and Confidence-Aware Regularization (CAR). Theoretical analysis confirms that VSD increases expected acceptance length and speedup. Extensive experiments across LLMs and MLLMs show that VSD achieves up to a 9.6% speedup over EAGLE-3 and 7.9% over ViSpec, significantly improving decoding efficiency.
Comment: Optimizes speculative draft training for sequence acceptance through a variational objective and weighted EM updates.
Topic Match: The training objective directly improves speculative-decoding efficiency by increasing accepted draft length.
Relevance: 8 Novelty: 7
7. LoRAQuant: Mixed-Precision Quantization of LoRA to Ultra-Low Bits
ArXiv ID: 2510.26690
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Amir Reza Mirzaei, Yuqiao Wen, Yanshuai Cao, Lili Mou
Abstract: Low-Rank Adaptation (LoRA) has become a popular technique for parameter-efficient fine-tuning of large language models (LLMs). In many real-world scenarios, multiple adapters are loaded simultaneously to enable LLM customization for personalized user experiences or to support a diverse range of tasks. Although each adapter is lightweight in isolation, their aggregate cost becomes substantial at scale. To address this, we propose LoRAQuant, a mixed-precision post-training quantization method tailored to LoRA. Specifically, LoRAQuant reparameterizes each adapter by singular value decomposition (SVD) to concentrate the most important information into specific rows and columns. This makes it possible to quantize the important components to higher precision, while quantizing the rest to ultra-low bitwidth. We conduct comprehensive experiments with LLaMA 2-7B, LLaMA 2-13B, and Mistral 7B models on mathematical reasoning, coding, and summarization tasks. Results show that our LoRAQuant uses significantly lower bits than other quantization methods, but achieves comparable or even higher performance.
Comment: Uses SVD reparameterization to concentrate important LoRA components for mixed-precision quantization.
Topic Match: Adapter compression is the core mechanism, directly reducing the aggregate storage cost of multiple LLM adapters.
Relevance: 8 Novelty: 6
8. LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration
ArXiv ID: 2608.12032
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Enhuai Liu, Yunke Wang, Yutong Wang, Changming Sun, Chang Xu
Abstract: Video diffusion transformers are costly to sample: every denoising step applies self-attention over a long 3D token sequence, a quadratic cost that dominates as resolution and duration grow. Sparse attention reduces this cost without retraining, but existing methods pursue aggressive sparsity, where further speedup costs disproportionately more attention fidelity. We target the opposite end of this trade-off: fix near-lossless fidelity by construction, and remove as much computation as this constraint permits. Two observations make this regime practical: roughly 40% of block interactions can be removed while retaining 99% of the attention mass, and the high-mass support remains stable across denoising steps. We propose LoSA, a training-free sparse-attention method that fixes a retained-mass threshold of 99% rather than a sparsity ratio: it measures exact block attention masses at one early dense step, keeps, for each head and query block, the smallest key/value block set meeting the threshold, and reuses the frozen block indices for all remaining steps. On Wan2.1-1.3B, LoSA alone gives a $1.36\times$ speedup with a 0.06-point VBench Overall drop. The benefit is largest under composition: combined with feature caching, LoSA reaches a $3.2\times$ speedup on HunyuanVideo at a 0.02-point drop, versus 0.32 points for the strongest sparse baseline at comparable speed. Across three video diffusion transformers and speedups up to $3.2\times$, LoSA consistently achieves the best training-free speed-quality trade-off.
Comment: Attention-mass thresholds select sparse blocks whose indices are reused across denoising steps, reducing computation with minimal quality loss.
Topic Match: The core contribution is a sparse-attention selection and reuse mechanism with measured speedups on large video diffusion transformers.
Relevance: 8 Novelty: 6
9. OrderMoE: An expert similarity driven distributed edge MoE inference
ArXiv ID: 2607.17154
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Xin Yuan, Ning Li, Quan Chen, Wenchao Xu, Song Guo
Abstract: Although mixture-of-experts, MoE, models have been increasingly adopted to scale large language models with moderate computation cost, it remains challenging to deploy MoE inference over resource-constrained and bandwidth-limited edge infrastructures. Existing distributed MoE serving methods mainly rely on exact expert placement, caching, replication, or communication scheduling, while overlooking the functional similarity among experts, which provides an opportunity to reduce cross-server token transmission. Therefore, this paper introduces a similarity-aware expert allocation and distributed deployment framework, dubbed OrderMoE, which aims to accelerate edge MoE inference while balancing inference latency, communication overhead, server workload, and inference quality. OrderMoE first constructs an expert similarity model based on router-induced logits representations and partitions experts in each MoE layer into multiple similarity groups. Then, it develops a similarity-aware expert grouping and deployment strategy to improve local similarity coverage across edge servers. Since reducing remote expert invocation and preserving exact inference quality are conflicting objectives, OrderMoE further designs a quality-aware and trajectory-aware runtime server-expert selection algorithm to decide whether a token should invoke its remote target expert or use a feasible local substitute expert. Experimental results on a real distributed edge testbed show that OrderMoE significantly reduces average latency, tail latency, cross-server traffic, and remote expert invocation ratio, while introducing only small and controllable inference quality degradation.
Comment: Similarity-based local expert substitution reduces cross-server token traffic during MoE inference.
Topic Match: Quality-aware expert substitution introduces an approximation mechanism for cheaper inference, making model execution efficiency the best fit despite the narrower edge-serving setting.
Relevance: 7 Novelty: 7
10. Accelerating Time Series Foundation Models with Speculative Decoding
ArXiv ID: 2511.18191
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Pranav Subbaraman, Fang Sun, Jinxi Yu, Yue Yao, Huacong Tang, Xiao Luo, Yizhou Sun
Abstract: Time series forecasting drives operational decisions under tight latency budgets, and autoregressive time series foundation models (TSFMs) increasingly deliver the most accurate forecasts. That accuracy is paid for at inference, since a horizon of $H$ steps takes $\lceil H / P\rceil$ sequential forward passes of a large model, so latency grows with exactly the long horizons these models are prized for. Yet a far cheaper model predicts most next patches nearly as well as the large one, and causal models can verify a block of future patches in one parallel pass even though they generate them one at a time. These are precisely the conditions under which speculative decoding thrives in LLMs, but its ingredients are all defined over discrete vocabularies. We therefore develop speculative decoding for continuous patch autoregression. A cheap draft proposes $K$ future patches, and the target verifies all of them in a single causal pass, accepting each by a log-domain Gaussian likelihood-ratio test and correcting the first rejection with its own prediction. We prove that the accelerated output stays within a squared-error radius of target-only decoding set by an acceptance temperature, and that throughput follows a capped-geometric law that makes speedups predictable before deployment. The method delivers up to $3.0 \times$ inference speedup at accuracy between target and draft across five TSFM families, and we characterize which architectures admit single-pass verification and when speculation does not pay.
Comment: Extends speculative decoding to continuous patches with Gaussian likelihood-ratio acceptance and a target-relative error bound.
Topic Match: The continuous-output acceleration algorithm and its error and throughput analysis establish an efficiency contribution, despite time-series-specific validation.
Relevance: 7 Novelty: 7
11. Trie Automata for Constrained Decoding over Large Finite Sets
ArXiv ID: 2608.12574
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Xingzi Xu, Karim Bouyarmane
Abstract: Large language models increasingly need to generate structured outputs that conform to predefined schemas, with one common constraint being selection from a finite set of valid strings. Current constrained decoding systems handle this through general-purpose grammar compilation, which becomes prohibitively slow as the number of valid values grows into the thousands, a cardinality wall. We introduce the trie automaton, a specialized mechanism that exploits finite-set structure (shared prefixes, bounded depth, known cardinality) via Aho-Corasick multi-pattern matching to precompute per-node token masks. The trie achieves 7X faster per-step valid-token computation (0.65 us vs. 5.8 us) compared to XGrammar, one of the primary backends in vLLM and SGLang, and 2--6.5X faster compilation at K >= 300. Because precomputed masks enable a stateless serving path that bypasses the guided decoding pipeline, this advantage compounds in batch serving: end-to-end vLLM throughput reaches 219 req/s vs. XGrammar's 7.5 req/s at batch size 256 (29X). The 29X combines the algorithmic speedup with integration-path savings that only precomputed masks can unlock. Across seven tokenizer families (32K--262K vocabulary), the trie maintains sub-100ms compilation up to K = 10,000 and flat per-step cost regardless of set size, while guaranteeing 100% output validity.
Comment: Prefix-sharing trie automata precompute token masks to accelerate finite-set constrained LLM decoding.
Topic Match: The specialized decoding algorithm reduces LLM execution overhead, with benefits confined to finite sets of allowed outputs.
Relevance: 7 Novelty: 6
12. QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving
ArXiv ID: 2608.12121
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Yilin Liu, Rui Meng, Wangze Ni, Jianxin Yan, Heng Cao, Libin Zheng, Peng Cheng, Jinfei Liu
Abstract: Retrieval-Augmented Generation (RAG) repeatedly prefills identical text chunks across queries, incurring redundant computations. Position-Independent Caching (PIC) mitigates it by reusing precomputed Key-Value (KV) across positions, but its efficiency is constrained by the large volume of text tokens. Rendering text chunks as images can compress the text into fewer visual tokens, but the rendered-image PIC suffers more severe quality degradation than the text PIC. This representation-specific gap primarily arises from contextual mismatches across independently compiled caches and the loss of fine-grained textual evidence during visual compression. Existing PIC repair methods mainly address the former through selective recomputation, but they incur online computation and cannot recover lost textual details. We propose QV-PIC, a query-aware dual-resolution PIC reuse framework guided by model-native templates. Offline, QV-PIC compiles visual caches under the model's native chat-template prefix, improving PIC quality without online recomputation. Online, it preserves global context with low resolution and restores fine-grained textual evidence within a high-resolution budget by cumulative query relevance scores, retaining the efficiency benefit of visual compression. Across six tasks, QV-PIC improves average F1 by 21.6 points over vanilla rendered-image PIC, closes the gap to vanilla text PIC, and surpasses optimized text PIC by 2.58 F1 while reducing TTFT by 17.2\%. Relative to full prefill, it cuts TTFT by 83.8%.
Comment: Uses query-aware dual-resolution visual KV caches to reduce repeated prefill computation.
Topic Match: The core contribution is a cache-reuse and visual-compression mechanism that reduces LLM prefill cost, although its scope is specialized to multimodal RAG serving.
Relevance: 7 Novelty: 6
13. XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication
ArXiv ID: 2608.11676
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Wooseong Yang, Wei-Chieh Huang, Weizhi Zhang, Yu Wang, Philip S. Yu, Junhyun Lee
Abstract: Heterogeneous multi-agent LLM systems, where agents are powered by different model families, can outperform homogeneous configurations by reducing redundant reasoning patterns. Yet existing communication protocols either operate through text, discarding the sender's internal representations, or require architectural homogeneity for latent-level transfer. We identify the entity grounding problem in cross-architecture communication: cross-attention bridges that transfer continuous representations across different LLM families suffer from rare-token compression collapse, where entity identity is lost in the continuous bottleneck (bridge-only F1 ~30%). We propose XBRIDGE, a decode-free communication protocol that addresses this through two mechanisms. Lexical Anchor Mapping (LAM) maps the sender's original context tokens to the receiver's vocabulary, providing discrete entity anchors. A Latent Enrichment Bridge (LEB) lets the receiver query the sender's hidden states for contextual enrichment. The entity anchors ground the bridge's contextual signals to specific entities through the receiver's own self-attention. Across three model families (Llama, Qwen, and Mistral), seven benchmarks, and both communication directions, XBRIDGE outperforms text-based communication on all seven tasks for each model pair while achieving 11x lower latency, and in a same-architecture setting it also exceeds a KV-sharing baseline on six of seven tasks. LEB requires only 264M trainable parameters (3.8% of the receiver), is trained on a small balanced sample set, and adds negligible inference overhead.
Comment: Combines lexical anchors with latent transfer to avoid text decoding during inter-model communication.
Topic Match: Decode-free communication introduces a concrete inference-cost mechanism, with relevance concentrated in heterogeneous multi-agent systems.
Relevance: 6 Novelty: 7
Paper Selection Prompt
System Prompt
You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.
User Prompt
Instructions
Respond in JSONL. Output exactly one JSON object per paper, one per line:
{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}
Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only:
daily_hot,new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return[]. -daily_hotmeans the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. -new_frontiermeans the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.Scoring Criteria
Relevance and Novelty are independent axes. Score both from 1 to 10.
Relevance Scoring
- 9-10: directly centered on the target foundational topics; highest when the core contribution is clearly within them.
- 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
- 5-6: touches the target topics, but the main contribution is elsewhere.
- 3-4: largely outside the target topics, often application-focused or domain-specific.
- 1-2: unrelated.
Important: Broad frontier relevance, major launch status, or daily buzz is not enough for a high Relevance score here. Those cases belong in the hotspot digest unless the paper strongly matches the specialized paper topics.
Novelty Scoring
- 9-10: new paradigm, theory, or major methodological breakthrough.
- 7-8: substantial methodological advance or strong new insight.
- 5-6: meaningful but incremental extension or refinement.
- 3-4: minor, narrow, or mostly engineering or domain-specific improvement.
- 1-2: little originality; mainly standard application of existing methods.
Topic Registry
Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - moe_training: MoE Training - Mixture-of-Experts training: routing, load balancing, expert granularity and shared experts, training stability, MoE parallelism and communication, and MoE kernels. - training_systems: Large-Scale Training Systems and Efficiency - Large-scale training systems and efficiency: distributed training algorithms, optimizers, and communication. - architecture_training: Architecture and Training Dynamics - Core architectural or computational mechanisms, dynamic computation, and training-stability dynamics. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, memory or cache efficiency, and large-scale training systems that materially change cost or behavior.
Papers
[PAPER LIST HERE]
Relevant Topics
This is a training-side feed. The centre of gravity is Mixture-of-Experts training and the large-scale training systems it sits on; architecture and efficiency work is kept when it bears on how a large model is trained or what it costs to train.
Keep a paper when its CORE CONTRIBUTION falls in one of the four topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.
MoE Training (primary) - Keep: expert routing and gating mechanisms; load balancing, with or without auxiliary losses; expert granularity, fine-grained and shared experts; MoE training stability (loss spikes, router collapse, z-loss and related fixes); expert and hybrid parallelism; all-to-all dispatch/combine and the communication schedules MoE needs; MoE kernels and grouped-GEMM style implementations; expert capacity, token dropping, and upcycling dense checkpoints into MoE. - Filter: papers that merely train on top of a MoE model without a new routing, balancing, stability, parallelism, or kernel contribution; "mixture of experts" in the classical ensemble or recommender sense.
Large-Scale Training Systems - Keep: distributed training algorithms (data, tensor, pipeline, sequence, and expert parallelism, FSDP and ZeRO-style sharding); optimizers and preconditioners for large-scale pretraining; communication algorithms, overlap, and collective design; throughput and MFU work that materially changes what a training run costs or whether it converges; scaling-law work that informs how such runs are configured. - Filter: routine infrastructure tuning with no new algorithmic or systems idea.
Architecture and Training Dynamics - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, dynamic or modular computation); optimisation and training-dynamics analysis that explains why large models train the way they do. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight.
Efficiency and Compression - Keep: quantization, sparsity, pruning, low-rank adaptation, KV-cache and memory-efficient designs, and other work that materially changes the cost of running or training a large model, especially when the mechanism is new rather than a tuned variant. - Filter: straightforward application of standard efficiency methods, and deployment or serving engineering with no new idea.
Leave these to the hotspot digest unless the core contribution clearly falls in one of the four topics above: - major model or product releases - agent, tooling, and RAG launches - benchmark, leaderboard, or evaluation-only papers - representation-learning theory, interpretability, and feature-formation analysis - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning) - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains