Previous Day 2026-09-04
Monthly Overview 2026-09
Next Day 2026-09-08

This is a remedial run for missed papers from 09/04/2026 to 09/06/2026.

Results generated on 09/14/2026.

Personalized Daily ArXiv Papers 2026-09-07

Model Metric Usage Papers
Prompt Completion Total Total arXiv Scanned Relevant
local agent CLI Tokens not reported not reported not reported 1045 1045 65
Cost not reported not reported not reported

Token counts are not reported for this run. 45 of 45 model calls succeeded, 7,045s of model wall clock.

Topic Coverage:

TopicPapers
MoE Training1
Large-Scale Training Systems and Efficiency4
Architecture and Training Dynamics30
Efficiency, Compression, and Large-Scale Training30

Table of contents by topic:

MoE Training (1)

  1. Multi-Modal Time Series Prediction via Mixture of Modulated Experts Authors: Lige Zhang, Ali Maatouk, Jialin Chen, Karthik Charan Konduri, Leandros Tassiulas, Rex Ying

Large-Scale Training Systems and Efficiency (4)

  1. RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training Authors: Haoran Sun, Yongjian Guo, Zhong Guan, Shuai Di, Xiaodong Bai, Jing Long, Tianyun Zhao, Mingxi Luo, Hongke Zhao, Likang Wu, Xiaotie Deng, Xu Chu, Xi Xiao, Sheng Wen, Yicheng Gong, Junwu Xiong

  2. FedSubMuon: Communication-Efficient Federated LLM Fine-Tuning via Structured Subspace Muon Authors: Shaolong Chen, Youming Tao, Shuzhen Chen, Falko Dressler, Qingqing Ye, Di Wang

  3. Limits of Reliability and Scaling in Language Models Authors: Subhabrata Majumdar

  4. Resilience Beyond Stationary Client Unavailability: Unlocking Efficient and Unbiased Federated Learning Authors: Ming Xiang, Stratis Ioannidis, Edmund Yeh, Carlee Joe-Wong, Lili Su

Architecture and Training Dynamics (30)

  1. Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biases Authors: Jingwen Liu, Ezra Edelman, Surbhi Goel, Bingbin Liu

  2. Bigger Text Encoders Can Hurt CLIP Zero-Shot Performance Authors: Samir Char, Carles Domingo-Enrich, Randall Balestriero

  3. How Does mHC Use Its Residual Streams? Selective Routing and Near-Identity Mixing Authors: Pengxiang Zhao, Xing Li, Xianzhi Yu, Wei Guo, Zhenhua Dong

  4. Convergent Stochastic Training of Multi-Headed Attention and Understanding LoRA Authors: Zhengkai Sun, Dibyakanti Kumar, Alejandro F Frangi, Anirbit Mukherjee, Mingfei Sun

  5. Observability conditions for neural state-space models with eigenvalues and their roots of unity Authors: Andrew Gracyk

  6. DART: Distributional Adversarial Recurrent Training for Algorithm Learning Authors: Hieu Tran Bao, Phung Thanh Dang, Pham Quang Nhat Minh, Hoang Thanh Tung

  7. Dynamics of Gradient Descent with Large Step Size Near a Manifold of Flat Minima Authors: Lachlan Ewen MacDonald, René Vidal

  8. Parity, Sensitivity, and Transformers Authors: Alexander Kozachinskiy, Tomasz Steifer, Przemysław Wałȩga

  9. Branch Geometry and Finite-Radius Sensitivity of Hard-ReLU Training Authors: Xiaoyang Li, Runni Zhou, Xinghao Yan

  10. Can We Change the Stroke Size for Easier Diffusion? Authors: Yunwei Bai, Ying Kiat Tan, Yao Shu, Tsuhan Chen

  11. Fractal and Chaotic Activation Functions in Echo State Networks: Preprocessing Topology Governs the Echo State Property Authors: Rae Chipera, Jenny Du, Irene Tsapara

  12. CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation Authors: Seonghyun Jin, Youngmin Kim, Sunwoo Park, Jong Chul Ye

  13. State-Dependent Lyapunov Analysis of Rank-1 Matrix Factorization Authors: Jaehong Moon

  14. Stability and Generalization of Straight-Through Estimators for Training Two-Layer Quantized Neural Networks Authors: Yiming Ying

  15. When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation Authors: Jiabin Shen, Guang Chen, Chengjun Mao

  16. Euclidean Fourier Neural Operators Authors: Nathanael Bosch, Niklas Frederik Schmitz, Michael F. Herbst

  17. Not Just Oversmoothing: Detecting the Echo Chamber Effect in Graph Neural Networks Authors: Asela Hevapathige, Ahad N. Zehmakan, Asiri Wijesinghe, Saman Halgamuge

  18. DeltaGNN: Graph Neural Network with Information Flow Control Authors: Kevin Mancini, Islem Rekik

  19. Canalization Before Generalization: Grokking as a Dynamical Probe Authors: Yiming Lin

  20. Diffusion-Inspired Reconfiguration of Transformers for Uncertainty Calibration Authors: Manh Cuong Dao, Quang Hung Pham, Phi Le Nguyen, Thao Nguyen Truong, Bryan Kian Hsiang Low, Trong Nghia Hoang

  21. RoPE attention is an exact forward-pass gradient step with softmax intact Authors: Julie Huang, Maggie Chlon, Leon Chlon

  22. ModularPhaseNet: Finite-Cyclic Phase Geometry for Computable Semantic Hierarchy, Direction, and Context Consistency in Standard Transformers Authors: Kiyotaka Kasubuchi, Kazuo Fukiya

  23. DODR: Deterministic Operator-Driven Reasoning in Latent Space Authors: Weicai Huang

  24. A Thermodynamic Theory of Learning Part II: History-Dependent Reachability and Continual Learning Authors: Daisuke Okanohara

  25. Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models Authors: Frederik Berenz

  26. Modal Logic Neural Networks Authors: Antonin Sulc, Noor Naddour

  27. Towards Unified Multimodal Graph Foundation Model: A Bridge-Router-Adapter Based Approach Authors: Sirui Zhang, Yubing Zhou, Xunkai Li, Zekai Chen, Shumeng Li, Wang Luo, Yinlin Zhu, Yujin Gao, Rong-Hua Li

  28. SAGE: Shape-Adapting Gated Experts for Adaptive Histopathology Image Segmentation Authors: Gia Huy Thai, Hoang-Nguyen Vu, Anh-Minh Phan, Quang-Thinh Ly, Tram Dinh, Thi-Ngoc-Truc Nguyen, Nhat Ho

  29. Cross-Task Generalization Between Understanding and Generation in Unified Vision-Language Models: A Controlled Study Authors: Jihai Zhang, Tianle Li, Linjie Li, Zhengyuan Yang, Yu Cheng

  30. Monotonic Kolmogorov-Arnold Networks: A Theoretical and Empirical Study of Monotonicity as an Inductive Bias Authors: Mikhail Krasnov, Blaž Bertalanič, Carolina Fortuna

Efficiency, Compression, and Large-Scale Training (30)

  1. LGQ: Learnable Geometric Quantization for Image Tokenization Authors: Idil Bilge Altun, Mert Onur Cakiroglu, Elham Buxton, Mehmet Dalkilic, Hasan Kurban

  2. Celty: SpMSpV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM Inference Authors: Ruokai Yin, Priyadarshini Panda

  3. BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference Authors: Janghyeon Kim, Minsoo Kim, Kyuhong Shim, Jungwook Choi

  4. All for 1-Bit: Towards Genuine 1-Bit Post-Training Quantization for LLMs Authors: Zhixiong Zhao, Zukang Xu, Guangyu Sun, Lifeng Liu, Dawei Yang

  5. Linear Algebra Foundations of Efficient Attention: A Phase Reversal in Rank Collapse Under SVD Compression Authors: Anjaneya Teja Sarma Kalvakolanu

  6. ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs Authors: Zukang Xu, Zhixiong Zhao, Xing Hu, Jiangyong Yu, Houji Wen, Jun Li, Zhe Jiang, Dawei Yang

  7. Accelerating Diffusion Transformers with Gaussian Process Rectified Feature Cache Authors: Zhirong Shen, Rui Huang, Chang Zou, Shikang Zheng, Jiacheng Liu, Peiliang Cai, Zhengyi Shi, Yaosong Du, Liang Feng, Xiaobing Tu, Jinkui Ren, Xiantao Zhang, Linfeng Zhang

  8. ACE: Adapter Consolidation across Experts for Parameter-Efficient Fine-Tuning of MoE LLMs Authors: Ahin Lee, Sehyun Yun, Joonha Park, Taesik Gong

  9. A Nuclear-Norm Lower Bound for Dithered Scalar Quantization of Matrix Products Authors: Piyush Sao, Narasinga Miniskar, Pedro Valero-Lara, Keita Teranishi, Sudip Seal

  10. Hidden in Plain Sight: The Overlooked Significance of Canonical Elements for Extreme LLM Sparsity Authors: Hyeondo Jang, Kwanhee Lee, Dongyeop Lee, Namhoon Lee

  11. ECOKV: Geometry-Aware KV Cache Eviction via Complementary Diversity Metrics Authors: Chin Ting Hsu, Yu-Syuan Xu, Ling Zou, Hsien-Kai Kuo, Wen-Huang Cheng

  12. One Rate Is Not Enough: Adaptive Anisotropic Learning Rates for LoRA Fine-Tuning Authors: Huiyi Wang, Daijiao Liu, Lina Yao, Dong Gong

  13. An AI4AI Framework for Visual Token Pruning Authors: Zhen Liu, Wenli Huang, Wei Song, Yuhan Liu, Zhiqin Yang, Jingwen Fu

  14. LookThere! Sparse Vision by Reinforced Selection Authors: Sreehari Rammohan, Yousef Yassin, Anthony Fuller, Junfeng Wen, Carl Vondrick, Evan Shelhamer

  15. EStream: Fast and Memory-Efficient MoE Prefill through Expert Virtualization on Mobile NPUs Authors: Junming Zhang, Zhenzhe Zheng, Fan Wu, Xiaoyao Huang, Jie Wu

  16. Interface-Aware KV Cache Quantization for Dense On-Chip NVM in Long-Context LLM Decoding Authors: Jiahao Zheng, Yifan Qin, Xiaobo Sharon Hu, Yiyu Shi

  17. Generating Pretraining Tokens from Organic Data for Data-Bound Scaling Authors: Zichun Yu, Chenyan Xiong

  18. PLUME: Parameter-Efficient Personalization of Large Language Models via Low-Rank User Modulation in Shared Subspaces Authors: Xinyu Li, Hao Zhou, Jianfeng Zhu, Julina Maharjan, Ruixin Guo, Feodor Dragan, Ruoming Jin

  19. STAR-Pro: Stage-Wise Token Adaptive Reduction with Progressive Refinement for Efficient Large Vision-Language Models Authors: Yichen Guo, Tinghao Wang, Qizhe Zhang, Lingbei Meng, Yuan Zhang, Jiajun Cao, Hao Jiang, Chenwei Wu, Jixian Wu, Sixiang Chen, Tao Luo, Hongyang Cheng, Kai Tang, Chenxi Li, Renyuan Li, Xiande Huang, Wenya Wang, Shanghang Zhang

  20. Cost-Optimal LLM Routing with Limited User Feedback under User Satisfaction Guarantees Authors: Herbert Woisetschläger, Arastun Mammadli, Ryan Zhang, Shiqiang Wang

  21. KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU Authors: Di Chai, Leye Wang, Zeshen Su, Zhiguo Xia, Zhihang Yu

  22. SplitLite: Low-Rank Residual Compression for Split Learning Authors: Tao Li, Yulin Tang, Qi Guo, Xianhao Chen

  23. Aplaud: Adaptive Personalized Low-Rank Decomposition for User-Specific LLM Authors: Xinyu Li, Ruoming Jin, Jianfeng Zhu, Ruixin Guo, Zhi Liu

  24. From Deep to Shallow: Unconstrained and Efficient Layer Merging Strategy Authors: Petro Shulzhenko, Gabriele Spadaro, Enzo Tartaglione

  25. RePro: Training Language Models to Faithfully Recycle the Web for Pretraining Authors: Zichun Yu, Chenyan Xiong

  26. SPD: Single Pass Decoding for Generative Reranking Authors: Emil Laftchiev, Prachi Agrawal, Moe Kayali, Bixing Yan, Qi Xu, Zijie Lei, Chen Qiu, Zhi Hua, Ke Li, Luke Simon

  27. Neural Field Tokenizations with Hierarchy and Spatial Locality Priors Authors: Alonso Urbano, David W. Romero, Max Zimmer, Sebastian Pokutta

  28. Coarse-Graining Hidden Representations: Unsupervised Neuron Selection via Mapping Entropy Authors: Margherita Mele, Andrea Castagna, Roberto Menichetti, Raffaello Potestio, Alessandro Ingrosso

  29. Subspace Optimization for Backpropagation-Free Continual Test-Time Adaptation Authors: Damian Sójka, Sebastian Cygert, Marc Masana

  30. InKAN: B-Spline KANs via Truncated Power Form Authors: Naveen Mysore


MoE Training (1)

1. Multi-Modal Time Series Prediction via Mixture of Modulated Experts

ArXiv ID: 2601.21547

Primary Topic: MoE Training

Also Matches: Architecture and Training Dynamics

Authors: Lige Zhang, Ali Maatouk, Jialin Chen, Karthik Charan Konduri, Leandros Tassiulas, Rex Ying

Abstract: Real-world time series exhibit complex and evolving dynamics, making accurate forecasting extremely challenging. Recent multi-modal forecasting methods leverage textual information such as news reports to improve prediction, but most rely on token-level fusion that mixes temporal patches with language tokens in a shared embedding space. However, such fusion can be ill-suited when high-quality time-text pairs are scarce and when time series exhibit substantial variation in characteristics, thus complicating cross-modal alignment. In parallel, mixture-of-experts (MoE) architectures have proven effective for both time series modeling and multi-modal learning, yet many existing MoE-based modality integration methods still depend on token-level fusion. To address this, we propose Expert Modulation, a new mechanism for multi-modal time series prediction that conditions both routing and expert computation on textual signals, enabling direct and efficient cross-modal control over expert behavior. Through theoretical analysis and experiments, our proposed method demonstrates strong improvements in multi-modal time series prediction. The current code implementation is available at https://github.com/BruceZhangReve/MoME

Comment: Expert Modulation conditions both expert routing and expert computation on textual signals.

Topic Match: The proposed mechanism changes routing and expert computation directly, supporting a MoE match; its forecasting focus limits relevance to foundational training.

Relevance: 7 Novelty: 6


Large-Scale Training Systems and Efficiency (4)

1. RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training

ArXiv ID: 2602.05765

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Haoran Sun, Yongjian Guo, Zhong Guan, Shuai Di, Xiaodong Bai, Jing Long, Tianyun Zhao, Mingxi Luo, Hongke Zhao, Likang Wu, Xiaotie Deng, Xu Chu, Xi Xiao, Sheng Wen, Yicheng Gong, Junwu Xiong

Abstract: Reinforcement learning (RL) has emerged as a critical paradigm for post-training Vision-Language-Action (VLA) models, enabling embodied agents to adapt and improve through environmental interaction. However, existing RL frameworks for VLAs inherit synchronous design principles from traditional LLM training, treating entire rollouts as indivisible units and alternating strictly between data collection and policy optimization. This fundamentally mismatches the unique characteristics of VLA training, as physical simulators introduce highly variable, resource-intensive latencies. To address this, we introduce RL-VLA$^3$, a fully asynchronous distributed RL framework that enables fine-grained asynchronous interaction between simulation, inference, and training components through dynamic batching schedulers and flexible environment sharding strategies. Extensive experiments across diverse simulation backends, VLA architectures, and RL algorithms demonstrate that RL-VLA$^3$ achieves throughput improvements of up to 85.2\% over synchronous baselines while maintaining identical sample efficiency, with scalability validated from 8 to 256 GPUs. To our knowledge, RL-VLA$^3$ is the first fully asynchronous RL training framework tailored specifically for the system-level challenges of VLA training.

Comment: Fine-grained asynchronous scheduling and environment sharding decouple simulation, inference, and distributed training.

Topic Match: The core contribution is distributed training orchestration with measured throughput improvements and scaling to 256 GPUs, specialized to VLA reinforcement learning.

Relevance: 8 Novelty: 7


2. FedSubMuon: Communication-Efficient Federated LLM Fine-Tuning via Structured Subspace Muon

ArXiv ID: 2609.06073

Primary Topic: Large-Scale Training Systems and Efficiency

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Shaolong Chen, Youming Tao, Shuzhen Chen, Falko Dressler, Qingqing Ye, Di Wang

Abstract: Federated fine-tuning adapts large language models (LLMs) to decentralized client data, but its scalability in cross-device training is often limited by the high communication cost. Muon is an optimizer that improves optimization performance by orthogonalizing momentum for matrix-valued parameters. Existing federated Muon methods demonstrate the benefit of matrix-aware optimization in federated learning, but still require transmitting full layer-size updates and optimizer state. A natural way to reduce communication is to directly apply Muon to LoRA factors, but this changes the optimized object and weakens Muon's matrix-aware update geometry. We propose FedSubMuon, a communication-efficient federated Muon fine-tuning method that optimizes compact coefficient matrices within shared structured subspaces. This design keeps Muon on a single matrix-valued trainable object, while reducing the client upload to compact coefficient matrices. We further introduce FedSubMuon-GT, an accuracy-oriented extension that uses projected gradients to adapt tracked subspace bases toward task-relevant gradient directions. Experiments on instruction tuning and mathematical reasoning show that FedSubMuon-GT achieves the best overall accuracy on four of five dataset-model pairs, while FedSubMuon performs best under all matched communication budgets. On Dolly-15K, the closest communication baseline requires 5.5 times and 1.4 times more total communication on Llama-1B and Qwen-4B, respectively.

Comment: Structured-subspace Muon reduces communicated update size while retaining matrix-valued optimization in federated LLM fine-tuning.

Topic Match: The optimizer and communication scheme are the central contributions; compact trainable subspaces also provide parameter efficiency.

Relevance: 8 Novelty: 7


3. Limits of Reliability and Scaling in Language Models

ArXiv ID: 2607.14112

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Subhabrata Majumdar

Abstract: Large language models (LLMs) are trained and evaluated as though perfect reliability is achievable for any task given sufficient scale. We show that this assumption is information-theoretically unjustified. Every generative task has a reliability ceiling that no model can exceed, determined by how much output uncertainty is resolvable from observable context. The gap decomposes into a resolvable component closable with additional context and a subjective component inherent to task ambiguity. Autoregressive generation further degrades this ceiling at a rate governed by the task's dependency kernel, which quantifies inter-token correlations in the output. From these two primitives, we derive a first-principles scaling law where LLM performance is bottlenecked by the scarcer resource: training data or model capacity. This law recovers the Chinchilla scaling law as a special case and provides a structural account of when scaling improves reliability. Beyond scaling, our framework unifies diverse practical phenomena, such as the benefits of retrieval-augmentation and the spectral mechanics of catastrophic forgetting. Our work formalizes the resource-complexity tradeoffs that govern model performance across domains, offering a unified theory of performance limits in generative language models.

Comment: Proposes a data-versus-capacity scaling law that recovers Chinchilla as a special case.

Topic Match: The scaling-law component addresses pretraining resource tradeoffs within a broader theory of generation reliability.

Relevance: 7 Novelty: 8


4. Resilience Beyond Stationary Client Unavailability: Unlocking Efficient and Unbiased Federated Learning

ArXiv ID: 2609.04763

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Ming Xiang, Stratis Ioannidis, Edmund Yeh, Carlee Joe-Wong, Lili Su

Abstract: Due to resource constraints or external and internal uncertainties, clients in real-world federated learning systems are often intermittently available edge devices. In highly dynamic environments, the parameter server lacks prior real-time knowledge of clients' availability, making it challenging to adapt traditional federated learning algorithms to be resilient to uncertainties in client availability. If not carefully addressed, complex client availability can introduce significant bias, potentially harming the performance of the trained model. Most prior work either fails to account for non-stationary client availability dynamics or demands significant memory and computational overhead. This paper aims to develop efficient federated learning algorithms that are provably resilient to heterogeneous and non-stationary stochastic client availability. We propose FedSWE, which admits novel algorithmic structures to (i) compensate for missed computations, (ii) stabilize and diffuse the global updates over rounds, and (iii) evenly mix the local updates through implicit gossiping, despite being agnostic to non-stationary dynamics. Compared with the standard FedAvg, FedSWE introduces light additional memory and computation overhead. We show that FedSWE converges to a stationary point of non-convex objectives while achieving the desired linear speedup property in certain special cases. We corroborate our analysis with numerical experiments over diversified client unavailability dynamics on real-world data sets.

Comment: Compensated updates and implicit gossiping enable convergent distributed optimization under non-stationary client availability.

Topic Match: The core contribution is a distributed training algorithm with convergence and conditional linear-speedup guarantees, although edge federated learning is peripheral to large-model pretraining.

Relevance: 7 Novelty: 7


Architecture and Training Dynamics (30)

1. Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biases

ArXiv ID: 2605.20314

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Jingwen Liu, Ezra Edelman, Surbhi Goel, Bingbin Liu

Abstract: This work investigates the ``small-vs-large gap'', where repeating on fewer samples can lead to compute saving during training compared to using a larger dataset. This is observed across algorithmic tasks, architectures and optimizers and cannot be explained using prior theory. We argue that the speedup comes from appropriate layer-wise growth enabled by sampling biases, which is more pronounced when the dataset size is smaller. We provide both theoretical analysis and empirical evidence from various interventions. Our results suggest that using a smaller dataset with more repetitions is not just a fallback strategy under data scarcity, but can be proactively leveraged as a favorable inductive biases for optimization, particularly in reasoning tasks.

Comment: Explains faster training on smaller repeated datasets through sampling biases that enable favorable layer-wise growth.

Topic Match: Its central contribution explains an optimization mechanism through theory and interventions; the resulting compute savings also match training efficiency.

Relevance: 9 Novelty: 8


2. Bigger Text Encoders Can Hurt CLIP Zero-Shot Performance

ArXiv ID: 2609.05730

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Samir Char, Carles Domingo-Enrich, Randall Balestriero

Abstract: Contrastive Language-Image Pretraining (CLIP) is a building block of many machine learning applications. Scaling laws have guided resource allocation for large-scale training, yet prior work treats total CLIP model size as a single variable, without exploring how the capacity split between encoders impacts downstream performance. Here, we train multiple CLIP models with different vision and text encoder sizes, revealing that for most vision encoders, there is an optimal text encoder size beyond which zero-shot performance degrades---even as total parameter count increases. Exploiting this behavior yields efficient configurations that match the zero-shot performance of the standard ViT-B/16 architecture with up to 55% fewer parameters. We further show that this degradation stems from overfitting induced by the oversized text encoder, and that using modality-specific weight decay coefficients not only recovers but improves performance across all degraded configurations. A geometric analysis reveals a trade-off in which scaling the text encoder improves embedding uniformity but worsens cross-modal alignment; we further show that these metrics are predictive of zero-shot performance. We hope these findings motivate CLIP architectures and training methods that counteract this degradation, a prerequisite for scaling CLIP reliably and efficiently.

Comment: Identifies and remedies overfitting caused by oversized text encoders during CLIP scaling.

Topic Match: Encoder-capacity allocation and modality-specific regularization explain pretraining behavior while improving parameter efficiency.

Relevance: 9 Novelty: 7


3. How Does mHC Use Its Residual Streams? Selective Routing and Near-Identity Mixing

ArXiv ID: 2609.05309

Primary Topic: Architecture and Training Dynamics

Authors: Pengxiang Zhao, Xing Li, Xianzhi Yu, Wei Guo, Zhenhua Dong

Abstract: Hyper-Connections and their manifold-constrained variant mHC widen a residual pathway from one stream to n, yet how trained models use this capacity remains unclear: how broadly blocks read and write, how strongly the residual pathway mixes streams, and whether the streams carry distinct representations. We examine these properties in the four-stream residual pathway of DeepSeek-V4-Flash using effective stream counts, cross-stream residual weights, and inter-stream cosine similarity. Read/write routing is concentrated but varies across depth: a typical attention or FFN site effectively uses about two streams, while the dominant stream changes across layers and the representations remain directionally distinct. Residual mixing is modest and occurs primarily in early layers; in layers 22-42, the pathway mostly carries each stream forward separately. Targeted interventions establish the functional significance of these patterns. Replacing the late mixers by identity increases C4 perplexity by only 1.9% and preserves the six-task average score, whereas replacing the early mixers increases perplexity by 41%. Fixing each early mixer to its C4 diagnostic mean increases perplexity by only 0.2% and reduces the average score by 0.25 percentage points, showing that its site-specific structure matters more than its token-wise variation on the evaluated metrics. Likewise, retaining the three largest routing weights per token at every site increases perplexity by at most 2.7% and changes the average score by at most 0.4 points. Thus, the studied model realizes only part of the flexibility afforded by four-stream mHC: individual blocks rarely require all four streams, and late residual mixing provides little measured benefit.

Comment: Residual-pathway interventions show early mHC mixing is important while late mixing is largely dispensable.

Topic Match: Direct mechanistic analysis of how a large model uses multi-stream residual connections, with implications for residual architecture design.

Relevance: 9 Novelty: 7


4. Convergent Stochastic Training of Multi-Headed Attention and Understanding LoRA

ArXiv ID: 2605.07959

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Zhengkai Sun, Dibyakanti Kumar, Alejandro F Frangi, Anirbit Mukherjee, Mingfei Sun

Abstract: Transformers have revolutionized machine learning and deploying attention layers in the model is increasingly standard across a myriad of applications. Further, for large models, it is common to implement Low Rank Adaptation (LoRA), whereby a factorized parameterization of them is trained, to achieve a surprisingly beneficial accuracy-size trade-off. In this work, via a unified framework we rigorously establish trainability of such models under stochastic methods. We prove that for a class of mild regularizations, the empirical regression loss on a attention layer and LoRA on a shallow neural net, both induce Poincaré inequality for the corresponding Gibbs' measure. Crucially, we show that the Poincaré constant is free of the data dimension for LoRA and for multi-head attention it is free of the head dimensions, Then it follows, via invoking recent results, that a certain SDE, which mimics the SGD, minimizes the corresponding losses. In both the cases, our first-of-its-kind results of trainability on attention and nets, do not use any assumptions on the data or the model size.

Comment: Head-dimension-independent Poincaré bounds support convergence of regularized stochastic attention training.

Topic Match: The core contribution is attention and LoRA optimization theory, with guarantees limited to SGD-like diffusion on attention layers and shallow networks.

Relevance: 8 Novelty: 8


5. Observability conditions for neural state-space models with eigenvalues and their roots of unity

ArXiv ID: 2504.15758

Primary Topic: Architecture and Training Dynamics

Authors: Andrew Gracyk

Abstract: We operate through the lens of ordinary differential equations and control theory to study the concept of observability in the context of neural state-space models and the Mamba architecture. We develop strategies to enforce observability, which are tailored to a learning context, specifically where the hidden states are learnable at initial time, in conjunction to over its continuum, and high-dimensional. We also highlight our methods emphasize eigenvalues, roots of unity, or both. Our methods effectuate computational efficiency when enforcing observability, sometimes at great scale. We formulate observability conditions in machine learning based on classical control theory and discuss their computational complexity. Our nontrivial results are fivefold. We discuss observability through the use of permutations in neural applications with learnable matrices without high precision. We present two results built upon the Fourier transform that effect observability with high probability up to the randomness in the learning. These results are worked with the interplay of representations in Fourier space and their eigenstructure, nonlinear mappings, and the observability matrix. We present a result for Mamba that is similar to a Hautus-type condition, but instead employs an argument using a Vandermonde matrix instead of eigenvectors. Our final result is a shared-parameter construction of the Mamba system, which is computationally efficient in high exponentiation. We develop a training algorithm with this coupling, showing it satisfies a Robbins-Monro condition under certain orthogonality, while a more classical training procedure fails to satisfy a contraction with high Lipschitz constant.

Comment: Derives efficient observability constraints for neural state-space models and a shared-parameter Mamba training construction.

Topic Match: The core contribution analyzes state-space architectural mechanisms and their training behavior, including conditional Robbins-Monro properties of the coupled parameterization.

Relevance: 8 Novelty: 7


6. DART: Distributional Adversarial Recurrent Training for Algorithm Learning

ArXiv ID: 2609.05988

Primary Topic: Architecture and Training Dynamics

Authors: Hieu Tran Bao, Phung Thanh Dang, Pham Quang Nhat Minh, Hoang Thanh Tung

Abstract: Recurrent reasoning models (RRMs) can solve structured problems, achieving easy-to-hard generalization through iterative computation in hidden space. These models are typically trained with instance-level supervision, which becomes increasingly problematic as task difficulty grows: valid solutions occupy a tiny region of the solution space, while invalid solutions proliferate rapidly. We propose Distributional Adversarial Recurrent Training (DART), a training framework that replaces single-point supervision with a local target distribution around the ground-truth solution and aligns model outputs with this distribution through an adversarial objective. DART provides a richer learning signal and encourages more stable iterative trajectories toward valid solutions. When evaluated on Maze, Chess, and masked Sudoku with multiple RRMs, including Deep Thinking Systems and Tiny Recursive Models, DART improves solution quality, stability, and robustness under the evaluated distribution shifts. Comparisons with label smoothing, Gaussian softened targets, and progressive training show that DART is not explained by target softening alone and is complementary to training schemes that stabilize long-horizon recurrence. These results identify DART as a promising approach for improving robustness across the evaluated recurrent reasoning models.

Comment: Distributional adversarial supervision improves the stability of iterative recurrent computation.

Topic Match: The central contribution changes recurrent-model learning signals and studies their effects on iteration stability and generalization.

Relevance: 8 Novelty: 7


7. Dynamics of Gradient Descent with Large Step Size Near a Manifold of Flat Minima

ArXiv ID: 2607.08380

Primary Topic: Architecture and Training Dynamics

Authors: Lachlan Ewen MacDonald, René Vidal

Abstract: An important quantity in the theory of gradient descent (GD) is the \emph{sharpness}, defined as the largest eigenvalue of the objective Hessian. Classical analyses typically require the step size to be uniformly smaller than twice the reciprocal of the sharpness, but this condition is frequently violated in the training of deep neural networks. Recent work bridges this gap in the setting of overparametrised least-squares with a \emph{single scalar output}, providing a normal form for large-step GD in a neighbourhood of an \emph{isolated} flat minimum and establishing three corresponding convergence results. In this paper, we extend this theory in two directions: (1) to overparametrised least-squares with \emph{vector-valued outputs} (including regression with arbitrarily many observations), and (2) to a neighbourhood of a \emph{manifold} of flat minima (which we show is essential for applications such as matrix factorisation). We generalise both the normal form and all three convergence theorems of \cite{macdonaldeos} to this broader setting, overcoming several technical challenges. We further show that our framework applies to deep matrix factorisation under mild assumptions, yielding several new structural results. In particular, we prove that the set of flat minima forms a fibre bundle over a product of spheres, and that the sharpness is Morse-Bott along this manifold.

Comment: Extends large-step gradient-descent convergence theory to vector outputs and manifolds of flat minima.

Topic Match: The core results explain optimization stability and sharpness near flat minima, including applications to deep matrix factorization.

Relevance: 8 Novelty: 7


8. Parity, Sensitivity, and Transformers

ArXiv ID: 2602.05896

Primary Topic: Architecture and Training Dynamics

Authors: Alexander Kozachinskiy, Tomasz Steifer, Przemysław Wałȩga

Abstract: Understanding what neural architectures can and cannot compute is a central challenge in the theory of AI. One of the fundamental problems in this context is the PARITY task, which asks whether the number of 1s in a binary input sequence is even or odd. PARITY is one of the central tasks studied in the theory of computation, yet it remains surprisingly unclear under which conditions transformers can or cannot solve it. In this paper, we show that the minimal number of layers a transformer needs to compute PARITY is two. In particular, we solve the open problem asking whether a one-layer transformer can compute PARITY. We answer it negatively by showing that average sensitivity of a one-layer transformer grows slower than that of PARITY. Furthermore, we show a new construction for transformer that computes PARITY, which improves on the existing constructions by removing a number of impractical assumptions. In particular, the existing transformers for PARITY rely on such impractical assumptions as length-dependent positional encoding, hardmax, layernorm without a regularisation parameter, or incompatibility with causal masking. We show that these assumptions can be removed, at the cost of increasing the number of layers from two to four. Specifically, we show that PARITY can be computed by a four-layer transformer using softmax attention, length-independent and polynomially bounded positional encoding, no layernorm, and compatible with both causal and non-causal masking.

Comment: Establishes transformer depth requirements for PARITY, including a four-layer construction using softmax attention and causal masking.

Topic Match: The results isolate how depth and attention assumptions determine architectural capabilities, with indirect implications for training.

Relevance: 7 Novelty: 8


9. Branch Geometry and Finite-Radius Sensitivity of Hard-ReLU Training

ArXiv ID: 2608.30960

Primary Topic: Architecture and Training Dynamics

Authors: Xiaoyang Li, Runni Zhou, Xinghao Yan

Abstract: Outer-learning algorithms use infinitesimal sensitivities to propose finite changes to initialization or training parameters. For hard-ReLU training, the derivative of the finite program and the derivative of its flow limit do not by themselves specify the response at a chosen radius. We characterize the intervening regime in which the perturbation radius is proportional to the GD step. Integer event rounding then survives at leading order: smooth Euler bias shifts each discrete phase, and upstream rounding moves downstream branch boundaries. We derive the crossing indices and a uniform endpoint expansion for finitely many separated transverse events in piecewise-$C^2$ dynamics, away from recursive phase boundaries. In contractive affine regions, an explicit remainder and complete branch verification certify finite candidate comparisons. Scalar phase frequencies and a coupled feedback ablation test the mechanism; frozen nonlinear-network experiments show radius-dependent prediction accuracy, including incomplete branch matches and failed-word tails. Together with local AD and uniform flow consistency, the result identifies sufficient response regimes: differentiating training is a choice of perturbation resolution as well as a choice of derivative.

Comment: Characterizes how finite perturbation radius and discrete event rounding alter gradients through hard-ReLU training trajectories.

Topic Match: It provides mechanistic training-dynamics analysis for nonsmooth neural networks.

Relevance: 7 Novelty: 8


10. Can We Change the Stroke Size for Easier Diffusion?

ArXiv ID: 2603.26783

Primary Topic: Architecture and Training Dynamics

Authors: Yunwei Bai, Ying Kiat Tan, Yao Shu, Tsuhan Chen

Abstract: Diffusion models can be challenged in the low signal-to-noise regime, where they have to make pixel-level predictions despite the presence of high noise. The geometric intuition is akin to using the finest stroke for oil painting throughout, which may be ineffective. We therefore study \emph{stroke-size control} as a controlled intervention that changes the roughness of the supervised target, predictions and perturbations across timesteps, in an attempt to ease the low signal-to-noise challenge via the prediction target simplification.

Comment: Controls timestep-dependent supervision granularity to investigate low-SNR diffusion training.

Topic Match: Directly examines a training mechanism: how target roughness interacts with noise level and prediction difficulty.

Relevance: 8 Novelty: 6


11. Fractal and Chaotic Activation Functions in Echo State Networks: Preprocessing Topology Governs the Echo State Property

ArXiv ID: 2512.14675

Primary Topic: Architecture and Training Dynamics

Authors: Rae Chipera, Jenny Du, Irene Tsapara

Abstract: Contemporary reservoir computing relies heavily on globally Lipschitz, well-behaved activation functions, limiting applications in defense, disaster response, and pharmaceutical modeling where robust operation under extreme conditions is critical. We systematically investigate non-smooth activation functions, including chaotic, stochastic, and fractal variants, in echo state networks. Through parameter sweeps across 36,610 reservoir configurations, we demonstrate that several non-smooth functions not only maintain behavior consistent with the Echo State Property (ESP) but outperform traditional smooth activations in convergence speed and spectral radius tolerance. Notably, the Cantor function (continuous everywhere, zero derivative almost everywhere) maintains ESP-consistent behavior up to spectral radii of rho = 10, an order of magnitude beyond typical bounds for traditional functions, while achieving 2.6x faster convergence than tanh and ReLU. We introduce a theoretical framework for quantized activation functions, defining a Degenerate Echo State Property (d-ESP) capturing stability for discrete-output functions, and prove that d-ESP implies traditional ESP. We conjecture a critical crowding ratio Q=N/k (reservoir size / quantization levels) predicting failure thresholds for discrete activations. Our analysis reveals that preprocessing topology, rather than continuity, determines stability: monotone, compressive preprocessing maintains ESP across scales, while dispersive or discontinuous preprocessing triggers sharp failures. Our findings challenge assumptions about activation function design in reservoir computing; the exceptional performance of certain fractal functions is only partially explained by the effective-gain analysis presented here, suggesting fundamental gaps in our understanding of how geometric properties of activation functions influence reservoir dynamics.

Comment: Analyzes how activation preprocessing controls recurrent-state stability beyond conventional spectral-radius regimes.

Topic Match: Activation mechanisms and recurrent-state stability are the core contribution, although the evidence concerns reservoir networks rather than large pretrained models.

Relevance: 7 Novelty: 7


12. CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation

ArXiv ID: 2605.12938

Primary Topic: Architecture and Training Dynamics

Authors: Seonghyun Jin, Youngmin Kim, Sunwoo Park, Jong Chul Ye

Abstract: Video world models should predict future appearance in a way that remains consistent with 3D scene structure, camera motion, and lens geometry. Existing attention-level camera encodings, however, either describe each token only by its viewing ray---without locating scene content along that ray---or assume pinhole projection, limiting camera control under wide-angle and fisheye lenses. We introduce Curved Ray Expectation Positional Encoding (CRePE), which represents each image token as a depth-aware distribution along its Unified Camera Model (UCM) ray and integrates the expected rotary positional phasor along the curved path this distribution traces when projected into each query view. CRePE is realized through a lightweight Geometric Attention Adapter on a frozen video diffusion transformer, with pseudo radial-distance supervision from a monocular geometry foundation model serving as a stabilizing anchor rather than an inference-time input. CRePE improves camera-control, lens, and orientation fidelity across pinhole, wide-angle, and fisheye settings, and transfers zero-shot to unseen real fisheye and diverse pinhole videos. Through Radial MixForcing, the same positional pathway further accepts externally supplied radial maps, enabling scene-geometry-conditioned generation and source-video motion transfer that follow the supplied geometry more faithfully than dedicated depth-conditioned baselines. CRePE thus offers a compact interface that unifies camera control, implicit 3D scene state, and external geometry control for video world models.

Comment: Introduces depth-aware rotary positional encoding through expected phasors along curved camera rays.

Topic Match: The core contribution is a new positional-attention mechanism, although its scope is specialized to camera-conditioned video generation.

Relevance: 7 Novelty: 7


13. State-Dependent Lyapunov Analysis of Rank-1 Matrix Factorization

ArXiv ID: 2604.26993

Primary Topic: Architecture and Training Dynamics

Authors: Jaehong Moon

Abstract: We develop a state-dependent Lyapunov framework for gradient descent on rank-1 matrix factorization. A parameterized quadratic certificate $I(δ;\,\cdot)$ generates strictly nested sublevel sets. Their ordering assigns each point a state $δ$, while a boundary-inward property makes this state monotone along gradient-descent trajectories. Together with internal chain transitivity, this geometry identifies the limiting dynamics even when all relevant stationary points are unstable. For scalar and rank-1 factorization, it yields convergence to a global minimizer below the stability threshold and to a balanced period-$2$ orbit in a post-critical interval, for almost every initialization in an explicit region. We formalize the mechanism through structural and dynamical axioms. Within the origin-centered, exchange-symmetric quadratic class, the axioms determine the ordered level-set geometry uniquely up to a monotone relabeling of the state, recovering the scalar geometry identified by Liang--Mont{ú}far. Additional analytic and numerical examples suggest broader applicability.

Comment: Uses state-dependent Lyapunov certificates to establish gradient-descent convergence and post-critical period-two dynamics.

Topic Match: Directly analyzes optimization dynamics beyond fixed-point stability, with applicability currently established for scalar and rank-1 factorization.

Relevance: 7 Novelty: 7


14. Stability and Generalization of Straight-Through Estimators for Training Two-Layer Quantized Neural Networks

ArXiv ID: 2609.06430

Primary Topic: Architecture and Training Dynamics

Authors: Yiming Ying

Abstract: We study the identity straight-through estimator (STE) for training a two-layer binary-activation network with hinge loss from the perspective of Statistical Learning Theory (SLT). Our central question is whether algorithmic stability can explain the statistical generalization of the estimator produced by the discontinuous STE training rule. In the saturated-output regime, the zero-initialized samplewise STE recursion is exactly the stochastic subgradient descent on the convex latent loss $(-yu^\top x)_+$. This representation makes a stability analysis possible. We derive an exact distance identity for two coupled updates and prove approximate non-expansiveness of the common-example map, with a quadratic defect only when the two latent margins straddle zero. We then obtain explicit $\ell_2$ on-average model-stability and generalization bounds, transferring stability isometrically from the latent vector to the full first-layer matrix. Combining stability with a standard optimization bound yields an explicit excess induced-risk guarantee and the rate $O(n^{-1/2})$ when $T=n^2$. Under margin separability, a complementary argument gives the optimal-order $O(R^2/(γ^2n))$ expected excess misclassification error for a randomized one-pass STE iterate and a corresponding majority-vote bound.

Comment: Shows that restricted binary-network STE updates equal convex latent subgradient descent, enabling stability and generalization guarantees.

Topic Match: The update equivalence provides a mechanistic analysis of quantized training, although the guarantees cover a restricted two-layer setting.

Relevance: 7 Novelty: 7


15. When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation

ArXiv ID: 2607.07050

Primary Topic: Architecture and Training Dynamics

Authors: Jiabin Shen, Guang Chen, Chengjun Mao

Abstract: Top-$K$ teacher logits make on-policy distillation tractable, but retained teacher mass does not certify student-relative gradient fidelity. We study routed, two-teacher tool-use distillation. In a frozen Qwen3.5-9B audit, the tool teacher ranks the entry token first on 500 tool prompts, and top-32 reinforces it in every matched pair. The response teacher's top-32 instead retains 99.99\% mass yet contains the same token on only 0.4\% of 500 response prompts; student-aware support restores the coordinate in all matched pairs and nearly matches full-vocabulary descent. On the tool route, top-32 preserves the reinforcing direction but substantially attenuates its full-vocabulary magnitude despite mass displayed as 1.000000. Matched restoration connects the response-side omission to behavior: restoring the entry coordinate at every supervised response position lowers full-generation over-calling from $14.2 \pm 2.1\%$ to $3.7 \pm 0.5\%$ across three seeds, but also lowers call recall by 12.4 points. A non-tool placebo changes over-calling by only 0.95 points. A synthetic output-logit-gradient-matched reroute also fails to reproduce exact restoration. We compare three optimization layers: student-aware support changes distilled coordinates, loss shaping changes retained-signal strength, and validation-tuned entry bias shifts inference without retraining. They expose distinct trade-offs in correction scope, system access, and required-call retention. Llama-3.1-8B reproduces the directional support asymmetry under its native JSON protocol. The results causally implicate decision-critical support omission as one contributor in the primary Qwen setting and show why both support fidelity and deployment costs must be audited. Code and aggregate artifacts are available at https://github.com/shen-jiabin/decision-support-opd.

Comment: Student-aware logit support restores decision-critical distillation gradients omitted by high-mass top-K truncation.

Topic Match: The causal analysis of truncation-induced gradient distortion fits training dynamics, with evidence concentrated on tool-use on-policy distillation.

Relevance: 7 Novelty: 7


16. Euclidean Fourier Neural Operators

ArXiv ID: 2608.28425

Primary Topic: Architecture and Training Dynamics

Authors: Nathanael Bosch, Niklas Frederik Schmitz, Michael F. Herbst

Abstract: Fourier neural operators (FNOs) provide an efficient framework for learning mappings between function spaces as they are, by construction, independent of the grid resolution at which they are trained and evaluated. However, FNOs are not independent of the periodic domain they are applied to: their discrete spectral weights are indexed by integer Fourier mode numbers, which correspond to physical wavevectors. When applied to a different domain, the same trained weights act at different wavevectors, and the FNO silently represents a different operator. This makes FNOs unsuitable for tasks where transfer across domains is crucial. We propose Euclidean Fourier neural operators~(EFNOs) as a domain-independent alternative to FNOs. By parameterizing the spectral kernel as a continuous function of the physical wavevector, the EFNO can learn operators that act consistently across periodic domains of varying shape and size. We evaluate the EFNO on a simple heat equation and on a practically relevant materials science task of learning exchange-correlation potentials across different crystal structures, and demonstrate that the EFNO is able to generalize to unseen grid sizes and domains.

Comment: Continuous physical-wavevector kernels make Fourier neural operators consistent across domains of different sizes and shapes.

Topic Match: The contribution is a general spectral parameterization mechanism with an explanation of domain-transfer failures, although its scope is narrower than large language model training.

Relevance: 7 Novelty: 7


17. Not Just Oversmoothing: Detecting the Echo Chamber Effect in Graph Neural Networks

ArXiv ID: 2609.06521

Primary Topic: Architecture and Training Dynamics

Authors: Asela Hevapathige, Ahad N. Zehmakan, Asiri Wijesinghe, Saman Halgamuge

Abstract: Oversmoothing is a well-known failure mode of Graph Neural Networks (GNNs). However, most existing diagnostics rely on global aggregation measures that fail to capture the heterogeneous dynamics of message passing. Real-world graphs exhibit pronounced community structure, and message passing operates on two timescales, with representations collapsing rapidly within communities and slowly across them. This creates a critical gap in which intra-community representations can become indistinguishable while inter community separation persists, a failure mode that we refer to as the Echo Chamber Effect. To quantify this effect, we introduce the Echo Chamber Index (ECI), which stratifies pairwise distances by community membership and reveals when global energy diminishes while inter-community separation persists. ECI further shows that feature retention mechanisms can preserve the echo chamber under the conditions of our theoretical analysis. The consequences depend on label structure: when communities align with classes, the echo chamber can sharpen node classification, whereas when they do not, the same collapse makes classification provably harder. Motivated by this analysis, we propose Community-Aware Split Propagation (CASP), a lightweight plugin that decouples intra- and inter-community aggregation and learns their balance from label structure. CASP improves diverse backbone GNNs across most evaluated homophilic and heterophilic settings.

Comment: Community-aware split propagation separates intra- and inter-community aggregation to control representation collapse.

Topic Match: The mechanistic analysis and modified propagation operator fit architectural dynamics, with scope limited to GNNs.

Relevance: 7 Novelty: 7


18. DeltaGNN: Graph Neural Network with Information Flow Control

ArXiv ID: 2501.06002

Primary Topic: Architecture and Training Dynamics

Authors: Kevin Mancini, Islem Rekik

Abstract: Graph Neural Networks (GNNs) are popular deep learning models designed to process graph-structured data through recursive neighborhood aggregations in the message passing process. When applied to semi-supervised node classification, the message-passing enables GNNs to understand short-range spatial interactions, but also causes them to suffer from over-smoothing and over-squashing. These challenges hinder model expressiveness and prevent the use of deeper models to capture long-range node interactions (LRIs) within the graph. Popular solutions for LRIs detection are either too expensive to process large graphs due to high time complexity or fail to generalize across diverse graph structures. To address these limitations, we propose a mechanism called \emph{information flow control}, which leverages a novel connectivity measure, called \emph{information flow score}, to address over-smoothing and over-squashing with linear computational overhead, supported by theoretical evidence. Building on this mechanism, we introduce DeltaGNN, to the best of our knowledge among the first \textit{scalable} (featuring linear computational and memory complexity overhead) and \textit{generalizable} (capable of effectively handling graphs with diverse homophily, density, and topology) architectures for long-range and short-range interaction detection. We benchmark our model across 10 real-world datasets, including graphs with varying sizes, topologies, densities, and homophilic ratios, showing superior performance with limited computational complexity. The implementation of the proposed methods are publicly available at https://github.com/basiralab/DeltaGNN.

Comment: An information-flow controller addresses oversmoothing and oversquashing with linear computational and memory overhead.

Topic Match: The contribution changes the message-passing mechanism and analyzes its behavior; its training relevance is specific to GNN architectures.

Relevance: 7 Novelty: 7


19. Canalization Before Generalization: Grokking as a Dynamical Probe

ArXiv ID: 2608.25813

Primary Topic: Architecture and Training Dynamics

Authors: Yiming Lin

Abstract: For overparameterized neural networks, many solutions can fit the training data equally well while behaving very differently on unseen samples. Grokking separates training fit from visible generalization, providing a window for studying how this selection develops during training. We sweep short, fixed-duration weight-decay (WD) perturbations across the pre-generalization plateau and measure how they shift later generalization time. Across three grokking tasks, these shifts are unordered early in the plateau but later form a stable dose ordering, with stronger WD increases leading to earlier generalization and stronger WD decreases leading to later generalization. This ordering emerges before visible generalization in all three tasks. Meanwhile, test-loss barriers between perturbed and baseline generalization checkpoints collapse toward zero while the ordered timing effects persist. Drawing on Waddington's developmental landscape as an analogy, we call this combination of increasingly constrained solution selection and persistent dose-ordered shifts in generalization timing the canalization of grokking solution selection.

Comment: Timed weight-decay perturbations expose ordered shifts in generalization timing before grokking becomes visible.

Topic Match: Interventions on optimization directly probe solution-selection dynamics during training, although evidence is confined to three grokking tasks.

Relevance: 7 Novelty: 7


20. Diffusion-Inspired Reconfiguration of Transformers for Uncertainty Calibration

ArXiv ID: 2602.08920

Primary Topic: Architecture and Training Dynamics

Authors: Manh Cuong Dao, Quang Hung Pham, Phi Le Nguyen, Thao Nguyen Truong, Bryan Kian Hsiang Low, Trong Nghia Hoang

Abstract: Uncertainty calibration in pre-trained transformers is critical for their reliable deployment in risk-sensitive applications. Yet, most existing pre-trained transformers do not have a principled mechanism for uncertainty propagation through their feature transformation stack. In this work, we propose a diffusion-inspired reconfiguration of transformers in which each feature transformation block is modeled as a probabilistic mapping. Composing these probabilistic mappings reveals a probability path that mimics the structure of a diffusion process, transporting data mass from the input distribution to the pre-trained feature distribution. This probability path can then be recompiled on a diffusion process with a unified transition model to enable principled propagation of representation uncertainty throughout the pre-trained model's architecture while maintaining its original predictive performance. Empirical results across a variety of vision and language benchmarks demonstrate that our method achieves superior calibration and predictive accuracy compared to existing uncertainty-aware transformers.

Comment: Reconstructs Transformer feature transformations as a unified diffusion process for propagating representation uncertainty.

Topic Match: Probabilistic reconfiguration of Transformer blocks supplies a general architectural mechanism, although the objective is calibration and training-cost implications are undeveloped.

Relevance: 7 Novelty: 7


21. RoPE attention is an exact forward-pass gradient step with softmax intact

ArXiv ID: 2609.06685

Primary Topic: Architecture and Training Dynamics

Authors: Julie Huang, Maggie Chlon, Leon Chlon

Abstract: We derive an exact gradient-step representation of the RoPE-softmax forward pass. For every deterministic RoPE-softmax attention head with arbitrary affine projection weights, we construct a query-dependent effective matrix $ΔM_i$ satisfying $y_i = μ_i + u_i^\top ΔM_i$, where $μ_i$ is the uniform mean of the attended values and $u_i$ is the augmented query input. The construction applies the classical exponential divided difference $ρ= ϕ_1$ to retain the softmax exactly. Its positive coefficients give a unit gradient-step representation on a query-conditioned quadratic objective. The same function connects the RoPE generator to exact positional finite differences. We derive a tokenwise formula for the error of reusing one query's matrix and prove that a nonconstant finite-cache head cannot admit a globally exact affine query readout. Reconstruction checks and frozen-reuse calibration on one pretrained Qwen2.5-0.5B layer verify the representation and quantify the correction required when one query's matrix is reused.

Comment: Derives an exact query-conditioned gradient-step representation of RoPE attention while retaining softmax.

Topic Match: Directly analyzes an attention mechanism and limits on reusing its effective matrices; implications for actual training dynamics remain indirect.

Relevance: 7 Novelty: 7


22. ModularPhaseNet: Finite-Cyclic Phase Geometry for Computable Semantic Hierarchy, Direction, and Context Consistency in Standard Transformers

ArXiv ID: 2609.06000

Primary Topic: Architecture and Training Dynamics

Authors: Kiyotaka Kasubuchi, Kazuo Fukiya

Abstract: We propose ModularPhaseNet, a classical and integer-computable discretization of the continuous complex phase geometry introduced in QuantumPhaseNet. The real-valued hidden states of a standard Transformer are retained, while only an auxiliary phase channel is quantized into a cyclic subgroup G = of order q | (p-1) in the multiplicative group of F_p. A continuous phase e^{i phi} is represented by z = g^a mod p; phase composition becomes group multiplication, relative phase becomes group division, conceptual hierarchy is induced by a filtration of cyclic quotients, semantic direction is represented by oriented relative group elements, and contextual consistency is measured by gauge-invariant cycle holonomy. The method introduces three components into an otherwise standard Transformer: a finite-phase encoder, a quotient-filtration hierarchy module, and a group-valued connection module. Their outputs enter self-attention as real-valued bias terms. Training uses distributions in the real group algebra or straight-through Gumbel-Softmax, whereas inference uses exact modular exponentiation and precomputed tables. No quantum hardware, complex-valued matrix multiplication, or discrete-logarithm computation is required. We prove quantization-distortion bounds, nesting of quotient-induced partitions, gauge invariance, a discrete integrability result for flat connections, and boundedness of the resulting attention output. The central empirical hypothesis is that these exact discrete invariants improve hierarchy recovery, discourse alignment, contradiction detection, and calibrated hallucination-risk prediction under a controlled compute budget. This paper reports the theory together with a pre-registered evaluation plan; the experiments described in Section 14 have not yet been carried out, and no empirical result is claimed here.

Comment: Adds finite-cyclic phase channels and group-valued connections as Transformer attention biases.

Topic Match: The proposed attention mechanism is an architectural contribution; its effects on training and computational efficiency remain experimentally untested.

Relevance: 7 Novelty: 6


23. DODR: Deterministic Operator-Driven Reasoning in Latent Space

ArXiv ID: 2609.04782

Primary Topic: Architecture and Training Dynamics

Authors: Weicai Huang

Abstract: Autoregressive (AR) large language models formulate reasoning as token-level probabilistic sampling, which induces three fundamental defects in complex logical reasoning: error accumulation, probability substituting necessity, and the linear-chain information bottleneck. This paper proposes the Deterministic Operator-Driven Reasoning in Latent Space architecture (DODR), which reconstructs reasoning as reasoning-graph computation in a high-dimensional linear-algebraic space. Reasoning states are represented as snapshot vectors whose primitives are semantic units (phrases or sentences) rather than tokens, and each inference step is a deterministic matrix operation with no token sampling. Peirce's three inference types are formalized as three trainable matrix operators: a rank-deficient deduction operator (information collapse), a full-rank induction operator (information expansion), and an abduction operator defined as the Moore-Penrose pseudo-inverse of deduction (information hypothesizing). We prove that the operator set is minimal and complete given Peirce's trichotomy, that no single "super-operator" can realize all three types (a rank obstruction), and that reasoning graphs are Turing-complete with contractive backflow converging by Banach's fixed-point theorem. Experiments on 503 sample records (420 deduplicated samples) across dedicated and end-to-end settings show: deduction loss converges to 1.40e-05; induction achieves 0.9996 generalization coverage with 20/20 hard vetoes on counterexamples; abduction solutions exceed the random baseline by 28x with judgment accuracies of 72.5% (58/80, Wilson 95% CI [61.9%, 81.1%]) and 81.7% (49/60, CI [70.1%, 89.4%]); frozen operators attain 100% (60/60) on unseen cross-domain deduction. The architecture provides a structural zero-hallucination guarantee and a three-layer continual-learning mechanism. All data and code are released.

Comment: Trainable deduction, induction, and pseudoinverse abduction operators define a modular latent reasoning architecture.

Topic Match: The core contribution changes computational primitives and reasoning-graph execution, although its connection to large-model training remains untested in the small experiments described.

Relevance: 7 Novelty: 6


24. A Thermodynamic Theory of Learning Part II: History-Dependent Reachability and Continual Learning

ArXiv ID: 2602.07950

Primary Topic: Architecture and Training Dynamics

Authors: Daisuke Okanohara

Abstract: We formulate plasticity as target-dependent, finite-horizon reachability under history-dependent dynamics, using standard minimum-energy control theory. An extended state includes parameters and internal variables that affect future updates. Around a reference trajectory, the input-to-endpoint response and its controllability Gramian determine the minimum input energy for each reachable displacement. Minimizing this energy over the joint task target defines a task-conditioned adaptation cost, exact for the arbitrary-input linear model. A solvable learning-retention example shows that this cost can increase without any rank loss, and distinguishes small directional gain from an inaccessible direction. Wasserstein transport supplies a complementary global lower bound, while learning-map Jacobians describe deformation of initial perturbations rather than response to future inputs. The resulting learning-retention frontier depends on the current extended state, admissible updates, target, horizon, and cost metric; nonlinear and constrained-update applications require additional control of these modeling choices.

Comment: Finite-horizon controllability explains how update history changes plasticity and learning-retention tradeoffs.

Topic Match: History-dependent learning dynamics are central, although the guarantees concern idealized control models without large-model validation.

Relevance: 7 Novelty: 6


25. Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models

ArXiv ID: 2608.27367

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Frederik Berenz

Abstract: Joint-Embedding Predictive Architectures (JEPAs) for world modeling typically employ fixed-size Vision Transformer encoders that are over-provisioned for simple tasks and under-provisioned for complex ones, with significant redundancy across attention heads. We propose Successive Capacity Growth (SCG), a method that starts from a minimal encoder (1 head, 2 layers, 283K parameters) and grows incrementally in width (adding attention heads for low-level semantic capacity) or depth (adding transformer blocks for higher-order semantic abstraction), driven by a task-agnostic test-and-verify mechanism that exploits function-preserving expansion to safely trial architectural changes and roll back if they do not improve prediction loss. The Sketched Isotropic Gaussian Regularizer (SIGReg) ensures that all learned semantic dimensions remain statistically independent and aligned with the predictive objective, preventing collapse even as the architecture grows. On a 60-dimensional multi-object dynamics task, SCG naturally triggers depth expansion, improving prediction loss by 20.3% over the fixed small baseline with 56 times greater parameter efficiency than scaling to the fixed large model; on a 2D navigation task, a single width expansion yields even an 23% improvement over the fixed large model. Across all three tested environments of increasing complexity, the adaptive encoder matches or exceeds the fixed small baseline, with zero false-positive expansions and bit-exact function preservation (ratio = 1.0, absolute difference = 0.0). The take-away is that JEPA world model encoders need not be pre-allocated at maximum capacity - they can grow successively as the task demands, achieving significant compute and data efficiency while maintaining representation quality.

Comment: Function-preserving width and depth expansion adapts Transformer capacity during training.

Topic Match: Adaptive architecture growth is the central mechanism, with parameter-efficiency evidence currently limited to small JEPA world-model encoders.

Relevance: 7 Novelty: 6


26. Modal Logic Neural Networks

ArXiv ID: 2512.03491

Primary Topic: Architecture and Training Dynamics

Authors: Antonin Sulc, Noor Naddour

Abstract: Neural Networks are indispensable to natural sciences and society. Their impact extends from applications in public health to workforce productivity. Here, we introduce Modal Logic Neural Networks (MLNNs) -- an end-to-end differentiable logical neural network realisation of modal logic which evaluates a learnable truth function across possible-world semantics. This neural architecture handles para-consistency and inconsistency via a learnable world accessibility relation and valuation function. Because the modality is fixed by which frame axioms the relation satisfies rather than by the operator, one differentiable engine covers the epistemic, doxastic, deontic and temporal readings, with applications from verification of reactive and distributed systems to legal discourse and microeconomic utility models. In this paper, we introduce a model of differentiable Kripke semantics, and establish their soundness, convergence, and structural guarantees. We show four applications, in which the learned relation reads as a trust matrix, an operating-regime embedding with safety bounds, a temporal precedence order, and a recovered constraint graph.

Comment: Learnable accessibility relations implement end-to-end differentiable modal-logic computation.

Topic Match: Differentiable logical operators constitute an architectural contribution, though its relevance to large-model training dynamics and cost remains limited.

Relevance: 6 Novelty: 7


27. Towards Unified Multimodal Graph Foundation Model: A Bridge-Router-Adapter Based Approach

ArXiv ID: 2609.06668

Primary Topic: Architecture and Training Dynamics

Authors: Sirui Zhang, Yubing Zhou, Xunkai Li, Zekai Chen, Shumeng Li, Wang Luo, Yinlin Zhu, Yujin Gao, Rong-Hua Li

Abstract: Multimodal graphs couple node attributes in different modalities, such as text and images, with relational structure, enabling topological structure and cross-modality attributes to be modeled jointly. Multimodal graph foundation models seek unified representations from such data that transfer across different graph domains and downstream tasks. However, existing methods exhibit two fundamental limitations. (1) Cross-Scope Context Entanglement. They merge scope-specific graph contexts into a unified representation, obscuring their distinctions during multimodal construction. (2) Scope-Ignorant Modality Routing. They route modalities within a fixed graph scope, overlooking how modality relevance varies across neighborhood ranges. To address these challenges, we propose BRAIN, a unified model that focuses on graph context that combines neighborhood scope with modality composition. BRAIN comprises a scope-conditioned Bridge that combines structural information spanning local-to-global neighborhood scopes with different modality compositions; a hierarchical Router that estimates the relevance between the scope and the task, and selects compositions separately within each scope, allowing modality utility to vary with graph range; and a lightweight residual Adapter that further specializes the routed embedding for downstream prediction. BRAIN is trained through multi-graph pretraining followed by task-specific adaptation. Experiments across nine datasets and four task families demonstrate its broad effectiveness, improving node-classification and link-prediction performance by up to 4.73% relative to the strongest baseline, while achieving an average relative improvement of 14.72% across four graph-to-text and two graph-to-image metrics.

Comment: Introduces hierarchical, scope-conditioned routing for dynamically selecting modality compositions.

Topic Match: Its central contribution is a dynamic modular-computation architecture.

Relevance: 6 Novelty: 6


28. SAGE: Shape-Adapting Gated Experts for Adaptive Histopathology Image Segmentation

ArXiv ID: 2511.18493

Primary Topic: Architecture and Training Dynamics

Authors: Gia Huy Thai, Hoang-Nguyen Vu, Anh-Minh Phan, Quang-Thinh Ly, Tram Dinh, Thi-Ngoc-Truc Nguyen, Nhat Ho

Abstract: The significant variability in cell size and shape continues to pose a major obstacle in computer-assisted cancer detection on gigapixel Whole Slide Images (WSIs), due to cellular heterogeneity. Current CNN-Transformer hybrids use static computation graphs with fixed routing. This leads to extra computation and makes it harder to adapt to changes in input. We propose Shape-Adapting Gated Experts (SAGE), an input-adaptive framework that enables dynamic expert routing in heterogeneous visual networks. SAGE reconfigures static backbones into dynamically routed expert architectures via a dual-path design with hierarchical gating and a Shape-Adapting Hub (SA-Hub) that harmonizes feature representations across convolutional and transformer modules. Embodied as SAGE with ConvNeXt and Vision Transformer UNet (SAGE-ConvNeXt+ViT-UNet), our model achieves a Dice score of 95.23% on EBHI, DSC scores of 92.78% and 91.42% on GlaS Test A and Test B, respectively, and 91.26% DSC at the WSI level on DigestPath, while exhibiting robust generalization under distribution shifts by adaptively balancing local refinement and global context. SAGE establishes a scalable foundation for dynamic expert routing in visual networks, thereby facilitating flexible visual reasoning. Project page: https://oxyzgiahuy.github.io/sage/

Comment: Reconfigures heterogeneous vision backbones with hierarchical gates and input-adaptive expert routing.

Topic Match: Dynamic computation across convolutional and transformer experts is the primary mechanism.

Relevance: 6 Novelty: 6


29. Cross-Task Generalization Between Understanding and Generation in Unified Vision-Language Models: A Controlled Study

ArXiv ID: 2505.23043

Primary Topic: Architecture and Training Dynamics

Authors: Jihai Zhang, Tianle Li, Linjie Li, Zhengyuan Yang, Yu Cheng

Abstract: Unified vision-language models (VLMs) aim to support both visual understanding and generation within a single framework, but it remains unclear when mixed training benefits both capabilities and when it introduces conflicts. This paper presents a controlled empirical study of cross-task generalization between understanding and generation in unified VLMs. We construct two controllable image-text benchmarks, SmartWatch and modified CelebA, with paired VQA, captioning, and text-to-image generation tasks, and evaluate multiple LLM-based unified architectures built from SigLIP and VQ-VAE visual spaces. Our experiments show that mixed understanding-generation training can improve both tasks over task-specific training, but the benefit depends strongly on the relation between vision input and output spaces. Unified models with better aligned visual spaces exhibit stronger cross-task transfer, while reversible affine distortions of the input visual space substantially weaken this effect and can turn mutual benefits into conflicts. We further find that increasing data from one task can initially improve the other, but excessive imbalance between understanding and generation data may degrade the complementary task. By controlling attribute frequencies, we show that generation supervision can help recover underrepresented visual concepts for understanding. Adapter analyses suggest that this transfer is not primarily caused by richer visual adapter features, but by the base language model learning relationships that generalize across aligned visual spaces. A real-case experiment on LLaVA provides additional evidence that mixed understanding-generation training can benefit visual understanding beyond controlled benchmarks.

Comment: Isolates how visual input-output alignment governs benefits and conflicts in mixed understanding-generation training.

Topic Match: Mixed-training behavior connects it to training analysis, but the main findings concern representation alignment and cross-task generalization.

Relevance: 6 Novelty: 6


30. Monotonic Kolmogorov-Arnold Networks: A Theoretical and Empirical Study of Monotonicity as an Inductive Bias

ArXiv ID: 2606.17886

Primary Topic: Architecture and Training Dynamics

Authors: Mikhail Krasnov, Blaž Bertalanič, Carolina Fortuna

Abstract: Monotonicity has been a long-running architectural inductive bias for neural networks, motivated by tabular, scientific, and economic settings where outputs are known to respond monotonically to certain inputs. Existing approaches are MLP- or flow-based and lack per-edge functional transparency; the only Kolmogorov--Arnold Network (KAN) variant with monotonicity, MonoKAN, enforces the constraint only on a restricted parameter subset and requires a projection-style training procedure. We close this gap with \textbf{MKAN}, a KAN with hard monotonicity guaranteed for \emph{all} parameter values via exponential reparameterization of B-spline coefficients, positive edge weights, and a monotone base activation. Training reduces to standard unconstrained gradient descent. Our headline theoretical contribution is a \emph{representation-cost} theorem: any $C^K, K >0$ feature extractor inducing a ball-shaped semantic-neighborhood partition admits a monotone realization of the equivalent neighborhood structure at $N' = N^ + k \le 2N^$, where $k$ is the number of non-monotone coordinates of the original. The bound is architecture-agnostic and gives a principled sizing rule for monotone encoders. Empirically, MKAN is competitive with state-of-the-art monotone NNs on the SMM/ICML-2024 benchmark while being the only method that combines hard unconstrained monotonicity with KAN's per-edge functional transparency; the $2N^*$ prediction is validated in a self-supervised feature-size sweep on four real datasets, and on a controlled monotone-generative dataset MKAN recovers ground-truth factors with substantially higher Spearman alignment than KAN, MLP, and linear baselines.

Comment: Enforces hard monotonicity through parameterization that permits unconstrained gradient descent.

Topic Match: The architectural parameterization changes how constrained networks train, although the main evidence concerns KANs and representation-size theory rather than large models.

Relevance: 6 Novelty: 6


Efficiency, Compression, and Large-Scale Training (30)

1. LGQ: Learnable Geometric Quantization for Image Tokenization

ArXiv ID: 2602.16086

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Idil Bilge Altun, Mert Onur Cakiroglu, Elham Buxton, Mehmet Dalkilic, Hasan Kurban

Abstract: Recent collapse-free quantizers such as FSQ achieve stable training by replacing the learnable codebook with an engineered geometry: a fixed scalar grid whose structure is dictated by the codebook size K. We show this trade-off is unnecessary. We introduce Learnable Geometric Quantization (LGQ), which retains a learnable codebook of codes and performs soft-to-hard assignment via temperature annealing, regularized by two cheap terms: a diversity term scaled by codebook size that penalizes concentrated batch-average usage is the primary driver of collapse resistance, complemented by a peakedness term that sharpens each token's soft-assignment toward one-hot; together they prevent codebook collapse without EMA, reset heuristics, or codebook reparameterization. Under a fixed VQ-GAN backbone, we benchmark LGQ against RotVQ, FSQ, LFQ, SimVQ, and IBQ on ImageNet 256x256 at K = 16,384, and sweep LGQ over K in {4096, ..., 65,536} without any per-K hyperparameter tuning. LGQ attains the best reconstruction FID at K = 16,384 while maintaining 100% codebook utilization, and continues to improve as the codebook grows to K = 65,536, holding 100% utilization at every K. Training MaskGIT on the frozen tokenizers, LGQ further attains the best class-conditional generation among the compared quantizers, leading on reconstruction and generation alike. Code is available at https://anonymous.4open.science/r/lgq-anon-E12C/.

Comment: Prevents learnable-tokenizer codebook collapse using diversity and peakedness regularization without EMA or resets.

Topic Match: Learnable quantization is primary, while its collapse-resistant training dynamics provide a secondary architectural match.

Relevance: 9 Novelty: 8


2. Celty: SpMSpV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM Inference

ArXiv ID: 2608.01536

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Ruokai Yin, Priyadarshini Panda

Abstract: Large Language Models (LLMs) increasingly rely on sparsity to cut inference cost, but most prior work exploits a single sparsity source and targets batched multi-user inference. Dual-sparsity, which pairs unstructured weight pruning with runtime activation sparsity, offers a compelling size-accuracy-latency tradeoff for single-user decoding, but forms a Sparse Matrix-Sparse Vector (SpMSpV) workload that existing GPU kernels handle poorly. We propose Celty, a co-designed sparse format, GPU kernel, and SIMT microarchitecture for SpMSpV in LLM inference. Celty's Run-Length Compressed CSC (RLC-CSC) format enables vectorized loading of compressed weight columns and exploits both sparsity sources to skip memory accesses, accumulating partial products in shared memory. The Celty Sparse SIMT Core then adds a pipelined RLC decoder that removes software index reconstruction and repurposes local register files for conflict-free accumulation, operating on the same compressed representation. The kernel alone achieves up to 2.8x over cuBLAS; with the Sparse SIMT Core, speedup reaches 5.3x over cuBLAS at 70% dual-sparsity.

Comment: A co-designed sparse format and SpMSpV kernel exploit weight and activation sparsity to accelerate LLM decoding.

Topic Match: New compressed storage, GPU kernels, and SIMT support directly target large-model inference cost; the abstract distinguishes kernel-only gains from gains requiring hardware changes.

Relevance: 9 Novelty: 8


3. BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference

ArXiv ID: 2609.04971

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Janghyeon Kim, Minsoo Kim, Kyuhong Shim, Jungwook Choi

Abstract: Large Reasoning Models (LRMs) achieve superior problem-solving through extended Chain-of-Thought (CoT) generation, but the resulting key-value (KV) cache grows linearly with sequence length and creates severe memory bottlenecks, often exceeding GPU capacity for long reasoning traces. Existing KV cache compression methods rely on recent queries to estimate future token importance, implicitly assuming these serve as reliable proxies for future attention patterns. We demonstrate that this assumption fails in long-horizon reasoning: certain decoding steps generate Thought Revisiting Tokens (TRT) that re-attend to distant previous context, such as task-solving plans formulated early in the trace. Through systematic analysis, we discover that queries corresponding to the TRT cluster into a small number of similarity groups in the embedding space. Based on this insight, we propose BeaconKV, a training-free KV cache compression method that maintains beacon queries, compact representatives for each global query cluster, to anticipate which KV pairs will be revisited without storing the entire query history. Across four open-source LRMs and diverse reasoning benchmarks, BeaconKV generally outperforms existing compression methods, achieving up to $5.8\times$ memory reduction while nearly preserving full cache accuracy and improving throughput by over $4.3\times$.

Comment: Cluster-representative beacon queries preserve KV entries likely to be revisited during long reasoning traces.

Topic Match: Introduces a concrete KV-cache retention mechanism with substantial reported memory savings and throughput gains, directly matching the cache-efficiency criterion.

Relevance: 9 Novelty: 8


4. All for 1-Bit: Towards Genuine 1-Bit Post-Training Quantization for LLMs

ArXiv ID: 2609.06161

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Zhixiong Zhao, Zukang Xu, Guangyu Sun, Lifeng Liu, Dawei Yang

Abstract: Large language models (LLMs) have achieved remarkable progress, yet their massive storage and memory-bandwidth demands still hinder efficient deployment. Weight binarization is a promising solution, but existing binarization-based post-training quantization (PTQ) methods usually far exceed the nominal 1-bit storage target due to hidden overhead. To address this gap, we propose All for 1-Bit (AF1), a genuine 1-bit PTQ framework for LLMs. AF1 comprises two complementary components: (1) Null-space-Aware Binary Factorization (NABF) for improving binary reconstruction through Hessian-aware surrogate reparameterization, null-space-aware binary factorization, and scale-only global reconstruction; and (2) Hierarchical Shapley Allocation (HiSA) for assigning structural capacity using hierarchical Shapley sensitivity. Together, they preserve model accuracy under a strict 1.0-BPW budget in the PTQ setting. Experiments on LLaMA, Qwen, and Gemma families show that AF1 consistently outperforms existing binarization-based PTQ methods in perplexity and zero-shot accuracy. Compared with BF16, AF1 achieves an average 2.5 times inference speedup and over 90% memory reduction across evaluated models, providing a practical path toward deployable genuine 1-bit compression for LLMs. The code for reproducibility is available at https://github.com/Kishon-zzx/AF1.

Comment: Achieves strict 1.0-bit weight compression through Hessian-aware binary factorization and sensitivity-based capacity allocation.

Topic Match: Genuine one-bit post-training quantization directly targets LLM memory bandwidth, storage, and inference cost.

Relevance: 9 Novelty: 8


5. Linear Algebra Foundations of Efficient Attention: A Phase Reversal in Rank Collapse Under SVD Compression

ArXiv ID: 2609.06341

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Anjaneya Teja Sarma Kalvakolanu

Abstract: Linear algebra provides the framework of concepts (matrix rank, singular value decomposition (SVD), and eigendecomposition) that modern artificial intelligence employs to encode, compress, and propagate information through neural networks. This paper unifies fourteen separate peer-reviewed works analyzing the usage of these techniques in the context of transformer-based foundation model research, focusing on three areas of the topic: derivations and properties of self-attention matrices' output rank, compression methods that purposefully utilize this phenomenon, and the low-rank key-value (KV) cache projection and its semiseparable-matrix duality to linear attention and state-space structured models. We were motivated to conduct this work after observing an open problem in this literature: the interplay of the mentioned compression methods with natural rank collapse of the network. With this paper, we report an original finding that using SVD compression of attention projections actually has the opposite effect on the rank collapse of the network: while it strongly suppresses it at initialization, it accelerates on pretrained models (for GPT-2 124M, GPT-2 Medium 355M, and Pythia-160M) with minimal risk of object aliasing artifacts appearing (verified on all compression ratios) and is consistent across four rank estimation methods. A controlled causal decomposition of the effect in both settings showed that the reason for this behavior can be explained by the choice of the subspace SVD makes when compressing the matrix better than the reduction of the operator norm it achieves, explaining roughly 76% of the effect at initialization and 83% on the pretrained weights, providing a refinement to the calibration-aware compression viewpoint and an explanation of why it outperformed naive SVD truncation.

Comment: Explains why SVD compression suppresses attention rank collapse at initialization but accelerates it after pretraining.

Topic Match: The original contribution analyzes attention compression through subspace selection, with a mechanistic connection to training-dependent rank dynamics.

Relevance: 9 Novelty: 7


6. ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs

ArXiv ID: 2609.05228

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Zukang Xu, Zhixiong Zhao, Xing Hu, Jiangyong Yu, Houji Wen, Jun Li, Zhe Jiang, Dawei Yang

Abstract: Mixture-of-Experts (MoE) architectures provide an efficient paradigm for scaling large language models (LLMs), yet fixed top-k routing activates the same number of expert slots for every token, causing substantial redundant computation. Existing expert-skipping methods often rely on router confidence, calibration data, or additional training, and therefore cannot reliably estimate the actual contribution of routed experts. To this end, we propose ACE, a training-free, calibration-free, and checkpoint-preserving framework for token-adaptive expert skipping in MoE-based LLMs. ACE contains two complementary components: 1) Global Spectral Proxy (GSP), which estimates global transformation capacity from the coupled gate, up, and down projections together with RMSNorm scaling; and 2) Router-Conditioned Refinement (RCR), which constructs expert-specific direction prototypes from centered router weights and evaluates expert responses along routing-preferred directions. During inference, ACE combines both estimates with runtime router gates and skips an expert slot only when both views identify it as low-contribution, while always retaining the top-1 expert. All expert statistics are computed offline, leaving only table lookups and lightweight scalar operations online. Extensive experiments across three MoE-based LLMs and eight benchmarks demonstrate that ACE consistently outperforms existing static and dynamic baselines, with increasingly pronounced advantages under aggressive expert skipping. For instance, at a 50% skipping ratio on Qwen3.6-35B-A3B, ACE reduces WikiText-2 perplexity by 7.96% and improves average downstream accuracy by 4.15 percentage points over the strongest competing method.

Comment: Spectral and router-conditioned contribution estimates enable calibration-free, token-adaptive expert skipping.

Topic Match: The core contribution is a new mechanism for reducing active MoE computation during inference through selective expert skipping.

Relevance: 9 Novelty: 7


7. Accelerating Diffusion Transformers with Gaussian Process Rectified Feature Cache

ArXiv ID: 2609.05981

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Zhirong Shen, Rui Huang, Chang Zou, Shikang Zheng, Jiacheng Liu, Peiliang Cai, Zhengyi Shi, Yaosong Du, Liang Feng, Xiaobing Tu, Jinkui Ren, Xiantao Zhang, Linfeng Zhang

Abstract: Diffusion Transformers have become the dominant paradigm in generative AI, but their high computational costs severely hinder real-time applications. Prediction-based feature caching is widely used to accelerate diffusion transformers; however, as the number of steps increases, the deviation between its predictions and the reference full-compute trajectory gradually grows. An intuitive idea is to use an online regression model to dynamically correct this deviation, but it faces the issue of label data being unavailable during the acceleration process. This paper presents a statistical observation that the residuals between the features of full computation steps using caching methods and reference full-compute trajectory locally exhibit a zero-mean Gaussian distribution. By treating the features of full computation steps as noisy observations of reference features, the data acquisition problem is resolved. Based on this observation, a plug-and-play GP-Refiner correction framework is proposed. This method utilizes Gaussian Process Regression for correction and, leveraging the properties of GPR, introduces an uncertainty-adaptive computation strategy that triggers necessary full-computation calibration by monitoring the posterior variance in real time. Experiments demonstrate significant improvements across different models when combined with various state-of-the-art methods. Integrating the proposed framework with TaylorSeer reduces the computational load by 19.3% while improving PSNR by 0.9 dB and reducing LPIPS from 0.46 to 0.29. Code is available in https://github.com/Aredstone/GP-Refiner.

Comment: Gaussian-process correction and posterior-variance triggers adapt full-computation refreshes during diffusion feature caching.

Topic Match: The core contribution is a new cache-correction and refresh mechanism that reduces diffusion-transformer computation while controlling accumulated error.

Relevance: 9 Novelty: 7


8. ACE: Adapter Consolidation across Experts for Parameter-Efficient Fine-Tuning of MoE LLMs

ArXiv ID: 2609.06072

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: MoE Training

Authors: Ahin Lee, Sehyun Yun, Joonha Park, Taesik Gong

Abstract: Parameter-efficient fine-tuning (PEFT) of mixture-of-experts (MoE) models commonly attaches a separate low-rank adapter to each expert. This expert-wise design fragments adaptation in three ways: capacity is split across narrow low-rank updates, gradient supervision becomes sparse and imbalanced under sparse routing, and execution is decomposed into many small GEMMs. We find that such expert-wise separation is often unnecessary, as subsets of LoRA adapters become functionally similar during fine-tuning, revealing redundancy among expert-specific adapters. Based on this redundancy, we propose ACE (Adapter Consolidation across Experts), which groups redundant experts and replaces their expert-specific adapters with group-shared higher-rank LoRA modules under the same PEFT budget. ACE further introduces grouped adapter execution, which consolidates fragmented expert-wise adapter computations into fewer, larger group-level GEMMs. Across evaluations covering 12 datasets and four MoE backbones, ACE achieves the highest observed mean accuracy among the parameter-matched PEFT methods on the three backbones with complete baseline coverage, while providing $1.31\times$ to $1.48\times$ wall-clock training speedup over expert-wise LoRA without increasing peak memory. Our code is available at https://github.com/UbiquitousAILab/ACE.

Comment: Group-shared adapters consolidate fragmented MoE adapter GEMMs, delivering reported 1.31-1.48x wall-clock training speedups.

Topic Match: Adapter consolidation under a fixed PEFT budget makes efficiency primary; the new MoE-specific grouped execution also directly matches MoE training implementations.

Relevance: 9 Novelty: 7


9. A Nuclear-Norm Lower Bound for Dithered Scalar Quantization of Matrix Products

ArXiv ID: 2609.05641

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Piyush Sao, Narasinga Miniskar, Pedro Valero-Lara, Keita Teranishi, Sudip Seal

Abstract: We consider the problem of minimizing error in quantized matrix multiplication $C=AB$. Scalar quantization of the factors introduces rounding errors whose scale depends on the maximum absolute entries -- the ranges -- of their rows and columns. These ranges determine the quantization grid steps. To reduce the error, we optimize over product-preserving transformations that alter the factor ranges and grid steps without changing $C$. Specifically, we seek the smallest leading expected squared error over invertible inner changes of basis and orthogonal outer rotations. Under independent, zero-mean subtractive dither noise on an unbounded lattice, we prove the output-only bound $E_{\rm lead} \ge (c_A+c_B)/K \Vert AB\Vert_^2$, where $K$ is the inner dimension, $c_A$ and $c_B$ are normalized noise variances, and $\Vert AB\Vert_$ is the nuclear norm. The bound is tight: an SVD-aligned Hadamard construction attains the infimum whenever a Hadamard matrix of order $K$ exists, including every power of two, while an SVD-aligned DCT construction is within a factor of two for every $K$. Without outer rotations, Gram-matrix balancing minimizes factorization energy, and finite-set flattening achieves the bound within $C\log(K(m+n))$. For power-of-two $K$, conditional expectations deterministically select the Hadamard signs in $O((m+n)K^2)$ exact-real operations. Synthetic experiments verify both constructions and illustrate the tradeoff between regularization and conditioning. These results characterize the full-gauge optimum and quantify the cost of preserving row and column indices.

Comment: Establishes tight quantized-matmul error bounds and constructive product-preserving transformations that attain them.

Topic Match: Quantization-error minimization is the core contribution, with optimal basis constructions under idealized dither assumptions; practical large-model gains remain untested.

Relevance: 8 Novelty: 8


10. Hidden in Plain Sight: The Overlooked Significance of Canonical Elements for Extreme LLM Sparsity

ArXiv ID: 2609.06557

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Hyeondo Jang, Kwanhee Lee, Dongyeop Lee, Namhoon Lee

Abstract: Large language models (LLMs) are often considered fragile under aggressive sparsification, and maintaining reliable performance typically requires sticking to moderate sparsity levels. However, recent studies suggest that LLMs are more resilient to high sparsity than previously thought, reframing the problem as a design challenge rather than a fundamental limitation. In this work, we challenge the perceived limits of unstructured post-training LLM pruning by revisiting elementary pruning strategies that have remained relatively underexplored at this scale. Through a progressive sparsification framework with second-order saliency and continued training coordinated with sparsity progression, we show that pretrained LLMs can retain strong performance far beyond commonly studied sparsity regimes. Across LLaMA-2 and Qwen-3 model families, our approach improves perplexity and downstream accuracy up to 99\% sparsity, surpassing both the current state-of-the-art and representative baselines. Precisely, on LLaMA-2-7B, our approach achieves WikiText-2 perplexities of 13.48 and 19.67 at 95\% and 99\% sparsity, respectively, while delivering 3.23$\times$ decoding speedup and 6.21$\times$ memory savings at 95\% sparsity. Taken together, our results show that LLMs can be pushed into extreme sparsity while retaining strong performance, providing a foundation for further improving sparse models in this regime.

Comment: Second-order pruning coordinated with progressive sparsification and continued training preserves LLM quality at 95–99% sparsity.

Topic Match: Extreme LLM compression is the core contribution, combining sparsity-aware recovery training with demonstrated memory savings and decoding speedup.

Relevance: 9 Novelty: 6


11. ECOKV: Geometry-Aware KV Cache Eviction via Complementary Diversity Metrics

ArXiv ID: 2609.06663

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Chin Ting Hsu, Yu-Syuan Xu, Ling Zou, Hsien-Kai Kuo, Wen-Huang Cheng

Abstract: Although multimodal Large Language Models (MLLMs) excel in diverse tasks, their scalability remains limited by the memory and computational overhead of KV cache storage. Recent KV cache eviction approaches incorporate a cosine similarity-based diversity metric with importance metrics to selectively retain critical key-value pairs. However, cosine similarity involves normalization that discards magnitude information, and it often yields uniformly high similarity values across layers due to the anisotropy property of hidden representations. In our study ECOKV, we rigorously deconstruct the capabilities of existing diversity metrics. Moving beyond simple measurement, we propose a geometry-aware composite metric that jointly leverages Euclidean distance and cosine similarity to capture token diversity from complementary perspectives. Furthermore, we use these two metrics to estimate the redundancy level of each attention head, allowing adaptive weighting between diversity and importance scores during token selection. Finally, we demonstrate that the observation window commonly employed to preserve recent tokens can be substantially reduced, thereby allocating more cache capacity to informative tokens and yielding consistent improvements. Extensive experiments demonstrate that ECOKV achieves state-of-the-art performance under various compression ratios and can be seamlessly integrated with existing KV cache eviction methods. We further analyze the relationship between importance and diversity, and examine redundancy patterns across layers and attention heads.

Comment: Compresses KV caches using complementary Euclidean and cosine diversity metrics with head-adaptive importance weighting.

Topic Match: KV-cache memory efficiency is the central contribution, with a revised eviction mechanism that adapts token selection to attention-head redundancy.

Relevance: 9 Novelty: 6


12. One Rate Is Not Enough: Adaptive Anisotropic Learning Rates for LoRA Fine-Tuning

ArXiv ID: 2609.05885

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Huiyi Wang, Daijiao Liu, Lina Yao, Dong Gong

Abstract: Low-rank adaptation (LoRA) has become the standard for parameter-efficient fine-tuning of large language models. Most LoRA variants follow a uniform-LR convention, applying a single global learning rate across every rank-one component of every adapter. We show that this convention overlooks substantial within-module heterogeneity, where the rank-one components of a LoRA adapter update at highly uneven rates and low-velocity modules converge to concentrated singular spectra that underutilize the nominal rank budget. To address this, we propose an adaptive anisotropic learning-rate model that assigns each rank-one component its own effective learning rate, computed online from training-time signals and mean-normalized per module to preserve the global LR budget. AnLR-LoRA instantiates this model with two signals available during AdamW optimization, namely function-space velocity and Adam SNR, as a lightweight scheme with no extra trainable parameters. Across commonsense reasoning, natural language generation and visual instruction-tuning benchmarks, AnLR-LoRA consistently improves over LoRA while encouraging broader use of rank capacity, with gains that remain robust across a wide range of global learning rates and transfer cleanly to other LoRA variants.

Comment: Assigns rank-specific LoRA learning rates to counter uneven update speeds and underused adapter rank capacity.

Topic Match: Parameter-efficient adaptation is the primary contribution; analysis of update velocity and singular-spectrum concentration also directly connects to training dynamics.

Relevance: 9 Novelty: 6


13. An AI4AI Framework for Visual Token Pruning

ArXiv ID: 2608.07193

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Zhen Liu, Wenli Huang, Wei Song, Yuhan Liu, Zhiqin Yang, Jingwen Fu

Abstract: Visual-token pruning can substantially reduce the inference cost of multimodal large language models (MLLMs), yet existing methods largely rely on fixed, handcrafted heuristics and costly expert trial and error. As pruning objectives, budgets, and model architectures diversify, manually navigating the expanding design space becomes increasingly difficult. This paper aims to build an AI4AI framework for visual-token pruning by addressing a natural question: Can large language models automatically design effective visual-token reduction algorithms? Although LLMs possess broad algorithmic knowledge and strong reasoning capabilities, translating such general knowledge into effective solutions for a specialized task remains nontrivial. We argue that the key lies in designing an appropriate search-state representation that connects the internal knowledge of LLMs with the structural requirements and constraints of visual-token pruning. Based on this insight, we propose AutoPrune, a training-free framework for LLM-driven visual-token pruning policy design. At its core, AutoPrune introduces a Token Pruning Domain-Specific Language (TPDSL) comprising 131 reusable atoms for budget control, token scoring, selection constraints, and token reassembly. A key property of TPDSL is that it represents each search state as a residual modification of a strong base policy. This residual formulation narrows the search space and directs the LLM's attention toward the policy components that are most consequential for performance. Experiments on 14 multimodal benchmarks and three MLLM backbones demonstrate the effectiveness, efficiency, and transferability of AutoPrune. Even when removing 94.4% of visual tokens, AutoPrune preserves more than 99% of full-token performance while reducing FLOPs by 9.9x and prefill latency by 6.4x.

Comment: A residual-policy search language enables automated discovery of visual-token pruning policies that reduce MLLM prefill cost.

Topic Match: The central contribution is a pruning-policy search mechanism with substantial token, FLOP, and prefill reductions across multiple backbones.

Relevance: 8 Novelty: 7


14. LookThere! Sparse Vision by Reinforced Selection

ArXiv ID: 2609.04698

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Sreehari Rammohan, Yousef Yassin, Anthony Fuller, Junfeng Wen, Carl Vondrick, Evan Shelhamer

Abstract: Vision transformers typically treat every image token as equally important, yet for most tasks in computer vision only a fraction are needed. Adaptive computation methods accelerate inference by choosing which tokens to process, but existing methods struggle at extreme sparsity and require heuristics that may not generalize like token diversity and attention scores. We address these limitations with LookThere, achieving a new pareto frontier in performance-compute trade-offs through an end-to-end reinforcement learning framework that jointly trains a shallow input selector and a deep representation extractor. The selector learns where to look and the extractor learns what to see, together saving computation by selecting only what is worth processing for a given task without relying on auxiliary signals. We show that LookThere only selects the task-specific input, excelling at sparse recognition in high-resolution settings (traffic signs, billiards), and maintaining accuracy with as little as 0.2% of the input. It generalizes across tasks and models, including global recognition (ImageNet classification), local recognition (ADE20K segmentation), zero-shot classification (by distillation), and regression (counting). Across all settings, LookThere surpasses state-of-the-art selection to provide a general and scalable framework for specialized and efficient adaptive computation.

Comment: Jointly trains an input selector and representation extractor to enable extremely sparse transformer computation without heuristic token scores.

Topic Match: Learned input sparsity directly reduces transformer compute; the selector-extractor mechanism also contributes adaptive computation across multiple vision tasks.

Relevance: 8 Novelty: 7


15. EStream: Fast and Memory-Efficient MoE Prefill through Expert Virtualization on Mobile NPUs

ArXiv ID: 2609.06551

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Junming Zhang, Zhenzhe Zheng, Fan Wu, Xiaoyao Huang, Jie Wu

Abstract: Mobile vendors and application developers increasingly deploy LLMs on smartphones for diverse prefill-only services. Yet current systems rely mainly on dense models whose regular computation maps efficiently to mobile NPUs, leaving more capable MoEs underused. MoE prefill does not fit mobile NPUs: NPU graphs are fixed at compile time, yet MoE picks experts at runtime; and one request touches most experts, more than a phone can hold in memory. We present EStream, which resolves both by separating what the NPU must fix from what MoE decides at runtime. A single compiled expert graph serves every expert, with each expert's routed tokens and weight address bound at call time, so dynamic MoE execution runs entirely on the NPU without padding or CPU/GPU fallback. Expert virtualization keeps the expert pool in UFS flash storage and pages it through a fixed-size NPU-addressable arena, group by group, with loading hidden behind computation, so memory is bounded by the arena rather than by the model. It further introduces a hardware-aware configuration algorithm that automatically configures the UFS--NPU pipeline and maximizes loading--computation overlap. Across 18 comparative settings covering three 7B--16B MoEs and 256--4,096-token prompts, we evaluate EStream on a commercial Snapdragon smartphone. Compared to the fastest baseline at each setting, EStream achieves a 2.25--27.57X pure-prefill TTFT speedup and reduces peak physical memory by 1.19--12.29X. EStream further scales to MoE models with up to 46.7B parameters.

Comment: Runtime-bound expert graphs and flash-backed expert virtualization reduce MoE prefill memory requirements and latency.

Topic Match: The substantive contribution is a new execution and memory design for mobile MoE inference.

Relevance: 8 Novelty: 7


16. Interface-Aware KV Cache Quantization for Dense On-Chip NVM in Long-Context LLM Decoding

ArXiv ID: 2609.05764

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Jiahao Zheng, Yifan Qin, Xiaobo Sharon Hu, Yiyu Shi

Abstract: The key-value (KV) cache is the dominant memory bottleneck in long-context large language model (LLM) decoding: every step reads it entirely, so decoding is memory-bandwidth bound. Holding a quantized KV cache in dense on-chip non-volatile memory (NVM) removes the off-chip transfer. Existing KV quantization methods, however, were designed for GPU-style memory systems: KIVI attaches per-group metadata, adding about 25% to the stored KV cache; KVQuant keeps sparse full-precision outliers that a dense array cannot hold in place. This paper examines what these structures cost when the KV cache resides in NVM behind fixed-range converters, and designs a quantization scheme matched to that interface. The architecture stores the quantized KV cache in dense on-chip NVM, uses a small static analog crossbar only for the fixed rotation, and keeps attention in on-chip digital logic. A randomized rotation and per-vector normalization give every coordinate the same range, so one fixed codebook for keys and one for values, each shared across all tokens of the corresponding tensor type, serve the entire KV cache. The codebook thresholds are programmed once as the read converter's reference levels, enabling fixed-range digitization with no per-token converter reconfiguration. Dequantization is a sixteen-entry lookup and one norm multiply; the only per-vector metadata is one scalar, about 3%. Across models from 3B to 14B and contexts to 32k tokens, the four-bit KV cache maintains accuracy under storage and crossbar noise simulated at realistic device levels. KIVI and KVQuant remain more accurate in software; the advantage of our format lies at the memory interface: 3.1-3.6x lower KV read energy than both mapped to the same NVM, and 8x lower metadata overhead than KIVI. The contribution is a KV quantization co-designed with the NVM memory interface rather than a new accuracy record.

Comment: Fixed converter codebooks and per-vector normalization reduce KV-cache metadata and on-chip NVM read energy.

Topic Match: Interface-aware KV quantization directly changes long-context inference memory and energy costs, with benefits specific to NVM hardware.

Relevance: 8 Novelty: 7


17. Generating Pretraining Tokens from Organic Data for Data-Bound Scaling

ArXiv ID: 2605.17849

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Zichun Yu, Chenyan Xiong

Abstract: LLM pretraining is shifting from a compute-bound to a data-bound regime, where available human (organic) text falls far short of scaling demands. However, reaching the data-bound regime does not mean the model has fully utilized its organic corpus. In this paper, we introduce SynPro, a synthetic data generation framework that helps LLMs more thoroughly learn from limited organic data. SynPro applies two operations, rephrasing and reformatting, that present the same organic source in diverse forms to facilitate deeper learning without introducing external information. Both generators are optimized via reinforcement learning with quality, faithfulness, and data influence rewards, and are continuously updated as pretraining plateaus to target content the model has yet to absorb. We pretrain 400M, 1.1B, and 2B models with 10% of their Chinchilla-optimal tokens (0.8B, 2.2B, and 4B) from DCLM-Baseline, reflecting a realistic data-bound regime in frontier pretraining. Our results reveal that organic data is significantly underutilized by standard repetition: SynPro unlocks 3.4--5.2x the effective tokens of repetition, even surpassing the non-data-bound oracle that trains on equivalent unique data at the 1.1B and 2B scales. Analyses confirm that faithful, model-aware synthesis sustains data-bound scaling without causing distribution collapse. We open-source our code at https://github.com/cxcscmu/SynPro.

Comment: Model-aware rephrasing and reformatting increase effective pretraining-data yield under limited organic-text budgets.

Topic Match: The central contribution improves pretraining data efficiency and sustains scaling when unique organic text is scarce.

Relevance: 8 Novelty: 7


18. PLUME: Parameter-Efficient Personalization of Large Language Models via Low-Rank User Modulation in Shared Subspaces

ArXiv ID: 2609.04715

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Xinyu Li, Hao Zhou, Jianfeng Zhu, Julina Maharjan, Ruixin Guo, Feodor Dragan, Ruoming Jin

Abstract: Personalizing large language models (LLMs) is essential for delivering AI assistance that aligns with individual users' styles, intents, and preferences. While per-user fine-tuning can substantially enhance personalization quality, it introduces significant parameter and storage overhead, limiting scalability to large user populations. We propose PLUME (Personalized Low-Rank Adaptation through User Modulation and Shared Subspace), a lightweight framework that achieves efficient and expressive per-user adaptation by leveraging a shared task-specific subspace. Specifically, PLUME first learns a global task subspace from aggregated user data. Personalization is then achieved by training only a lightweight small square matrix within this subspace, enabling each user to obtain a tailored model while keeping shared components fixed. Cross-layer shared parameters and rank-1 residual terms are further introduced to significantly reduce redundancy while maintaining expressiveness. Experiments on multiple personalized text generation benchmarks demonstrate that PLUME achieves comparable or superior performance to strong baselines, while reducing per-user parameters by over 95%. These results establish shared-subspace modulation with minimal residuals as a scalable and semantically grounded approach to LLM personalization.

Comment: Shared low-rank task subspaces and small user-specific modulation matrices reduce adaptation parameter storage.

Topic Match: The shared-subspace parameterization, cross-layer sharing, and minimal residuals constitute an efficiency mechanism beyond simply applying standard LoRA.

Relevance: 8 Novelty: 6


19. STAR-Pro: Stage-Wise Token Adaptive Reduction with Progressive Refinement for Efficient Large Vision-Language Models

ArXiv ID: 2609.05916

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Yichen Guo, Tinghao Wang, Qizhe Zhang, Lingbei Meng, Yuan Zhang, Jiajun Cao, Hao Jiang, Chenwei Wu, Jixian Wu, Sixiang Chen, Tao Luo, Hongyang Cheng, Kai Tang, Chenxi Li, Renyuan Li, Xiande Huang, Wenya Wang, Shanghang Zhang

Abstract: Large vision-language models (LVLMs) achieve strong multimodal understanding, but the hundreds to thousands of visual tokens they process impose substantial computational overhead, motivating training-free visual token pruning. In this work, we conduct two complementary analyses of visual token pruning. First, we measure the feature-space coverage of tokens retained before cross-modal fusion and find that aggressive pruning discards substantial visual information. Second, we track text-to-visual attention across decoder layers and find that the visual tokens considered important change substantially with depth, making one-shot pruning decisions unreliable. Together, these findings show that effective pruning should preserve broad visual coverage before fusion and progressively refine the retained tokens as cross-modal evidence evolves during fusion. We therefore propose STAR-Pro (STage-Wise Adaptive Token Reduction with Progressive Refinement), a training-free two-stage framework. Its Adaptive Stage applies pivoted QR to construct an over-budget feature-coverage candidate pool, while its Progressive Stage uses evolving text-to-visual attention at selected decoder layers to prune a nested survivor set under a target layer-average token budget. Extensive experiments across seven LVLMs spanning multiple architectures and 18 image and video benchmarks demonstrate the effectiveness of STAR-Pro under aggressive pruning. On LLaVA-Video-7B, STAR-Pro reduces visual tokens by 90.5%, retains 92.7% of baseline performance, and achieves a $2.24\times$ measured inference speedup. Code is available at https://github.com/EasonAI-5589/starpro.

Comment: QR-based coverage selection and progressive attention-based pruning reduce visual-token computation across decoder layers.

Topic Match: Its core is a token-reduction mechanism that lowers LVLM inference cost across architectures, with measured speedup under aggressive pruning.

Relevance: 8 Novelty: 6


20. Cost-Optimal LLM Routing with Limited User Feedback under User Satisfaction Guarantees

ArXiv ID: 2606.19376

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Herbert Woisetschläger, Arastun Mammadli, Ryan Zhang, Shiqiang Wang

Abstract: Inference costs for large language model (LLM) applications are rapidly growing, driven by surging demand and rising infrastructure cost. Users expect high-quality responses, and in commercial settings this is formally codified in Service Level Agreements (SLAs), creating a fundamental tension between cost and quality. Recent progress on cost-aware LLM request routing has shown potential to resolve this tension, but existing approaches rely on complete feedback signals, offline training, extensive per-workload tuning, and most lack SLA guarantees or inference-time adaptivity. We introduce SLARouter, an online routing algorithm that learns a cost-optimal policy from the sparse, one-sided user feedback available in production systems. SLARouter provides theoretical guarantees for both cost optimality and strict SLA compliance. Experiments across a wide range of LLM benchmarks show that SLARouter satisfies SLA constraints without the need for per-benchmark tuning, reducing operating cost by up to 2.2x over existing baselines.

Comment: An online model-selection algorithm reduces inference cost using sparse feedback while providing satisfaction guarantees.

Topic Match: The novel cost-optimization algorithm fits inference efficiency, although request-level serving policies are peripheral to this training-centered feed.

Relevance: 7 Novelty: 7


21. KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU

ArXiv ID: 2609.04852

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Di Chai, Leye Wang, Zeshen Su, Zhiguo Xia, Zhihang Yu

Abstract: Modern LLM agents operate in persistent workspaces whose accumulated history can exceed both GPU KV capacity and the model's native context window. Existing systems typically compact older context into summaries or retrieve it later as text, either losing fine-grained execution evidence or repeatedly prefilling content that the model has already processed. We present KVMem, a KV-context virtualization system that preserves overflowed workspace history as paged KV state across GPU memory, host memory, and NVMe. KVMem uses lightweight, model-native attention-space indexes to select relevant historical blocks and materializes a query-dependent execution view bounded by the model's native context window. Extensive evaluations on long-context agent benchmarks spanning histories up to one million tokens, including LongMemEval, MemoryAgentBench, and AgentLongBench, show that KVMem generally achieves higher task utility and greater inference efficiency than compaction-based approaches, the de facto standard for handling context overflow. In the DeepSWE long-context test with Qwen3.8-27B, KVMem improves task success from 43.8% with compaction-only context management to 48.4%. In our local-deployment evaluation, KVMem runs Qwen3.6/3.8-27B NVFP4 with MTP on an off-the-shelf laptop equipped with a 24\,GB RTX 5090 Laptop GPU, virtualizing agent workspaces of up to 1M tokens-four times the model's native 256K-token context window. In a single-session setting, KVMem generates $\sim$50 tokens/s, providing interactive responsiveness for local agent execution. More broadly, by decoupling addressable workspace size from the LLM's native context window, KVMem provides a practical path toward long-running agents whose workspaces can grow beyond that window.

Comment: Reuses historical KV blocks across GPU, host memory, and NVMe through query-selected execution views.

Topic Match: KV reuse and attention-indexed materialization provide a concrete memory-efficiency mechanism, although the system primarily targets persistent agent workspaces.

Relevance: 7 Novelty: 7


22. SplitLite: Low-Rank Residual Compression for Split Learning

ArXiv ID: 2608.23018

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Large-Scale Training Systems and Efficiency

Authors: Tao Li, Yulin Tang, Qi Guo, Xianhao Chen

Abstract: Federated fine-tuning of on-device large language models (LLMs) faces a significant computing burden. To overcome this limitation, split learning (SL) has emerged as a promising solution, which offloads the primary training workload to a powerful server. However, SL requires exchanging high-dimensional activations and gradients between clients and the server, resulting in prohibitive communication costs. To overcome this challenge, we propose SplitLite, a communication-efficient split federated LoRA fine-tuning method that exploits the low effective rank structure of consecutive-epoch activation and gradient residuals. Our key finding is that, when LoRA uses rank $r$ updates in parameter space, the activation and gradient residuals of the same data sample between adjacent epochs also exhibit effective rank-$2r$ and rank-$4r$ structures, respectively. By revealing this property, SplitLite transmits only quantized truncated singular value decomposition (SVD) residual factors, thereby significantly reducing both activation uplink and gradient downlink traffic. Extensive experiments on the GLUE benchmark across a series of advanced on-device LLMs demonstrate that our method reduces activation uplink communication costs by up to 93.5\% and total communication costs by up to 83.7\%, without performance degradation.

Comment: Derives low-rank activation and gradient residual structures to compress communication during split LoRA fine-tuning.

Topic Match: Low-rank residual transmission directly reduces distributed fine-tuning communication cost.

Relevance: 7 Novelty: 7


23. Aplaud: Adaptive Personalized Low-Rank Decomposition for User-Specific LLM

ArXiv ID: 2609.04738

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Xinyu Li, Ruoming Jin, Jianfeng Zhu, Ruixin Guo, Zhi Liu

Abstract: In this paper, we study the problem of personalized survey response prediction using fine-tuned large language models (LLMs). This task poses unique challenges: limited per-user training data, scalability of model storage, and the need to exploit shared structure across survey questions. To address these issues, we propose Aplaud (Adaptive Personalized Low-rank and User-specific Nested Decomposition), a lightweight and scalable framework for LLM personalization. Aplaud extends the LoRA paradigm by separating adaptation into a frozen, shared low-rank basis and a compact user-specific correction, augmented with a rank-one residual for finer personalization. To further reduce per-user parameter cost and mitigate overfitting, the correction matrix can be factorized into an even lower-rank form. Empirical results demonstrate that Aplaud achieves efficient, scalable personalization across users while outperforming state-of-the-art LoRA-based personalized LLM approaches in both generalization and inference efficiency.

Comment: A frozen shared LoRA basis with nested user-specific low-rank corrections reduces per-user adapter storage.

Topic Match: The central contribution changes LLM adapter parameterization and per-user memory cost; its personalization focus makes the match narrower.

Relevance: 7 Novelty: 6


24. From Deep to Shallow: Unconstrained and Efficient Layer Merging Strategy

ArXiv ID: 2609.04881

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Petro Shulzhenko, Gabriele Spadaro, Enzo Tartaglione

Abstract: Although Deep Neural Networks have become foundational in many areas of Machine Learning, high computational demands limit their application in resource-constrained environments. To address this issue, depth compression methods have been proposed to identify and linearize redundant activation functions, thereby allowing for the merging of layers without intermediate non-linearities. However, these methods face two key challenges: they cannot be directly applied to convolutions with padding due to the absence of an analytical solution for merging these layers, and they typically increase the kernel size of merged layers, thus limiting speed-up gains. To overcome these limitations, we propose an efficient strategy that enables merging of layers without an existing analytical solution, and also without increasing kernel size. We validate our approach across multiple architectures and datasets, and measure inference speed-up gains on real embedded platforms. We publicly released the code at https://github.com/ShulzhenkoPetr/deep-to-shallow.

Comment: Layer merging handles padded convolutions without increasing kernel size or requiring an analytical fusion rule.

Topic Match: Depth compression is the core methodological contribution, although its demonstrated scope emphasizes convolutional networks and embedded inference.

Relevance: 7 Novelty: 6


25. RePro: Training Language Models to Faithfully Recycle the Web for Pretraining

ArXiv ID: 2510.10681

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Zichun Yu, Chenyan Xiong

Abstract: High-quality data is a cornerstone of large language model (LLM) pretraining, yet its growth has not kept pace with the needs of frontier models. In this paper, we introduce RePro, a novel web recycling method that trains a relatively small LM with reinforcement learning to generate effective and faithful rephrasings of pretraining data. Specifically, we design one quality reward and three faithfulness rewards, optimizing the LM rephraser to convert organic data into high-quality rephrasings while maintaining its core semantics and structure. In our experiment, we train a rephraser as small as 1B parameters to recycle 72B tokens sampled from DCLM-RefinedWeb. Pretraining results on 400M, 1.4B, and 2.8B models demonstrate that RePro delivers 3.7%-14.5% relative accuracy gains over organic-only baseline on 22 downstream tasks, doubling the performance gains achieved by the state-of-the-art web recycling method that prompts a 70B rephraser. Experiments with different amounts of recycled data highlight that RePro improves organic data efficiency by 2-3x. Individual and distributional analyses validate that RePro preserves more critical information and faithfully reflects the characteristics of organic data compared to prompting-based methods. Together, these results show that RePro provides an efficient and controllable path to effectively recycle organic data for pretraining. Our code is available at https://github.com/cxcscmu/RePro.

Comment: Uses a 1B learned rephraser to improve pretraining-data efficiency while replacing a prompted 70B data generator.

Topic Match: The strongest fit is pretraining data efficiency, supported by reported 2-3x improvements in organic-data utilization and a substantially smaller rephrasing model.

Relevance: 7 Novelty: 6


26. SPD: Single Pass Decoding for Generative Reranking

ArXiv ID: 2609.01807

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Emil Laftchiev, Prachi Agrawal, Moe Kayali, Bixing Yan, Qi Xu, Zijie Lei, Chen Qiu, Zhi Hua, Ke Li, Luke Simon

Abstract: Large language models (LLMs) achieve state-of-the-art generative ranking quality, but the ranking they produce must be decoded, and autoregressive decoding spends one sequential forward pass per emitted token. We observe that the only tokens a ranker must emit are the $N$ ordinal values naming the items in ranked order, and that this narrow, permutation-structured output format admits decoding strategies which are much more efficient than left-to-right generation. We introduce SPD (Single Forward Pass), a format-specialized decoding strategy that decodes all $N$ ordinals in $O(1)$ forward passes. SPD reads an $N \times K$ item-position score matrix off the LLM's prefill hidden states with a lightweight self-attention head, then decodes the ordinals as the optimal bipartite assignment of that matrix via the Hungarian algorithm, yielding a valid permutation by construction rather than by repair. Through a systematic study of training signals and backbone adaptation, we show that LoRA-based fine-tuning combined with auto-regressive LLM ranking distillation reaches 28 ms end-to-end inference, a speed-up of 64x while maintaining ranking quality on par with the teacher. We provide a complete ablation decomposing the contributions of architecture, training signal, and backbone adaptation. Our framework connects generative ranking to combinatorial optimization, opening a path toward other $O(1)$-decode mechanisms for real-time ranking.

Comment: One-pass item-position scoring followed by bipartite assignment replaces sequential LLM ranking generation.

Topic Match: The core mechanism removes repeated backbone decoding passes, making it a substantive inference-efficiency contribution, although restricted to permutation-structured ranking outputs.

Relevance: 7 Novelty: 6


27. Neural Field Tokenizations with Hierarchy and Spatial Locality Priors

ArXiv ID: 2606.08204

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Alonso Urbano, David W. Romero, Max Zimmer, Sebastian Pokutta

Abstract: Neural fields parameterize data as functions from coordinates to values, providing a unified framework for representation learning across modalities. Existing approaches are dominated by per-sample meta-learning, which scales poorly due to memory-intensive inner-loop optimization. The natural alternative -- feed-forward encoding -- typically introduces modality-specific assumptions, sacrificing the generality that makes learning with neural fields attractive. We argue that locality and hierarchy are useful priors for learning field representations that can be injected without compromising modality-agnosticism. We propose LH-NeF, a framework to learn general-purpose tokenized representations of continuous signals. A locality-preserving hierarchical encoder maps raw coordinate-value field observations to structured tokens, from which the field is reconstructed during training. By replacing meta-learning's inner loop with a single forward pass, LH-NeF uses 42x less memory and supports 133x larger batches than the strongest modality-agnostic baseline. Across images, 3D shapes, and climate fields, our learned representations match or exceed performance of modality-agnostic, modality-specific, and specialized generative neural field baselines on both reconstruction and downstream tasks.

Comment: Replaces memory-intensive per-sample meta-learning with a locality-preserving feed-forward tokenizer.

Topic Match: The headline contribution is a large reduction in representation-training memory and batch cost.

Relevance: 6 Novelty: 7


28. Coarse-Graining Hidden Representations: Unsupervised Neuron Selection via Mapping Entropy

ArXiv ID: 2609.05126

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Margherita Mele, Andrea Castagna, Roberto Menichetti, Raffaello Potestio, Alessandro Ingrosso

Abstract: Overparameterized neural networks carry far more hidden units than a task nominally requires, raising the question of which neurons are essential and whether that distinction is legible in the representation itself, without labels or gradients. We cast neuron selection as the problem of coarse-graining the hidden layer by retaining a subset of its neurons, and score each putative selection by the mapping entropy (ME). This quantity measures the loss of discriminatory power inherent in discarding part of the network neurons, and the selection that minimises the ME is taken as particularly informative. This criterion is fully unsupervised, in that it depends only on hidden-activation statistics. In teacher-student networks, ME optimisation recovers the minimal teacher-consistent representation and retains extra units in proportion to the hidden layer's residual variability; in a non-linear Gaussian process task, it selects coherent functional-class mappings whose preferred class shifts across training. On this task and on translation-augmented MNIST, ME-selected subnetworks outperform random subsets of equal size, most clearly under strong compression - linking configurational distinguishability to predictive performance.

Comment: Mapping entropy supplies an activation-only criterion for selecting neurons under compression.

Topic Match: Neuron selection directly concerns compression, although the evidence covers small networks and does not establish large-model savings.

Relevance: 6 Novelty: 7


29. Subspace Optimization for Backpropagation-Free Continual Test-Time Adaptation

ArXiv ID: 2603.28678

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Damian Sójka, Sebastian Cygert, Marc Masana

Abstract: We introduce PACE, a backpropagation-free continual test-time adaptation system that directly optimizes the affine parameters of normalization layers. Existing derivative-free approaches struggle to balance runtime efficiency with learning capacity, as they either restrict updates to input prompts or require continuous, resource-intensive adaptation regardless of domain stability. To address these limitations, PACE leverages the Covariance Matrix Adaptation Evolution Strategy with the Fastfood projection to optimize high-dimensional affine parameters within a low-dimensional subspace, leading to superior adaptive performance. Furthermore, we enhance the runtime efficiency by incorporating an adaptation stopping criterion and a domain-specialized vector bank to eliminate redundant computation. Our framework achieves state-of-the-art accuracy across multiple benchmarks under continual distribution shifts, reducing runtime by over 50% compared to existing backpropagation-free methods.

Comment: Reduces adaptation compute through derivative-free optimization of normalization parameters in a low-dimensional subspace.

Topic Match: Adaptation runtime provides a meaningful efficiency connection, but the contribution combines established search and projection methods for continual test-time adaptation, with limited evidence of large-model training relevance.

Relevance: 6 Novelty: 6


30. InKAN: B-Spline KANs via Truncated Power Form

ArXiv ID: 2609.01956

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Naveen Mysore

Abstract: Kolmogorov-Arnold Networks (KANs) place learnable B-spline activations on network edges rather than fixed activations on nodes. The standard Cox-de Boor recursion evaluates these activations through $k$ sequential passes for degree-$k$ splines, consuming over 90% of forward-pass time. InKAN replaces this recursion with the truncated power form, a classical result from approximation theory that expresses each uniform cubic B-spline as five $(x)_+^3$ terms at shifted knot positions. The resulting expression computes exact B-spline basis values: the same mathematical function as the Cox-de Boor recursion, evaluated without sequential passes. This paper documents three contributions: (1) an implementation structured for torch . compile fusion, eliminating recursion, span lookup, and scatter-gather operations; (2) a bounded-coordinate evaluation that clamps the normalized input to $[0, k{+}1]$, preventing the growth of cancellation error at large off-support coordinates; and (3) an open-source package (pip install inkan). In the tested configurations, InKAN has 2.8--3.5$\times$ lower forward-pass latency than the Cox-de Boor recursion. Partition-of-unity errors remain below $10^{-5}$ for grid sizes up to 200.

Comment: Reformulates exact B-spline evaluation to eliminate sequential recursion and enable compiler fusion.

Topic Match: The new evaluation path materially lowers the computational cost of a learned activation architecture.

Relevance: 6 Novelty: 6


Paper Selection Prompt

System Prompt

You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.

User Prompt

Instructions

Respond in JSONL. Output exactly one JSON object per paper, one per line:

{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}

Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only: daily_hot, new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return []. - daily_hot means the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. - new_frontier means the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.

Scoring Criteria

Relevance and Novelty are independent axes. Score both from 1 to 10.

Relevance Scoring

  • 9-10: directly centered on the target foundational topics; highest when the core contribution is clearly within them.
  • 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
  • 5-6: touches the target topics, but the main contribution is elsewhere.
  • 3-4: largely outside the target topics, often application-focused or domain-specific.
  • 1-2: unrelated.

Important: Broad frontier relevance, major launch status, or daily buzz is not enough for a high Relevance score here. Those cases belong in the hotspot digest unless the paper strongly matches the specialized paper topics.

Novelty Scoring

  • 9-10: new paradigm, theory, or major methodological breakthrough.
  • 7-8: substantial methodological advance or strong new insight.
  • 5-6: meaningful but incremental extension or refinement.
  • 3-4: minor, narrow, or mostly engineering or domain-specific improvement.
  • 1-2: little originality; mainly standard application of existing methods.

Topic Registry

Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - moe_training: MoE Training - Mixture-of-Experts training: routing, load balancing, expert granularity and shared experts, training stability, MoE parallelism and communication, and MoE kernels. - training_systems: Large-Scale Training Systems and Efficiency - Large-scale training systems and efficiency: distributed training algorithms, optimizers, and communication. - architecture_training: Architecture and Training Dynamics - Core architectural or computational mechanisms, dynamic computation, and training-stability dynamics. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, memory or cache efficiency, and large-scale training systems that materially change cost or behavior.

Papers

[PAPER LIST HERE]

Relevant Topics

This is a training-side feed. The centre of gravity is Mixture-of-Experts training and the large-scale training systems it sits on; architecture and efficiency work is kept when it bears on how a large model is trained or what it costs to train.

Keep a paper when its CORE CONTRIBUTION falls in one of the four topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.

  1. MoE Training (primary) - Keep: expert routing and gating mechanisms; load balancing, with or without auxiliary losses; expert granularity, fine-grained and shared experts; MoE training stability (loss spikes, router collapse, z-loss and related fixes); expert and hybrid parallelism; all-to-all dispatch/combine and the communication schedules MoE needs; MoE kernels and grouped-GEMM style implementations; expert capacity, token dropping, and upcycling dense checkpoints into MoE. - Filter: papers that merely train on top of a MoE model without a new routing, balancing, stability, parallelism, or kernel contribution; "mixture of experts" in the classical ensemble or recommender sense.

  2. Large-Scale Training Systems - Keep: distributed training algorithms (data, tensor, pipeline, sequence, and expert parallelism, FSDP and ZeRO-style sharding); optimizers and preconditioners for large-scale pretraining; communication algorithms, overlap, and collective design; throughput and MFU work that materially changes what a training run costs or whether it converges; scaling-law work that informs how such runs are configured. - Filter: routine infrastructure tuning with no new algorithmic or systems idea.

  3. Architecture and Training Dynamics - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, dynamic or modular computation); optimisation and training-dynamics analysis that explains why large models train the way they do. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight.

  4. Efficiency and Compression - Keep: quantization, sparsity, pruning, low-rank adaptation, KV-cache and memory-efficient designs, and other work that materially changes the cost of running or training a large model, especially when the mechanism is new rather than a tuned variant. - Filter: straightforward application of standard efficiency methods, and deployment or serving engineering with no new idea.

Leave these to the hotspot digest unless the core contribution clearly falls in one of the four topics above: - major model or product releases - agent, tooling, and RAG launches - benchmark, leaderboard, or evaluation-only papers - representation-learning theory, interpretability, and feature-formation analysis - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning) - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains