Previous Day 2026-03-31
Monthly Overview 2026-04
Next Day 2026-04-02

Personalized Daily ArXiv Papers 2026-04-01

Model Metric Usage Papers
Prompt Completion Total Total arXiv Scanned Relevant
gpt-5.4 Tokens 165268 21187 186455 595 335 15
Cost $0.41 $0.32 $0.73

Papers on this page were re-scored with local agent CLI; the usage above is from the original run.

Topic Coverage:

TopicPapers
Architecture and Training Dynamics21
Training Algorithms That Change What Is Possible1
MoE Where It Changes the Design Space1
Efficiency, Compression, and Large-Scale Training3

Table of contents by topic:

Architecture and Training Dynamics (21)

  1. Tucker Attention: A generalization of approximate attention mechanisms Authors: Timon Klein, Jonas Kusch, Sebastian Sager, Stefan Schnake, Steffen Schotth\"ofer

  2. OptiMer: Optimal Distribution Vector Merging Is Better than Data Mixing for Continual Pre-Training Authors: Haiyue Song, Masao Utiyama

  3. The Spectral Edge Thesis: A Mathematical Framework for Intra-Signal Phase Transitions in Neural Network Training Authors: Yongzhong Xu

  4. From Density Matrices to Phase Transitions in Deep Learning: Spectral Early Warnings and Interpretability Authors: Max Hennick, Guillaume Corlouer

  5. Metriplector: From Field Theory to Neural Architecture Authors: Dan Oprisa, Peter Toth

  6. A Comprehensive Information-Decomposition Analysis of Large Vision-Language Models Authors: Lixin Xiu, Xufang Luo, Hideki Nakayama

  7. ASI-Evolve: AI Accelerates AI Authors: Weixian Xu, Tiantian Mi, Yixiu Liu, Yang Nan, Zhimeng Zhou, Lyumanshan Ye, Lin Zhang, Yu Qiao, Pengfei Liu

  8. Think Anywhere in Code Generation Authors: Xue Jiang, Tianyu Zhang, Ge Li, Mengyang Liu, Taozhi Chen, Zhenhua Xu, Binhua Li, Wenpin Jiao, Zhi Jin, Yongbin Li, Yihong Dong

  9. On the Mirage of Long-Range Dependency, with an Application to Integer Multiplication Authors: Zichao Wei

  10. Grokking From Abstraction to Intelligence Authors: Junjie Zhang, Zhen Shen, Gang Xiong, Xisong Dong

  11. Is the Modality Gap a Bug or a Feature? A Robustness Perspective Authors: Rhea Chowers, Oshri Naparstek, Udi Barzelay, Yair Weiss

  12. DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Authors: Yi Chen, Yuying Ge, Hui Zhou, Mingyu Ding, Yixiao Ge, Xihui Liu

  13. APEX-EM: Non-Parametric Online Learning for Autonomous Agents via Structured Procedural-Episodic Experience Replay Authors: Pratyay Banerjee, Masud Moshtaghi, Ankit Chadha

  14. LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning Authors: Haihong Hao, Lei Chen, Mingfei Han, Changlin Li, Dong An, Yuqiang Yang, Zhihui Li, Xiaojun Chang

  15. Tracking Equivalent Mechanistic Interpretations Across Neural Networks Authors: Alan Sun, Mariya Toneva

  16. OccSim: Multi-kilometer Simulation with Long-horizon Occupancy World Models Authors: Tianran Liu, Shengwen Zhao, Mozhgan Pourkeshavarz, Weican Li, Nicholas Rhinehart

  17. Concept frustration: Aligning human concepts and machine representations Authors: Enrico Parisini, Christopher J. Soelistyo, Ahab Isaac, Alessandro Barp, Christopher R. S. Banerji

  18. Lie Generator Networks for Nonlinear Partial Differential Equations Authors: Shafayeth Jamil, Rehan Kapadia

  19. A Pontryagin Method of Model-based Reinforcement Learning via Hamiltonian Actor-Critic Authors: Chengyang Gu, Yuxin Pan, Hui Xiong, Yize Chen

  20. Minimum Norm Interpolation via The Local Theory of Banach Spaces: The Role of $2$-Uniform Convexity Authors: Gil Kur, Pierre Bizeul

  21. GENIE: Gram-Eigenmode INR Editing with Closed-Form Geometry Updates Authors: Samundra Karki, Adarsh Krishnamurthy, Baskar Ganapathysubramanian

Training Algorithms That Change What Is Possible (1)

  1. Byzantine-Robust and Communication-Efficient Distributed Training: Compressive and Cyclic Gradient Coding Authors: Chengxi Li, Youssef Allouah, Rachid Guerraoui, Mikael Skoglund, Ming Xiao

MoE Where It Changes the Design Space (1)

  1. Training-Free Dynamic Upcycling of Expert Language Models Authors: Eros Fan`i, O\u{g}uzhan Ersoy

Efficiency, Compression, and Large-Scale Training (3)

  1. Baby Scale: Investigating Models Trained on Individual Children's Language Input Authors: Steven Y. Feng, Alvin W. M. Tan, Michael C. Frank

  2. Big2Small: A Unifying Neural Network Framework for Model Compression Authors: Jing-Xiao Liao, Haoran Wang, Tao Li, Daoming Lyu, Yi Zhang, Chengjun Cai, Feng-Lei Fan

  3. Nonnegative Matrix Factorization in the Component-Wise L1 Norm for Sparse Data Authors: Giovanni Seraghiti, K\'evin Dubrulle, Arnaud Vandaele, Nicolas Gillis


Architecture and Training Dynamics (21)

1. Tucker Attention: A generalization of approximate attention mechanisms

ArXiv ID: 2603.30033

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Timon Klein, Jonas Kusch, Sebastian Sager, Stefan Schnake, Steffen Schotth\"ofer

Abstract: The pursuit of reducing the memory footprint of the self-attention mechanism in multi-headed self attention (MHA) spawned a rich portfolio of methods, e.g., group-query attention (GQA) and multi-head latent attention (MLA). The methods leverage specialized low-rank factorizations across embedding dimensions or attention heads. From the point of view of classical low-rank approximation, these methods are unconventional and raise questions of which objects they really approximate and how to interpret the low-rank behavior of the resulting representations. To answer these questions, this work proposes a generalized view on the weight objects in the self-attention layer and a factorization strategy, which allows us to construct a parameter efficient scheme, called Tucker Attention. Tucker Attention requires an order of magnitude fewer parameters for comparable validation metrics, compared to GQA and MLA, as evaluated in LLM and ViT test cases. Additionally, Tucker Attention~encompasses GQA, MLA, MHA as special cases and is fully compatible with flash-attention and rotary position embeddings (RoPE). This generalization strategy yields insights of the actual ranks achieved by MHA, GQA, and MLA, and further enables simplifications for MLA.

Comment: Unifies MHA, GQA and MLA through tensor factorization and derives a lower-parameter attention scheme.

Topic Match: Redesigning attention parameterization is the primary contribution; compression follows from the factorization and is evaluated against GQA and MLA.

Relevance: 10 Novelty: 8


2. OptiMer: Optimal Distribution Vector Merging Is Better than Data Mixing for Continual Pre-Training

ArXiv ID: 2603.28858

Primary Topic: Architecture and Training Dynamics

Authors: Haiyue Song, Masao Utiyama

Abstract: Continual pre-training is widely used to adapt LLMs to target languages and domains, yet the mixture ratio of training data remains a sensitive hyperparameter that is expensive to tune: they must be fixed before training begins, and a suboptimal choice can waste weeks of compute. In this work, we propose OptiMer, which decouples ratio selection from training: we train one CPT model per dataset, extract each model's distribution vector, which represents the parameter shift induced by that dataset, and search for optimal composition weights post-hoc via Bayesian optimization. Experiments on Gemma 3 27B across languages (Japanese, Chinese) and domains (Math, Code) show that OptiMer consistently outperforms data mixture and model averaging baselines with 15-35 times lower search cost. Key findings reveal that 1) the optimized weights can be interpreted as data mixture ratios, and retraining with these ratios improves data mixture CPT, and 2) the same vector pool can be re-optimized for a given objective without any retraining, producing target-tailored models on demand. Our work establishes that data mixture ratio selection, traditionally a pre-training decision, can be reformulated as a post-hoc optimization over distribution vectors, offering a more flexible paradigm for continual pre-training.

Comment: Moves data-mixture selection after training by varying composition weights over a fixed pool of dataset-specific parameter shifts.

Topic Match: It changes the continual-pretraining recipe by separating model training from mixture selection, conditional on constructing one CPT model per dataset.

Relevance: 9 Novelty: 8


3. The Spectral Edge Thesis: A Mathematical Framework for Intra-Signal Phase Transitions in Neural Network Training

ArXiv ID: 2603.28964

Primary Topic: Architecture and Training Dynamics

Authors: Yongzhong Xu

Abstract: We develop the spectral edge thesis: phase transitions in neural network training -- grokking, capability gains, loss plateaus -- are controlled by the spectral gap of the rolling-window Gram matrix of parameter updates. In the extreme aspect ratio regime (parameters $P \sim 10^8$, window $W \sim 10$), the classical BBP detection threshold is vacuous; the operative structure is the intra-signal gap separating dominant from subdominant modes at position $k^ = \mathrm{argmax}\, \sigma_j/\sigma_{j+1}$. From three axioms we derive: (i) gap dynamics governed by a Dyson-type ODE with curvature asymmetry, damping, and gradient driving; (ii) a spectral loss decomposition linking each mode's learning contribution to its Davis--Kahan stability coefficient; (iii) the Gap Maximality Principle, showing that $k^$ is the unique dynamically privileged position -- its collapse is the only one that disrupts learning, and it sustains itself through an $\alpha$-feedback loop requiring no assumption on the optimizer. The adiabatic parameter $\mathcal{A} = |\Delta G|_F / (\eta\, g^2)$ controls circuit stability: $\mathcal{A} \ll 1$ (plateau), $\mathcal{A} \sim 1$ (phase transition), $\mathcal{A} \gg 1$ (forgetting). Tested across six model families (150K--124M parameters): gap dynamics precede every grokking event (24/24 with weight decay, 0/24 without), the gap position is optimizer-dependent (Muon: $k^=1$, AdamW: $k^=2$ on the same model), and 19/20 quantitative predictions are confirmed. The framework is consistent with the edge of stability, Tensor Programs, Dyson Brownian motion, the Lottery Ticket Hypothesis, and neural scaling laws.

Comment: Predicts training transitions from update-spectrum gaps, including a falsifiable optimizer-dependent gap position.

Topic Match: The core is a mechanism-level theory of optimization and circuit stability; experiments reach 124M parameters, but language-model datasets are unspecified.

Relevance: 8 Novelty: 9


4. From Density Matrices to Phase Transitions in Deep Learning: Spectral Early Warnings and Interpretability

ArXiv ID: 2603.29805

Primary Topic: Architecture and Training Dynamics

Authors: Max Hennick, Guillaume Corlouer

Abstract: A key problem in the modern study of AI is predicting and understanding emergent capabilities in models during training. Inspired by methods for studying reactions in quantum chemistry, we present the ``2-datapoint reduced density matrix". We show that this object provides a computationally efficient, unified observable of phase transitions during training. By tracking the eigenvalue statistics of the 2RDM over a sliding window, we derive two complementary signals: the spectral heat capacity, which we prove provides early warning of second-order phase transitions via critical slowing down, and the participation ratio, which reveals the dimensionality of the underlying reorganization. Remarkably, the top eigenvectors of the 2RDM are directly interpretable making it straightforward to study the nature of the transitions. We validate across four settings distinct settings: deep linear networks, induction head formation, grokking, and emergent misalignment. We then discuss directions for future work using the 2RDM.

Comment: Derives spectral early-warning signals for training transitions and measures the dimension of representation reorganization.

Topic Match: The central contribution concerns training dynamics and representation transitions, including induction-head formation, rather than merely cataloguing interpretable features.

Relevance: 8 Novelty: 7


5. Metriplector: From Field Theory to Neural Architecture

ArXiv ID: 2603.29496

Primary Topic: Architecture and Training Dynamics

Authors: Dan Oprisa, Peter Toth

Abstract: We present Metriplector, a neural architecture primitive in which the input configures an abstract physical system--fields, sources, and operators--and the dynamics of that system is the computation. Multiple fields evolve via coupled metriplectic dynamics, and the stress-energy tensor T^{{\mu}{\nu}}, derived from Noether's theorem, provides the readout. The metriplectic formulation admits a natural spectrum of instantiations: the dissipative branch alone yields a screened Poisson equation solved exactly via conjugate gradient; activating the full structure--including the antisymmetric Poisson bracket--gives field dynamics for image recognition and language modeling. We evaluate Metriplector across four domains, each using a task-specific architecture built from this shared primitive with progressively richer physics: F1=1.0 on maze pathfinding, generalizing from 15x15 training grids to unseen 39x39 grids; 97.2% exact Sudoku solve rate with zero structural injection; 81.03% on CIFAR-100 with 2.26M parameters; and 1.182 bits/byte on language modeling with 3.6x fewer training tokens than a GPT baseline.

Comment: Uses coupled physical-field dynamics and a stress-energy readout as a shared neural computation primitive.

Topic Match: A new computational primitive is the core contribution, including a language-model experiment whose model scale and dataset are unspecified.

Relevance: 8 Novelty: 8


6. A Comprehensive Information-Decomposition Analysis of Large Vision-Language Models

ArXiv ID: 2603.29676

Primary Topic: Architecture and Training Dynamics

Authors: Lixin Xiu, Xufang Luo, Hideki Nakayama

Abstract: Large vision-language models (LVLMs) achieve impressive performance, yet their internal decision-making processes remain opaque, making it difficult to determine if the success stems from true multimodal fusion or from reliance on unimodal priors. To address this attribution gap, we introduce a novel framework using partial information decomposition (PID) to quantitatively measure the "information spectrum" of LVLMs -- decomposing a model's decision-relevant information into redundant, unique, and synergistic components. By adapting a scalable estimator to modern LVLM outputs, our model-agnostic pipeline profiles 26 LVLMs on four datasets across three dimensions -- breadth (cross-model & cross-task), depth (layer-wise information dynamics), and time (learning dynamics across training). Our analysis reveals two key results: (i) two task regimes (synergy-driven vs. knowledge-driven) and (ii) two stable, contrasting family-level strategies (fusion-centric vs. language-centric). We also uncover a consistent three-phase pattern in layer-wise processing and identify visual instruction tuning as the key stage where fusion is learned. Together, these contributions provide a quantitative lens beyond accuracy-only evaluation and offer insights for analyzing and designing the next generation of LVLMs. Code and data are available at https://github.com/RiiShin/pid-lvlm-analysis .

Comment: Tracks how cross-modal fusion develops across layers and training stages, locating its emergence in visual instruction tuning.

Topic Match: Analysis of layer-wise computation and training-stage reorganization connects directly to architecture_training, though the findings remain primarily descriptive.

Relevance: 7 Novelty: 6


7. ASI-Evolve: AI Accelerates AI

ArXiv ID: 2603.29640

Primary Topic: Architecture and Training Dynamics

Authors: Weixian Xu, Tiantian Mi, Yixiu Liu, Yang Nan, Zhimeng Zhou, Lyumanshan Ye, Lin Zhang, Yu Qiao, Pengfei Liu

Abstract: Can AI accelerate the development of AI itself? While recent agentic systems have shown strong performance on well-scoped tasks with rapid feedback, it remains unclear whether they can tackle the costly, long-horizon, and weakly supervised research loops that drive real AI progress. We present ASI-Evolve, an agentic framework for AI-for-AI research that closes this loop through a learn-design-experiment-analyze cycle. ASI-Evolve augments standard evolutionary agents with two key components: a cognition base that injects accumulated human priors into each round of exploration, and a dedicated analyzer that distills complex experimental outcomes into reusable insights for future iterations. To our knowledge, ASI-Evolve is the first unified framework to demonstrate AI-driven discovery across three central components of AI development: data, architectures, and learning algorithms. In neural architecture design, it discovered 105 SOTA linear attention architectures, with the best discovered model surpassing DeltaNet by +0.97 points, nearly 3x the gain of recent human-designed improvements. In pretraining data curation, the evolved pipeline improves average benchmark performance by +3.96 points, with gains exceeding 18 points on MMLU. In reinforcement learning algorithm design, discovered algorithms outperform GRPO by up to +12.5 points on AMC32, +11.67 points on AIME24, and +5.04 points on OlympiadBench. We further provide initial evidence that this AI-for-AI paradigm can transfer beyond the AI stack through experiments in mathematics and biomedicine. Together, these results suggest that ASI-Evolve represents a promising step toward enabling AI to accelerate AI across the foundational stages of development, offering early evidence for the feasibility of closed-loop AI research.

Comment: Automates experimental search over linear-attention architectures, reporting an improvement over DeltaNet.

Topic Match: Architecture search provides the connection, but the core contribution is a research-agent framework rather than an architectural mechanism.

Relevance: 5 Novelty: 6


8. Think Anywhere in Code Generation

ArXiv ID: 2603.29957

Primary Topic: Architecture and Training Dynamics

Authors: Xue Jiang, Tianyu Zhang, Ge Li, Mengyang Liu, Taozhi Chen, Zhenhua Xu, Binhua Li, Wenpin Jiao, Zhi Jin, Yongbin Li, Yihong Dong

Abstract: Recent advances in reasoning Large Language Models (LLMs) have primarily relied on upfront thinking, where reasoning occurs before final answer. However, this approach suffers from critical limitations in code generation, where upfront thinking is often insufficient as problems' full complexity only reveals itself during code implementation. Moreover, it cannot adaptively allocate reasoning effort throughout the code generation process where difficulty varies significantly. In this paper, we propose Think-Anywhere, a novel reasoning mechanism that enables LLMs to invoke thinking on-demand at any token position during code generation. We achieve Think-Anywhere by first teaching LLMs to imitate the reasoning patterns through cold-start training, then leveraging outcome-based RL rewards to drive the model's autonomous exploration of when and where to invoke reasoning. Extensive experiments on four mainstream code generation benchmarks (i.e., LeetCode, LiveCodeBench, HumanEval, and MBPP) show that Think-Anywhere achieves state-of-the-art performance over both existing reasoning methods and recent post-training approaches, while demonstrating consistent generalization across diverse LLMs. Our analysis further reveals that Think-Anywhere enables the model to adaptively invoke reasoning at high-entropy positions, providing enhanced interpretability.

Comment: Learns when to insert reasoning segments at arbitrary positions during code generation.

Topic Match: Dynamic computation is the nearest category, but the contribution is post-training for code reasoning.

Relevance: 3 Novelty: 6


9. On the Mirage of Long-Range Dependency, with an Application to Integer Multiplication

ArXiv ID: 2603.29069

Primary Topic: Architecture and Training Dynamics

Authors: Zichao Wei

Abstract: Integer multiplication has long been considered a hard problem for neural networks, with the difficulty widely attributed to the O(n) long-range dependency induced by carry chains. We argue that this diagnosis is wrong: long-range dependency is not an intrinsic property of multiplication, but a mirage produced by the choice of computational spacetime. We formalize the notion of mirage and provide a constructive proof: when two n-bit binary integers are laid out as a 2D outer-product grid, every step of long multiplication collapses into a $3 \times 3$ local neighborhood operation. Under this representation, a neural cellular automaton with only 321 learnable parameters achieves perfect length generalization up to $683\times$ the training range. Five alternative architectures -- including Transformer (6,625 params), Transformer+RoPE, and Mamba -- all fail under the same representation. We further analyze how partial successes locked the community into an incorrect diagnosis, and argue that any task diagnosed as requiring long-range dependency should first be examined for whether the dependency is intrinsic to the task or induced by the computational spacetime.

Comment: Constructively turns multiplication's apparent long-range dependency into local operations by changing its computational layout.

Topic Match: Computational layout connects to architecture_training, but the demonstrated scope is binary multiplication with small networks rather than language-model training.

Relevance: 4 Novelty: 9


10. Grokking From Abstraction to Intelligence

ArXiv ID: 2603.29262

Primary Topic: Architecture and Training Dynamics

Authors: Junjie Zhang, Zhen Shen, Gang Xiong, Xisong Dong

Abstract: Grokking in modular arithmetic has established itself as the quintessential fruit fly experiment, serving as a critical domain for investigating the mechanistic origins of model generalization. Despite its significance, existing research remains narrowly focused on specific local circuits or optimization tuning, largely overlooking the global structural evolution that fundamentally drives this phenomenon. We propose that grokking originates from a spontaneous simplification of internal model structures governed by the principle of parsimony. We integrate causal, spectral, and algorithmic complexity measures alongside Singular Learning Theory to reveal that the transition from memorization to generalization corresponds to the physical collapse of redundant manifolds and deep information compression, offering a novel perspective for understanding the mechanisms of model overfitting and generalization.

Comment: Attributes grokking to global structural simplification and compression of redundant representations.

Topic Match: Structural changes during optimization are closest to architecture_training; the abstract supplies modular-arithmetic framing without an identified language-model setting.

Relevance: 5 Novelty: 6


11. Is the Modality Gap a Bug or a Feature? A Robustness Perspective

ArXiv ID: 2603.29080

Primary Topic: Architecture and Training Dynamics

Authors: Rhea Chowers, Oshri Naparstek, Udi Barzelay, Yair Weiss

Abstract: Many modern multi-modal models (e.g. CLIP) seek an embedding space in which the two modalities are aligned. Somewhat surprisingly, almost all existing models show a strong modality gap: the distribution of images is well-separated from the distribution of texts in the shared embedding space. Despite a series of recent papers on this topic, it is still not clear why this gap exists nor whether closing the gap in post-processing will lead to better performance on downstream tasks. In this paper we show that under certain conditions, minimizing the contrastive loss yields a representation in which the two modalities are separated by a global gap vector that is orthogonal to their embeddings. We also show that under these conditions the modality gap is monotonically related to robustness: decreasing the gap does not change the clean accuracy of the models but makes it less likely that a model will change its output when the embeddings are perturbed. Our experiments show that for many real-world VLMs we can significantly increase robustness by a simple post-processing step that moves one modality towards the mean of the other modality, without any loss of clean accuracy.

Comment: Relates modality separation to perturbation robustness while holding clean accuracy fixed under stated conditions.

Topic Match: Contrastive-training geometry is closest to architecture_training; the subject is multimodal embedding alignment rather than language-model training.

Relevance: 3 Novelty: 8


12. DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA

ArXiv ID: 2603.29844

Primary Topic: Architecture and Training Dynamics

Authors: Yi Chen, Yuying Ge, Hui Zhou, Mingyu Ding, Yixiao Ge, Xihui Liu

Abstract: The development of Vision-Language-Action (VLA) models has been significantly accelerated by pre-trained Vision-Language Models (VLMs). However, most existing end-to-end VLAs treat the VLM primarily as a multimodal encoder, directly mapping vision-language features to low-level actions. This paradigm underutilizes the VLM's potential in high-level decision making and introduces training instability, frequently degrading its rich semantic representations. To address these limitations, we introduce DIAL, a framework bridging high-level decision making and low-level motor execution through a differentiable latent intent bottleneck. Specifically, a VLM-based System-2 performs latent world modeling by synthesizing latent visual foresight within the VLM's native feature space; this foresight explicitly encodes intent and serves as the structural bottleneck. A lightweight System-1 policy then decodes this predicted intent together with the current observation into precise robot actions via latent inverse dynamics. To ensure optimization stability, we employ a two-stage training paradigm: a decoupled warmup phase where System-2 learns to predict latent futures while System-1 learns motor control under ground-truth future guidance within a unified feature space, followed by seamless end-to-end joint optimization. This enables action-aware gradients to refine the VLM backbone in a controlled manner, preserving pre-trained knowledge. Extensive experiments on the RoboCasa GR1 Tabletop benchmark show that DIAL establishes a new state-of-the-art, achieving superior performance with 10x fewer demonstrations than prior methods. Furthermore, by leveraging heterogeneous human demonstrations, DIAL learns physically grounded manipulation priors and exhibits robust zero-shot generalization to unseen objects and novel configurations during real-world deployment on a humanoid robot.

Comment: A differentiable intent bottleneck separates high-level prediction from motor control during joint training.

Topic Match: The bottleneck and staged optimization are architectural choices, but the core contribution and evidence concern robot-control training.

Relevance: 3 Novelty: 7


13. APEX-EM: Non-Parametric Online Learning for Autonomous Agents via Structured Procedural-Episodic Experience Replay

ArXiv ID: 2603.29093

Primary Topic: Architecture and Training Dynamics

Authors: Pratyay Banerjee, Masud Moshtaghi, Ankit Chadha

Abstract: LLM-based autonomous agents lack persistent procedural memory: they re-derive solutions from scratch even when structurally identical tasks have been solved before. We present \textbf{APEX-EM}, a non-parametric online learning framework that accumulates, retrieves, and reuses structured procedural plans without modifying model weights. APEX-EM introduces: (1) a \emph{structured experience representation} encoding the full procedural-episodic trace of each execution -- planning steps, artifacts, iteration history with error analysis, and quality scores; (2) a \emph{Plan-Retrieve-Generate-Iterate-Ingest} (PRGII) workflow with Task Verifiers providing multi-dimensional reward signals; and (3) a \emph{dual-outcome Experience Memory} with hybrid retrieval combining semantic search, structural signature matching, and plan DAG traversal -- enabling cross-domain transfer between tasks sharing no lexical overlap but analogous operational structure. Successful experiences serve as positive in-context examples; failures as negative examples with structured error annotations. We evaluate on BigCodeBench~\cite{zhuo2025bigcodebench}, KGQAGen-10k~\cite{zhang2025kgqagen}, and Humanity's Last Exam~\cite{phan2025hle} using Claude Sonnet 4.5 and Opus 4.5. On KGQAGen-10k, APEX-EM achieves 89.6\% accuracy versus 41.3\% without memory (+48.3pp), surpassing the oracle-retrieval upper bound (84.9\%). On BigCodeBench, it reaches 83.3\% SR from a 53.9\% baseline (+29.4pp), exceeding MemRL's~\cite{memrl2025} +11.0pp gain under comparable frozen-backbone conditions (noting backbone differences controlled for in our analysis). On HLE, entity graph retrieval reaches 48.0\% from 25.2\% (+22.8pp). Ablations show component value is task-dependent: rich judge feedback is negligible for code generation but critical for structured queries (+10.3pp), while binary-signal iteration partially compensates for weaker feedback.

Comment: Reuses structured execution plans and failure traces through procedural experience replay.

Topic Match: Structured computation is the nearest architectural connection; learning occurs in external agent memory while model weights remain fixed.

Relevance: 2 Novelty: 5


14. LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning

ArXiv ID: 2603.29165

Primary Topic: Architecture and Training Dynamics

Authors: Haihong Hao, Lei Chen, Mingfei Han, Changlin Li, Dong An, Yuqiang Yang, Zhihui Li, Xiaojun Chang

Abstract: Existing vision-and-language navigation (VLN) models primarily reason over past and current visual observations, while largely ignoring the future visual dynamics induced by actions. As a result, they often lack an effective understanding of the causal relationship between actions and how the visual world changes, limiting robust decision-making. Humans, in contrast, can imagine the near future by leveraging action-dynamics causality, which improves both environmental understanding and navigation choices. Inspired by this capability, we propose LatentPilot, a new paradigm that exploits future observations during training as a valuable data source to learn action-conditioned visual dynamics, while requiring no access to future frames at inference. Concretely, we propose a flywheel-style training mechanism that iteratively collects on-policy trajectories and retrains the model to better match the agent's behavior distribution, with an expert takeover triggered when the agent deviates excessively. LatentPilot further learns visual latent tokens without explicit supervision; these latent tokens attend globally in a continuous latent space and are carried across steps, serving as both the current output and the next input, thereby enabling the agent to dream ahead and reason about how actions will affect subsequent observations. Experiments on R2R-CE, RxR-CE, and R2R-PE benchmarks achieve new SOTA results, and real-robot tests across diverse environments demonstrate LatentPilot's superior understanding of environment-action dynamics in scene. Project page:https://abdd.top/latentpilot/

Comment: Learns action-conditioned latent foresight tokens that carry predicted visual dynamics between navigation steps.

Topic Match: The recurrent latent computation is the nearest architectural connection; its core purpose is embodied navigation.

Relevance: 2 Novelty: 6


15. Tracking Equivalent Mechanistic Interpretations Across Neural Networks

ArXiv ID: 2603.30002

Primary Topic: Architecture and Training Dynamics

Authors: Alan Sun, Mariya Toneva

Abstract: Mechanistic interpretability (MI) is an emerging framework for interpreting neural networks. Given a task and model, MI aims to discover a succinct algorithmic process, an interpretation, that explains the model's decision process on that task. However, MI is difficult to scale and generalize. This stems in part from two key challenges: there is no precise notion of a valid interpretation; and, generating interpretations is often an ad hoc process. In this paper, we address these challenges by defining and studying the problem of interpretive equivalence: determining whether two different models share a common interpretation, without requiring an explicit description of what that interpretation is. At the core of our approach, we propose and formalize the principle that two interpretations of a model are equivalent if all of their possible implementations are also equivalent. We develop an algorithm to estimate interpretive equivalence and case study its use on Transformer-based models. To analyze our algorithm, we introduce necessary and sufficient conditions for interpretive equivalence based on models' representation similarity. We provide guarantees that simultaneously relate a model's algorithmic interpretations, circuits, and representations. Our framework lays a foundation for the development of more rigorous evaluation methods of MI and automated, generalizable interpretation discovery methods.

Comment: Formalizes when two networks admit a common algorithmic interpretation using representation-based equivalence conditions.

Topic Match: Circuits and representations are the nearest architectural connection; the core result is interpretability methodology rather than an explanation of a specific learned mechanism.

Relevance: 3 Novelty: 7


16. OccSim: Multi-kilometer Simulation with Long-horizon Occupancy World Models

ArXiv ID: 2603.28887

Primary Topic: Architecture and Training Dynamics

Authors: Tianran Liu, Shengwen Zhao, Mozhgan Pourkeshavarz, Weican Li, Nicholas Rhinehart

Abstract: Data-driven autonomous driving simulation has long been constrained by its heavy reliance on pre-recorded driving logs or spatial priors, such as HD maps. This fundamental dependency severely limits scalability, restricting open-ended generation capabilities to the finite scale of existing collected datasets. To break this bottleneck, we present OccSim, the first occupancy world model-driven 3D simulator. OccSim obviates the requirement for continuous logs or HD maps; conditioned only on a single initial frame and a sequence of future ego-actions, it can stably generate over 3,000 continuous frames, enabling the continuous construction of large-scale 3D occupancy maps spanning over 4 kilometers for simulation. This represents an >80x improvement in stable generation length over previous state-of-the-art occupancy world models. OccSim is powered by two modules: W-DiT based static occupancy world model and the Layout Generator. W-DiT handles the ultra-long-horizon generation of static environments by explicitly introducing known rigid transformations in architecture design, while the Layout Generator populates the dynamic foreground with reactive agents based on the synthesized road topology. With these designs, OccSim can synthesize massive, diverse simulation streams. Extensive experiments demonstrate its downstream utility: data collected directly from OccSim can pre-train 4D semantic occupancy forecasting models to achieve up to 67% zero-shot performance on unseen data, outperforming previous asset-based simulator by 11%. When scaling the OccSim dataset to 5x the size, the zero-shot performance increases to about 74%, while the improvement over asset-based simulators expands to 22.1%.

Comment: Uses action-conditioned occupancy dynamics to generate driving scenes without continuous logs or HD maps.

Topic Match: Its generative architecture is the nearest registry connection, but the trained system is an autonomous-driving world model.

Relevance: 1 Novelty: 7


17. Concept frustration: Aligning human concepts and machine representations

ArXiv ID: 2603.29654

Primary Topic: Architecture and Training Dynamics

Authors: Enrico Parisini, Christopher J. Soelistyo, Ahab Isaac, Alessandro Barp, Christopher R. S. Banerji

Abstract: Aligning human-interpretable concepts with the internal representations learned by modern machine learning systems remains a central challenge for interpretable AI. We introduce a geometric framework for comparing supervised human concepts with unsupervised intermediate representations extracted from foundation model embeddings. Motivated by the role of conceptual leaps in scientific discovery, we formalise the notion of concept frustration: a contradiction that arises when an unobserved concept induces relationships between known concepts that cannot be made consistent within an existing ontology. We develop task-aligned similarity measures that detect concept frustration between supervised concept-based models and unsupervised representations derived from foundation models, and show that the phenomenon is detectable in task-aligned geometry while conventional Euclidean comparisons fail. Under a linear-Gaussian generative model we derive a closed-form expression for Bayes-optimal concept-based classifier accuracy, decomposing predictive signal into known-known, known-unknown and unknown-unknown contributions and identifying analytically where frustration affects performance. Experiments on synthetic data and real language and vision tasks demonstrate that frustration can be detected in foundation model representations and that incorporating a frustrating concept into an interpretable model reorganises the geometry of learned concept representations, to better align human and machine reasoning. These results suggest a principled framework for diagnosing incomplete concept ontologies and aligning human and machine conceptual reasoning, with implications for the development and validation of safe interpretable AI for high-risk applications.

Comment: Defines task-aligned geometric tests for contradictions caused by missing concepts in an ontology.

Topic Match: Representation geometry is the nearest architectural connection; the core contribution concerns concept-ontology interpretability.

Relevance: 3 Novelty: 7


18. Lie Generator Networks for Nonlinear Partial Differential Equations

ArXiv ID: 2603.29264

Primary Topic: Architecture and Training Dynamics

Authors: Shafayeth Jamil, Rehan Kapadia

Abstract: Linear dynamical systems are fully characterized by their eigenspectra, accessible directly from the generator of the dynamics. For nonlinear systems governed by partial differential equations, no equivalent theory exists. We introduce Lie Generator Network--Koopman (LGN-KM), a neural operator that lifts nonlinear dynamics into a linear latent space and learns the continuous-time Koopman generator ($L_k$) through a decomposition $L_k = S - D_k$, where $S$ is skew-symmetric representing conservative inter-modal coupling, and $D_k$ is a positive-definite diagonal encoding modal dissipation. This architectural decomposition enforces stability and enables interpretability through direct spectral access to the learned dynamics. On two-dimensional Navier--Stokes turbulence, the generator recovers the known dissipation scaling and a complete multi-branch dispersion relation from trajectory data alone with no physics supervision. Independently trained models at different flow regimes recover matched gauge-invariant spectral structure, exposing a gauge freedom in the Koopman lifting. Because the generator is provably stable, it enables guaranteed long-horizon stability, continuous-time evaluation at arbitrary time, and physics-informed cross-viscosity model transfer.

Comment: Decomposes a learned Koopman generator into conservative and dissipative parts to guarantee stable dynamics.

Topic Match: The stability-enforcing parameterization is architectural, but it trains a neural operator for fluid dynamics.

Relevance: 2 Novelty: 7


19. A Pontryagin Method of Model-based Reinforcement Learning via Hamiltonian Actor-Critic

ArXiv ID: 2603.28971

Primary Topic: Architecture and Training Dynamics

Authors: Chengyang Gu, Yuxin Pan, Hui Xiong, Yize Chen

Abstract: Model-based reinforcement learning (MBRL) improves sample efficiency by leveraging learned dynamics models for policy optimization. However, the effectiveness of methods such as actor-critic is often limited by compounding model errors, which degrade long-horizon value estimation. Existing approaches, such as Model-Based Value Expansion (MVE), partially mitigate this issue through multi-step rollouts, but remain sensitive to rollout horizon selection and residual model bias. Motivated by the Pontryagin Maximum Principle (PMP), we propose Hamiltonian Actor-Critic (HAC), a model-based approach that eliminates explicit value function learning by directly optimizing a Hamiltonian defined over the learned dynamics and reward for deterministic systems. By avoiding value approximation, HAC reduces sensitivity to model errors while admitting convergence guarantees. Extensive experiments on continuous control benchmarks, in both online and offline RL settings, demonstrate that HAC outperforms model-free and MVE-based baselines in control performance, convergence speed, and robustness to distributional shift, including out-of-distribution (OOD) scenarios. In offline settings with limited data, HAC matches or exceeds state-of-the-art methods, highlighting its strong sample efficiency.

Comment: Optimizes a Hamiltonian over learned dynamics and rewards to eliminate explicit value-function fitting.

Topic Match: Replacing an optimization component is the nearest architectural connection; the method trains control policies for deterministic systems.

Relevance: 2 Novelty: 8


20. Minimum Norm Interpolation via The Local Theory of Banach Spaces: The Role of $2$-Uniform Convexity

ArXiv ID: 2603.28956

Primary Topic: Architecture and Training Dynamics

Authors: Gil Kur, Pierre Bizeul

Abstract: The minimum-norm interpolator (MNI) framework has recently attracted considerable attention as a tool for understanding generalization in overparameterized models, such as neural networks. In this work, we study the MNI under a $2$-uniform convexity assumption, which is weaker than requiring the norm to be induced by an inner product, and it typically does not admit a closed-form solution. At a high level, we show that this condition yields an upper bound on the MNI bias in both linear and nonlinear models. We further show that this bound is sharp for overparameterized linear regression when the unit ball of the norm is in isotropic (or John's) position, and the covariates are isotropic, symmetric, i.i.d. sub-Gaussian, such as vectors with i.i.d. Bernoulli entries. Finally, under the same assumption on the covariates, we prove sharp generalization bounds for the $\ell_p$-MNI when $p \in \bigl(1 + C/\log d, 2\bigr]$. To the best of our knowledge, this is the first work to establish sharp bounds for non-Gaussian covariates in linear models when the norm is not induced by an inner product. This work is deeply inspired by classical works on $K$-convexity, and more modern work on the geometry of 2-uniform and isotropic convex bodies.

Comment: Extends sharp minimum-norm generalization bounds beyond inner-product norms and Gaussian covariates.

Topic Match: Optimization geometry is the nearest connection to architecture_training; the contribution is interpolation theory without a language-model training result.

Relevance: 2 Novelty: 6


21. GENIE: Gram-Eigenmode INR Editing with Closed-Form Geometry Updates

ArXiv ID: 2603.29860

Primary Topic: Architecture and Training Dynamics

Authors: Samundra Karki, Adarsh Krishnamurthy, Baskar Ganapathysubramanian

Abstract: Implicit Neural Representations (INRs) provide compact models of geometry, but it is unclear when their learned shapes can be edited without retraining. We show that the Gram operator induced by the INR's penultimate features admits deformation eigenmodes that parameterize a family of realizable edits of the SDF zero level set. A key finding is that these modes are not intrinsic to the geometry alone: they are reliably recoverable only when the Gram operator is estimated from sufficiently rich sampling distributions. We derive a single closed-form update that performs geometric edits to the INR without optimization by leveraging the deformation modes. We characterize theoretically the precise set of deformations that are feasible under this one-shot update, and show that editing is well-posed exactly within the span of these deformation modes.

Comment: Derives closed-form neural-representation edits from Gram-operator deformation modes.

Topic Match: Architectural analysis is the nearest category, but the model represents geometry rather than language.

Relevance: 1 Novelty: 7


Training Algorithms That Change What Is Possible (1)

1. Byzantine-Robust and Communication-Efficient Distributed Training: Compressive and Cyclic Gradient Coding

ArXiv ID: 2603.28780

Primary Topic: Training Algorithms That Change What Is Possible

Authors: Chengxi Li, Youssef Allouah, Rachid Guerraoui, Mikael Skoglund, Ming Xiao

Abstract: In this paper, we study the problem of distributed training (DT) under Byzantine attacks with communication constraints. While prior work has developed various robust aggregation rules at the server to enhance robustness to Byzantine attacks, the existing methods suffer from a critical limitation in that the solution error does not diminish when the local gradients sent by different devices vary considerably, as a result of data heterogeneity among the subsets held by different devices. To overcome this limitation, we propose a novel DT method, cyclic gradient coding-based DT (LAD). In LAD, the server allocates the entire training dataset to the devices before training begins. In each iteration, it assigns computational tasks redundantly to the devices using cyclic gradient coding. Each honest device then computes local gradients on a fixed number of data subsets and encodes the local gradients before transmitting to the server. The server aggregates the coded vectors from the honest devices and the potentially incorrect messages from Byzantine devices using a robust aggregation rule. Leveraging the redundancy of computation across devices, the convergence performance of LAD is analytically characterized, demonstrating improved robustness against Byzantine attacks and significantly lower solution error. Furthermore, we extend LAD to a communication-efficient variant, compressive and cyclic gradient coding-based DT (Com-LAD), which further reduces communication overhead under constrained settings. Numerical results validate the effectiveness of the proposed methods in enhancing both Byzantine resilience and communication efficiency.

Comment: Redundant cyclic gradient coding changes robustness guarantees under heterogeneous data and Byzantine workers.

Topic Match: The contribution is a distributed optimization algorithm, making training_systems the closest label; no language-model training setting is established.

Relevance: 4 Novelty: 7


MoE Where It Changes the Design Space (1)

1. Training-Free Dynamic Upcycling of Expert Language Models

ArXiv ID: 2603.29765

Primary Topic: MoE Where It Changes the Design Space

Authors: Eros Fan`i, O\u{g}uzhan Ersoy

Abstract: Large Language Models (LLMs) have achieved remarkable performance on a wide range of specialized tasks, exhibiting strong problem-solving capabilities. However, training these models is prohibitively expensive, and they often lack domain-specific expertise because they rely on general knowledge datasets. Expertise finetuning can address this issue; however, it often leads to overspecialization, and developing a single multi-domain expert remains difficult due to diverging objectives. Furthermore, multitask training is challenging due to interference and catastrophic forgetting. Existing work proposes combining the expertise of dense models within a Mixture of Experts (MoE) architecture, although this approach still requires multitask finetuning. To address these issues, we introduce Dynamic Upcycling MoE (DUME), a novel approach that reuses dense experts trained on different domains to construct a unified MoE model. Our method builds a single multitask model that preserves the capabilities of the original dense experts without requiring additional training. DUME is both cost-efficient and scalable: by leveraging the closed-form solution of ridge regression, it eliminates the need for further optimization and enables experts to be added dynamically while maintaining the model's original performance. We demonstrate that DUME consistently outperforms baseline approaches in both causal language modeling and reasoning settings. Finally, we also show that the DUME model can be fine-tuned to further improve performance. We show that, in the causal language modeling setting, DUME can retain up to 97.6% of a dense expert model specialized in one particular domain, and that it can also surpass it in the reasoning setting, where it can achieve 102.1% of the dense expert performance. Our code is available at: github.com/gensyn-ai/dume.

Comment: Replaces additional multitask upcycling optimization with a closed-form ridge solution for combining existing dense experts.

Topic Match: Training-free dense-to-MoE conversion directly changes expert composition; model scale and the cost of growing the expert pool remain unspecified.

Relevance: 9 Novelty: 9


Efficiency, Compression, and Large-Scale Training (3)

1. Baby Scale: Investigating Models Trained on Individual Children's Language Input

ArXiv ID: 2603.29522

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Steven Y. Feng, Alvin W. M. Tan, Michael C. Frank

Abstract: Modern language models (LMs) must be trained on many orders of magnitude more words of training data than human children receive before they begin to produce useful behavior. Assessing the nature and origins of this "data gap" requires benchmarking LMs on human-scale datasets to understand how linguistic knowledge emerges from children's natural training data. Using transcripts from the BabyView dataset (videos from children ages 6-36 months), we investigate (1) scaling performance at child-scale data regimes, (2) variability in model performance across datasets from different children's experiences and linguistic predictors of dataset quality, and (3) relationships between model and child language learning outcomes. LMs trained on child data show acceptable scaling for grammar tasks, but lower scaling on semantic and world knowledge tasks than models trained on synthetic data; we also observe substantial variability on data from different children. Beyond dataset size, performance is most associated with a combination of distributional and interactional linguistic features, broadly consistent with what makes high-quality input for child language development. Finally, model likelihoods for individual words correlate with children's learning of those words, suggesting that properties of child-directed input may influence both model learning and human language development. Overall, understanding what properties make language data efficient for learning can enable more powerful small-scale language models while also shedding light on human language acquisition.

Comment: Attributes differences in learning efficiency to the distributional and interactional properties of training data.

Topic Match: Training-data attribution is the strongest fit, demonstrated in a narrow child-scale data regime.

Relevance: 7 Novelty: 5


2. Big2Small: A Unifying Neural Network Framework for Model Compression

ArXiv ID: 2603.29768

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Jing-Xiao Liao, Haoran Wang, Tao Li, Daoming Lyu, Yi Zhang, Chengjun Cai, Feng-Lei Fan

Abstract: With the development of foundational models, model compression has become a critical requirement. Various model compression approaches have been proposed such as low-rank decomposition, pruning, quantization, ergodic dynamic systems, and knowledge distillation, which are based on different heuristics. To elevate the field from fragmentation to a principled discipline, we construct a unifying mathematical framework for model compression grounded in measure theory. We further demonstrate that each model compression technique is mathematically equivalent to a neural network subject to a regularization. Building upon this mathematical and structural equivalence, we propose an experimentally-verified data-free model compression framework, termed \textit{Big2Small}, which translates Implicit Neural Representations (INRs) from data domain to the domain of network parameters. \textit{Big2Small} trains compact INRs to encode the weights of larger models and reconstruct the weights during inference. To enhance reconstruction fidelity, we introduce Outlier-Aware Preprocessing to handle extreme weight values and a Frequency-Aware Loss function to preserve high-frequency details. Experiments on image classification and segmentation demonstrate that \textit{Big2Small} achieves competitive accuracy and compression ratios compared to state-of-the-art baselines.

Comment: Encodes network weights in compact implicit neural representations and reconstructs them for inference.

Topic Match: Weight-space compression is closest to efficiency_scaling; the reported evaluations concern image classification and segmentation.

Relevance: 3 Novelty: 7


3. Nonnegative Matrix Factorization in the Component-Wise L1 Norm for Sparse Data

ArXiv ID: 2603.29715

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Giovanni Seraghiti, K\'evin Dubrulle, Arnaud Vandaele, Nicolas Gillis

Abstract: Nonnegative matrix factorization (NMF) approximates a nonnegative matrix, $X$, by the product of two nonnegative factors, $WH$, where $W$ has $r$ columns and $H$ has $r$ rows. In this paper, we consider NMF using the component-wise L1 norm as the error measure (L1-NMF), which is suited for data corrupted by heavy-tailed noise, such as Laplace noise or salt and pepper noise, or in the presence of outliers. Our first contribution is an NP-hardness proof for L1-NMF, even when $r=1$, in contrast to the standard NMF that uses least squares. Our second contribution is to show that L1-NMF strongly enforces sparsity in the factors for sparse input matrices, thereby favoring interpretability. However, if the data is affected by false zeros, too sparse solutions might degrade the model. Our third contribution is a new, more general, L1-NMF model for sparse data, dubbed weighted L1-NMF (wL1-NMF), where the sparsity of the factorization is controlled by adding a penalization parameter to the entries of $WH$ associated with zeros in the data. The fourth contribution is a new coordinate descent (CD) approach for wL1-NMF, denoted as sparse CD (sCD), where each subproblem is solved by a weighted median algorithm. To the best of our knowledge, sCD is the first algorithm for L1-NMF whose complexity scales with the number of nonzero entries in the data, making it efficient in handling large-scale, sparse data. We perform extensive numerical experiments on synthetic and real-world data to show the effectiveness of our new proposed model (wL1-NMF) and algorithm (sCD).

Comment: Makes L1 factorization updates scale with input nonzeros using weighted-median coordinate descent.

Topic Match: Sparse factorization is closest to efficiency_scaling, but the paper studies standalone matrix optimization rather than neural-model compression or training.

Relevance: 1 Novelty: 6


Paper Selection Prompt

System Prompt

You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.

User Prompt

Relevant Topics

This is a training-side feed for someone who builds and modifies model internals. The centre of gravity is what a frontier model is made of and how it is trained: architecture, training dynamics, and the decisions a lab actually made. MoE is still in scope, but only when the contribution changes the design space, not when it tunes an existing MoE.

Keep a paper when its CORE CONTRIBUTION falls in one of the five topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.

THE SUBJECT IS LANGUAGE-MODEL TRAINING. Every topic below is scoped to it. A technique named in one of them -- decentralised or asynchronous training, discrete diffusion, byte-level modelling, 1-bit weights, mixture-of-experts -- earns nothing when the model being trained is an image generator, a recommender, a scientific surrogate, or a vision classifier. Match on what is being trained, not on the vocabulary in the abstract.

THE ONE QUESTION THAT DECIDES A HIGH SCORE, asked before the topic list:

Does this paper REMOVE or REPLACE a component that everyone downstream inherits without thinking, or DECOUPLE two quantities the field currently conflates -- and does it say which end is held fixed while the other moves?

A paper that ADDS a mechanism on top of existing primitives is ordinary work, however good the numbers. A paper that takes one away, or splits one knob into two, is the reason this feed exists. Papers whose gains are CONDITIONAL (only above some batch size, only in one budget regime, only with an extra training stage) are ordinary work too, even when the mechanism is new: a conditional win is a new knob, not a removed one.

  1. Frontier Model Releases and Technical Reports (primary) - Keep: model and technical reports that DISCLOSE architecture or training decisions -- layer and attention design, normalization and residual choices, hybrid attention/state-space stacks, sparsity and expert layout, tokenizer and vocabulary decisions, data mixture and curriculum, optimizer and schedule, precision and numerics, stability fixes and what broke; open-weight releases whose report explains a choice rather than only reporting it. - Filter: releases that are a scorecard -- benchmark tables, a capability announcement, a product or API launch, or a report that names its recipe without saying why it was chosen. A frontier name in the title earns nothing on its own.

  2. Architecture and Training Dynamics (primary) - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, positional schemes, dynamic or modular computation, tokenizer-free and byte-level or learned-chunking stacks, non-autoregressive and discrete-diffusion generation); optimizers, preconditioners, and parameterisations, especially hyperparameter transfer across scale; training-dynamics and stability analysis that explains why large models train the way they do; work that removes a standard component (normalization, weight decay, positional encoding, the tokenizer, left-to-right decoding) and shows the model still trains. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight; an architectural tweak evaluated only at toy scale with no account of what it costs.

  3. Large-Scale Training Systems - Keep: distributed training ALGORITHMS that change what is possible, not only what is fast -- asynchronous, low-communication and decentralised optimisation, sharding and parallelism schemes with a new invariant, numerics and low-precision training (FP8, FP4, ternary and 1-bit) when the claim is about training stability rather than a kernel; scaling-law work that informs how a run is configured, including data-constrained and repeated-data regimes and critical batch size. - Filter: kernel engineering, communication scheduling, overlap, caching, serving and inference throughput. This is a real exclusion, not a soft one: work whose contribution is "the same model, faster" belongs to someone else's feed even when it is excellent and even when it is about MoE.

  4. MoE Where It Changes the Design Space - Keep: routing that is newly differentiable or newly stable; the training/inference router gap; expert collapse and what actually prevents it; expert granularity, shared experts and layout when a new axis is opened rather than tuned; dense-to-MoE conversion; MoE scaling laws that fix one quantity and vary another; analyses that carry a criterion which could have failed (a random-guess baseline, an ablation with a predicted sign) rather than similarity heatmaps. - Filter: MoE SYSTEMS and inference work (expert parallelism, all-to-all schedules, grouped GEMM, offload, prefetching, capacity and scheduling tricks, compression for serving); MoE surveys; papers that train on top of a MoE without a routing, balancing, stability, or structural contribution; "mixture of experts" in the classical ensemble or recommender sense. - The exclusion above is about MAKING MoE RUN FASTER, and it does not reach routing and expert design. A paper on how tokens are assigned to experts, how experts specialise, how balance is enforced, or how a dense model becomes sparse stays fully in scope and scores on its merits, even when the mechanism is a small one.

  5. Efficiency and Compression When the Mechanism Is New - Keep: quantization, sparsity, pruning and low-rank work whose mechanism is new and whose gain is unconditional; compression that changes what can be trained, not only what can be served; attribution of a model's behaviour to its data. - Filter: tuned variants of standard efficiency methods; deployment and serving engineering; anything whose headline is a compression ratio with no account of what was given up.

Also keep, even though they read like evaluation work, because the field has no one checking them: - benchmark validity itself: contamination, saturation, LLM-as-judge reliability, whether a benchmark measures what it claims. A paper AUDITING a benchmark is in scope; a paper PROPOSING one is not.

Subjects He Has Never Once Kept

Measured, not guessed. Over nine months where he was actively curating, 4953 papers passed this filter and 274 made his list -- a keep rate of 5.2%. Each subject below appeared in the rejected pile the number of times shown and appeared ZERO times in his entire 508-entry list.

Filter a paper whose CORE SUBJECT is one of these. The count is the evidence; where a paper only mentions the term in passing while contributing somewhere in the five topics above, keep it.

variational methods and inference (52) unlearning (41) long-context methods (40) neural operators (35) time series (35) vision transformers (35) parameter-efficient fine-tuning (28) stochastic gradient analysis (21) LLM safety (20) LLM reasoning as a subject (20) hallucination (20) sequence modelling as a subject (19) differential equations (19) spiking neural networks (17) post-training quantization (16) mechanistic interpretability as a framing (15) test-time scaling (13)

Two more he kept exactly once each, so suppress rather than drop: graph neural networks (1 of 42), knowledge distillation (1 of 39).

Do NOT extend this list by analogy. Four subjects that look like they belong here were checked and do not: diffusion models (8.1% kept), prompting (10.0%), benchmarks (9.1%) and neural architecture search (13.3%) all sit ABOVE his 5.2% base rate.

Note the one distinction that matters: he rejects papers FRAMED as interpretability, while analysis of how a trained model's computation is organised is among his highest-rated work. The subject is the framing, not the act of analysing.

Leave these to the hotspot digest unless the core contribution clearly falls in one of the five topics above: - agent, tooling, and RAG launches, and agent framework papers - new benchmarks, leaderboards, and evaluation-only papers - interpretability that stops at cataloguing features, with nothing said about the mechanism that produced them. Analysis of how a trained model's computation is organised -- what drives expert selection, what a circuit computes, how representations reorganise across training -- is IN scope, not here. - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning), EXCEPT where the subject is how post-training destabilises the architecture itself - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains

Scoring Criteria

Score five independent axes from 1 to 10: Relevance, Novelty, Evidence, Load, Proximity.

They answer different questions and must not be collapsed. Relevance: is this the kind of work this feed is for. Evidence: did the paper earn its claim. Load: what does it cost to get the usable result out. Proximity: how close is it to what the reader is working on right now (see Active Lines below).

Novelty is scored for continuity with the archive and for the hotspot spotlight cutoffs, and it does NOT order the feed. Measured against 274 hand-tiered papers it carried no signal at all: papers scoring 8 or more were his must-reads exactly as often as papers scoring 5 or less, both at the pool's base rate. Every abstract claims novelty, so the axis measures the claiming rather than the work. Score it honestly and do not let it influence the other four.

A strong claim with thin evidence is a worse read than a modest claim that holds. Score each axis on its own; do not let a high one pull up a low one.

Relevance Scoring

  • 9-10: directly centered on the target topics; highest when the core contribution is clearly within them and the paper is about how a model is BUILT or TRAINED.
  • 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
  • 5-6: touches the target topics, but the main contribution is elsewhere.
  • 3-4: largely outside the target topics, often application-focused or domain-specific.
  • 1-2: unrelated.

Important: Daily buzz, a frontier lab's name, or a large headline number is not enough for a high Relevance score. A model release scores high only when the report explains an architecture or training decision; a release that only reports results is hotspot material, not this feed.

Work whose contribution is "the same model, faster" -- kernels, communication schedules, overlap, caching, serving -- caps at Relevance 5 even when it is excellent and even when it is about MoE. Making a known design run faster is a different job from changing the design.

That cap is about SPEED, not about size. Compression that is lossless, or that changes what can be trained or held in memory at all rather than how quickly it runs, is not capped: it changes what is possible, which is the thing this feed is for.

Analysis of how a trained model's computation is actually organised -- what a circuit computes, what drives expert selection, how representations reorganise over training -- IS in scope and scores like any other work on mechanism. Only interpretability that stops at describing features, with nothing said about the mechanism that produced them, drops out of the feed.

Novelty Scoring

Novelty means a change to the design space, not a change to a number. Score against this ladder:

  • 9-10: removes or replaces a component the whole field inherits without thinking (normalization, weight decay, the tokenizer, positional encoding, left-to-right decoding, the standard optimizer, a numerical format), or decouples two quantities everyone conflates -- AND states which end is held fixed while the other moves. The claim is unconditional: it does not require a particular scale, budget regime, or extra stage. An ANALYSIS paper reaches this band when it settles a mechanism-level question with a criterion that could have come out the other way.
  • 7-8: a substantial new mechanism, or a decoupling whose gain is real but CONDITIONAL (holds above some batch size, in one budget band, or with an added training phase). A conditional win is a new knob, not a removed one, and belongs here rather than above.
  • 5-6: meaningful but incremental extension or refinement of an existing primitive; a tuned variant; a well-executed combination of known parts.
  • 3-4: minor, narrow, or mostly engineering or domain-specific improvement; a survey or overview, however thorough.
  • 1-2: little originality; mainly standard application of existing methods.

Score the claim as the paper states it here. Whether the paper supports that claim is Evidence, and it is scored separately -- do not discount Novelty for weak experiments, and do not raise it because the framing is confident. Phrases like "we remove", "without X", "holding Y fixed" are what the paper SAYS; they earn a high Novelty only if the thing named really is a component the field inherits by default, not a component this paper introduced two sentences earlier.

Evidence Scoring

Does the demonstration reach as far as the claim?

  • 9-10: shown at the scale and breadth the claim asserts -- language-model pretraining at billion-parameter or hundred-billion-token scale, or across several model families, sizes, or domains -- with ablations that isolate the named mechanism and could have failed.
  • 7-8: one credible setting that matches the claim's stated scope, with the baselines a skeptic would ask for.
  • 5-6: a single small setting, or a claim stated more broadly than the demonstration reaches.
  • 3-4: the claim is about transformers or language models in general, but the demonstration lives in one narrow domain or one small benchmark -- image classification alone, a single toy task, one dataset; or the comparisons a reader needs in order to believe it are missing.
  • 1-2: the central claim is asserted, illustrated, or supported only by plots that could not have come out the other way.

A structural claim demonstrated only outside the setting it claims is capped at 4, however striking the claim. "We removed a component everyone uses" shown only on small vision models is a result about small vision models.

When the abstract names no scale, no dataset, and no baseline at all, score exactly 5. Do not infer rigor from confident writing. Unstated is not the same as strong, and two thirds of abstracts say nothing here: guessing on those was measured to make this axis worse.

Adjustments: subtract 1 when no code is released and the procedure is not reproducible from the paper alone; subtract 1 when the method's cost grows in the number of components it adds (one loss per pair of experts, one module per domain) and the paper does not account for that growth; add 1 when the paper states its own limitation precisely enough that a reader could design the experiment that breaks it.

Load Scoring

What does it cost to get the usable result out? This is about the reader's effort, not quality.

  • 9-10: the result cannot be taken without following the derivation -- a new theoretical framework, unfamiliar mathematical machinery, or a proof that IS the contribution.
  • 7-8: substantial theory or an involved formalism, but the operational result is stated plainly somewhere.
  • 4-6: ordinary methods paper; the recipe is legible from the paper's own description.
  • 1-3: the takeaway is a single decision or number a reader can act on immediately.

A high Load is not a criticism. It only says the paper is a project rather than a read.

Proximity Scoring

How close is this to the Active Lines stated below? This axis is about the reader, not the paper, and a paper can be excellent and distant at the same time.

  • 9-10: squarely on a centre line -- the paper's core contribution is the thing the reader is working on, and a result here changes what they would do next.
  • 7-8: on a centre line but from a direction they are not working from, or on an adjacent line where the result carries over directly.
  • 5-6: adjacent: same stack, different layer; they would want to know it happened.
  • 3-4: far: recognisable as the same field, but nothing here reaches their work.
  • 1-2: another area entirely.

Judge distance from the reader's stated centre, NOT from whatever is currently prominent in the field. A paper everyone is discussing is not thereby close.

Topic Registry

Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - frontier_models: Frontier Model Releases and Technical Reports - Frontier model and technical reports that disclose an architecture or training decision: layer and attention design, normalization and residual choices, hybrid stacks, sparsity layout, tokenizer and data mixture, optimizer and schedule, precision, and the stability problems that had to be solved. - architecture_training: Architecture and Training Dynamics - Architectural and optimisation mechanisms, and the training dynamics that explain them: attention and normalization design, positional schemes, state-space and recurrent stacks, tokenizer-free and learned-chunking models, non-autoregressive and discrete-diffusion generation, optimizers and parameterisations including hyperparameter transfer across scale, and work that removes a standard component and shows the model still trains. - training_systems: Training Algorithms That Change What Is Possible - Distributed training algorithms that change what can be trained rather than how fast it runs: asynchronous, low-communication and decentralised optimisation, parallelism schemes with a new invariant, low-precision and 1-bit training when the claim is stability, and scaling laws including data-constrained regimes and critical batch size. - moe_training: MoE Where It Changes the Design Space - Mixture-of-Experts work that opens or closes a design axis: differentiable and stable routing, the training/inference router gap, expert collapse, granularity and layout as a new axis, dense-to-MoE conversion, and MoE scaling laws that fix one quantity and vary another. MoE systems and inference work is out. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, quantization, pruning, low-rank and memory or cache efficiency whose mechanism is new and whose gain is unconditional.

Papers

[PAPER LIST HERE]

Instructions

Respond in JSONL. Output exactly one JSON object per paper, one per line:

{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"EVIDENCE":0,"LOAD":0,"PROXIMITY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}

Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10, scoring the claim as stated. - EVIDENCE: integer from 1 to 10, scoring whether the demonstration reaches as far as the claim. - LOAD: integer from 1 to 10, scoring what it costs the reader to extract the usable result. - PROXIMITY: integer from 1 to 10, scoring distance from the stated Active Lines. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only: daily_hot, new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return []. - daily_hot means the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. - new_frontier means the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.