Personalized Daily ArXiv Papers 2026-09-29
| Model | Metric | Usage | Papers | ||||
|---|---|---|---|---|---|---|---|
| Prompt | Completion | Total | Total arXiv | Scanned | Relevant | ||
local agent CLI |
Tokens | not reported | not reported | not reported | 663 | 403 | 12 |
| Cost | not reported | not reported | not reported | ||||
Topic Coverage:
| Topic | Papers |
|---|---|
| Architecture and Training Dynamics | 11 |
| Efficiency, Compression, and Large-Scale Training | 1 |
Table of contents by topic:
Architecture and Training Dynamics (11)
-
Programs-of-Layers in LLMs through the Lens of Cortical Areas Authors: Justus Westerhoff, Stephan Olbrich, Hatem Oraby, Matthew Evan Larkum, Felix Alexander Gers
-
Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling Authors: Narges Mokhtari, Farzan Haddadi, Ebrahim Rezaii
-
MoSAR: Mixture of Semantic Attention Regimes for Learning Adaptive and Approximable Attention Geometries Authors: Michele Paolicelli, Alessandro Petruzzelli, Alessandro Franceso Maria Martina, Cataldo Musto, Giovanni Semeraro
-
Low-Rank Friction for Memory-Efficient Transformer Pretraining Authors: Rajit Rajpal, Benedict Leimkuhler
-
FTB Graph: Determining and Validating First-token Broadcasters and Language-Identity Head Circuits in Multilingual Language Models Authors: Arjun Pillai, Christian Hoang, Anjelo Laroza
-
Convergence guarantees for Muon: New parameter regimes and generalizations Authors: Arthur C. B. de Oliveira, Dhruv D. Jatkar, Guilherme S. Vicinansa, Eduardo D. Sontag
-
Decodable In-Context State and Model Output Across Training Authors: Manas Venkata Sai Ravulapalli, Samrath Singh Chadha
-
Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline Authors: Haowei Liu, Hsin-Tai Wu, Yi Fang
-
When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess Authors: Salma Roshdy Aly, Hussein Assaf, Ziad Kobti
-
DIAL: Position-Debiased LLM Judges with Adaptive Human Preference Calibration Authors: Zesheng Cai, Yingqi Fan, Sichang Chen, Jin-Hong Du
-
CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production Authors: Mukul Chhabra, Shail Patel, Luigi Medrano
Efficiency, Compression, and Large-Scale Training (1)
- Accounting for Bias Enables Sustainable LLM Evaluation Authors: Harshita Katoch, David Antony Selby, Gerrit Gro{\ss}mann, Sebastian Vollmer
Architecture and Training Dynamics (11)
1. Programs-of-Layers in LLMs through the Lens of Cortical Areas
ArXiv ID: 2609.31360
Primary Topic: Architecture and Training Dynamics
Authors: Justus Westerhoff, Stephan Olbrich, Hatem Oraby, Matthew Evan Larkum, Felix Alexander Gers
Abstract: Inference in LLMs is conventionally a fixed-depth, fixed-order forward pass through every layer, regardless of how difficult the input is. The human brain does not work this way: using the thalamus as a central hub, it routes information flexibly to all regions of the cortex according to demand. Li et al. (2026) recently showed, with a system they call program-of-layers (PoLar), that transformers can be given an analogous flexibility if their layers are treated as a library of functions rather than a fixed sequence. Performance improves over the standard forward pass when each input is dynamically routed through an adaptive sequence of skipped or repeated contiguous layer blocks. We reconstructed PoLar's diagnostic MCTS in more detail than the original paper and applied it across 5 models. We reproduced several of PoLar's findings: skipping outperformed the standard pass, repeating outperformed skipping, and combining both outperformed either alone. Shorter programs sufficed for easier questions, while harder questions required more layer repeats. However, we failed to replicate the main claim regarding their learned router for single-shot inference: its top-ranked prediction consistently collapsed back to the standard pass, even though its top-k predicted programs, taken together, did show a real accuracy gain. Beyond reproduction, we find that a small number of generic programs are enough to solve most of the questions. We also provide a much deeper analysis of these programs' structure and robustness: for example, we found that programs that correct errors are highly brittle: undoing even a single edit inside a program typically breaks the correction. Connecting this to the brain's routing mechanisms, PoLar mirrors principles of thalamo-cortical coordination between cortical-area-like transformer layers. We publicly release the code at https://datexis.github.io/RE-PoLar/
Comment: Tests learned replacement of fixed layer order, exposing top-ranked router collapse and brittle corrective layer programs.
Topic Match: Controlled layer skipping and repetition directly test how transformer computation is organized and whether adaptive routing delivers its claimed benefit.
Relevance: 9 Novelty: 6
2. Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling
ArXiv ID: 2609.30288
Primary Topic: Architecture and Training Dynamics
Authors: Narges Mokhtari, Farzan Haddadi, Ebrahim Rezaii
Abstract: In Transformer-based masked language models, attention is the primary mechanism for context mixing, but there are other ways to mix data across tokens. Recent attention-free mixers replace attention with fixed or hypernetwork-generated MLPs, alternating their dynamic, content-dependent weighting for computational simplicity. We build an alternative that gets the same property from a low-rank bottleneck autoencoder. We replace attention with a stack of autoencoder-based mixing modules, one operating over local neighborhoods, one over the full sequence, and one across attention heads, each compressing and reconstructing its input through a bottleneck, and its width is a hyperparameter rather than a training effect. In masked positions, we introduce an iterative refinement procedure that has two distinct steps. A pulling step that pulls an embedding representation toward a weighted average of its neighbors, and a correcting step that projects the result back to the learned manifold via an autoencoder. Our architecture achieves a significant portion of attention's performance at about $1.9 \times$ fewer FLOPs when pretrained on C4 and evaluated with parameter-matched BERT baselines. Our model equals parameter-matched BERT and TinyBERT baselines on the rarest-token frequency bucket using a frequency-aware training schedule that samples rare tokens more than uniformly for the masking tasks.
Comment: Replaces self-attention with bottleneck autoencoder mixers and compares against BERT at fixed parameter count.
Topic Match: Attention replacement is directly central to layer design. Overall quality remains below attention baselines, and rare-token parity additionally depends on frequency-aware masking.
Relevance: 9 Novelty: 7
3. MoSAR: Mixture of Semantic Attention Regimes for Learning Adaptive and Approximable Attention Geometries
ArXiv ID: 2609.31261
Primary Topic: Architecture and Training Dynamics
Authors: Michele Paolicelli, Alessandro Petruzzelli, Alessandro Franceso Maria Martina, Cataldo Musto, Giovanni Semeraro
Abstract: The quadratic complexity of dense self-attention remains a central bottleneck for long-context language modeling. Many efficient alternatives address this cost by deciding in advance where attention should be sparse or local. We argue that attention approximation should instead be approached as a geometric problem, with the relevant interaction geometry learned from data: natural-language dependencies are input-dependent and difficult to prescribe in advance, so the model should learn where positional relevance can decay and where broader interactions must be preserved. We introduce Mixture of Semantic Attention Regimes (MoSAR), which learns such an adaptive, controlled-decay geometry over query--key interactions. Input-conditioned query and key routers, applied after positional encoding, select mixtures over short, medium, and global regimes, inducing a continuous distance-dependent attention field rather than a fixed sparsity pattern. This geometry is learned during training and can subsequently be discretized through top-1 routing. In controlled pre-training experiments with matched 500M-parameter models, MoSAR learns a substantially lower-reach attention geometry without degrading language-modeling quality, improving perplexity over dense RoPE at the training context length. Under length extrapolation, MoSAR achieves the best perplexity among all evaluated variants, including strong baselines such as ALiBi. Moreover, the learned geometry remains stable under deterministic top-1 discretization, suggesting that it is not only adaptive, but also amenable to low-cost approximation at inference time.
Comment: Learns input-dependent query-key decay regimes during pretraining, with matched 500M-model comparisons and stable top-1 discretization.
Topic Match: The core contribution changes the trainable attention geometry and improves training-context perplexity; length extrapolation provides additional validation of that mechanism.
Relevance: 9 Novelty: 7
4. Low-Rank Friction for Memory-Efficient Transformer Pretraining
ArXiv ID: 2609.30342
Primary Topic: Architecture and Training Dynamics
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Rajit Rajpal, Benedict Leimkuhler
Abstract: iKFAD is a recently proposed optimiser that replaces adaptive learning rates with adaptive friction in the momentum dynamics, yet performs as well as Adam. Its limitation is that the full friction tensor $\xi\in\mathbb{R}^{m\times n}$ carries the same $\mathcal{O}(mn)$ memory overhead per layer as Adam's second-moment buffer. Here we replace iKFAD's friction tensor $\xi$ with a rank-1 outer-product factorisation built from row and column momentum statistics, resulting in Rank-1 iKFAD (R-iKFAD). This reduces the friction memory footprint from $\mathcal{O}(mn)$ to $\mathcal{O}(m+n)$ per layer, which approximately halves iKFAD's total optimiser state. Despite this reduction, R-iKFAD maintains parity in performance with iKFAD: experiments on GPT2-Nano, TinyViT, DistilBERT and GPT2-S confirm that it matches or exceeds iKFAD while nearly halving the memory footprint and remaining comparably robust to hyperparameters. We analyse the continuous-time dynamics in two damping regimes. For linear damping ($\gamma>0$) we prove exponential convergence under strong convexity. For $\gamma=0$, the preferred option in our experiments, the friction is generated entirely from past momentum and switches off as the momentum vanishes, so geometric convergence cannot be shown. We nonetheless prove convergence to the minimiser, together with matching upper and lower bounds on the energy: of order $t^{-1}$ when the regularisation scale $\epsilon_{\mathrm{stab}}$ is zero, and of order $t^{-1/2}$ when it is positive. To our knowledge this is the first convergence rate for a rank-1 factored optimiser in continuous time, and the first such result that does not require positive damping.
Comment: Factorizes an optimizer's adaptive friction into row and column statistics, reducing state memory while testing pretraining performance against the unfactored optimizer.
Topic Match: The primary contribution changes optimizer parameterization and dynamics; its reduction in optimizer-state memory also matches training efficiency.
Relevance: 9 Novelty: 6
5. FTB Graph: Determining and Validating First-token Broadcasters and Language-Identity Head Circuits in Multilingual Language Models
ArXiv ID: 2609.30954
Primary Topic: Architecture and Training Dynamics
Authors: Arjun Pillai, Christian Hoang, Anjelo Laroza
Abstract: Large language models operating in multilingual contexts must resolve target response languages early in generation, yet the causal circuitry governing first-token language identity decisions remains poorly mapped. We present an end-to-end structural circuit analysis across six model architectures spanning four families: GPT-2, BLOOM-560M, Pythia-1B/2.8B, and Qwen2.5-1.5B Base/Instruct. Using Edge Attribution Patching (EAP) with FP16 active clamping, followed by exact activation patching verification with a 2,000-candidate-edge search ceiling, we extract directed acyclic graphs driving first-token language broadcasting. Across the standalone models, we observe deep or mid-to-deep broadcasting hubs, though the evidence is strongest for Pythia-2.8B and BLOOM-560M because GPT-2 and Pythia-1B leave few out-of-graph heads for comparison, while both Qwen2.5-1.5B variants invert the necessity check. Scaling from Pythia-1B to 2.8B expands node participation while maintaining a similar verified edge budget, producing sparser topology. The Qwen2.5-1.5B base and instruct circuits retain 84.7% Jaccard similarity, including the Layer 27 hub, indicating that first-token routing is largely established during pretraining and preserved by instruction tuning. Finally, EAP scores correlate weakly with exact patching deltas across most models, showing that linear gradient approximations can diverge from causal interventions in FP16 and motivating exact-patching verification for reliable circuit discovery.
Comment: Tests first-token language-selection circuits with exact activation patching and model-specific necessity checks.
Topic Match: Analyzes what attention-head circuits compute through causal interventions; failed necessity checks limit the cross-model conclusion.
Relevance: 8 Novelty: 7
6. Convergence guarantees for Muon: New parameter regimes and generalizations
ArXiv ID: 2609.30546
Primary Topic: Architecture and Training Dynamics
Authors: Arthur C. B. de Oliveira, Dhruv D. Jatkar, Guilherme S. Vicinansa, Eduardo D. Sontag
Abstract: In this paper, we establish the first asymptotic convergence guarantees for the Muon algorithm through a more accurate proxy for the Newton-Schultz iteration than the typical matrix sign function. We prove that, for appropriate choices of hyperparameters, the iterates satisfy $\lim_{k\to\infty}|\nabla f(x_k)|=0$, and, under a global Polyak-\L{}ojasiewicz condition, that the sequence of function values converges linearly. The key insight is that the regularization, implicit in Muon's Newton-Schulz implementation, induces a bounded preconditioner, exposing Muon as a \emph{preconditioned Polyak heavy-ball} method and enabling a classical Lyapunov analysis. This observation naturally motivates applying the same preconditioning structure to the Nesterov gradient evaluation. We formalize this idea by introducing \emph{Muesterov}, a Nesterov-based variant of Muon, and prove that it enjoys the same convergence guarantees, extending the theoretical framework beyond the heavy-ball setting. Numerical experiments on a scalar cross-entropy problem corroborate the theory and illuminate the joint role of the learning rate and the Newton-Schulz regularizer in controlling convergence. Preliminary numerical simulations training the nanoGPT dataset provide intuition regarding the relevance of the observations in this paper to practical applications.
Comment: Explains Muon's Newton-Schulz regularization as a bounded preconditioner and derives convergence regimes.
Topic Match: Directly analyzes an optimizer's implemented mechanism and motivates a Nesterov variant; practical language-model training evidence remains preliminary.
Relevance: 8 Novelty: 7
7. Decodable In-Context State and Model Output Across Training
ArXiv ID: 2609.31401
Primary Topic: Architecture and Training Dynamics
Authors: Manas Venkata Sai Ravulapalli, Samrath Singh Chadha
Abstract: Prior work established that a probe can decode an in-context binding on model errors and that probe-guided steering can repair some of them. We follow probe accuracy, model output, and steering response across public pretraining and post-training checkpoints. Probe accuracy rises during Pythia pretraining, while probe-guided steering moves from negligible all-trial benefit to a larger benefit at two model sizes. Saved scores distinguish probe-correct errors with low and above-uniform model probability for the correct candidate. Oracle-target steering already repairs many early errors, but saved aggregates cannot separate target quality from intervention sensitivity. A held-out comparison of decoders trained on the final state or candidate logits finds no detected final-state advantage on late-checkpoint model errors. An information-theoretic counterexample explains why decodability on errors alone cannot establish discarded output information. The connection to downstream omissions remains open.
Comment: Tracks changes in decodable state and steering efficacy across pretraining, comparing final-state decoders with candidate-logit decoders.
Topic Match: Checkpoint comparisons connect representation development to output behavior, although the causal source of changing steering efficacy remains unresolved.
Relevance: 7 Novelty: 6
8. Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline
ArXiv ID: 2609.30290
Primary Topic: Architecture and Training Dynamics
Authors: Haowei Liu, Hsin-Tai Wu, Yi Fang
Abstract: Production text-to-SQL pipelines often end with an LLM-as-judge whose agreement with human annotators has never actually been measured. When we checked ours, the deployed gpt-4o-mini judge agreed with two-author gold at only Cohen's kappa = 0.04 on a disagreement-enriched set and 0.42 on a uniform-random spot-check, over-flagging 77.1% of the human-FAITHFUL cases in the enriched set. Most of its over-flags trace back to a single mechanism we call GRADE-HALLUCINATION. A self-hosted Qwen3.6-27B replacement (kappa = 0.72) lands in the same range as Claude Opus 4.7 (kappa = 0.71); the head-to-head is underpowered at n = 96, but for the deployment decision that hardly matters, since Qwen costs roughly 1/300 as much per call. Ensembling does not help for free. Pairing the weak judge with a stronger one degrades agreement, whereas three strong judges under unanimity routing reach kappa = 0.79 at 89.7% auto-coverage. Applied out-of-domain, the same audit recipe flags 25.5% of BIRD-financial's expert-authored gold SQLs as candidate gold-SQL issues under our annotation protocol. Code and pre-registration are at https://github.com/JamesL404/synca-audit.
Comment: Audits deployed-judge agreement against human annotations, exposing systematic over-flagging and near-chance agreement on the enriched sample.
Topic Match: Retained under the explicit judge-reliability audit exception. The registry lacks an audit category, so architecture_training is the nearest available label.
Relevance: 8 Novelty: 6
9. When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess
ArXiv ID: 2609.30328
Primary Topic: Architecture and Training Dynamics
Authors: Salma Roshdy Aly, Hussein Assaf, Ziad Kobti
Abstract: When one language model judges whether another's code is correct, it does not report the absence of evidence. It returns a confident verdict with reasoning attached, indistinguishable from a verdict it had grounds for. Multi-agent verification, which decomposes a judgment into checkable claims and verifies each against evidence, is a promising response and works well when the evidence is a set of retrieved documents. We argue such methods require two things of their evidence: it must be independent of the answer under review, and it must differ between the two candidates being compared. The second condition holds automatically with retrieved documents and stops holding in code judging. Running MARCH, a published framework unmodified over 80 condition-by-cell measurements on two code judging benchmarks, we find it declares both solutions equally good on 78 to 95% of comparisons, reaching 4.4% accuracy where the same model asked directly reaches 43.7%. Neither easier problems nor a larger judge changes this. Two measurements taken from the pipeline's own logs explain it without needing labels. Gating on one of them, the pipeline declines the comparisons it cannot make and raises its accuracy from 20.7 to 36.9% while still answering half of all comparisons. The contribution is not a more accurate judge, but a label-free way to tell when a judge has no basis for its answer.
Comment: Audits LLM-as-judge reliability by measuring whether a code-verification pipeline has evidence that distinguishes candidates.
Topic Match: Qualifies through the explicit judge-reliability audit exception; architecture_training is a required fallback because the registry has no audit topic.
Relevance: 8 Novelty: 7
10. DIAL: Position-Debiased LLM Judges with Adaptive Human Preference Calibration
ArXiv ID: 2609.31215
Primary Topic: Architecture and Training Dynamics
Authors: Zesheng Cai, Yingqi Fan, Sichang Chen, Jin-Hong Du
Abstract: Large language models (LLMs) as a judge enable scalable evaluation, but their judgments can be sensitive to response order and, even after removing such position effects, can still diverge systematically from human preferences.We introduce DIAL, a unified framework that combines abundant LLM comparisons with limited human comparisons to separate judge-specific position effects, learn shared structure in position-debiased LLM preferences, and adaptively calibrate that structure toward the human preference target. Theoretically, we study three aspects of DIAL: (i) identification of latent LLM preferences, position effects, and human calibration; (ii) adaptive estimation that balances LLM anchoring against limited human evidence; and (iii) fixed-weight uncertainty quantification for the calibrated human preference. Empirically, we evaluate position debiasing and human alignment separately in controlled simulations and on three human-preference benchmarks, showing that DIAL remains robust to unbalanced response order, achieves strong human-aligned rankings with limited labels, and adapts toward human evidence when LLM information is imperfect. Our real-data study collects over 410K judgments from 21 LLM judges in both display orders, providing a resource for future studies of LLM-judge bias, heterogeneity, and human alignment.
Comment: Separates LLM-judge position bias from disagreement with human preferences using paired-order judgments.
Topic Match: Qualifies under the explicit LLM-judge reliability exception; architecture_training is a required registry fallback because evaluation validity has no topic ID.
Relevance: 8 Novelty: 7
11. CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production
ArXiv ID: 2609.30471
Primary Topic: Architecture and Training Dynamics
Authors: Mukul Chhabra, Shail Patel, Luigi Medrano
Abstract: Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support cases, assets, accounts), the closest available reference typically applies the correct procedure to a different entity, so a literal judge penalizes different identifiers, dates, and statuses as errors or hallucinations. We name this failure mode reference-instance divergence (RID). We propose CARGO, a framework that (i) treats retrieved references as procedural exemplars and grounds factual judgments in the live instance's observed context, (ii) assigns each claim a three-way status (supported, contradicted, unverifiable) and penalizes only contradictions, and (iii) gates evaluation by retrieval confidence, casting production evaluation as selective prediction. We introduce CARGO-Bench, a perturbation-based diagnostic suite with ground truth by construction that separates leniency from discrimination. On CARGO-Bench (246 items, two judge models, 7,872 judgments), the standard reference-based judge penalizes 100% of correct entity-transplanted answers and is uninformative (discrimination index DI ~ 0); supplying the live facts without reframing changes nothing. CARGO eliminates these false penalties (0/50) while retaining near-complete contradiction recall (50/50 and 49/50), raising DI to 0.58 [0.48, 0.68]; a rubric-swap control attributes most of the effect to context-grounded dimension definitions. CARGO also exposes a limitation of its own design: the leniency that protects entity values suppresses detection of procedural corruptions (20% recall). A post-hoc fix does not close the gap, and an LLM-as-annotator study with written guidelines and adjudication shows the same blind spot. We release a preregistered protocol for extending the evaluation to expert agreement, risk-coverage, and cost on production traffic.
Comment: Uses entity-swap and rubric-swap controls to expose invalid judge penalties caused by mismatched reference instances.
Topic Match: The controlled judge-failure audit fits the explicit reliability exception, despite the accompanying benchmark proposal. architecture_training is the nearest available registry label.
Relevance: 7 Novelty: 6
Efficiency, Compression, and Large-Scale Training (1)
1. Accounting for Bias Enables Sustainable LLM Evaluation
ArXiv ID: 2609.31184
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Harshita Katoch, David Antony Selby, Gerrit Gro{\ss}mann, Sebastian Vollmer
Abstract: LLM-as-a-judge has become the de facto standard for scalable, subjective evaluation, yet current leaderboards compensate for systematic measurement bias by running ever more comparisons, an approach that is both statistically unsound and computationally wasteful. The root cause is an incomplete measurement model, treating LLM judges as neutral, interchangeable instruments ignores documented biases like position bias, verbosity bias, judge severity, and self-enhancement, that no volume of additional data can eliminate. We propose a unified latent variable framework that jointly models pairwise and ordinal data while explicitly correcting for these confounders, recovering reliable rankings from substantially fewer comparisons. Because fitting this model costs negligible compute relative to a single round of LLM inference, bias correction is not only more statistically rigorous but also a more sustainable approach to trustworthy evaluation.
Comment: Corrects position, verbosity and judge-specific biases in LLM-as-judge rankings.
Topic Match: Kept through the explicit judge-reliability exception. Comparison efficiency supplies the closest required registry label; the abstract names no empirical validation setting.
Relevance: 8 Novelty: 6
Paper Selection Prompt
System Prompt
You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.
User Prompt
Relevant Topics
This is a training-side feed for someone who builds and modifies model internals. The centre of gravity is what a frontier model is made of and how it is trained: architecture, training dynamics, and the decisions a lab actually made. MoE is still in scope, but only when the contribution changes the design space, not when it tunes an existing MoE.
Keep a paper when its CORE CONTRIBUTION falls in one of the five topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.
THE SUBJECT IS LANGUAGE-MODEL TRAINING. Every topic below is scoped to it. A technique named in one of them -- decentralised or asynchronous training, discrete diffusion, byte-level modelling, 1-bit weights, mixture-of-experts -- earns nothing when the model being trained is an image generator, a recommender, a scientific surrogate, or a vision classifier. Match on what is being trained, not on the vocabulary in the abstract.
THE ONE QUESTION THAT DECIDES A HIGH SCORE, asked before the topic list:
Does this paper REMOVE or REPLACE a component that everyone downstream inherits without thinking, or DECOUPLE two quantities the field currently conflates -- and does it say which end is held fixed while the other moves?
A paper that ADDS a mechanism on top of existing primitives is ordinary work, however good the numbers. A paper that takes one away, or splits one knob into two, is the reason this feed exists. Papers whose gains are CONDITIONAL (only above some batch size, only in one budget regime, only with an extra training stage) are ordinary work too, even when the mechanism is new: a conditional win is a new knob, not a removed one.
Frontier Model Releases and Technical Reports (primary) - Keep: model and technical reports that DISCLOSE architecture or training decisions -- layer and attention design, normalization and residual choices, hybrid attention/state-space stacks, sparsity and expert layout, tokenizer and vocabulary decisions, data mixture and curriculum, optimizer and schedule, precision and numerics, stability fixes and what broke; open-weight releases whose report explains a choice rather than only reporting it. - Filter: releases that are a scorecard -- benchmark tables, a capability announcement, a product or API launch, or a report that names its recipe without saying why it was chosen. A frontier name in the title earns nothing on its own.
Architecture and Training Dynamics (primary) - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, positional schemes, dynamic or modular computation, tokenizer-free and byte-level or learned-chunking stacks, non-autoregressive and discrete-diffusion generation); optimizers, preconditioners, and parameterisations, especially hyperparameter transfer across scale; training-dynamics and stability analysis that explains why large models train the way they do; work that removes a standard component (normalization, weight decay, positional encoding, the tokenizer, left-to-right decoding) and shows the model still trains. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight; an architectural tweak evaluated only at toy scale with no account of what it costs.
Large-Scale Training Systems - Keep: distributed training ALGORITHMS that change what is possible, not only what is fast -- asynchronous, low-communication and decentralised optimisation, sharding and parallelism schemes with a new invariant, numerics and low-precision training (FP8, FP4, ternary and 1-bit) when the claim is about training stability rather than a kernel; scaling-law work that informs how a run is configured, including data-constrained and repeated-data regimes and critical batch size. - Filter: kernel engineering, communication scheduling, overlap, caching, serving and inference throughput. This is a real exclusion, not a soft one: work whose contribution is "the same model, faster" belongs to someone else's feed even when it is excellent and even when it is about MoE.
MoE Where It Changes the Design Space - Keep: routing that is newly differentiable or newly stable; the training/inference router gap; expert collapse and what actually prevents it; expert granularity, shared experts and layout when a new axis is opened rather than tuned; dense-to-MoE conversion; MoE scaling laws that fix one quantity and vary another; analyses that carry a criterion which could have failed (a random-guess baseline, an ablation with a predicted sign) rather than similarity heatmaps. - Filter: MoE SYSTEMS and inference work (expert parallelism, all-to-all schedules, grouped GEMM, offload, prefetching, capacity and scheduling tricks, compression for serving); MoE surveys; papers that train on top of a MoE without a routing, balancing, stability, or structural contribution; "mixture of experts" in the classical ensemble or recommender sense. - The exclusion above is about MAKING MoE RUN FASTER, and it does not reach routing and expert design. A paper on how tokens are assigned to experts, how experts specialise, how balance is enforced, or how a dense model becomes sparse stays fully in scope and scores on its merits, even when the mechanism is a small one.
Efficiency and Compression When the Mechanism Is New - Keep: quantization, sparsity, pruning and low-rank work whose mechanism is new and whose gain is unconditional; compression that changes what can be trained, not only what can be served; attribution of a model's behaviour to its data. - Filter: tuned variants of standard efficiency methods; deployment and serving engineering; anything whose headline is a compression ratio with no account of what was given up.
Also keep, even though they read like evaluation work, because the field has no one checking them: - benchmark validity itself: contamination, saturation, LLM-as-judge reliability, whether a benchmark measures what it claims. A paper AUDITING a benchmark is in scope; a paper PROPOSING one is not.
Subjects He Has Never Once Kept
Measured, not guessed. Over nine months where he was actively curating, 4953 papers passed this filter and 274 made his list -- a keep rate of 5.2%. Each subject below appeared in the rejected pile the number of times shown and appeared ZERO times in his entire 508-entry list.
Filter a paper whose CORE SUBJECT is one of these. The count is the evidence; where a paper only mentions the term in passing while contributing somewhere in the five topics above, keep it.
variational methods and inference (52) unlearning (41) long-context methods (40) neural operators (35) time series (35) vision transformers (35) parameter-efficient fine-tuning (28) stochastic gradient analysis (21) LLM safety (20) LLM reasoning as a subject (20) hallucination (20) sequence modelling as a subject (19) differential equations (19) spiking neural networks (17) post-training quantization (16) mechanistic interpretability as a framing (15) test-time scaling (13)
Two more he kept exactly once each, so suppress rather than drop: graph neural networks (1 of 42), knowledge distillation (1 of 39).
Do NOT extend this list by analogy. Four subjects that look like they belong here were checked and do not: diffusion models (8.1% kept), prompting (10.0%), benchmarks (9.1%) and neural architecture search (13.3%) all sit ABOVE his 5.2% base rate.
Note the one distinction that matters: he rejects papers FRAMED as interpretability, while analysis of how a trained model's computation is organised is among his highest-rated work. The subject is the framing, not the act of analysing.
Leave these to the hotspot digest unless the core contribution clearly falls in one of the five topics above: - agent, tooling, and RAG launches, and agent framework papers - new benchmarks, leaderboards, and evaluation-only papers - interpretability that stops at cataloguing features, with nothing said about the mechanism that produced them. Analysis of how a trained model's computation is organised -- what drives expert selection, what a circuit computes, how representations reorganise across training -- is IN scope, not here. - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning), EXCEPT where the subject is how post-training destabilises the architecture itself - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains
Scoring Criteria
Score five independent axes from 1 to 10: Relevance, Novelty, Evidence, Load, Proximity.
They answer different questions and must not be collapsed. Relevance: is this the kind of work this feed is for. Evidence: did the paper earn its claim. Load: what does it cost to get the usable result out. Proximity: how close is it to what the reader is working on right now (see Active Lines below).
Novelty is scored for continuity with the archive and for the hotspot spotlight cutoffs, and it does NOT order the feed. Measured against 274 hand-tiered papers it carried no signal at all: papers scoring 8 or more were his must-reads exactly as often as papers scoring 5 or less, both at the pool's base rate. Every abstract claims novelty, so the axis measures the claiming rather than the work. Score it honestly and do not let it influence the other four.
A strong claim with thin evidence is a worse read than a modest claim that holds. Score each axis on its own; do not let a high one pull up a low one.
Relevance Scoring
- 9-10: directly centered on the target topics; highest when the core contribution is clearly within them and the paper is about how a model is BUILT or TRAINED.
- 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
- 5-6: touches the target topics, but the main contribution is elsewhere.
- 3-4: largely outside the target topics, often application-focused or domain-specific.
- 1-2: unrelated.
Important: Daily buzz, a frontier lab's name, or a large headline number is not enough for a high Relevance score. A model release scores high only when the report explains an architecture or training decision; a release that only reports results is hotspot material, not this feed.
Work whose contribution is "the same model, faster" -- kernels, communication schedules, overlap, caching, serving -- caps at Relevance 5 even when it is excellent and even when it is about MoE. Making a known design run faster is a different job from changing the design.
That cap is about SPEED, not about size. Compression that is lossless, or that changes what can be trained or held in memory at all rather than how quickly it runs, is not capped: it changes what is possible, which is the thing this feed is for.
Analysis of how a trained model's computation is actually organised -- what a circuit computes, what drives expert selection, how representations reorganise over training -- IS in scope and scores like any other work on mechanism. Only interpretability that stops at describing features, with nothing said about the mechanism that produced them, drops out of the feed.
Novelty Scoring
Novelty means a change to the design space, not a change to a number. Score against this ladder:
- 9-10: removes or replaces a component the whole field inherits without thinking (normalization, weight decay, the tokenizer, positional encoding, left-to-right decoding, the standard optimizer, a numerical format), or decouples two quantities everyone conflates -- AND states which end is held fixed while the other moves. The claim is unconditional: it does not require a particular scale, budget regime, or extra stage. An ANALYSIS paper reaches this band when it settles a mechanism-level question with a criterion that could have come out the other way.
- 7-8: a substantial new mechanism, or a decoupling whose gain is real but CONDITIONAL (holds above some batch size, in one budget band, or with an added training phase). A conditional win is a new knob, not a removed one, and belongs here rather than above.
- 5-6: meaningful but incremental extension or refinement of an existing primitive; a tuned variant; a well-executed combination of known parts.
- 3-4: minor, narrow, or mostly engineering or domain-specific improvement; a survey or overview, however thorough.
- 1-2: little originality; mainly standard application of existing methods.
Score the claim as the paper states it here. Whether the paper supports that claim is Evidence, and it is scored separately -- do not discount Novelty for weak experiments, and do not raise it because the framing is confident. Phrases like "we remove", "without X", "holding Y fixed" are what the paper SAYS; they earn a high Novelty only if the thing named really is a component the field inherits by default, not a component this paper introduced two sentences earlier.
Evidence Scoring
Does the demonstration reach as far as the claim?
- 9-10: shown at the scale and breadth the claim asserts -- language-model pretraining at billion-parameter or hundred-billion-token scale, or across several model families, sizes, or domains -- with ablations that isolate the named mechanism and could have failed.
- 7-8: one credible setting that matches the claim's stated scope, with the baselines a skeptic would ask for.
- 5-6: a single small setting, or a claim stated more broadly than the demonstration reaches.
- 3-4: the claim is about transformers or language models in general, but the demonstration lives in one narrow domain or one small benchmark -- image classification alone, a single toy task, one dataset; or the comparisons a reader needs in order to believe it are missing.
- 1-2: the central claim is asserted, illustrated, or supported only by plots that could not have come out the other way.
A structural claim demonstrated only outside the setting it claims is capped at 4, however striking the claim. "We removed a component everyone uses" shown only on small vision models is a result about small vision models.
When the abstract names no scale, no dataset, and no baseline at all, score exactly 5. Do not infer rigor from confident writing. Unstated is not the same as strong, and two thirds of abstracts say nothing here: guessing on those was measured to make this axis worse.
Adjustments: subtract 1 when no code is released and the procedure is not reproducible from the paper alone; subtract 1 when the method's cost grows in the number of components it adds (one loss per pair of experts, one module per domain) and the paper does not account for that growth; add 1 when the paper states its own limitation precisely enough that a reader could design the experiment that breaks it.
Load Scoring
What does it cost to get the usable result out? This is about the reader's effort, not quality.
- 9-10: the result cannot be taken without following the derivation -- a new theoretical framework, unfamiliar mathematical machinery, or a proof that IS the contribution.
- 7-8: substantial theory or an involved formalism, but the operational result is stated plainly somewhere.
- 4-6: ordinary methods paper; the recipe is legible from the paper's own description.
- 1-3: the takeaway is a single decision or number a reader can act on immediately.
A high Load is not a criticism. It only says the paper is a project rather than a read.
Proximity Scoring
How close is this to the Active Lines stated below? This axis is about the reader, not the paper, and a paper can be excellent and distant at the same time.
- 9-10: squarely on a centre line -- the paper's core contribution is the thing the reader is working on, and a result here changes what they would do next.
- 7-8: on a centre line but from a direction they are not working from, or on an adjacent line where the result carries over directly.
- 5-6: adjacent: same stack, different layer; they would want to know it happened.
- 3-4: far: recognisable as the same field, but nothing here reaches their work.
- 1-2: another area entirely.
Judge distance from the reader's stated centre, NOT from whatever is currently prominent in the field. A paper everyone is discussing is not thereby close.
Active Lines
What the reader is working on right now. This is the target the Proximity axis measures against, and it is the ONLY part of the prompt that is expected to change as the work changes. Everything else in the scoring prompt is about the paper; this is about the reader.
Why it exists: over two years of hand-tiered papers, the reader's three middle tiers turned out not to be a quality ladder at all. Measured within a single period, the share of papers sitting on the then-active line fell monotonically across them -- 57%, 36%, 18% -- while the quality axes were flat to three decimal places. Must-read, better-to-read and not-now is a distance, and without a statement of where the centre is there is nothing to measure distance from.
Centre
- Frontier model releases and technical reports, read for the architecture and training decisions they disclose rather than for the results they report.
- Architecture and training dynamics: what a layer is made of, what can be removed from it, and what happens to optimisation when you do.
Adjacent
- Training algorithms that change what can be trained: asynchronous and low-communication optimisation, parallelism with a new invariant, low-precision pretraining, scaling laws that configure a run.
- MoE where the contribution is routing, balance, specialisation, or dense-to-sparse structure.
- Compression and sparsity whose mechanism is new and whose gain is unconditional.
Far
- Making a known design run faster: kernels, communication schedules, serving, inference.
- Post-training, agents, benchmarks, interpretability that stops at cataloguing features.
- Everything outside language-model training.
Topic Registry
Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - frontier_models: Frontier Model Releases and Technical Reports - Frontier model and technical reports that disclose an architecture or training decision: layer and attention design, normalization and residual choices, hybrid stacks, sparsity layout, tokenizer and data mixture, optimizer and schedule, precision, and the stability problems that had to be solved. - architecture_training: Architecture and Training Dynamics - Architectural and optimisation mechanisms, and the training dynamics that explain them: attention and normalization design, positional schemes, state-space and recurrent stacks, tokenizer-free and learned-chunking models, non-autoregressive and discrete-diffusion generation, optimizers and parameterisations including hyperparameter transfer across scale, and work that removes a standard component and shows the model still trains. - training_systems: Training Algorithms That Change What Is Possible - Distributed training algorithms that change what can be trained rather than how fast it runs: asynchronous, low-communication and decentralised optimisation, parallelism schemes with a new invariant, low-precision and 1-bit training when the claim is stability, and scaling laws including data-constrained regimes and critical batch size. - moe_training: MoE Where It Changes the Design Space - Mixture-of-Experts work that opens or closes a design axis: differentiable and stable routing, the training/inference router gap, expert collapse, granularity and layout as a new axis, dense-to-MoE conversion, and MoE scaling laws that fix one quantity and vary another. MoE systems and inference work is out. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, quantization, pruning, low-rank and memory or cache efficiency whose mechanism is new and whose gain is unconditional.
Papers
[PAPER LIST HERE]
Instructions
Respond in JSONL. Output exactly one JSON object per paper, one per line:
{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"EVIDENCE":0,"LOAD":0,"PROXIMITY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}
Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10, scoring the claim as stated. - EVIDENCE: integer from 1 to 10, scoring whether the demonstration reaches as far as the claim. - LOAD: integer from 1 to 10, scoring what it costs the reader to extract the usable result. - PROXIMITY: integer from 1 to 10, scoring distance from the stated Active Lines. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only:
daily_hot,new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return[]. -daily_hotmeans the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. -new_frontiermeans the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.