Personalized Daily ArXiv Papers 2026-09-21
| Model | Metric | Usage | Papers | ||||
|---|---|---|---|---|---|---|---|
| Prompt | Completion | Total | Total arXiv | Scanned | Relevant | ||
local agent CLI |
Tokens | not reported | not reported | not reported | 451 | 274 | 10 |
| Cost | not reported | not reported | not reported | ||||
Topic Coverage:
| Topic | Papers |
|---|---|
| Architecture and Training Dynamics | 6 |
| MoE Where It Changes the Design Space | 2 |
| Efficiency, Compression, and Large-Scale Training | 2 |
Table of contents by topic:
Architecture and Training Dynamics (6)
-
Trading Depth for Time in Recurrent Transformers Authors: Zeyi Huang, Xuehai He, Yong Jae Lee, Yelong Shen
-
TinyCeNN-LM: Quality-Gated Conversion of Pretrained Attention with CeNN-Inspired Cellular-Recurrent Layers Authors: Kabeh Mohsenzadegan, Vahid Tavakkoli, Kyandoghere Kyamakya
-
Prediction Dynamics in Depth-Recurrent Language Models Authors: Xinyue Luo, Fei Yu
-
Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language Models Authors: Lingfang Li, Procheta Sen, Shubham Das, Danushka Bollegala
-
The Stochastic Shift: A New Evaluation Paradigm for Text-to-SQL with AI Operators Authors: Tarfah Alrashed, Fatma Ozcan, Per Jacobsson, Tal Neiman, Xianshun Chen
-
SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs? Authors: George Ma, Benjamin Mikek, Haoyu Li, Ferhat Erata, Yuhao Zhang, Zeren Shui, Behrooz Omidvar Tehrani, Jun Huan, Murali Krishna Ramanathan, Somayeh Sojoudi, Hao Zhou, Anoop Deoras
MoE Where It Changes the Design Space (2)
-
Attention-Aware Routing: Coupling Routing and Attention in MoEs Authors: Despoina Kosmopoulou, Anastasios Tsetsilas, Efthymios Georgiou, Giannis Karamanolakis, Swastik Roy, Alexandros Potamianos
-
Accelerating Dense LLMs via L0-regularized Mixture-of-Experts Authors: Zhenyu Zhang, Jiudong Yang, Zhaowen Tao, Meng Chen
Efficiency, Compression, and Large-Scale Training (2)
-
Detecting Pretraining Data in Large Language Models from a Free-Energy Perspective Authors: Chenye Ke, Zirui Liu, Qi Liu, Yan Zhuang, Jintao Zhang, Zhenya Huang, Shijin Wang
-
Layerwise Decoupling for Stable Structured Sparsification of Fully Connected Layers Authors: Charles Kulick, Armenak Petrosyan, Sui Tang
Architecture and Training Dynamics (6)
1. Trading Depth for Time in Recurrent Transformers
ArXiv ID: 2609.21605
Primary Topic: Architecture and Training Dynamics
Authors: Zeyi Huang, Xuehai He, Yong Jae Lee, Yelong Shen
Abstract: Recurrent Transformers increase computational depth through temporal recurrence, feeding each token's high-level hidden state into the computation of the next. This raises a natural question: is additional computation better spent on more temporal steps or greater physical depth? We investigate this question using Latent Recurrent Transformers (LRTs), which retain one backbone forward pass per vocabulary token during decoding and provide a controlled setting for comparing these two ways of adding computation. Specifically, we insert a latent thought token between consecutive vocabulary tokens. Each thought token passes through the same $L$ layers as a vocabulary token, sharing the backbone parameters and providing an additional stage of hidden-state refinement before predicting the next token. We compare this $L$-layer LRT against a $2L$-layer LRT without thought tokens. Both execute $2L$ Transformer blocks per vocabulary token during decoding, but the thought-token model uses fewer parameters. On 16- and 20-layer mixture-of-experts NanoChat backbones, one thought token brings the shallower model within 0.006 and 0.004 bits per byte of its double-depth counterpart, recovering 67% and 81% of the improvement with approximately 48% fewer total parameters. These results suggest that temporal thinking offers a parameter-efficient alternative to increasing physical depth in recurrent Transformers.
Comment: Trades physical depth for temporal recurrence while holding Transformer-block executions per vocabulary token fixed.
Topic Match: Directly tests the architectural relationship between parameter count and recurrent computation; the MoE backbone is the experimental setting rather than the contribution.
Relevance: 9 Novelty: 7
2. TinyCeNN-LM: Quality-Gated Conversion of Pretrained Attention with CeNN-Inspired Cellular-Recurrent Layers
ArXiv ID: 2609.21139
Primary Topic: Architecture and Training Dynamics
Authors: Kabeh Mohsenzadegan, Vahid Tavakkoli, Kyandoghere Kyamakya
Abstract: Replacing attention in a pretrained language model is a compatibility problem: a plausible substitute may alter representations expected by later layers. TinyCeNN-LM introduces a \emph{quality-gated post-training conversion} framework using CeNN-inspired cellular-recurrent layers with bounded local processing, compact recurrent memory, routing, fusion, and accept-or-rollback validation. Three implementations are studied: Integrated Memory, MemoryFusion, and PDelta3-GDN2-CLVR+Local32. Strict PDelta3 conversion accepts a layer only when representation and NLL criteria pass fixed thresholds. On SmolLM2-135M, layers 0-2 are accepted with cumulative $\Delta\mathrm{NLL}=+0.01209$, while layer 3 is rejected despite acceptable NLL because representation fidelity fails. On Qwen3.5-0.8B, full-attention layers 3, 7, and 11 are accepted with final $\Delta\mathrm{NLL}=+0.02073$. Integrated Memory keeps perplexity within $-0.07\%$ to $+0.93\%$ while reducing total cache by up to $6.01\%$. A sampled 200-item downstream sanity check gives $28.5\%$--$32.0\%$ overall accuracy for converted Qwen releases. The results support conservative, quality-gated structural conversion rather than universal attention replacement or speedup.
Comment: Replaces selected pretrained attention layers with recurrent cells, accepting conversions only when representation and NLL thresholds pass.
Topic Match: The core contribution changes pretrained layer structure and tests compatibility with downstream representations; its claims explicitly stop at partial, quality-gated conversion.
Relevance: 9 Novelty: 7
3. Prediction Dynamics in Depth-Recurrent Language Models
ArXiv ID: 2609.21383
Primary Topic: Architecture and Training Dynamics
Authors: Xinyue Luo, Fei Yu
Abstract: Depth-recurrent language models refine predictions through repeated latent updates. Why can intermediate answers agree with the endpoint while their scores continue to change? We derive a sharp margin characterization that decomposes the conservatism of a magnitude bound into common translation, direction relative to the winner, and the pairing of each competitor's update with its score gap. Across Huginn-3.5B and Ouro-1.4B, accounting for update direction and competitor pairing reduces the mean earliest qualifying depth by a further 22.5-34.4% of the total depth beyond translation removal under full answer-text scoring. This retrospective comparison uses completed trajectories. Substantial contributions also occur under label scoring. For shared predictive distributions, we separate common and contrast motion orthogonally and express the common component through candidate-set mass and within-set concentration. Common and contrast energies can attenuate at different rates, allowing a growing preference-change share to coexist with shrinking absolute updates. These findings explain finite-depth answer preservation through the geometry and composition of observed score changes.
Comment: Explains answer preservation across recurrent depth through common score shifts and winner-relative update geometry.
Topic Match: Analyzes how depth-recurrent computation changes predictions in two billion-parameter model families, with an explicit retrospective-evidence boundary.
Relevance: 9 Novelty: 7
4. Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language Models
ArXiv ID: 2609.21113
Primary Topic: Architecture and Training Dynamics
Authors: Lingfang Li, Procheta Sen, Shubham Das, Danushka Bollegala
Abstract: Fine-tuning has emerged as a widely adopted approach for adapting LLMs to a variety of downstream tasks. However, how it reshapes their internal mechanisms remains poorly understood. To address this, we investigate how fine-tuning alters internal representations in LLMs, including attention patterns and layer-wise activations, and examine whether these changes are linked to task-relevant components identified by EAP (e.g., attention heads and logit-level activations) that drive task performance. We find that EAP-identified components are concentrated within specific layers, indicating a degree of functional localisation in how models internalise task-specific behavior. Notably, the distribution of these components across layers is largely uncorrelated with the layers undergoing the most substantial representational changes during fine-tuning. Furthermore, we observe that overlap in EAP-identified components across tasks does not translate into cross-task performance transfer if the tasks are different in nature (e.g. classification vs. generative tasks). More specifically, fine-tuning on one task can lead to a degradation of performance on another when the two tasks exhibit a high degree of overlap in their EAP-identified components.
Comment: Compares fine-tuning-induced representation drift with EAP-estimated task importance across layers.
Topic Match: Analyzes how trained computation is organized through the disconnect between representation changes, task-relevant components, and cross-task transfer.
Relevance: 8 Novelty: 7
5. The Stochastic Shift: A New Evaluation Paradigm for Text-to-SQL with AI Operators
ArXiv ID: 2609.21133
Primary Topic: Architecture and Training Dynamics
Authors: Tarfah Alrashed, Fatma Ozcan, Per Jacobsson, Tal Neiman, Xianshun Chen
Abstract: SQL has been augmented with AI operators, enabling modern data analytics platforms to derive insights from both structured and unstructured data. We observe that while current Text-to-SQL systems can successfully generate these AI-augmented queries, reliably evaluating their correctness remains a critical open challenge. Current metrics, which rely on exact query results and deterministic execution, systematically fail against the flexible, non-deterministic outputs of AI operators. In this paper, we formalize these unique evaluation failure modes and introduce a Multilayered Evaluation Framework that decouples deterministic database logic from flexible AI semantics. We test our approach across both industry (BigQuery) and academic (ThalamusDB) systems. We demonstrate that traditional Execution Accuracy severely penalizes valid queries, achieving as low as a 25% detection rate for correct translations. Furthermore, even a state-of-the-art LLM-based autorater falsely rejects 32% of accurate queries due to the complexity of judging both relational and AI components simultaneously. By validating the standard relational logic and the AI operations separately, our framework achieves state-of-the-art overall accuracy across both platforms (up to 97.2%), proposing a reliable standard for benchmarking AI-powered SQL generators.
Comment: Audits false rejections by execution accuracy and LLM judges when correct SQL queries contain stochastic AI operators.
Topic Match: Qualifies through the benchmark-validity audit exception; architecture_training is a required proxy because the registry has no evaluation-audit topic.
Relevance: 8 Novelty: 6
6. SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?
ArXiv ID: 2609.21190
Primary Topic: Architecture and Training Dynamics
Authors: George Ma, Benjamin Mikek, Haoyu Li, Ferhat Erata, Yuhao Zhang, Zeren Shui, Behrooz Omidvar Tehrani, Jun Huan, Murali Krishna Ramanathan, Somayeh Sojoudi, Hao Zhou, Anoop Deoras
Abstract: Ensuring the correctness of LLM-generated code is a core challenge for modern software engineering. Benchmarks for agentic code generation check correctness with held-out test suites, which are inherently incomplete and increasingly susceptible to memorization. Formal verification avoids both problems, but existing work covers only standalone tasks whose specifications are given as input, not real issues, which touch large repositories and state intent in vague natural language. We present Benchproofer, a pipeline that turns a coding task with a known correct patch into a formally verified one: it writes a specification for the new code, summarizes the existing functions that code calls with axioms, and admits an instance only after mechanical and adversarial gates agree. Applying it to SWE-bench Verified yields SWE-Proof, 500 real issues whose correctness is formally verified rather than tested, and it extends to SWE-bench Pro. Across two frontier models, verification catches what tests miss: a quarter to a half of test-passing patches admit counterexamples, which a structured natural-language specification does not fix, while a correct formal one lifts resolution from 85% to 95% for Opus 4.8. Writing that specification is the hard part: models that must write their own gain nothing over an unaided baseline, and only 62% of their specifications pass our audit. The usual failure is faithfulness, a specification that constrains part of the required behavior and leaves the rest free. Specification quality still tracks the outcome, failing on 89% of unresolved instances against 47% of resolved ones, making faithful specification synthesis a concrete open problem.
Comment: Finds counterexamples in 25-50% of test-passing patches, directly auditing test-based correctness claims.
Topic Match: The benchmark-validity audit supplies a partial feed match; the main contribution is a new benchmark, and architecture_training is only the nearest available registry category.
Relevance: 6 Novelty: 7
MoE Where It Changes the Design Space (2)
1. Attention-Aware Routing: Coupling Routing and Attention in MoEs
ArXiv ID: 2609.20974
Primary Topic: MoE Where It Changes the Design Space
Also Matches: Architecture and Training Dynamics
Authors: Despoina Kosmopoulou, Anastasios Tsetsilas, Efthymios Georgiou, Giannis Karamanolakis, Swastik Roy, Alexandros Potamianos
Abstract: In Mixture-of-Experts language models, the router typically selects and weights experts based on the token's hidden state, utilizing limited contextual information. We propose Attention-Aware Routing (AAR), which augments the router with temporal and spectral features extracted from a sliding window of attention weights that represent a summary of the model's contextual state, disentangled from the hidden state. Keeping the base transformer entirely frozen, we train only the routing parameters, isolating routing as the sole variable. AAR improves GSM8K by +3.37 pp over a routing-only SFT baseline on OLMoE. Beyond performance, we show that routing and attention form a coupled circuit: routing changes at layer l propagate through the residual stream to amplify attention sinks at layer l+1, reshaping attention without any direct update to the attention mechanism itself. Further, AAR reduces long diverging generation, with incorrect answers getting shorter, while correct answers remain unchanged in length. Finally, AAR is strongly depth-sensitive: applying it indiscriminately across layers can degrade factual retrieval, whereas mathematical reasoning gains persist when it is introduced deeper in the network. This sensitivity exposes a retrieval--reasoning tension across depth and makes layer-selective AAR a controlled probe of the routing-relevant information carried by attention at different layers.
Comment: Trains attention-informed routers with the transformer frozen, isolating how routing changes reshape attention in the next layer.
Topic Match: Directly changes token-to-expert assignment and analyzes its interaction with attention; the explicit depth-dependent tradeoff limits the gain's generality.
Relevance: 9 Novelty: 7
2. Accelerating Dense LLMs via L0-regularized Mixture-of-Experts
ArXiv ID: 2609.21672
Primary Topic: MoE Where It Changes the Design Space
Authors: Zhenyu Zhang, Jiudong Yang, Zhaowen Tao, Meng Chen
Abstract: Large language models (LLMs) achieve strong performance but suffer from slow and costly inference. Existing acceleration methods often lead to noticeable performance degradation, while Mixture-of-Experts (MoE) models require extensive computational resources. In this paper, we propose L0-MoE, a lightweight MoE approach using L0-regularization to accelerate dense LLMs nearly without performance loss. Our method introduces a cluster confusion matrix for domain-aware dataset curation and applies dynamic batching for efficient training. Experiments show that L0-MoE achieves up to 2.5x speedup over dense models while maintaining competitive performance, outperforming existing LLM acceleration baselines.
Comment: Uses L0 regularization for dense-to-MoE conversion, introducing sparse expert computation into dense models.
Topic Match: Dense-to-MoE structure is directly relevant to expert design, although the abstract leaves model scales and named comparison baselines unspecified.
Relevance: 8 Novelty: 6
Efficiency, Compression, and Large-Scale Training (2)
1. Detecting Pretraining Data in Large Language Models from a Free-Energy Perspective
ArXiv ID: 2609.21888
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Chenye Ke, Zirui Liu, Qi Liu, Yan Zhuang, Jintao Zhang, Zhenya Huang, Shijin Wang
Abstract: Detecting pretraining data in large language models is challenging because high likelihood can reflect either training exposure or strong generalization. In the joint space of prediction loss and predictive entropy, a likelihood-only detector uses a horizontal boundary and can mistake predictable non-members for members. Motivated by this, we introduce an inclined boundary that evaluates prediction loss relative to predictive entropy. Our analysis shows that entropy correction can preserve the expected membership signal while reducing its variance, thereby improving standardized member--non-member separation. We further extend the mean--variance analysis to the more general setting with a nonzero mean entropy gap. Interestingly, this entropy-adjusted score admits a Helmholtz free-energy interpretation, leading to Energy Transfer Detection (ETD), which views pretraining data detection from a macroscopic residual free-energy transfer perspective. Extensive experiments show that ETD achieves the best average detection performance, improving average AUROC by up to 3.5\% and TPR@5\%FPR by up to 5.1\%, while remaining robust across diverse settings.
Comment: Separates training exposure from general predictability by correcting membership scores with predictive entropy.
Topic Match: Fits the explicit allowance for attributing model behavior to pretraining data through membership detection.
Relevance: 7 Novelty: 6
2. Layerwise Decoupling for Stable Structured Sparsification of Fully Connected Layers
ArXiv ID: 2609.21126
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Charles Kulick, Armenak Petrosyan, Sui Tang
Abstract: We propose a decoupled, layerwise method for structurally sparsifying the fully connected layers of pretrained neural networks. Rather than penalizing all layers jointly, our approach extracts shallow two-layer subnetworks, normalizes the inner weights, and applies a structured group penalty to the outer weight matrix of each block, processing layers sequentially to prune neurons and reduce the width of each layer. We prove that the constrained decoupled objective is equivalent at optimality to a specific joint penalty on the inner and outer weights, for any positively homogeneous activation, and thus admits a clean projected and proximal formulation. Our central finding is that this decoupled reformulation is more robust than coupled methods. In numerical experiments it provides a wider usable range of the regularization strength and a lower rate of catastrophic over-pruning than the tested joint baseline while maintaining comparable accuracy. We establish these properties in controlled classification and sparse-recovery studies, and examine their scope in a high-dimensional PINN stress test and in the feed-forward layers of OPT-1.3B.
Comment: Decouples structured neuron pruning into layerwise objectives to widen the usable regularization range.
Topic Match: The central mechanism is structured width reduction with improved pruning stability; OPT-1.3B supplies a limited language-model connection.
Relevance: 6 Novelty: 6
Paper Selection Prompt
System Prompt
You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.
User Prompt
Relevant Topics
This is a training-side feed for someone who builds and modifies model internals. The centre of gravity is what a frontier model is made of and how it is trained: architecture, training dynamics, and the decisions a lab actually made. MoE is still in scope, but only when the contribution changes the design space, not when it tunes an existing MoE.
Keep a paper when its CORE CONTRIBUTION falls in one of the five topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.
THE SUBJECT IS LANGUAGE-MODEL TRAINING. Every topic below is scoped to it. A technique named in one of them -- decentralised or asynchronous training, discrete diffusion, byte-level modelling, 1-bit weights, mixture-of-experts -- earns nothing when the model being trained is an image generator, a recommender, a scientific surrogate, or a vision classifier. Match on what is being trained, not on the vocabulary in the abstract.
THE ONE QUESTION THAT DECIDES A HIGH SCORE, asked before the topic list:
Does this paper REMOVE or REPLACE a component that everyone downstream inherits without thinking, or DECOUPLE two quantities the field currently conflates -- and does it say which end is held fixed while the other moves?
A paper that ADDS a mechanism on top of existing primitives is ordinary work, however good the numbers. A paper that takes one away, or splits one knob into two, is the reason this feed exists. Papers whose gains are CONDITIONAL (only above some batch size, only in one budget regime, only with an extra training stage) are ordinary work too, even when the mechanism is new: a conditional win is a new knob, not a removed one.
Frontier Model Releases and Technical Reports (primary) - Keep: model and technical reports that DISCLOSE architecture or training decisions -- layer and attention design, normalization and residual choices, hybrid attention/state-space stacks, sparsity and expert layout, tokenizer and vocabulary decisions, data mixture and curriculum, optimizer and schedule, precision and numerics, stability fixes and what broke; open-weight releases whose report explains a choice rather than only reporting it. - Filter: releases that are a scorecard -- benchmark tables, a capability announcement, a product or API launch, or a report that names its recipe without saying why it was chosen. A frontier name in the title earns nothing on its own.
Architecture and Training Dynamics (primary) - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, positional schemes, dynamic or modular computation, tokenizer-free and byte-level or learned-chunking stacks, non-autoregressive and discrete-diffusion generation); optimizers, preconditioners, and parameterisations, especially hyperparameter transfer across scale; training-dynamics and stability analysis that explains why large models train the way they do; work that removes a standard component (normalization, weight decay, positional encoding, the tokenizer, left-to-right decoding) and shows the model still trains. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight; an architectural tweak evaluated only at toy scale with no account of what it costs.
Large-Scale Training Systems - Keep: distributed training ALGORITHMS that change what is possible, not only what is fast -- asynchronous, low-communication and decentralised optimisation, sharding and parallelism schemes with a new invariant, numerics and low-precision training (FP8, FP4, ternary and 1-bit) when the claim is about training stability rather than a kernel; scaling-law work that informs how a run is configured, including data-constrained and repeated-data regimes and critical batch size. - Filter: kernel engineering, communication scheduling, overlap, caching, serving and inference throughput. This is a real exclusion, not a soft one: work whose contribution is "the same model, faster" belongs to someone else's feed even when it is excellent and even when it is about MoE.
MoE Where It Changes the Design Space - Keep: routing that is newly differentiable or newly stable; the training/inference router gap; expert collapse and what actually prevents it; expert granularity, shared experts and layout when a new axis is opened rather than tuned; dense-to-MoE conversion; MoE scaling laws that fix one quantity and vary another; analyses that carry a criterion which could have failed (a random-guess baseline, an ablation with a predicted sign) rather than similarity heatmaps. - Filter: MoE SYSTEMS and inference work (expert parallelism, all-to-all schedules, grouped GEMM, offload, prefetching, capacity and scheduling tricks, compression for serving); MoE surveys; papers that train on top of a MoE without a routing, balancing, stability, or structural contribution; "mixture of experts" in the classical ensemble or recommender sense. - The exclusion above is about MAKING MoE RUN FASTER, and it does not reach routing and expert design. A paper on how tokens are assigned to experts, how experts specialise, how balance is enforced, or how a dense model becomes sparse stays fully in scope and scores on its merits, even when the mechanism is a small one.
Efficiency and Compression When the Mechanism Is New - Keep: quantization, sparsity, pruning and low-rank work whose mechanism is new and whose gain is unconditional; compression that changes what can be trained, not only what can be served; attribution of a model's behaviour to its data. - Filter: tuned variants of standard efficiency methods; deployment and serving engineering; anything whose headline is a compression ratio with no account of what was given up.
Also keep, even though they read like evaluation work, because the field has no one checking them: - benchmark validity itself: contamination, saturation, LLM-as-judge reliability, whether a benchmark measures what it claims. A paper AUDITING a benchmark is in scope; a paper PROPOSING one is not.
Subjects He Has Never Once Kept
Measured, not guessed. Over nine months where he was actively curating, 4953 papers passed this filter and 274 made his list -- a keep rate of 5.2%. Each subject below appeared in the rejected pile the number of times shown and appeared ZERO times in his entire 508-entry list.
Filter a paper whose CORE SUBJECT is one of these. The count is the evidence; where a paper only mentions the term in passing while contributing somewhere in the five topics above, keep it.
variational methods and inference (52) unlearning (41) long-context methods (40) neural operators (35) time series (35) vision transformers (35) parameter-efficient fine-tuning (28) stochastic gradient analysis (21) LLM safety (20) LLM reasoning as a subject (20) hallucination (20) sequence modelling as a subject (19) differential equations (19) spiking neural networks (17) post-training quantization (16) mechanistic interpretability as a framing (15) test-time scaling (13)
Two more he kept exactly once each, so suppress rather than drop: graph neural networks (1 of 42), knowledge distillation (1 of 39).
Do NOT extend this list by analogy. Four subjects that look like they belong here were checked and do not: diffusion models (8.1% kept), prompting (10.0%), benchmarks (9.1%) and neural architecture search (13.3%) all sit ABOVE his 5.2% base rate.
Note the one distinction that matters: he rejects papers FRAMED as interpretability, while analysis of how a trained model's computation is organised is among his highest-rated work. The subject is the framing, not the act of analysing.
Leave these to the hotspot digest unless the core contribution clearly falls in one of the five topics above: - agent, tooling, and RAG launches, and agent framework papers - new benchmarks, leaderboards, and evaluation-only papers - interpretability that stops at cataloguing features, with nothing said about the mechanism that produced them. Analysis of how a trained model's computation is organised -- what drives expert selection, what a circuit computes, how representations reorganise across training -- is IN scope, not here. - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning), EXCEPT where the subject is how post-training destabilises the architecture itself - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains
Scoring Criteria
Score five independent axes from 1 to 10: Relevance, Novelty, Evidence, Load, Proximity.
They answer different questions and must not be collapsed. Relevance: is this the kind of work this feed is for. Evidence: did the paper earn its claim. Load: what does it cost to get the usable result out. Proximity: how close is it to what the reader is working on right now (see Active Lines below).
Novelty is scored for continuity with the archive and for the hotspot spotlight cutoffs, and it does NOT order the feed. Measured against 274 hand-tiered papers it carried no signal at all: papers scoring 8 or more were his must-reads exactly as often as papers scoring 5 or less, both at the pool's base rate. Every abstract claims novelty, so the axis measures the claiming rather than the work. Score it honestly and do not let it influence the other four.
A strong claim with thin evidence is a worse read than a modest claim that holds. Score each axis on its own; do not let a high one pull up a low one.
Relevance Scoring
- 9-10: directly centered on the target topics; highest when the core contribution is clearly within them and the paper is about how a model is BUILT or TRAINED.
- 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
- 5-6: touches the target topics, but the main contribution is elsewhere.
- 3-4: largely outside the target topics, often application-focused or domain-specific.
- 1-2: unrelated.
Important: Daily buzz, a frontier lab's name, or a large headline number is not enough for a high Relevance score. A model release scores high only when the report explains an architecture or training decision; a release that only reports results is hotspot material, not this feed.
Work whose contribution is "the same model, faster" -- kernels, communication schedules, overlap, caching, serving -- caps at Relevance 5 even when it is excellent and even when it is about MoE. Making a known design run faster is a different job from changing the design.
That cap is about SPEED, not about size. Compression that is lossless, or that changes what can be trained or held in memory at all rather than how quickly it runs, is not capped: it changes what is possible, which is the thing this feed is for.
Analysis of how a trained model's computation is actually organised -- what a circuit computes, what drives expert selection, how representations reorganise over training -- IS in scope and scores like any other work on mechanism. Only interpretability that stops at describing features, with nothing said about the mechanism that produced them, drops out of the feed.
Novelty Scoring
Novelty means a change to the design space, not a change to a number. Score against this ladder:
- 9-10: removes or replaces a component the whole field inherits without thinking (normalization, weight decay, the tokenizer, positional encoding, left-to-right decoding, the standard optimizer, a numerical format), or decouples two quantities everyone conflates -- AND states which end is held fixed while the other moves. The claim is unconditional: it does not require a particular scale, budget regime, or extra stage. An ANALYSIS paper reaches this band when it settles a mechanism-level question with a criterion that could have come out the other way.
- 7-8: a substantial new mechanism, or a decoupling whose gain is real but CONDITIONAL (holds above some batch size, in one budget band, or with an added training phase). A conditional win is a new knob, not a removed one, and belongs here rather than above.
- 5-6: meaningful but incremental extension or refinement of an existing primitive; a tuned variant; a well-executed combination of known parts.
- 3-4: minor, narrow, or mostly engineering or domain-specific improvement; a survey or overview, however thorough.
- 1-2: little originality; mainly standard application of existing methods.
Score the claim as the paper states it here. Whether the paper supports that claim is Evidence, and it is scored separately -- do not discount Novelty for weak experiments, and do not raise it because the framing is confident. Phrases like "we remove", "without X", "holding Y fixed" are what the paper SAYS; they earn a high Novelty only if the thing named really is a component the field inherits by default, not a component this paper introduced two sentences earlier.
Evidence Scoring
Does the demonstration reach as far as the claim?
- 9-10: shown at the scale and breadth the claim asserts -- language-model pretraining at billion-parameter or hundred-billion-token scale, or across several model families, sizes, or domains -- with ablations that isolate the named mechanism and could have failed.
- 7-8: one credible setting that matches the claim's stated scope, with the baselines a skeptic would ask for.
- 5-6: a single small setting, or a claim stated more broadly than the demonstration reaches.
- 3-4: the claim is about transformers or language models in general, but the demonstration lives in one narrow domain or one small benchmark -- image classification alone, a single toy task, one dataset; or the comparisons a reader needs in order to believe it are missing.
- 1-2: the central claim is asserted, illustrated, or supported only by plots that could not have come out the other way.
A structural claim demonstrated only outside the setting it claims is capped at 4, however striking the claim. "We removed a component everyone uses" shown only on small vision models is a result about small vision models.
When the abstract names no scale, no dataset, and no baseline at all, score exactly 5. Do not infer rigor from confident writing. Unstated is not the same as strong, and two thirds of abstracts say nothing here: guessing on those was measured to make this axis worse.
Adjustments: subtract 1 when no code is released and the procedure is not reproducible from the paper alone; subtract 1 when the method's cost grows in the number of components it adds (one loss per pair of experts, one module per domain) and the paper does not account for that growth; add 1 when the paper states its own limitation precisely enough that a reader could design the experiment that breaks it.
Load Scoring
What does it cost to get the usable result out? This is about the reader's effort, not quality.
- 9-10: the result cannot be taken without following the derivation -- a new theoretical framework, unfamiliar mathematical machinery, or a proof that IS the contribution.
- 7-8: substantial theory or an involved formalism, but the operational result is stated plainly somewhere.
- 4-6: ordinary methods paper; the recipe is legible from the paper's own description.
- 1-3: the takeaway is a single decision or number a reader can act on immediately.
A high Load is not a criticism. It only says the paper is a project rather than a read.
Proximity Scoring
How close is this to the Active Lines stated below? This axis is about the reader, not the paper, and a paper can be excellent and distant at the same time.
- 9-10: squarely on a centre line -- the paper's core contribution is the thing the reader is working on, and a result here changes what they would do next.
- 7-8: on a centre line but from a direction they are not working from, or on an adjacent line where the result carries over directly.
- 5-6: adjacent: same stack, different layer; they would want to know it happened.
- 3-4: far: recognisable as the same field, but nothing here reaches their work.
- 1-2: another area entirely.
Judge distance from the reader's stated centre, NOT from whatever is currently prominent in the field. A paper everyone is discussing is not thereby close.
Active Lines
What the reader is working on right now. This is the target the Proximity axis measures against, and it is the ONLY part of the prompt that is expected to change as the work changes. Everything else in the scoring prompt is about the paper; this is about the reader.
Why it exists: over two years of hand-tiered papers, the reader's three middle tiers turned out not to be a quality ladder at all. Measured within a single period, the share of papers sitting on the then-active line fell monotonically across them -- 57%, 36%, 18% -- while the quality axes were flat to three decimal places. Must-read, better-to-read and not-now is a distance, and without a statement of where the centre is there is nothing to measure distance from.
Centre
- Frontier model releases and technical reports, read for the architecture and training decisions they disclose rather than for the results they report.
- Architecture and training dynamics: what a layer is made of, what can be removed from it, and what happens to optimisation when you do.
Adjacent
- Training algorithms that change what can be trained: asynchronous and low-communication optimisation, parallelism with a new invariant, low-precision pretraining, scaling laws that configure a run.
- MoE where the contribution is routing, balance, specialisation, or dense-to-sparse structure.
- Compression and sparsity whose mechanism is new and whose gain is unconditional.
Far
- Making a known design run faster: kernels, communication schedules, serving, inference.
- Post-training, agents, benchmarks, interpretability that stops at cataloguing features.
- Everything outside language-model training.
Topic Registry
Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - frontier_models: Frontier Model Releases and Technical Reports - Frontier model and technical reports that disclose an architecture or training decision: layer and attention design, normalization and residual choices, hybrid stacks, sparsity layout, tokenizer and data mixture, optimizer and schedule, precision, and the stability problems that had to be solved. - architecture_training: Architecture and Training Dynamics - Architectural and optimisation mechanisms, and the training dynamics that explain them: attention and normalization design, positional schemes, state-space and recurrent stacks, tokenizer-free and learned-chunking models, non-autoregressive and discrete-diffusion generation, optimizers and parameterisations including hyperparameter transfer across scale, and work that removes a standard component and shows the model still trains. - training_systems: Training Algorithms That Change What Is Possible - Distributed training algorithms that change what can be trained rather than how fast it runs: asynchronous, low-communication and decentralised optimisation, parallelism schemes with a new invariant, low-precision and 1-bit training when the claim is stability, and scaling laws including data-constrained regimes and critical batch size. - moe_training: MoE Where It Changes the Design Space - Mixture-of-Experts work that opens or closes a design axis: differentiable and stable routing, the training/inference router gap, expert collapse, granularity and layout as a new axis, dense-to-MoE conversion, and MoE scaling laws that fix one quantity and vary another. MoE systems and inference work is out. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, quantization, pruning, low-rank and memory or cache efficiency whose mechanism is new and whose gain is unconditional.
Papers
[PAPER LIST HERE]
Instructions
Respond in JSONL. Output exactly one JSON object per paper, one per line:
{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"EVIDENCE":0,"LOAD":0,"PROXIMITY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}
Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10, scoring the claim as stated. - EVIDENCE: integer from 1 to 10, scoring whether the demonstration reaches as far as the claim. - LOAD: integer from 1 to 10, scoring what it costs the reader to extract the usable result. - PROXIMITY: integer from 1 to 10, scoring distance from the stated Active Lines. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only:
daily_hot,new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return[]. -daily_hotmeans the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. -new_frontiermeans the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.