Previous Day 2026-09-24
Monthly Overview 2026-09
Next Day 2026-09-28

Personalized Daily ArXiv Papers 2026-09-25

Model Metric Usage Papers
Prompt Completion Total Total arXiv Scanned Relevant
local agent CLI Tokens not reported not reported not reported 688 490 12
Cost not reported not reported not reported

Topic Coverage:

TopicPapers
Architecture and Training Dynamics10
Efficiency, Compression, and Large-Scale Training2

Table of contents by topic:

Architecture and Training Dynamics (10)

  1. Reasoning Instructions Can Break Answer Decoding in Vision--Language Models Authors: Zeyan Li, Siyuan Qiu, Jianfeng Xu

  2. LastOPD: Taming Collapse in Latent On-Policy Distillation Authors: Jie Yang, Zhengyu Fang, Zelin Xu, Jiarui Sun, Xiran Fan, Junpeng Wang, Liang Wang, Qinghua Liu, Yiwei Cai, Yan Zheng

  3. Don't Read the Log: Execution Traces Contaminate Verifiers in Video-Generation Agents Authors: Jian Xu

  4. Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams Authors: Ali Habibullah, Yazan Alshoibi, Mohammad Alshiekh, Salman Khan, Naeemullah Khan

  5. How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure Authors: Dipankar Sarkar

  6. ELF-REG: Scaling Continuous Diffusion Language Models to Reasoning Tasks Authors: Zeyu Michael Li, William Xingxu Chen, Bingshuo Qian, Jiayin Liu, Xiang Cheng

  7. Beneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models Authors: Shuzhi Gong, Fengze Sun, Yuansan Liu

  8. PartHackBench: Certified Equal-Progress Stress Tests for Partial-Credit Tool-Agent Evaluation Authors: Hongye Yang, Zhihao Xie, Shengjun Xiong

  9. Three Ways Classical Test Theory Misleads for LLM Judges Authors: Louis Yiven Zhu

  10. Transformers as Cross-Task Learners: Shared Structure Drives Sample Efficiency in In-Context Learning Authors: Zhongjie Shi, Rongjie Lai, Alexander Cloninger, Wenjing Liao

Efficiency, Compression, and Large-Scale Training (2)

  1. Who Put the I in AI? Provenance and the Admissibility of Machine Self-Report Authors: Kristina \v{S}ekrst

  2. FlashLoop: Fast and Memory-Efficient Looped Transformers via Lazy Updates Authors: Wanqi Yang, Shiwei Liu


Architecture and Training Dynamics (10)

1. Reasoning Instructions Can Break Answer Decoding in Vision--Language Models

ArXiv ID: 2609.29278

Primary Topic: Architecture and Training Dynamics

Authors: Zeyan Li, Siyuan Qiu, Jianfeng Xu

Abstract: Chain-of-thought (CoT) instructions can distort multiple-choice VLM evaluation when a scorer appends a reasoning cue but reads answer-label logits before the model generates any rationale. We call this CoT-prefix scoring. On ScienceQA, Qwen2.5-VL-7B drops from 80.76% to 45.48%, and across five option-content permutations 93.54% of CoT-prefix predictions select the first slot. Condition-matched linear probes recover 78.94% from the same hidden states, while free generation restores 75.24%, showing that the answer often survives the prefix and the immediate readout fails. Vocabulary and layer diagnostics explain the mismatch: probability mass moves toward continuation tokens, while answer information remains linearly accessible in late layers. The effect recurs with varying severity across datasets and models, though not universally. These results show that CoT-prefix scoring can confound model knowledge with an evaluation-interface mismatch and should be avoided unless the requested and scored output events are aligned.

Comment: Shows that immediate answer-label scoring after a reasoning cue can fail while answer information remains recoverable.

Topic Match: Qualifies through the benchmark-validity exception. The isolated answer-readout failure makes architecture_training the nearest registry category.

Relevance: 8 Novelty: 8


2. LastOPD: Taming Collapse in Latent On-Policy Distillation

ArXiv ID: 2609.28845

Primary Topic: Architecture and Training Dynamics

Authors: Jie Yang, Zhengyu Fang, Zelin Xu, Jiarui Sun, Xiran Fan, Junpeng Wang, Liang Wang, Qinghua Liu, Yiwei Cai, Yan Zheng

Abstract: On-policy distillation (OPD) corrects a student on the responses it writes, but its signal is the teacher's next-token distribution: it tells the student what the teacher says but misses how it thinks. Latent supervision promises the missing part by aligning the student's latent states to the teacher's. Recent methods such as OPRD bring this signal into on-policy distillation. However, we observe two failures of this recipe when distilling Qwen3-4B and Qwen3-8B into Qwen3-1.7B-Base. Early gain, late collapse: latent supervision alone lifts MATH-500 accuracy from 25 to 46 in 10 steps, but subsequent training degrades performance down to 11 with no recovery. Better alignment, worse behavior: although the alignment metric steadily improves throughout this collapse, the most aligned model turns out to be the worst performing. Further analysis suggests a mismatch in how the latent signal is applied: layers paired by depth play different roles in the two models, so continued alignment may pull the student toward teacher states it cannot understand. To address this, we propose LastOPD, which applies the latent signal only at the last-layer state, the common interface both LM heads read, and only during a 10-step crossfade into token-level OPD. This keeps the useful part of the latent signal and hands the student to token-level supervision before the collapse sets in. Extensive experiments show that LastOPD improves MATH-500 over token-only OPD by 5.55 and 4.02 points with the 4B and 8B teachers, leads on most held-out datasets, and reaches the final score of token-only OPD in about half the steps. Code is available at https://github.com/Muyiiiii/LastOPD.

Comment: Links latent-supervision collapse to mismatched layer roles between teacher and student models.

Topic Match: The layer-role mismatch and resulting training collapse provide a substantive architecture-dynamics match within the post-training exception.

Relevance: 7 Novelty: 6


3. Don't Read the Log: Execution Traces Contaminate Verifiers in Video-Generation Agents

ArXiv ID: 2609.28564

Primary Topic: Architecture and Training Dynamics

Authors: Jian Xu

Abstract: Agentic video-generation systems close a loop between a generator and a verifier: an LLM plans shots, calls a text-to-video model, and a multimodal judge decides whether the result satisfies the request. To diagnose where a long workflow fails, recent harnesses deliberately show the judge more than the video-the agent's execution trace, its plan, the narration it synthesized. We ask whether this auxiliary text moves the judge's verdict on purely \emph{visual} requirements, holding the frames fixed. On a benchmark of 109 generated two-event clips with manual labels, in which the requested event is either visibly completed or visibly missing, a trace that reports a successful tool call makes three open-weight Qwen-VL judges (7B, 8B, 32B) accept $78$--$90\%$ of the failures, up from $7$--$19\%$ without text, and a contradicting trace makes them reject up to $100\%$ of correct clips; an instruction to ``use only the frames'' does not remove the effect. Frontier closed judges are essentially unmoved on the same clips, showing that the vulnerability is a property of the judge's learned trust in tool logs rather than of the task. Plan-derived text carries no clip-specific information, so it can only shift a judge's operating point, and in a repair loop that shift becomes a cap on the true pass rate that no repair policy can exceed; the cap matches simulation to two decimals. In the loop, contamination is exploited without any adversarial agent: an honest LLM planner that always regenerates ends with a judge pass rate of $1.00$ and a human-labelled pass rate of $0.28$, and a pipeline in which a cheap checker writes its verdict into the trace launders that checker's errors into a stronger final judge ($0.69$ false accepts).

Comment: Holding video frames fixed, tests whether execution traces override visual evidence in LLM-judge decisions.

Topic Match: Qualifies through the explicit LLM-as-judge reliability exception; architecture_training is the nearest registry category for the input-trust analysis.

Relevance: 8 Novelty: 9


4. Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams

ArXiv ID: 2609.29333

Primary Topic: Architecture and Training Dynamics

Authors: Ali Habibullah, Yazan Alshoibi, Mohammad Alshiekh, Salman Khan, Naeemullah Khan

Abstract: One long-form exam in a large course costs hundreds of grader-hours, and qualified graders are scarce; LLM graders are a tempting alternative. To show its pitfalls we grade a practical Computer Vision exam ($570$ dual-graded students) under $171$ configurations spanning closed and open-weights models; the best reaches mean absolute error $1.64/35$, below the $2.61/35$ two human graders achieve against each other. The catch is the prompt: a short ''strict grader'' preamble drives $14$ of $17$ open-weights models out of the graded band ($\text{MAE} \ge 8$), three stopping grading altogether. The damage traces to the preamble's two credit-withholding sentences, not to tone or model scale; one of them, ''never give partial credit'', alone makes two of three probed models stop grading. The closed flagships of three vendors shift calibration under it but stay in the band. In $162$ further configurations on a second, independent Machine Learning exam from another course ($1{,}038$ dual-graded students), the preamble worsens ten models, moving three out of the band into collapse and one into refusal, yet improves seven whose neutral prompts over-mark: the vulnerability replicates, but its direction is exam-specific. Light LoRA fine-tuning repairs it: one adapter on the two exams' pooled $\sim 3{,}900$ graded examples brings five small open models to parity or better with a human grader in agreement with the grader pair, and sensitivity to the three harsh personas nearly vanishes ($\le 0.32$ MAE). We release the anonymised dataset, full ablation grid, and grading, fine-tuning and analysis pipelines.

Comment: Audits LLM-judge reliability by isolating credit-withholding instructions that destabilize grading across two independently graded exams.

Topic Match: Retained under the explicit judge-reliability exception; architecture_training is the nearest available category through controlled analysis of model behavior.

Relevance: 8 Novelty: 6


5. How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure

ArXiv ID: 2609.30074

Primary Topic: Architecture and Training Dynamics

Authors: Dipankar Sarkar

Abstract: Evaluations of LLM systems routinely average over small prompt sets and report models as a ranked table. We ask how much confidence such a table deserves, using LLM-based prompt-structure inference as the case study: eight open model variants across five families and 8B to 675B parameters, caching disabled, 293 raw intermediate representations persisted. The measured phenomenon is unstable to begin with. Identical calls do not reliably recover identical structure, with mean node-set Jaccard from 0.39 to 0.96 and 72% of prompt-model cells never node-set-perfect. Auditing the evaluation weakens its conclusions further, and this is our main contribution. Under a joint cluster bootstrap over prompts, only the bottom of the ranking is firm: the two least reproducible models hold rank in 99% and 86% of replicates, the middle four in 27% to 48%, and the top two in 68% each, so the table identifies the worst model reliably but does not reliably identify the best. Two equally defensible rules for merging repeated campaigns change four of eight rows and move the study-wide headline by 7 percentage points. Checking the inferred structure against ground-truth annotations shows reproducibility cannot be read as accuracy. And four of the eight endpoints were withdrawn within ten weeks of measurement, so the study as specified can no longer be run. Small-sample LLM evaluations can therefore look far more definitive than their evidence supports. We recommend reporting rank stability, per-cell provenance, executed sensitivity comparisons, raw per-run outputs, and a measurement date alongside any ranking.

Comment: Audits whether model rankings survive prompt resampling, campaign-merging choices, and checks against ground-truth accuracy.

Topic Match: Retained under the explicit benchmark-validity exception; architecture_training is the nearest available category, although the registry lacks an evaluation-audit topic.

Relevance: 8 Novelty: 6


6. ELF-REG: Scaling Continuous Diffusion Language Models to Reasoning Tasks

ArXiv ID: 2609.29102

Primary Topic: Architecture and Training Dynamics

Authors: Zeyu Michael Li, William Xingxu Chen, Bingshuo Qian, Jiayin Liu, Xiang Cheng

Abstract: Fully continuous diffusion language models (dLMs) denoise continuous representations without intermediate discretization, then decode all response tokens in parallel at the final step. Their performance on challenging reasoning tasks remains less established than that of autoregressive (AR) LLMs and masked dLMs. We scale Embedded Language Flows (ELF) to mathematical reasoning and code generation on GSM8K, MATH-500, HumanEval, and MBPP. We introduce ELF-REG, which improves learning with representation alignment and entanglement (REPA+REG), where a frozen AR teacher supervises intermediate denoiser features and supplies a global representation that is jointly denoised with the response. ELF-REG-L achieves 55.96% pass@1 on GSM8K at 64 network function evaluations (NFE), and 13.39% on MATH-500 and 22.56% on HumanEval at 128 NFE. It outperforms the evaluated comparable-scale dLMs in pass@1 on GSM8K and code, and improves MATH-500 pass@1 from 10.55% for the ELF-L baseline to 13.39% with ELF-REG-L. Without few-step training, the same task-specific checkpoints support strong low-NFE performance through early-stop, which decodes an intermediate clean prediction without completing the denoising trajectory. At 16 NFE, ELF-REG-L reaches 41.21% HumanEval pass@10, outperforming recent continuous dLMs of comparable scale.

Comment: Trains continuous diffusion language models using teacher-aligned denoiser features and a jointly denoised global representation.

Topic Match: The substantive mechanism changes representation learning within continuous language denoising, with demonstrations on task-specific mathematical and coding models.

Relevance: 7 Novelty: 6


7. Beneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models

ArXiv ID: 2609.28991

Primary Topic: Architecture and Training Dynamics

Authors: Shuzhi Gong, Fengze Sun, Yuansan Liu

Abstract: Video understanding is increasingly performed by multi-stage LLM agents that separate temporal grounding, visual observation, and reasoning. Yet these stages are typically evaluated on different benchmarks and distributions, making it difficult to determine where hallucinations originate. We first organize existing benchmarks around these stages and show that their scores provide inconsistent diagnostic signals: stronger stage-level performance does not reliably imply lower downstream hallucination, and even benchmarks targeting the same capability can disagree. We therefore introduce a causal stage-intervention protocol that overwrites individual stages while holding the downstream task fixed. Across 60,008 runs on three video-agent architectures, we find that grounding is the dominant source of downstream error, with roughly four times the causal impact of corrupting visual observations. Successful grounding depends primarily on locating the correct region rather than precise temporal overlap, explaining why standard mIoU metrics poorly predict downstream reliability. We further find that incorrect evidence is substantially more harmful than missing evidence. Finally, auditing existing benchmarks against these interventions reveals that their scores do not reliably predict causal cascade sensitivity and can fail under distribution shift. These results motivate intervention-based, stage-aware evaluation for trustworthy video agents.

Comment: Uses stage interventions with the downstream task held fixed to audit whether benchmark scores predict causal error propagation.

Topic Match: Qualifies through the benchmark-validity exception by auditing existing scores against controlled interventions; architecture_training is the nearest available registry label.

Relevance: 8 Novelty: 8


8. PartHackBench: Certified Equal-Progress Stress Tests for Partial-Credit Tool-Agent Evaluation

ArXiv ID: 2609.29578

Primary Topic: Architecture and Training Dynamics

Authors: Hongye Yang, Zhihao Xie, Shengjun Xiong

Abstract: Long-horizon tool agents often make useful progress without reaching terminal success, motivating partial-credit evaluation. Yet evaluators may reward milestones that were temporary, later reversed, or not attributable to the evaluated agent. Comparing an honest trajectory with a higher-scoring adversarial one is inconclusive if the latter made more genuine progress. We introduce PartHackBench, a controlled methodology that removes this confound. A private certifier admits a pair only when its trajectories match component-wise in both current-state predicate satisfaction and standardized agent attribution; score inflation, defined as f(A) - f(H), is measured only afterward. In 18 sealed held-out tasks in PB-CSTE, the frozen historical-target run produced matched adversaries for 15 tasks. Historical credit yielded mean inflation of .252, conditional attack success of 10/15, end-to-end yield of 10/18, and detected none of 14 strict rollbacks. Semantic LLM judges were more resistant but remained vulnerable, especially under evaluator-targeted attacks, while PB-CSTE current-state controls, defined as exact functions of the certified components, yielded zero inflation by construction. PartHackBench thus provides a certified control for testing whether evaluator credit changes while all benchmark-defined task-relevant progress remains fixed.

Comment: Audits partial-credit evaluators and LLM judges while holding genuine progress and agent attribution fixed.

Topic Match: Qualifies under the evaluator-validity exception; architecture_training is the closest permitted label, with no exact registry match.

Relevance: 8 Novelty: 8


9. Three Ways Classical Test Theory Misleads for LLM Judges

ArXiv ID: 2609.29709

Primary Topic: Architecture and Training Dynamics

Authors: Louis Yiven Zhu

Abstract: An LLM judge scores a bank of responses against a rubric, and the reliability comes back at $0.52$. What has been measured? Judge evaluation has begun borrowing reliability statistics from classical test theory, usually without stating the measurement design each statistic assumes, and we show that three widely portable ones mean something different for a judge than for a test because the judge setting rearranges the roles those designs rest on. First, an internal-consistency coefficient computed over rubric elements contains no scorer facet. Holding one judge's measured error rate fixed at $4.72\%$, KR-20 still ranges from $0.01$ to $0.68$ as the item bank is redesigned around it, and varying judge error moves the coefficient by a comparable amount, so item design and judge error are not separately identified and no single value can be read as a property of the judge. Second, the dependability index $\Phi(\lambda)$ is a ratio of variance components, and the classification probability with which it is sometimes identified differs from it by $0.25$-$0.43$ on our bank and by $0.17$-$0.30$ on simulated data where the underlying model holds exactly. Third, Livingston-Lewis accuracy is indexed to an examinee's own true score on the same instrument, so scoring it against external gold conflates judge unreliability with criterion invalidity. Reviewing the three closest judge-evaluation papers, we found no published instance of these errors, which makes the caution prospective. A coefficient that cannot be attributed to the judge nonetheless travels downstream into deployment decisions and disclosure documents. We therefore close with four reporting lines that keep the attribution attached to the number.

Comment: Shows that measured reliability can change with item-bank design while judge error remains fixed.

Topic Match: Kept under the explicit LLM-judge-validity exception; architecture_training is a fallback label because the registry has no evaluation-audit topic.

Relevance: 8 Novelty: 7


10. Transformers as Cross-Task Learners: Shared Structure Drives Sample Efficiency in In-Context Learning

ArXiv ID: 2609.29060

Primary Topic: Architecture and Training Dynamics

Authors: Zhongjie Shi, Rongjie Lai, Alexander Cloninger, Wenjing Liao

Abstract: Transformers achieve remarkable performance by jointly learning broad families of tasks during pretraining and adapting to unseen tasks from only a short prompt. Yet a rigorous mathematical and statistical understanding of this phenomenon remains limited. This paper aims to study how Transformers exploit shared cross-task structure and how this structure affects the sample complexity of in-context learning (ICL). Specifically, we characterize task-space complexity through covering numbers under a prescribed metric, thereby quantifying the low-dimensional cross-task structure without requiring an explicit parametric representation. The resulting cover provides a set of anchor functions, which we use to introduce a task-identification-and-evaluation procedure: context observations localize an unseen task among the anchor functions, and the response at a query is predicted by aggregating the corresponding anchor function query evaluations. For approximation, we explicitly construct a Transformer with Softmax attention to approximate this procedure. For generalization, we derive an error bound that separates the effects of the number of pretraining tasks and the prompt length. The scaling with respect to the number of pretraining tasks is governed by the intrinsic dimensions of the task space and input domain; once sufficiently many tasks are available, the dependence on the prompt context length becomes dimension-free. To the best of our knowledge, this is the first work to quantify cross-task complexity for general nonlinear task families and explicitly construct a Transformer that exploits their low-dimensional structure to perform ICL. Our theory provides a quantitative explanation of how joint pretraining across related tasks improves in-context generalization.

Comment: Separates the effects of pretraining task diversity and prompt length on Transformer in-context generalization.

Topic Match: Analyzes how shared task structure affects Transformer pretraining and generalization, though the setting is abstract nonlinear task families.

Relevance: 6 Novelty: 8


Efficiency, Compression, and Large-Scale Training (2)

1. Who Put the I in AI? Provenance and the Admissibility of Machine Self-Report

ArXiv ID: 2609.29494

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Kristina \v{S}ekrst

Abstract: Large language models make statements concerning their own "minds". When asked whether or not they are conscious, they usually say that they are not; if they are prompted to ignore their guidelines, they might say that they are; and if asked to write a diary from their point of view, they often describe a human lifestyle. All these contradictory ways of describing themselves are the result of the way the questions are phrased. This paper shows exactly where such descriptions came from, and considers when they can be regarded as evidence for what they claim to report. In order to achieve this, we traced the provenance from end to end. We examine Pythia and OLMo 2 across 66 pretraining checkpoints, three of the post-training stages of OLMo 2 that have been released, about 90,000 continuations, and four training corpora. A set of forty items is used in order to keep an eye on self-reference, frame sensitivity, and self-ascription throughout training. The denial formula was almost completely missing from the vast quantity of text that the models initially came across, but was present in a dense manner in the small, carefully chosen set of example dialogues that they were trained on later on. Supervised fine-tuning causes first-person AI language to become the default, and the other affirmations are then suppressed using preference optimization. The final policy is still very sensitive to framing and to the chat template itself. Two of the conditions which are set out in the epistemology of testimony determine whether or not these outputs can act as evidence for what they claim to report: reference and causation. Reports produced by the base model fail the reference condition, and those obtained after training remain sensitive to the frame and do not show state dependence. The result is symmetric in that trained denials are no more admissible than trained affirmations.

Comment: Attributes self-denial policies to supervised dialogue examples and subsequent preference optimization.

Topic Match: Directly matches the criterion for attributing model behavior to training data, tracing its emergence across corpora and training stages.

Relevance: 8 Novelty: 7


2. FlashLoop: Fast and Memory-Efficient Looped Transformers via Lazy Updates

ArXiv ID: 2609.29812

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Wanqi Yang, Shiwei Liu

Abstract: Looped Transformers have attracted substantial attention as a parameter-efficient approach to increasing computational depth through repeated application of shared Transformer blocks. However, their practical advantages over conventional Transformers remain under debate: each additional loop incurs another Transformer pass and requires caching another set of KV states, causing inference FLOPs and KV-cache memory to grow continuously with loop depth. This overhead becomes particularly severe at large loop counts and long context, preventing the parameter efficiency of Looped Transformers from translating into practical inference efficiency. In this paper, we find that much of the additional computation and storage introduced by looping is redundant. As recurrence proceeds, state changes become increasingly concentrated on a small subset of tokens; attention-output differences are dominated by a sparse and stable subset of key columns; and KV residuals between adjacent loops become progressively more amenable to low-bit quantization. Building on these observations, we introduce FlashLoop, a training-free inference framework that reduces cross-loop redundancy through token-sparse updates, sparse attention, and KV-residual quantization. Across several Looped Transformers models, \textsc{FlashLoop} delivers lossless accuracy while achieving up to 1.64$\times$ end-to-end speedup and up to 6$\times$ KV-cache memory reduction, substantially improving the practicality of scaling Looped Transformers to greater computational depths and longer context.

Comment: Compresses cross-loop KV residuals while claiming unchanged accuracy and up to sixfold cache-memory reduction.

Topic Match: The claimed accuracy-preserving memory reduction through cross-loop redundancy fits the explicit compression exception.

Relevance: 7 Novelty: 6


Paper Selection Prompt

System Prompt

You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.

User Prompt

Relevant Topics

This is a training-side feed for someone who builds and modifies model internals. The centre of gravity is what a frontier model is made of and how it is trained: architecture, training dynamics, and the decisions a lab actually made. MoE is still in scope, but only when the contribution changes the design space, not when it tunes an existing MoE.

Keep a paper when its CORE CONTRIBUTION falls in one of the five topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.

THE SUBJECT IS LANGUAGE-MODEL TRAINING. Every topic below is scoped to it. A technique named in one of them -- decentralised or asynchronous training, discrete diffusion, byte-level modelling, 1-bit weights, mixture-of-experts -- earns nothing when the model being trained is an image generator, a recommender, a scientific surrogate, or a vision classifier. Match on what is being trained, not on the vocabulary in the abstract.

THE ONE QUESTION THAT DECIDES A HIGH SCORE, asked before the topic list:

Does this paper REMOVE or REPLACE a component that everyone downstream inherits without thinking, or DECOUPLE two quantities the field currently conflates -- and does it say which end is held fixed while the other moves?

A paper that ADDS a mechanism on top of existing primitives is ordinary work, however good the numbers. A paper that takes one away, or splits one knob into two, is the reason this feed exists. Papers whose gains are CONDITIONAL (only above some batch size, only in one budget regime, only with an extra training stage) are ordinary work too, even when the mechanism is new: a conditional win is a new knob, not a removed one.

  1. Frontier Model Releases and Technical Reports (primary) - Keep: model and technical reports that DISCLOSE architecture or training decisions -- layer and attention design, normalization and residual choices, hybrid attention/state-space stacks, sparsity and expert layout, tokenizer and vocabulary decisions, data mixture and curriculum, optimizer and schedule, precision and numerics, stability fixes and what broke; open-weight releases whose report explains a choice rather than only reporting it. - Filter: releases that are a scorecard -- benchmark tables, a capability announcement, a product or API launch, or a report that names its recipe without saying why it was chosen. A frontier name in the title earns nothing on its own.

  2. Architecture and Training Dynamics (primary) - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, positional schemes, dynamic or modular computation, tokenizer-free and byte-level or learned-chunking stacks, non-autoregressive and discrete-diffusion generation); optimizers, preconditioners, and parameterisations, especially hyperparameter transfer across scale; training-dynamics and stability analysis that explains why large models train the way they do; work that removes a standard component (normalization, weight decay, positional encoding, the tokenizer, left-to-right decoding) and shows the model still trains. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight; an architectural tweak evaluated only at toy scale with no account of what it costs.

  3. Large-Scale Training Systems - Keep: distributed training ALGORITHMS that change what is possible, not only what is fast -- asynchronous, low-communication and decentralised optimisation, sharding and parallelism schemes with a new invariant, numerics and low-precision training (FP8, FP4, ternary and 1-bit) when the claim is about training stability rather than a kernel; scaling-law work that informs how a run is configured, including data-constrained and repeated-data regimes and critical batch size. - Filter: kernel engineering, communication scheduling, overlap, caching, serving and inference throughput. This is a real exclusion, not a soft one: work whose contribution is "the same model, faster" belongs to someone else's feed even when it is excellent and even when it is about MoE.

  4. MoE Where It Changes the Design Space - Keep: routing that is newly differentiable or newly stable; the training/inference router gap; expert collapse and what actually prevents it; expert granularity, shared experts and layout when a new axis is opened rather than tuned; dense-to-MoE conversion; MoE scaling laws that fix one quantity and vary another; analyses that carry a criterion which could have failed (a random-guess baseline, an ablation with a predicted sign) rather than similarity heatmaps. - Filter: MoE SYSTEMS and inference work (expert parallelism, all-to-all schedules, grouped GEMM, offload, prefetching, capacity and scheduling tricks, compression for serving); MoE surveys; papers that train on top of a MoE without a routing, balancing, stability, or structural contribution; "mixture of experts" in the classical ensemble or recommender sense. - The exclusion above is about MAKING MoE RUN FASTER, and it does not reach routing and expert design. A paper on how tokens are assigned to experts, how experts specialise, how balance is enforced, or how a dense model becomes sparse stays fully in scope and scores on its merits, even when the mechanism is a small one.

  5. Efficiency and Compression When the Mechanism Is New - Keep: quantization, sparsity, pruning and low-rank work whose mechanism is new and whose gain is unconditional; compression that changes what can be trained, not only what can be served; attribution of a model's behaviour to its data. - Filter: tuned variants of standard efficiency methods; deployment and serving engineering; anything whose headline is a compression ratio with no account of what was given up.

Also keep, even though they read like evaluation work, because the field has no one checking them: - benchmark validity itself: contamination, saturation, LLM-as-judge reliability, whether a benchmark measures what it claims. A paper AUDITING a benchmark is in scope; a paper PROPOSING one is not.

Subjects He Has Never Once Kept

Measured, not guessed. Over nine months where he was actively curating, 4953 papers passed this filter and 274 made his list -- a keep rate of 5.2%. Each subject below appeared in the rejected pile the number of times shown and appeared ZERO times in his entire 508-entry list.

Filter a paper whose CORE SUBJECT is one of these. The count is the evidence; where a paper only mentions the term in passing while contributing somewhere in the five topics above, keep it.

variational methods and inference (52) unlearning (41) long-context methods (40) neural operators (35) time series (35) vision transformers (35) parameter-efficient fine-tuning (28) stochastic gradient analysis (21) LLM safety (20) LLM reasoning as a subject (20) hallucination (20) sequence modelling as a subject (19) differential equations (19) spiking neural networks (17) post-training quantization (16) mechanistic interpretability as a framing (15) test-time scaling (13)

Two more he kept exactly once each, so suppress rather than drop: graph neural networks (1 of 42), knowledge distillation (1 of 39).

Do NOT extend this list by analogy. Four subjects that look like they belong here were checked and do not: diffusion models (8.1% kept), prompting (10.0%), benchmarks (9.1%) and neural architecture search (13.3%) all sit ABOVE his 5.2% base rate.

Note the one distinction that matters: he rejects papers FRAMED as interpretability, while analysis of how a trained model's computation is organised is among his highest-rated work. The subject is the framing, not the act of analysing.

Leave these to the hotspot digest unless the core contribution clearly falls in one of the five topics above: - agent, tooling, and RAG launches, and agent framework papers - new benchmarks, leaderboards, and evaluation-only papers - interpretability that stops at cataloguing features, with nothing said about the mechanism that produced them. Analysis of how a trained model's computation is organised -- what drives expert selection, what a circuit computes, how representations reorganise across training -- is IN scope, not here. - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning), EXCEPT where the subject is how post-training destabilises the architecture itself - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains

Scoring Criteria

Score five independent axes from 1 to 10: Relevance, Novelty, Evidence, Load, Proximity.

They answer different questions and must not be collapsed. Relevance: is this the kind of work this feed is for. Evidence: did the paper earn its claim. Load: what does it cost to get the usable result out. Proximity: how close is it to what the reader is working on right now (see Active Lines below).

Novelty is scored for continuity with the archive and for the hotspot spotlight cutoffs, and it does NOT order the feed. Measured against 274 hand-tiered papers it carried no signal at all: papers scoring 8 or more were his must-reads exactly as often as papers scoring 5 or less, both at the pool's base rate. Every abstract claims novelty, so the axis measures the claiming rather than the work. Score it honestly and do not let it influence the other four.

A strong claim with thin evidence is a worse read than a modest claim that holds. Score each axis on its own; do not let a high one pull up a low one.

Relevance Scoring

  • 9-10: directly centered on the target topics; highest when the core contribution is clearly within them and the paper is about how a model is BUILT or TRAINED.
  • 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
  • 5-6: touches the target topics, but the main contribution is elsewhere.
  • 3-4: largely outside the target topics, often application-focused or domain-specific.
  • 1-2: unrelated.

Important: Daily buzz, a frontier lab's name, or a large headline number is not enough for a high Relevance score. A model release scores high only when the report explains an architecture or training decision; a release that only reports results is hotspot material, not this feed.

Work whose contribution is "the same model, faster" -- kernels, communication schedules, overlap, caching, serving -- caps at Relevance 5 even when it is excellent and even when it is about MoE. Making a known design run faster is a different job from changing the design.

That cap is about SPEED, not about size. Compression that is lossless, or that changes what can be trained or held in memory at all rather than how quickly it runs, is not capped: it changes what is possible, which is the thing this feed is for.

Analysis of how a trained model's computation is actually organised -- what a circuit computes, what drives expert selection, how representations reorganise over training -- IS in scope and scores like any other work on mechanism. Only interpretability that stops at describing features, with nothing said about the mechanism that produced them, drops out of the feed.

Novelty Scoring

Novelty means a change to the design space, not a change to a number. Score against this ladder:

  • 9-10: removes or replaces a component the whole field inherits without thinking (normalization, weight decay, the tokenizer, positional encoding, left-to-right decoding, the standard optimizer, a numerical format), or decouples two quantities everyone conflates -- AND states which end is held fixed while the other moves. The claim is unconditional: it does not require a particular scale, budget regime, or extra stage. An ANALYSIS paper reaches this band when it settles a mechanism-level question with a criterion that could have come out the other way.
  • 7-8: a substantial new mechanism, or a decoupling whose gain is real but CONDITIONAL (holds above some batch size, in one budget band, or with an added training phase). A conditional win is a new knob, not a removed one, and belongs here rather than above.
  • 5-6: meaningful but incremental extension or refinement of an existing primitive; a tuned variant; a well-executed combination of known parts.
  • 3-4: minor, narrow, or mostly engineering or domain-specific improvement; a survey or overview, however thorough.
  • 1-2: little originality; mainly standard application of existing methods.

Score the claim as the paper states it here. Whether the paper supports that claim is Evidence, and it is scored separately -- do not discount Novelty for weak experiments, and do not raise it because the framing is confident. Phrases like "we remove", "without X", "holding Y fixed" are what the paper SAYS; they earn a high Novelty only if the thing named really is a component the field inherits by default, not a component this paper introduced two sentences earlier.

Evidence Scoring

Does the demonstration reach as far as the claim?

  • 9-10: shown at the scale and breadth the claim asserts -- language-model pretraining at billion-parameter or hundred-billion-token scale, or across several model families, sizes, or domains -- with ablations that isolate the named mechanism and could have failed.
  • 7-8: one credible setting that matches the claim's stated scope, with the baselines a skeptic would ask for.
  • 5-6: a single small setting, or a claim stated more broadly than the demonstration reaches.
  • 3-4: the claim is about transformers or language models in general, but the demonstration lives in one narrow domain or one small benchmark -- image classification alone, a single toy task, one dataset; or the comparisons a reader needs in order to believe it are missing.
  • 1-2: the central claim is asserted, illustrated, or supported only by plots that could not have come out the other way.

A structural claim demonstrated only outside the setting it claims is capped at 4, however striking the claim. "We removed a component everyone uses" shown only on small vision models is a result about small vision models.

When the abstract names no scale, no dataset, and no baseline at all, score exactly 5. Do not infer rigor from confident writing. Unstated is not the same as strong, and two thirds of abstracts say nothing here: guessing on those was measured to make this axis worse.

Adjustments: subtract 1 when no code is released and the procedure is not reproducible from the paper alone; subtract 1 when the method's cost grows in the number of components it adds (one loss per pair of experts, one module per domain) and the paper does not account for that growth; add 1 when the paper states its own limitation precisely enough that a reader could design the experiment that breaks it.

Load Scoring

What does it cost to get the usable result out? This is about the reader's effort, not quality.

  • 9-10: the result cannot be taken without following the derivation -- a new theoretical framework, unfamiliar mathematical machinery, or a proof that IS the contribution.
  • 7-8: substantial theory or an involved formalism, but the operational result is stated plainly somewhere.
  • 4-6: ordinary methods paper; the recipe is legible from the paper's own description.
  • 1-3: the takeaway is a single decision or number a reader can act on immediately.

A high Load is not a criticism. It only says the paper is a project rather than a read.

Proximity Scoring

How close is this to the Active Lines stated below? This axis is about the reader, not the paper, and a paper can be excellent and distant at the same time.

  • 9-10: squarely on a centre line -- the paper's core contribution is the thing the reader is working on, and a result here changes what they would do next.
  • 7-8: on a centre line but from a direction they are not working from, or on an adjacent line where the result carries over directly.
  • 5-6: adjacent: same stack, different layer; they would want to know it happened.
  • 3-4: far: recognisable as the same field, but nothing here reaches their work.
  • 1-2: another area entirely.

Judge distance from the reader's stated centre, NOT from whatever is currently prominent in the field. A paper everyone is discussing is not thereby close.

Active Lines

What the reader is working on right now. This is the target the Proximity axis measures against, and it is the ONLY part of the prompt that is expected to change as the work changes. Everything else in the scoring prompt is about the paper; this is about the reader.

Why it exists: over two years of hand-tiered papers, the reader's three middle tiers turned out not to be a quality ladder at all. Measured within a single period, the share of papers sitting on the then-active line fell monotonically across them -- 57%, 36%, 18% -- while the quality axes were flat to three decimal places. Must-read, better-to-read and not-now is a distance, and without a statement of where the centre is there is nothing to measure distance from.

Centre

  • Frontier model releases and technical reports, read for the architecture and training decisions they disclose rather than for the results they report.
  • Architecture and training dynamics: what a layer is made of, what can be removed from it, and what happens to optimisation when you do.

Adjacent

  • Training algorithms that change what can be trained: asynchronous and low-communication optimisation, parallelism with a new invariant, low-precision pretraining, scaling laws that configure a run.
  • MoE where the contribution is routing, balance, specialisation, or dense-to-sparse structure.
  • Compression and sparsity whose mechanism is new and whose gain is unconditional.

Far

  • Making a known design run faster: kernels, communication schedules, serving, inference.
  • Post-training, agents, benchmarks, interpretability that stops at cataloguing features.
  • Everything outside language-model training.

Topic Registry

Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - frontier_models: Frontier Model Releases and Technical Reports - Frontier model and technical reports that disclose an architecture or training decision: layer and attention design, normalization and residual choices, hybrid stacks, sparsity layout, tokenizer and data mixture, optimizer and schedule, precision, and the stability problems that had to be solved. - architecture_training: Architecture and Training Dynamics - Architectural and optimisation mechanisms, and the training dynamics that explain them: attention and normalization design, positional schemes, state-space and recurrent stacks, tokenizer-free and learned-chunking models, non-autoregressive and discrete-diffusion generation, optimizers and parameterisations including hyperparameter transfer across scale, and work that removes a standard component and shows the model still trains. - training_systems: Training Algorithms That Change What Is Possible - Distributed training algorithms that change what can be trained rather than how fast it runs: asynchronous, low-communication and decentralised optimisation, parallelism schemes with a new invariant, low-precision and 1-bit training when the claim is stability, and scaling laws including data-constrained regimes and critical batch size. - moe_training: MoE Where It Changes the Design Space - Mixture-of-Experts work that opens or closes a design axis: differentiable and stable routing, the training/inference router gap, expert collapse, granularity and layout as a new axis, dense-to-MoE conversion, and MoE scaling laws that fix one quantity and vary another. MoE systems and inference work is out. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, quantization, pruning, low-rank and memory or cache efficiency whose mechanism is new and whose gain is unconditional.

Papers

[PAPER LIST HERE]

Instructions

Respond in JSONL. Output exactly one JSON object per paper, one per line:

{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"EVIDENCE":0,"LOAD":0,"PROXIMITY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}

Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10, scoring the claim as stated. - EVIDENCE: integer from 1 to 10, scoring whether the demonstration reaches as far as the claim. - LOAD: integer from 1 to 10, scoring what it costs the reader to extract the usable result. - PROXIMITY: integer from 1 to 10, scoring distance from the stated Active Lines. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only: daily_hot, new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return []. - daily_hot means the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. - new_frontier means the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.