← Previous Summary | Monthly Overview | Next Summary →
2025-02 | 2025-03 | 2026-03

Personalized Monthly Topic Summary 2025/03

MetricValue
Total Papers18
Frontier Model Releases and Technical Reports1
Architecture and Training Dynamics12
Training Algorithms That Change What Is Possible2
MoE Where It Changes the Design Space2
Efficiency, Compression, and Large-Scale Training1
Representation Learning Theory and Structure0
Memory Structures and Agent Memory Systems0
World Models, Exploration, and Open-Ended Reinforcement Learning0

Frontier Model Releases and Technical Reports (1)

  1. RWKV-7 "Goose" with Expressive Dynamic State Evolution - Score: 18 (R=10, N=8) - Date: 2025-03-19 - Comment: Generalized delta updates with vector-valued gates expand recurrent state expressivity while preserving parallelizable training.

Architecture and Training Dynamics (12)

  1. Transformers without Normalization - Score: 20 (R=10, N=10) - Date: 2025-03-14 - Comment: Replaces normalization with an element-wise learned tanh operation while reporting preserved or improved Transformer performance.

  2. SuperBPE: Space Travel for Language Models - Score: 19 (R=10, N=9) - Date: 2025-03-18 - Comment: Removes word-boundary tokenization constraints while holding model size, vocabulary size, and pretraining compute fixed.

  3. Predictable Scale: Part I -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining - Score: 18 (R=10, N=8) - Date: 2025-03-07 - Comment: Maps optimal pretraining learning rate and batch size against parameter count and training-token budget.

  4. Forgetting Transformer: Softmax Attention with a Forget Gate - Score: 19 (R=10, N=9) - Date: 2025-03-05 - Comment: Data-dependent attention forgetting removes the need for positional embeddings.

  5. Computation Mechanism Behind LLM Position Generalization - Score: 18 (R=9, N=9) - Date: 2025-03-18 - Comment: Finds a learned additive separation of positional relevance and semantic importance in attention logits.

  6. Outlier dimensions favor frequent tokens in language model - Score: 19 (R=10, N=9) - Date: 2025-03-28 - Comment: Identifies final-layer outliers as a frequent-token prediction mechanism and explains the counterweights that suppress it.

  7. Generalized Interpolating Discrete Diffusion - Score: 18 (R=10, N=8) - Date: 2025-03-07 - Comment: Generalizes discrete-diffusion training beyond absorbing masks, allowing generated tokens to be revised.

  8. I Predict Therefore I Am: Is Next Token Prediction Enough to Learn Human-Interpretable Concepts from Data? - Score: 18 (R=9, N=9) - Date: 2025-03-13 - Comment: Derives when next-token training produces linear representations of latent-concept posteriors.

  9. Interpreting the Repeated Token Phenomenon in Large Language Models - Score: 18 (R=9, N=9) - Date: 2025-03-13 - Comment: Identifies the attention-sink circuit disrupted by repetition and tests a targeted repair.

  10. Strategy Coopetition Explains the Emergence and Transience of In-Context Learning - Score: 18 (R=9, N=9) - Date: 2025-03-10 - Comment: Shared subcircuits make context-constrained in-weights learning both enable and eventually displace in-context learning.

  11. (How) Do Language Models Track State? - Score: 18 (R=9, N=9) - Date: 2025-03-05 - Comment: Training interventions select between associative-scan and parity-assisted state-tracking algorithms.

  12. Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries - Score: 18 (R=9, N=9) - Date: 2025-03-03 - Comment: Identifies a promote-then-suppress computation shared by attention and MLPs, tested through token-specific knockouts.

Training Algorithms That Change What Is Possible (2)

  1. Compute Optimal Scaling of Skills: Knowledge vs Reasoning - Score: 17 (R=9, N=8) - Date: 2025-03-14 - Comment: Finds skill-dependent compute-optimal scaling after controlling for pretraining data mixture, changing how model size should be selected.

  2. Training LLMs with MXFP4 - Score: 17 (R=9, N=8) - Date: 2025-03-03 - Comment: Combines stochastic rounding with Hadamard transforms to control MXFP4 gradient variance during pretraining up to 6.7B parameters.

MoE Where It Changes the Design Space (2)

  1. Continual Pre-training of MoEs: How robust is your router? - Score: 18 (R=9, N=9) - Date: 2025-03-10 - Comment: Four large MoEs retain router balance through continual pretraining, including without replay.

  2. Mixture of Lookup Experts - Score: 18 (R=9, N=9) - Date: 2025-03-21 - Comment: Token-embedding-only experts become lookup tables, eliminating their inference-time FFN computation.

Efficiency, Compression, and Large-Scale Training (1)

  1. Language Models May Verbatim Complete TextThey Were Not Explicitly Trained On - Score: 18 (R=9, N=9) - Date: 2025-03-25 - Comment: Removal-and-retraining counterexamples show that verbatim completion does not establish n-gram-defined training membership.