Personalized Monthly Topic Summary 2025/03
| Metric | Value |
|---|---|
| Total Papers | 18 |
| Frontier Model Releases and Technical Reports | 1 |
| Architecture and Training Dynamics | 12 |
| Training Algorithms That Change What Is Possible | 2 |
| MoE Where It Changes the Design Space | 2 |
| Efficiency, Compression, and Large-Scale Training | 1 |
| Representation Learning Theory and Structure | 0 |
| Memory Structures and Agent Memory Systems | 0 |
| World Models, Exploration, and Open-Ended Reinforcement Learning | 0 |
Frontier Model Releases and Technical Reports (1)
- RWKV-7 "Goose" with Expressive Dynamic State Evolution - Score: 18 (R=10, N=8) - Date: 2025-03-19 - Comment: Generalized delta updates with vector-valued gates expand recurrent state expressivity while preserving parallelizable training.
Architecture and Training Dynamics (12)
-
Transformers without Normalization - Score: 20 (R=10, N=10) - Date: 2025-03-14 - Comment: Replaces normalization with an element-wise learned tanh operation while reporting preserved or improved Transformer performance.
-
SuperBPE: Space Travel for Language Models - Score: 19 (R=10, N=9) - Date: 2025-03-18 - Comment: Removes word-boundary tokenization constraints while holding model size, vocabulary size, and pretraining compute fixed.
-
Predictable Scale: Part I -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining - Score: 18 (R=10, N=8) - Date: 2025-03-07 - Comment: Maps optimal pretraining learning rate and batch size against parameter count and training-token budget.
-
Forgetting Transformer: Softmax Attention with a Forget Gate - Score: 19 (R=10, N=9) - Date: 2025-03-05 - Comment: Data-dependent attention forgetting removes the need for positional embeddings.
-
Computation Mechanism Behind LLM Position Generalization - Score: 18 (R=9, N=9) - Date: 2025-03-18 - Comment: Finds a learned additive separation of positional relevance and semantic importance in attention logits.
-
Outlier dimensions favor frequent tokens in language model - Score: 19 (R=10, N=9) - Date: 2025-03-28 - Comment: Identifies final-layer outliers as a frequent-token prediction mechanism and explains the counterweights that suppress it.
-
Generalized Interpolating Discrete Diffusion - Score: 18 (R=10, N=8) - Date: 2025-03-07 - Comment: Generalizes discrete-diffusion training beyond absorbing masks, allowing generated tokens to be revised.
-
I Predict Therefore I Am: Is Next Token Prediction Enough to Learn Human-Interpretable Concepts from Data? - Score: 18 (R=9, N=9) - Date: 2025-03-13 - Comment: Derives when next-token training produces linear representations of latent-concept posteriors.
-
Interpreting the Repeated Token Phenomenon in Large Language Models - Score: 18 (R=9, N=9) - Date: 2025-03-13 - Comment: Identifies the attention-sink circuit disrupted by repetition and tests a targeted repair.
-
Strategy Coopetition Explains the Emergence and Transience of In-Context Learning - Score: 18 (R=9, N=9) - Date: 2025-03-10 - Comment: Shared subcircuits make context-constrained in-weights learning both enable and eventually displace in-context learning.
-
(How) Do Language Models Track State? - Score: 18 (R=9, N=9) - Date: 2025-03-05 - Comment: Training interventions select between associative-scan and parity-assisted state-tracking algorithms.
-
Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries - Score: 18 (R=9, N=9) - Date: 2025-03-03 - Comment: Identifies a promote-then-suppress computation shared by attention and MLPs, tested through token-specific knockouts.
Training Algorithms That Change What Is Possible (2)
-
Compute Optimal Scaling of Skills: Knowledge vs Reasoning - Score: 17 (R=9, N=8) - Date: 2025-03-14 - Comment: Finds skill-dependent compute-optimal scaling after controlling for pretraining data mixture, changing how model size should be selected.
-
Training LLMs with MXFP4 - Score: 17 (R=9, N=8) - Date: 2025-03-03 - Comment: Combines stochastic rounding with Hadamard transforms to control MXFP4 gradient variance during pretraining up to 6.7B parameters.
MoE Where It Changes the Design Space (2)
-
Continual Pre-training of MoEs: How robust is your router? - Score: 18 (R=9, N=9) - Date: 2025-03-10 - Comment: Four large MoEs retain router balance through continual pretraining, including without replay.
-
Mixture of Lookup Experts - Score: 18 (R=9, N=9) - Date: 2025-03-21 - Comment: Token-embedding-only experts become lookup tables, eliminating their inference-time FFN computation.
Efficiency, Compression, and Large-Scale Training (1)
- Language Models May Verbatim Complete TextThey Were Not Explicitly Trained On - Score: 18 (R=9, N=9) - Date: 2025-03-25 - Comment: Removal-and-retraining counterexamples show that verbatim completion does not establish n-gram-defined training membership.