Personalized Monthly Topic Summary 2025/02
| Metric | Value |
|---|---|
| Total Papers | 19 |
| Frontier Model Releases and Technical Reports | 0 |
| Architecture and Training Dynamics | 12 |
| Training Algorithms That Change What Is Possible | 3 |
| MoE Where It Changes the Design Space | 3 |
| Efficiency, Compression, and Large-Scale Training | 1 |
| Representation Learning Theory and Structure | 0 |
| Memory Structures and Agent Memory Systems | 0 |
| World Models, Exploration, and Open-Ended Reinforcement Learning | 0 |
Architecture and Training Dynamics (12)
-
Reasoning with Latent Thoughts: On the Power of Looped Transformers - Score: 19 (R=10, N=9) - Date: 2025-02-25 - Comment: Holds effective depth at kL while replacing kL distinct layers with k layers reused L times.
-
The underlying structures of self-attention: symmetry, directionality, and emergent dynamics in Transformer training - Score: 19 (R=10, N=9) - Date: 2025-02-18 - Comment: Derives how training objectives produce symmetric versus directional structure in self-attention weights.
-
Which Attention Heads Matter for In-Context Learning? - Score: 19 (R=10, N=9) - Date: 2025-02-21 - Comment: Ablations across 12 language models distinguish function-vector and induction-head contributions and track their relationship during training.
-
Systematic Outliers in Large Language Models - Score: 19 (R=10, N=9) - Date: 2025-02-11 - Comment: Explains systematic outliers as implicit context-aware scaling induced by attention softmax.
-
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach - Score: 19 (R=10, N=9) - Date: 2025-02-10 - Comment: Decouples stored parameter count from computation depth by repeatedly applying a pretrained recurrent block.
-
MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections - Score: 18 (R=10, N=8) - Date: 2025-02-19 - Comment: Gives query, key, value, and residual streams separate input-dependent mixtures of earlier layers.
-
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations - Score: 18 (R=9, N=9) - Date: 2025-02-26 - Comment: Tests whether computational sparsity is learned by comparing pretrained transformers with randomized counterparts.
-
Neural Attention: A Novel Mechanism for Enhanced Expressive Power in Transformer Models - Score: 18 (R=9, N=9) - Date: 2025-02-25 - Comment: Replaces dot-product attention scoring with learned feed-forward networks at unchanged attention-matrix dimensions.
-
Continuous Diffusion Model for Language Modeling - Score: 18 (R=10, N=8) - Date: 2025-02-18 - Comment: Connects categorical diffusion to continuous manifold flow and derives simulation-free training.
-
Prediction hubs are context-informed frequent tokens in LLMs - Score: 18 (R=9, N=9) - Date: 2025-02-17 - Comment: Distinguishes frequency-driven prediction hubs from distance-concentration artifacts at the unembedding readout.
-
Understanding Why Adam Outperforms SGD: Gradient Heterogeneity in Transformers - Score: 17 (R=9, N=8) - Date: 2025-02-04 - Comment: Connects Adam's advantage to gradient-norm heterogeneity and its dependence on layer-normalization placement.
-
Norm Growth and Stability Challenges in Localized Sequential Knowledge Editing - Score: 17 (R=9, N=8) - Date: 2025-02-27 - Comment: Localized updates increase weight norms while shrinking and shifting activations, exposing instability in layer balance.
Training Algorithms That Change What Is Possible (3)
-
LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws - Score: 18 (R=9, N=9) - Date: 2025-02-18 - Comment: Finds that data and tokenizer determine loss-to-loss curves while architecture and optimizer choices largely leave them unchanged.
-
SkipPipe: Partial and Reordered Pipelining Framework for Training LLMs in Heterogeneous Networks - Score: 17 (R=9, N=8) - Date: 2025-02-28 - Comment: Derives convergence constraints that allow training microbatches to skip and reorder pipeline stages.
-
QuEST: Stable Training of LLMs with 1-Bit Weights and Activations - Score: 18 (R=10, N=8) - Date: 2025-02-10 - Comment: Introduces a trust gradient estimator that stabilizes LLM training with 1-bit weights and activations.
MoE Where It Changes the Design Space (3)
-
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient - Score: 18 (R=10, N=8) - Date: 2025-02-10 - Comment: Models active parameters, expert count and data jointly to configure MoE training under fixed memory and compute budgets.
-
Scaling Laws for Upcycling Mixture-of-Experts Language Models - Score: 17 (R=9, N=8) - Date: 2025-02-06 - Comment: Models the interaction between dense-pretraining and upcycled-MoE token budgets to identify when reuse beats training from scratch.
-
Mixture of Tunable Experts - Behavior Modification of DeepSeek-R1 at Inference Time - Score: 18 (R=9, N=9) - Date: 2025-02-18 - Comment: Random deactivation and forced activation test whether specific experts causally control localized behavior.
Efficiency, Compression, and Large-Scale Training (1)
- Pivoting Factorization: A Compact Meta Low-Rank Representation of Sparsity for Efficient Inference in Large Language Models - Score: 18 (R=9, N=9) - Date: 2025-02-03 - Comment: Losslessly removes redundancy from low-rank representations, reducing storage while holding the represented layer's function fixed.