Personalized Monthly Topic Summary 2025/01
| Metric | Value |
|---|---|
| Total Papers | 4 |
| Frontier Model Releases and Technical Reports | 0 |
| Architecture and Training Dynamics | 2 |
| Training Algorithms That Change What Is Possible | 0 |
| MoE Where It Changes the Design Space | 2 |
| Efficiency, Compression, and Large-Scale Training | 0 |
| Representation Learning Theory and Structure | 0 |
| Memory Structures and Agent Memory Systems | 0 |
| World Models, Exploration, and Open-Ended Reinforcement Learning | 0 |
Architecture and Training Dynamics (2)
-
Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling - Score: 20 (R=10, N=10) - Date: 2025-01-29 - Comment: Decouples input and output vocabularies, scaling multi-gram input capacity while keeping the output vocabulary fixed.
-
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models - Score: 19 (R=10, N=9) - Date: 2025-01-20 - Comment: Replaces the fixed subword vocabulary with character-to-word encoding and character decoding, demonstrated at up to 7B parameters.
MoE Where It Changes the Design Space (2)
-
Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models - Score: 17 (R=9, N=8) - Date: 2025-01-22 - Comment: Varies MoE sparsity under parameter or training-compute constraints to separate stored capacity from active compute.
-
Autonomy-of-Experts Models - Score: 19 (R=10, N=9) - Date: 2025-01-23 - Comment: Removes the standalone MoE router and selects experts using their own internal activation norms.