Personalized Monthly Topic Summary 2026/03
| Metric | Value |
|---|---|
| Total Papers | 22 |
| Frontier Model Releases and Technical Reports | 1 |
| Architecture and Training Dynamics | 12 |
| Training Algorithms That Change What Is Possible | 3 |
| MoE Where It Changes the Design Space | 5 |
| Efficiency, Compression, and Large-Scale Training | 1 |
| Representation Learning Theory and Structure | 0 |
| Memory Structures and Agent Memory Systems | 0 |
| World Models, Exploration, and Open-Ended Reinforcement Learning | 0 |
Frontier Model Releases and Technical Reports (1)
- A Family of LLMs Liberated from Static Vocabularies - Score: 19 (R=10, N=9) - Date: 2026-03-17 - Comment: Replaces fixed vocabulary embeddings and output tables with byte-to-word encoding and word-to-byte decoding.
Architecture and Training Dynamics (12)
-
Lost in the Middle at Birth: An Exact Theory of Transformer Position Bias - Score: 19 (R=10, N=9) - Date: 2026-03-12 - Comment: Uses initialization and RoPE controls to attribute position bias to causal masking and residual connections.
-
The Geometric Cost of Normalization: Affine Bounds on the Bayesian Complexity of Neural Networks - Score: 18 (R=9, N=9) - Date: 2026-03-31 - Comment: LayerNorm's centering reduces the next weight matrix's local learning coefficient by exactly m/2, while RMSNorm preserves it.
-
Rethinking Language Model Scaling under Transferable Hypersphere Optimization - Score: 19 (R=10, N=9) - Date: 2026-03-31 - Comment: HyperP transfers one base learning rate across width, depth, token budget, and MoE granularity under fixed-norm Muon optimization.
-
The Discrete Charm of the MLP: Binary Routing of Continuous Signals in Transformer Feed-Forward Layers - Score: 19 (R=10, N=9) - Date: 2026-03-12 - Comment: Binary-versus-continuous controls and conditional MLP ablations test whether feed-forward layers act as routing gates.
-
Rethinking Attention Output Projection: Structured Hadamard Transforms for Efficient Transformers - Score: 19 (R=10, N=9) - Date: 2026-03-10 - Comment: Replaces attention's learned dense output projection with a fixed Hadamard transform and diagonal affine map.
-
Beyond Masks: Efficient, Flexible Diffusion Language Models via Deletion-Insertion Processes - Score: 19 (R=10, N=9) - Date: 2026-03-27 - Comment: Replaces masked diffusion with deletion-insertion processes, removing mask and padding computation while supporting native variable-length generation.
-
Functional Component Ablation Reveals Specialization Patterns in Hybrid Language Model Architectures - Score: 19 (R=10, N=9) - Date: 2026-03-25 - Comment: Matched-control ablations identify recurrent components as the main computational backbone in two hybrid language models.
-
Attention Sinks Induce Gradient Sinks - Score: 19 (R=10, N=9) - Date: 2026-03-19 - Comment: V-scale preserves attention sinks while suppressing massive activations, testing a gradient-mediated training mechanism.
-
Ablate and Rescue: A Causal Analysis of Residual Stream Hyper-Connections - Score: 19 (R=10, N=9) - Date: 2026-03-17 - Comment: Causal stream ablation and rescue distinguish redundancy from asymmetric utilization in a 780M mHC language model.
-
The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks - Score: 19 (R=10, N=9) - Date: 2026-03-06 - Comment: Pre-norm ablation separates the global function of massive activations from the local function of attention sinks.
-
Taming Momentum: Rethinking Optimizer States Through Low-Rank Approximation - Score: 18 (R=10, N=8) - Date: 2026-03-02 - Comment: Recasts momentum EMA as online linear regression to construct a low-rank optimizer for Llama pretraining.
-
Half the Nonlinearity Is Wasted: Measuring and Reallocating the Transformer's MLP Budget - Score: 18 (R=10, N=8) - Date: 2026-03-05 - Comment: Tests when transformer MLP nonlinearity can be replaced by linear maps, using contextual routing and layerwise interventions.
Training Algorithms That Change What Is Possible (3)
-
Problems with Chinchilla Approach 2: Systematic Biases in IsoFLOP Parabola Fits - Score: 18 (R=9, N=9) - Date: 2026-03-25 - Comment: Identifies systematic parameter-allocation bias at fixed training compute and uses variable projection to fit the full scaling law.
-
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training - Score: 18 (R=10, N=8) - Date: 2026-03-12 - Comment: Identifies rank-one activation mean bias as a source of FP4 instability and counteracts it with mean subtraction.
-
Attn-QAT: 4-Bit Attention With Quantization-Aware Training - Score: 17 (R=9, N=8) - Date: 2026-03-03 - Comment: Stabilizes FP4 attention by matching backward recomputation precision and correcting gradient precision assumptions.
MoE Where It Changes the Design Space (5)
-
Holistic Scaling Laws for Optimal Mixture-of-Experts Architecture Optimization - Score: 19 (R=10, N=9) - Date: 2026-03-24 - Comment: Separates MoE capacity from compute through joint constraints on FLOPs per token, active parameters, and total parameters.
-
Path-Constrained Mixture-of-Experts - Score: 19 (R=10, N=9) - Date: 2026-03-19 - Comment: Shares routers across consecutive layers, constraining expert paths and removing auxiliary load-balancing losses.
-
Expert Threshold Routing for Autoregressive Language Modeling with Dynamic Computation Allocation and Load Balancing - Score: 19 (R=10, N=9) - Date: 2026-03-13 - Comment: Replaces fixed top-k routing and auxiliary balancing losses with causal per-expert thresholds shared by training and inference.
-
Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation - Score: 18 (R=9, N=9) - Date: 2026-03-06 - Comment: Reuses a layer-agnostic expert pool to trade depth for virtual width at a fixed per-token activation budget.
-
Optimal Expert-Attention Allocation in Mixture-of-Experts: A Scalable Law for Dynamic Model Design - Score: 17 (R=9, N=8) - Date: 2026-03-12 - Comment: Models optimal expert-versus-attention compute allocation under fixed total budgets while accounting for sparsity.
Efficiency, Compression, and Large-Scale Training (1)
- The Residual Stream Is All You Need: On the Redundancy of the KV Cache in Transformer Inference - Score: 18 (R=9, N=9) - Date: 2026-03-23 - Comment: Replaces stored KV tensors with residual checkpoints while claiming bit-identical outputs, reducing stored state at fixed decoding fidelity.