← Previous Summary | Monthly Overview | Next Summary →
2025-03 | 2026-03 | 2026-04

Personalized Monthly Topic Summary 2026/03

MetricValue
Total Papers22
Frontier Model Releases and Technical Reports1
Architecture and Training Dynamics12
Training Algorithms That Change What Is Possible3
MoE Where It Changes the Design Space5
Efficiency, Compression, and Large-Scale Training1
Representation Learning Theory and Structure0
Memory Structures and Agent Memory Systems0
World Models, Exploration, and Open-Ended Reinforcement Learning0

Frontier Model Releases and Technical Reports (1)

  1. A Family of LLMs Liberated from Static Vocabularies - Score: 19 (R=10, N=9) - Date: 2026-03-17 - Comment: Replaces fixed vocabulary embeddings and output tables with byte-to-word encoding and word-to-byte decoding.

Architecture and Training Dynamics (12)

  1. Lost in the Middle at Birth: An Exact Theory of Transformer Position Bias - Score: 19 (R=10, N=9) - Date: 2026-03-12 - Comment: Uses initialization and RoPE controls to attribute position bias to causal masking and residual connections.

  2. The Geometric Cost of Normalization: Affine Bounds on the Bayesian Complexity of Neural Networks - Score: 18 (R=9, N=9) - Date: 2026-03-31 - Comment: LayerNorm's centering reduces the next weight matrix's local learning coefficient by exactly m/2, while RMSNorm preserves it.

  3. Rethinking Language Model Scaling under Transferable Hypersphere Optimization - Score: 19 (R=10, N=9) - Date: 2026-03-31 - Comment: HyperP transfers one base learning rate across width, depth, token budget, and MoE granularity under fixed-norm Muon optimization.

  4. The Discrete Charm of the MLP: Binary Routing of Continuous Signals in Transformer Feed-Forward Layers - Score: 19 (R=10, N=9) - Date: 2026-03-12 - Comment: Binary-versus-continuous controls and conditional MLP ablations test whether feed-forward layers act as routing gates.

  5. Rethinking Attention Output Projection: Structured Hadamard Transforms for Efficient Transformers - Score: 19 (R=10, N=9) - Date: 2026-03-10 - Comment: Replaces attention's learned dense output projection with a fixed Hadamard transform and diagonal affine map.

  6. Beyond Masks: Efficient, Flexible Diffusion Language Models via Deletion-Insertion Processes - Score: 19 (R=10, N=9) - Date: 2026-03-27 - Comment: Replaces masked diffusion with deletion-insertion processes, removing mask and padding computation while supporting native variable-length generation.

  7. Functional Component Ablation Reveals Specialization Patterns in Hybrid Language Model Architectures - Score: 19 (R=10, N=9) - Date: 2026-03-25 - Comment: Matched-control ablations identify recurrent components as the main computational backbone in two hybrid language models.

  8. Attention Sinks Induce Gradient Sinks - Score: 19 (R=10, N=9) - Date: 2026-03-19 - Comment: V-scale preserves attention sinks while suppressing massive activations, testing a gradient-mediated training mechanism.

  9. Ablate and Rescue: A Causal Analysis of Residual Stream Hyper-Connections - Score: 19 (R=10, N=9) - Date: 2026-03-17 - Comment: Causal stream ablation and rescue distinguish redundancy from asymmetric utilization in a 780M mHC language model.

  10. The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks - Score: 19 (R=10, N=9) - Date: 2026-03-06 - Comment: Pre-norm ablation separates the global function of massive activations from the local function of attention sinks.

  11. Taming Momentum: Rethinking Optimizer States Through Low-Rank Approximation - Score: 18 (R=10, N=8) - Date: 2026-03-02 - Comment: Recasts momentum EMA as online linear regression to construct a low-rank optimizer for Llama pretraining.

  12. Half the Nonlinearity Is Wasted: Measuring and Reallocating the Transformer's MLP Budget - Score: 18 (R=10, N=8) - Date: 2026-03-05 - Comment: Tests when transformer MLP nonlinearity can be replaced by linear maps, using contextual routing and layerwise interventions.

Training Algorithms That Change What Is Possible (3)

  1. Problems with Chinchilla Approach 2: Systematic Biases in IsoFLOP Parabola Fits - Score: 18 (R=9, N=9) - Date: 2026-03-25 - Comment: Identifies systematic parameter-allocation bias at fixed training compute and uses variable projection to fit the full scaling law.

  2. The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training - Score: 18 (R=10, N=8) - Date: 2026-03-12 - Comment: Identifies rank-one activation mean bias as a source of FP4 instability and counteracts it with mean subtraction.

  3. Attn-QAT: 4-Bit Attention With Quantization-Aware Training - Score: 17 (R=9, N=8) - Date: 2026-03-03 - Comment: Stabilizes FP4 attention by matching backward recomputation precision and correcting gradient precision assumptions.

MoE Where It Changes the Design Space (5)

  1. Holistic Scaling Laws for Optimal Mixture-of-Experts Architecture Optimization - Score: 19 (R=10, N=9) - Date: 2026-03-24 - Comment: Separates MoE capacity from compute through joint constraints on FLOPs per token, active parameters, and total parameters.

  2. Path-Constrained Mixture-of-Experts - Score: 19 (R=10, N=9) - Date: 2026-03-19 - Comment: Shares routers across consecutive layers, constraining expert paths and removing auxiliary load-balancing losses.

  3. Expert Threshold Routing for Autoregressive Language Modeling with Dynamic Computation Allocation and Load Balancing - Score: 19 (R=10, N=9) - Date: 2026-03-13 - Comment: Replaces fixed top-k routing and auxiliary balancing losses with causal per-expert thresholds shared by training and inference.

  4. Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation - Score: 18 (R=9, N=9) - Date: 2026-03-06 - Comment: Reuses a layer-agnostic expert pool to trade depth for virtual width at a fixed per-token activation budget.

  5. Optimal Expert-Attention Allocation in Mixture-of-Experts: A Scalable Law for Dynamic Model Design - Score: 17 (R=9, N=8) - Date: 2026-03-12 - Comment: Models optimal expert-versus-attention compute allocation under fixed total budgets while accounting for sparsity.

Efficiency, Compression, and Large-Scale Training (1)

  1. The Residual Stream Is All You Need: On the Redundancy of the KV Cache in Transformer Inference - Score: 18 (R=9, N=9) - Date: 2026-03-23 - Comment: Replaces stored KV tensors with residual checkpoints while claiming bit-identical outputs, reducing stored state at fixed decoding fidelity.