← Previous Summary | Monthly Overview | Next Summary →
2026-04 | 2026-06 | 2026-07

Personalized Monthly Topic Summary 2026/06

MetricValue
Total Papers21
Frontier Model Releases and Technical Reports0
Architecture and Training Dynamics12
Training Algorithms That Change What Is Possible5
MoE Where It Changes the Design Space4
Efficiency, Compression, and Large-Scale Training0
Representation Learning Theory and Structure0
Memory Structures and Agent Memory Systems0
World Models, Exploration, and Open-Ended Reinforcement Learning0

Architecture and Training Dynamics (12)

  1. Spectral Scaling Laws of Muon - Score: 19 (R=10, N=9) - Date: 2026-06-03 - Comment: Relates momentum-spectrum scaling to when a fixed Newton–Schulz iteration count stops adequately orthogonalizing particular layers.

  2. Low-Rank Attention Residuals - Score: 19 (R=10, N=9) - Date: 2026-06-20 - Comment: Keeps residual values full-width while independently shrinking the keys used for depth routing.

  3. CascadeFormer: Depth-Tapered Transformers Motivated by Gradient Fan-in Asymmetry - Score: 19 (R=10, N=9) - Date: 2026-06-26 - Comment: Tests gradient fan-in causally: equalizing gradient norms fails to restore late-layer value, while increasing downstream paths restores it.

  4. Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers - Score: 19 (R=10, N=9) - Date: 2026-06-22 - Comment: Removes explicit key representations and halves cached attention state while reporting comparable model quality.

  5. On the Residual Scaling of Looped Transformers: Stability and Transferability - Score: 19 (R=10, N=9) - Date: 2026-06-17 - Comment: Separates loop count from unique depth so learning rates transfer across loop counts while unique depth stays fixed.

  6. How Linear Is a Transformer Feed-Forward Block? Per-Block Linear Recoverability Is Learned, Not Architectural - Score: 19 (R=10, N=9) - Date: 2026-06-15 - Comment: Exact linear recoverability tests expose learned differences in FFN computation even with activation type and width held fixed.

  7. Where does Absolute Position come from in decoder-only Transformers? - Score: 19 (R=10, N=9) - Date: 2026-06-05 - Comment: Identifies causal-mask normalization and the position-zero residual trajectory as sources of absolute-position signals under RoPE.

  8. Phase Transitions in Attention: A Bayesian Theory of Copy Head Emergence - Score: 18 (R=9, N=9) - Date: 2026-06-11 - Comment: Predicts first-order copy-head emergence with softmax versus second-order emergence with linear attention.

  9. Positional versus Symbolic Attention Heads: Learning Dynamics, RoPE Geometry, and Length Generalization - Score: 18 (R=9, N=9) - Date: 2026-06-01 - Comment: Links learned positional and symbolic head computations to quantitative, testable predictions derived from RoPE geometry.

  10. MultiHashFormer: Hash-based Generative Language Models - Score: 19 (R=10, N=9) - Date: 2026-06-29 - Comment: Expands multilingual vocabulary while holding parameter count fixed through hash-signature autoregression.

  11. Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining - Score: 19 (R=10, N=9) - Date: 2026-06-25 - Comment: Tests how corpus support frequency controls rule loss during pretraining, including asymmetric responses to destructive and restorative interventions.

  12. Massive Activations Are Architecturally Robust: A Controlled Scratch/Commitment Residual Stream Test - Score: 19 (R=10, N=9) - Date: 2026-06-18 - Comment: A matched-loss residual-stream split falsifies the shared read/write explanation for massive activations.

Training Algorithms That Change What Is Possible (5)

  1. Repetition Mismatch: Why Data Mixture Experiments Don't Scale and How to Fix Them - Score: 19 (R=10, N=9) - Date: 2026-06-01 - Comment: Holds target data repetition rates fixed while shrinking proxy training budgets, separating repetition effects from scale.

  2. Breaking the Bubble: Asynchronous Pipeline Parallel Training with Bounded Weight Inconsistency - Score: 17 (R=9, N=8) - Date: 2026-06-06 - Comment: Bounds asynchronous weight-version drift without parameter stashing while preserving synchronous-training peak memory.

  3. Decentralised AI Training and Inference with BlockTrain - Score: 18 (R=9, N=9) - Date: 2026-06-24 - Comment: Replaces end-to-end optimization with independently trained blocks sharing a global target and avoiding full-model optimizer state.

  4. Unifying Local Communications and Local Updates for LLM Pretraining - Score: 17 (R=9, N=8) - Date: 2026-06-10 - Comment: Unifies sparse peer gossip and local adaptive-optimizer steps without globally synchronized model replicas.

  5. When Data Is Scarce: Scaling Sparse Language Models with Repeated Training - Score: 17 (R=9, N=8) - Date: 2026-06-01 - Comment: Separates unique tokens, repetition, active parameters and sparsity to determine training choices under fixed data budgets.

MoE Where It Changes the Design Space (4)

  1. How Modular Is a Frontier Mixture-of-Experts? A Pre-registered Causal Test in Which Apparent Expert Modularity Mostly Dissolves - Score: 19 (R=10, N=9) - Date: 2026-06-24 - Comment: Tests expert specialization against size-matched random ablations using preregistered selectivity criteria.

  2. From Observation to Intervention: A Causal Audit of Expert Importance in Mixture-of-Experts Models - Score: 18 (R=9, N=9) - Date: 2026-06-10 - Comment: A matched positive control supports the finding that routing summaries fail to predict causal expert importance in redundant MoEs.

  3. cMoLLM at Scale: Horizontal Scaling Laws for Mixture-of-LLMs - Score: 17 (R=9, N=8) - Date: 2026-06-08 - Comment: Fully differentiable routing over entire model streams extends expert sparsity beyond FFNs.

  4. Geometric and Stochastic Analysis of Discontinuities in Sparse Mixture-of-Experts - Score: 17 (R=9, N=8) - Date: 2026-06-18 - Comment: Localized smoothing near Top-k switching surfaces directly addresses discontinuous expert routing.