Personalized Monthly Topic Summary 2026/06
| Metric | Value |
|---|---|
| Total Papers | 21 |
| Frontier Model Releases and Technical Reports | 0 |
| Architecture and Training Dynamics | 12 |
| Training Algorithms That Change What Is Possible | 5 |
| MoE Where It Changes the Design Space | 4 |
| Efficiency, Compression, and Large-Scale Training | 0 |
| Representation Learning Theory and Structure | 0 |
| Memory Structures and Agent Memory Systems | 0 |
| World Models, Exploration, and Open-Ended Reinforcement Learning | 0 |
Architecture and Training Dynamics (12)
-
Spectral Scaling Laws of Muon - Score: 19 (R=10, N=9) - Date: 2026-06-03 - Comment: Relates momentum-spectrum scaling to when a fixed Newton–Schulz iteration count stops adequately orthogonalizing particular layers.
-
Low-Rank Attention Residuals - Score: 19 (R=10, N=9) - Date: 2026-06-20 - Comment: Keeps residual values full-width while independently shrinking the keys used for depth routing.
-
CascadeFormer: Depth-Tapered Transformers Motivated by Gradient Fan-in Asymmetry - Score: 19 (R=10, N=9) - Date: 2026-06-26 - Comment: Tests gradient fan-in causally: equalizing gradient norms fails to restore late-layer value, while increasing downstream paths restores it.
-
Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers - Score: 19 (R=10, N=9) - Date: 2026-06-22 - Comment: Removes explicit key representations and halves cached attention state while reporting comparable model quality.
-
On the Residual Scaling of Looped Transformers: Stability and Transferability - Score: 19 (R=10, N=9) - Date: 2026-06-17 - Comment: Separates loop count from unique depth so learning rates transfer across loop counts while unique depth stays fixed.
-
How Linear Is a Transformer Feed-Forward Block? Per-Block Linear Recoverability Is Learned, Not Architectural - Score: 19 (R=10, N=9) - Date: 2026-06-15 - Comment: Exact linear recoverability tests expose learned differences in FFN computation even with activation type and width held fixed.
-
Where does Absolute Position come from in decoder-only Transformers? - Score: 19 (R=10, N=9) - Date: 2026-06-05 - Comment: Identifies causal-mask normalization and the position-zero residual trajectory as sources of absolute-position signals under RoPE.
-
Phase Transitions in Attention: A Bayesian Theory of Copy Head Emergence - Score: 18 (R=9, N=9) - Date: 2026-06-11 - Comment: Predicts first-order copy-head emergence with softmax versus second-order emergence with linear attention.
-
Positional versus Symbolic Attention Heads: Learning Dynamics, RoPE Geometry, and Length Generalization - Score: 18 (R=9, N=9) - Date: 2026-06-01 - Comment: Links learned positional and symbolic head computations to quantitative, testable predictions derived from RoPE geometry.
-
MultiHashFormer: Hash-based Generative Language Models - Score: 19 (R=10, N=9) - Date: 2026-06-29 - Comment: Expands multilingual vocabulary while holding parameter count fixed through hash-signature autoregression.
-
Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining - Score: 19 (R=10, N=9) - Date: 2026-06-25 - Comment: Tests how corpus support frequency controls rule loss during pretraining, including asymmetric responses to destructive and restorative interventions.
-
Massive Activations Are Architecturally Robust: A Controlled Scratch/Commitment Residual Stream Test - Score: 19 (R=10, N=9) - Date: 2026-06-18 - Comment: A matched-loss residual-stream split falsifies the shared read/write explanation for massive activations.
Training Algorithms That Change What Is Possible (5)
-
Repetition Mismatch: Why Data Mixture Experiments Don't Scale and How to Fix Them - Score: 19 (R=10, N=9) - Date: 2026-06-01 - Comment: Holds target data repetition rates fixed while shrinking proxy training budgets, separating repetition effects from scale.
-
Breaking the Bubble: Asynchronous Pipeline Parallel Training with Bounded Weight Inconsistency - Score: 17 (R=9, N=8) - Date: 2026-06-06 - Comment: Bounds asynchronous weight-version drift without parameter stashing while preserving synchronous-training peak memory.
-
Decentralised AI Training and Inference with BlockTrain - Score: 18 (R=9, N=9) - Date: 2026-06-24 - Comment: Replaces end-to-end optimization with independently trained blocks sharing a global target and avoiding full-model optimizer state.
-
Unifying Local Communications and Local Updates for LLM Pretraining - Score: 17 (R=9, N=8) - Date: 2026-06-10 - Comment: Unifies sparse peer gossip and local adaptive-optimizer steps without globally synchronized model replicas.
-
When Data Is Scarce: Scaling Sparse Language Models with Repeated Training - Score: 17 (R=9, N=8) - Date: 2026-06-01 - Comment: Separates unique tokens, repetition, active parameters and sparsity to determine training choices under fixed data budgets.
MoE Where It Changes the Design Space (4)
-
How Modular Is a Frontier Mixture-of-Experts? A Pre-registered Causal Test in Which Apparent Expert Modularity Mostly Dissolves - Score: 19 (R=10, N=9) - Date: 2026-06-24 - Comment: Tests expert specialization against size-matched random ablations using preregistered selectivity criteria.
-
From Observation to Intervention: A Causal Audit of Expert Importance in Mixture-of-Experts Models - Score: 18 (R=9, N=9) - Date: 2026-06-10 - Comment: A matched positive control supports the finding that routing summaries fail to predict causal expert importance in redundant MoEs.
-
cMoLLM at Scale: Horizontal Scaling Laws for Mixture-of-LLMs - Score: 17 (R=9, N=8) - Date: 2026-06-08 - Comment: Fully differentiable routing over entire model streams extends expert sparsity beyond FFNs.
-
Geometric and Stochastic Analysis of Discontinuities in Sparse Mixture-of-Experts - Score: 17 (R=9, N=8) - Date: 2026-06-18 - Comment: Localized smoothing near Top-k switching surfaces directly addresses discontinuous expert routing.