Personalized Monthly Topic Summary 2026/07
| Metric | Value |
|---|---|
| Total Papers | 15 |
| Frontier Model Releases and Technical Reports | 0 |
| Architecture and Training Dynamics | 12 |
| Training Algorithms That Change What Is Possible | 0 |
| MoE Where It Changes the Design Space | 1 |
| Efficiency, Compression, and Large-Scale Training | 2 |
| Representation Learning Theory and Structure | 0 |
| Memory Structures and Agent Memory Systems | 0 |
| World Models, Exploration, and Open-Ended Reinforcement Learning | 0 |
Architecture and Training Dynamics (12)
-
A Compositional Theory of Causally Masked Transformers - Score: 18 (R=9, N=9) - Date: 2026-07-30 - Comment: Derives how attention type and finite-precision evaluation order determine the memory and expressivity of position-free transformers.
-
Dynamic Parameterization Is Not Dynamic Inference - Score: 19 (R=10, N=9) - Date: 2026-07-29 - Comment: Freezes model weights and reassigns controller coefficients to distinguish learned layer profiles from genuine content-dependent computation.
-
Deep Delta Learning - Score: 19 (R=10, N=9) - Date: 2026-07-28 - Comment: Separates persistent residual-state capacity from backbone compute width through structured delta-rule updates.
-
Mean-to-Score Discrete Diffusion: Posterior-Mean Denoisers for Score Entropy - Score: 19 (R=10, N=9) - Date: 2026-07-24 - Comment: Replaces unconstrained discrete-diffusion score ratios with posterior-derived scores that guarantee Bayes realizability.
-
Muse: Representation Geometry of Muon Beyond Normalized Momentum - Score: 19 (R=10, N=9) - Date: 2026-07-17 - Comment: Holding momentum and the orthogonalization backend fixed isolates how parameter representation determines Muon's update geometry.
-
An Exact Instrument for State Usage in Selective State-Space Models, and the Input-Driven Migration It Reveals - Score: 19 (R=10, N=9) - Date: 2026-07-14 - Comment: Counterfactuals identify the input-dependent write map as the main driver of context-dependent state use.
-
Fingerprint, Not Blueprint: How Positional Schemes Set the Default Spectral Algebra of Attention - Score: 19 (R=10, N=9) - Date: 2026-07-08 - Comment: Constrained training preserves induction by rerouting around spectral bans, distinguishing positional defaults from necessary circuit structure.
-
How Transformers Learn to Plan via Multi-Token Prediction - Score: 18 (R=9, N=9) - Date: 2026-07-27 - Comment: Explains how multi-token gradients induce an end-to-start planning circuit.
-
Toward Manifest Relationality in Transformers via Symmetry Reduction - Score: 18 (R=9, N=9) - Date: 2026-07-22 - Comment: Invariant relational parameterization removes redundant coordinates from transformer representations, attention, and optimization.
-
Sparse Inter-Layer Dependencies of Transformer FFN Neurons - Score: 18 (R=9, N=9) - Date: 2026-07-14 - Comment: Tests sparse inter-layer FFN dependencies by applying neuron-specific masks throughout the network and measuring perplexity preservation.
-
LayerNorm as Implicit Gain Control in Looped Transformers - Score: 18 (R=9, N=9) - Date: 2026-07-13 - Comment: Identifies the carry as a stabilizer and the nonlinear recurrence as the main memory path in pre-LayerNorm looped transformers.
-
Adaptive Depth in Looped Transformers: Diagnosing Learned Halting Gates and Trajectory Readouts - Score: 18 (R=9, N=9) - Date: 2026-07-09 - Comment: Holds recurrent trajectories fixed while varying halting readouts to isolate failures caused by jointly learned depth supervision.
MoE Where It Changes the Design Space (1)
- Detecting a Route Flip Is Easier Than Knowing Whether to Fix It: Causal Route-Mediated Damage in Quantized Mixture-of-Experts - Score: 18 (R=9, N=9) - Date: 2026-07-15 - Comment: A causal decomposition and chance baseline test whether router statistics distinguish harmful from helpful expert flips.
Efficiency, Compression, and Large-Scale Training (2)
-
GaugeQuant: Online Learning of Quantization-Optimal Bases from LLM Symmetries - Score: 17 (R=9, N=8) - Date: 2026-07-27 - Comment: Learns quantization-friendly symmetry rotations while keeping the language-model objective unchanged.
-
Variable Bit-width Quantization: Learning Per-Group Precision for "Bigger-but-Smaller" Language Models - Score: 17 (R=9, N=8) - Date: 2026-07-06 - Comment: Learns per-group bit widths during pretraining to allocate a fixed storage budget unevenly across weights.