Personalized Monthly Topic Summary 2026/04
| Metric | Value |
|---|---|
| Total Papers | 14 |
| Frontier Model Releases and Technical Reports | 0 |
| Architecture and Training Dynamics | 12 |
| Training Algorithms That Change What Is Possible | 0 |
| MoE Where It Changes the Design Space | 1 |
| Efficiency, Compression, and Large-Scale Training | 1 |
| Representation Learning Theory and Structure | 0 |
| Memory Structures and Agent Memory Systems | 0 |
| World Models, Exploration, and Open-Ended Reinforcement Learning | 0 |
Architecture and Training Dynamics (12)
-
Graph Memory Transformer (GMT) - Score: 19 (R=10, N=9) - Date: 2026-04-29 - Comment: Replaces every dense decoder FFN with graph-mediated state transitions, demonstrated in an 82.2M-parameter language model.
-
Tucker Attention: A generalization of approximate attention mechanisms - Score: 18 (R=10, N=8) - Date: 2026-04-01 - Comment: Unifies MHA, GQA and MLA through tensor factorization and derives a lower-parameter attention scheme.
-
Dual Triangle Attention: Effective Bidirectional Attention Without Positional Embeddings - Score: 19 (R=10, N=9) - Date: 2026-04-22 - Comment: Replaces explicit bidirectional positional embeddings with complementary directional attention masks without adding parameters.
-
A Mechanistic Analysis of Looped Reasoning Language Models - Score: 19 (R=10, N=9) - Date: 2026-04-14 - Comment: Explains how normalization, input injection, and recurrent-block size influence layer-specific fixed points in looped language models.
-
Rethinking Token Prediction: Tree-Structured Diffusion Language Model - Score: 19 (R=10, N=9) - Date: 2026-04-07 - Comment: Replaces full-vocabulary diffusion prediction with a token tree, halving peak training memory at fixed parameter budget.
-
Does Your Optimizer Care How You Normalize? Normalization-Optimizer Coupling in LLM Training - Score: 19 (R=10, N=9) - Date: 2026-04-03 - Comment: Factorial controls isolate Derf's loss penalty under Muon and test fixes that preserve activation-scale information.
-
EMA Is Not All You Need: Mapping the Boundary Between Structure and Content in Recurrent Context - Score: 18 (R=9, N=9) - Date: 2026-04-13 - Comment: Holding EMA traces fixed while replacing the predictor isolates information loss in the recurrent representation.
-
Architecture Determines Observability in Transformers - Score: 18 (R=9, N=9) - Date: 2026-04-29 - Comment: Tracks training-induced erasure of confidence-independent decision signals across controlled layer and head configurations.
-
Sensitivity-Positional Co-Localization in GQA Transformers - Score: 18 (R=9, N=9) - Date: 2026-04-10 - Comment: Finds anti-localization between task-sensitive and RoPE-influential layers, with crossed interventions testing their interaction.
-
Nexus: Same Pretraining Loss, Better Downstream Generalization via Common Minima - Score: 19 (R=10, N=9) - Date: 2026-04-13 - Comment: Improves downstream generalization at the same pretraining loss by encouraging agreement between data-source gradients.
-
Neural Continuous-Time Markov Chain: Discrete Diffusion via Decoupled Jump Timing and Direction - Score: 18 (R=10, N=8) - Date: 2026-04-20 - Comment: Separately parameterizes reverse-diffusion jump timing and direction, yielding a factored training objective for text generation.
-
Representational Curvature Modulates Behavioral Uncertainty in Large Language Models - Score: 18 (R=9, N=9) - Date: 2026-04-29 - Comment: Links training-emergent trajectory curvature to token entropy using aligned interventions and geometrically misaligned controls.
MoE Where It Changes the Design Space (1)
- Training-Free Dynamic Upcycling of Expert Language Models - Score: 18 (R=9, N=9) - Date: 2026-04-01 - Comment: Replaces additional multitask upcycling optimization with a closed-form ridge solution for combining existing dense experts.
Efficiency, Compression, and Large-Scale Training (1)
- BASIS: Balanced Activation Sketching with Invariant Scalars for "Ghost Backpropagation" - Score: 18 (R=9, N=9) - Date: 2026-04-21 - Comment: At fixed sketch rank, claims activation memory independent of batch and sequence length while preserving exact error propagation.