← Previous Summary | Monthly Overview | Next Summary →
2026-03 | 2026-04 | 2026-06

Personalized Monthly Topic Summary 2026/04

MetricValue
Total Papers14
Frontier Model Releases and Technical Reports0
Architecture and Training Dynamics12
Training Algorithms That Change What Is Possible0
MoE Where It Changes the Design Space1
Efficiency, Compression, and Large-Scale Training1
Representation Learning Theory and Structure0
Memory Structures and Agent Memory Systems0
World Models, Exploration, and Open-Ended Reinforcement Learning0

Architecture and Training Dynamics (12)

  1. Graph Memory Transformer (GMT) - Score: 19 (R=10, N=9) - Date: 2026-04-29 - Comment: Replaces every dense decoder FFN with graph-mediated state transitions, demonstrated in an 82.2M-parameter language model.

  2. Tucker Attention: A generalization of approximate attention mechanisms - Score: 18 (R=10, N=8) - Date: 2026-04-01 - Comment: Unifies MHA, GQA and MLA through tensor factorization and derives a lower-parameter attention scheme.

  3. Dual Triangle Attention: Effective Bidirectional Attention Without Positional Embeddings - Score: 19 (R=10, N=9) - Date: 2026-04-22 - Comment: Replaces explicit bidirectional positional embeddings with complementary directional attention masks without adding parameters.

  4. A Mechanistic Analysis of Looped Reasoning Language Models - Score: 19 (R=10, N=9) - Date: 2026-04-14 - Comment: Explains how normalization, input injection, and recurrent-block size influence layer-specific fixed points in looped language models.

  5. Rethinking Token Prediction: Tree-Structured Diffusion Language Model - Score: 19 (R=10, N=9) - Date: 2026-04-07 - Comment: Replaces full-vocabulary diffusion prediction with a token tree, halving peak training memory at fixed parameter budget.

  6. Does Your Optimizer Care How You Normalize? Normalization-Optimizer Coupling in LLM Training - Score: 19 (R=10, N=9) - Date: 2026-04-03 - Comment: Factorial controls isolate Derf's loss penalty under Muon and test fixes that preserve activation-scale information.

  7. EMA Is Not All You Need: Mapping the Boundary Between Structure and Content in Recurrent Context - Score: 18 (R=9, N=9) - Date: 2026-04-13 - Comment: Holding EMA traces fixed while replacing the predictor isolates information loss in the recurrent representation.

  8. Architecture Determines Observability in Transformers - Score: 18 (R=9, N=9) - Date: 2026-04-29 - Comment: Tracks training-induced erasure of confidence-independent decision signals across controlled layer and head configurations.

  9. Sensitivity-Positional Co-Localization in GQA Transformers - Score: 18 (R=9, N=9) - Date: 2026-04-10 - Comment: Finds anti-localization between task-sensitive and RoPE-influential layers, with crossed interventions testing their interaction.

  10. Nexus: Same Pretraining Loss, Better Downstream Generalization via Common Minima - Score: 19 (R=10, N=9) - Date: 2026-04-13 - Comment: Improves downstream generalization at the same pretraining loss by encouraging agreement between data-source gradients.

  11. Neural Continuous-Time Markov Chain: Discrete Diffusion via Decoupled Jump Timing and Direction - Score: 18 (R=10, N=8) - Date: 2026-04-20 - Comment: Separately parameterizes reverse-diffusion jump timing and direction, yielding a factored training objective for text generation.

  12. Representational Curvature Modulates Behavioral Uncertainty in Large Language Models - Score: 18 (R=9, N=9) - Date: 2026-04-29 - Comment: Links training-emergent trajectory curvature to token entropy using aligned interventions and geometrically misaligned controls.

MoE Where It Changes the Design Space (1)

  1. Training-Free Dynamic Upcycling of Expert Language Models - Score: 18 (R=9, N=9) - Date: 2026-04-01 - Comment: Replaces additional multitask upcycling optimization with a closed-form ridge solution for combining existing dense experts.

Efficiency, Compression, and Large-Scale Training (1)

  1. BASIS: Balanced Activation Sketching with Invariant Scalars for "Ghost Backpropagation" - Score: 18 (R=9, N=9) - Date: 2026-04-21 - Comment: At fixed sketch rank, claims activation memory independent of batch and sequence length while preserving exact error propagation.