← Previous Summary | Monthly Overview
2026-07 | 2026-08

Personalized Monthly Topic Summary 2026/08

MetricValue
Total Papers17
Frontier Model Releases and Technical Reports0
Architecture and Training Dynamics12
Training Algorithms That Change What Is Possible2
MoE Where It Changes the Design Space2
Efficiency, Compression, and Large-Scale Training1
Representation Learning Theory and Structure0
Memory Structures and Agent Memory Systems0
World Models, Exploration, and Open-Ended Reinforcement Learning0

Architecture and Training Dynamics (12)

  1. Riemannian Attention Mechanisms for Transformers: A Theoretical Framework and Architecture Design - Score: 18 (R=9, N=9) - Date: 2026-08-27 - Comment: Replaces dot-product attention with token-specific Riemannian metrics whose scores resist an O(d)-width QK factorization.

  2. Why Post-Norm Transformers Collapse: Attention Amplification and Gradient Repair Failure - Score: 19 (R=10, N=9) - Date: 2026-08-12 - Comment: Explains Post-Norm collapse through causal-attention similarity amplification followed by residual-driven contraction of normalization gradients.

  3. Feed-Forward Steering in Transformer Residual Dynamics - Score: 19 (R=10, N=9) - Date: 2026-08-04 - Comment: Tangential-versus-radial FFN interventions isolate the component that steers residual directions and preserves model quality.

  4. Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference - Score: 18 (R=9, N=9) - Date: 2026-08-11 - Comment: Derives when input-generated attention collapses to a constant bilinear operator and tests that mechanism on a released checkpoint.

  5. MultiHashFormer: Hash-based Generative Language Models - Score: 19 (R=10, N=9) - Date: 2026-08-29 - Comment: Expands multilingual vocabulary while holding parameter count fixed through hash-signature autoregression.

  6. Vowel Signs Are Not Letters: A Pre-tokenization Ceiling on Multilingual Tokenizer Fertility - Score: 19 (R=10, N=9) - Date: 2026-08-27 - Comment: Separates a regex-imposed tokenization floor from data scarcity using matched tokenizers and fixed-compute pretraining.

  7. TANGO: Token-Aggregated Nonlinear Gating Operators for Natural and Formal Language Modeling - Score: 19 (R=10, N=9) - Date: 2026-08-24 - Comment: Replaces separate attention and per-token FFN sublayers with one cross-token gated residual update.

  8. Quantifying Depth Sufficiency in Residual Neural Networks: A First-Order Criterion - Score: 19 (R=10, N=9) - Date: 2026-08-03 - Comment: Gives an exact first-order criterion for useful additional depth under a fixed, function-preserving residual-insertion protocol.

  9. Disentangling the Expressivity of RoPE - Score: 18 (R=9, N=9) - Date: 2026-08-13 - Comment: Distinguishes periodic modular computation from conventional RoPE's precision-bounded local access.

  10. The Loss Does Not See the Basis, but Adam Does - Score: 18 (R=9, N=9) - Date: 2026-08-06 - Comment: Identifies gauge equivariance as a necessary condition for transferring gradient-flow low-rank bias to an optimizer.

  11. A Hamiltonian-Inspired Local-Operator Ansatz for Slimming Large Language Models - Score: 17 (R=9, N=8) - Date: 2026-08-03 - Comment: Replaces dense Transformer maps with executable sums of local tensor operators and identifies a sharp depth-dependent tolerance boundary.

  12. Allocating Recurrent Compute in Looped Language Models - Score: 18 (R=10, N=8) - Date: 2026-08-19 - Comment: Separates mixer recurrence from FFN recurrence and tests whether additional mixer passes contribute useful cross-position influence.

Training Algorithms That Change What Is Possible (2)

  1. CurveFP: Co-Designing Numerical Representation and Product Arithmetic for Language Models - Score: 18 (R=9, N=9) - Date: 2026-08-13 - Comment: Changes the training number format so nonzero products become exact sign and integer-index updates.

  2. Attn-QAT: 4-Bit Attention With Quantization-Aware Training - Score: 17 (R=9, N=8) - Date: 2026-08-10 - Comment: Stabilizes FP4 attention by matching backward recomputation precision and correcting gradient precision assumptions.

MoE Where It Changes the Design Space (2)

  1. Beyond Routing: Decoupling Expert Dispatch and Aggregation in Sparse Mixture-of-Experts - Score: 19 (R=10, N=9) - Date: 2026-08-10 - Comment: Separates expert dispatch from aggregation while holding selected experts, expert computation, and selected router mass fixed.

  2. Output Dilution: Redundant but Fragile Representations in MoE Models - Score: 18 (R=9, N=9) - Date: 2026-08-26 - Comment: Attributes MoE noise fragility to expert aggregation shrinking the residual contribution while routing remains stable.

Efficiency, Compression, and Large-Scale Training (1)

  1. Lossless Tensor Compression as Program Synthesis - Score: 17 (R=9, N=8) - Date: 2026-08-04 - Comment: Synthesizes reversible tensor programs that reduce checkpoint storage while preserving every source byte.