Previous Day 2026-08-11
Monthly Overview 2026-08
Next Day 2026-08-13

This is a remedial run for missed papers from 08/11/2026 to 08/11/2026.

Results generated on 09/13/2026.

Personalized Daily ArXiv Papers 2026-08-12

Model Metric Usage Papers
Prompt Completion Total Total arXiv Scanned Relevant
local agent CLI Tokens not reported not reported not reported 541 541 26
Cost not reported not reported not reported

Token counts are not reported for this run. 5 of 7 model calls succeeded, 3,660s of model wall clock.

Topic Coverage:

TopicPapers
MoE Training3
Large-Scale Training Systems and Efficiency3
Architecture and Training Dynamics7
Efficiency, Compression, and Large-Scale Training13

Table of contents by topic:

MoE Training (3)

  1. MoE Proxy Models for Low-Cost Failure Reproduction and Diagnosis in LLM RL Post-Training Authors: Yikai Wang, Chuansai Zhou, Yuhang Zhou, Weiqiang Wu, Cong Wu, Yue Deng, Ben Feng, Mingming Zhu, Beirong Zhou, Zhibin Wang, Sheng Zhong, Chen Tian, Wangze Zhang

  2. Policy-Masked Private Experts: Auditable and Reversible Capability Access Control in Sparse MoE Models Authors: Zhuoheng Huang, Mukesh Singh

  3. A Theoretical Framework for Modular Learning of Robust Generative Models Authors: Corinna Cortes, Mehryar Mohri, Yutao Zhong

Large-Scale Training Systems and Efficiency (3)

  1. SCOUT: Symmetric Consensus Outlier Detection for Failure Localization in LLM Pre-Training Authors: Zhuang Wang

  2. Scheduling Mixed RL Rollouts Beyond Prefix Locality Authors: Zetao Hong, Song Yuan, Yuanhao Ding, Yibo Zhu, Daxin Jiang, Zhibin Wang, Chen Tian

  3. Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks Authors: Zelei Cheng, Amritansh Mishra, Sambit Sahu, William Campbell

Architecture and Training Dynamics (7)

  1. Why Post-Norm Transformers Collapse: Attention Amplification and Gradient Repair Failure Authors: Xingjian Wang, Qingyu Han, Xiaodong Luo, Yin Zhang

  2. Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions Authors: David R. Wessels, Farhad Ramezanghorbani, Alireza Moradzadeh, David W. Romero, Olivia Viessmann, Maksim Zhdanov, John St. John, Ken Janik, David M Knigge, Yucheng Tang, Erik J Bekkers, Saee Gopal Paliwal

  3. Spherical Flows for Sampling Categorical Data Authors: Jannis Chemseddine, Gregor Kornhardt, Gabriele Steidl

  4. Diffract: Spectral View of LLM Domain Adaptation Authors: Nikita Borodin, Maria Krylova, Artem Zabolotnyi, Dmitry Aspisov, Egor Shikov, Nikita Tyuplyaev, Oleg Travkin, Roman Alferov, Dmitry Vinichenko

  5. Three Tokens Force Exponential Feature Rank in Nonnegative Kernel Attention Authors: Vicente Opazo

  6. Convergence of Sign-based Random Reshuffling Algorithms for Nonconvex Optimization Authors: Zhen Qin, Zhishuai Liu, Pan Xu

  7. Convergence Guarantees of Gradient Descent for Neural Networks via Generalized Lipschitz Smoothness Authors: Siqiao Mu, Diego Klabjan

Efficiency, Compression, and Large-Scale Training (13)

  1. SQuaT: Self-Supervised Knowledge Distillation via Student-Aware Quantized Teacher Features Authors: HyeonJun Lee, Hyeonsik Jo, Jinwoo Chung, Jangho Kim

  2. Hybrid Token Compression for Vision-Language Models Authors: Jusheng Zhang, Xiaoyang Guo, Tongyu Mo, Qinhan Lv, Wenhao Chai, Jian Wang, Keze Wang, Liang Lin

  3. Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference Authors: Soumil Mandal

  4. ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization Authors: He-Yen Hsieh, H. T. Kung

  5. Lifecycle-Optimal Tokenization: Vocabulary Size as a Deployment-Regime-Dependent Infrastructure Parameter Authors: Rima Mittal, Ankit Gubrani, Satyanarayana Kakollu

  6. Weightless Fine-Tuning: Personalizing LLMs via Logit-Space Transport Authors: Bohan Zhang, Anqi Ni, Yixin Wang, Paramveer S. Dhillon

  7. HoloQ-VLA: Uniform W4A4 Quantization of Vision-Language-Action Models Authors: Xinyu Wang, Mingze Li, Sicheng Lyu, Dongxiu Liu, Kaicheng Yang, Ziyu Zhao, Yufei Cui, Xiao-Wen Chang, Peng Lu

  8. Do LLMs Benefit From Their Own Words? Authors: Jenny Y. Huang, Leshem Choshen, Wei Sun, Omar Khattab, Ramón Fernandez Astudillo, Mehul Damani, Tamara Broderick, Jacob Andreas

  9. Putting Registers to Work: Task Registers for Token Pruning in Vision Transformers Authors: Hongsen Cao, Mona Jaber, Shanxin Yuan, Ahmed Sayed

  10. UPAIR: Diagnosing Reasoning States via Uncertainty-Progress Alignment for Selective Intervention Authors: Cheng Yan, Zhijun Fan, Guangyang Ye, Fan Xu, Xiang Xia, Yawei Wang, Wuyang Zhang

  11. GLAM: Efficient Continual Learning at Scale via Grouped LoRA Adapter Merging Authors: Irene Testa, Luigi Quarantiello, Eric Nuertey Coleman, Samrat Mukherjee, Julio Hurtado, Vincenzo Lomonaco

  12. Dynamic Context Adapters: Efficiently Infusing History into Vision-and-Language Models Authors: Yuhang Song, Bor-Jiun Lin, Jiaxu Liu, Te-Chuan Chiu, Anh Nguyen, Chun-Yi Lee

  13. TACTICL: Task-Aware Compression of Tabular ICL Models Authors: Mykhailo Koshil, Matthias Feurer, Katharina Eggensperger


MoE Training (3)

1. MoE Proxy Models for Low-Cost Failure Reproduction and Diagnosis in LLM RL Post-Training

ArXiv ID: 2608.10823

Primary Topic: MoE Training

Also Matches: Large-Scale Training Systems and Efficiency, Efficiency, Compression, and Large-Scale Training

Authors: Yikai Wang, Chuansai Zhou, Yuhang Zhou, Weiqiang Wu, Cong Wu, Yue Deng, Ben Feng, Mingming Zhu, Beirong Zhou, Zhibin Wang, Sheng Zhong, Chen Tian, Wangze Zhang

Abstract: Reinforcement learning (RL) post-training of large language models (LLMs) is computationally intensive and involves complex system pipelines with substantial debugging overhead. In practice, factors such as framework adaptation, numerical precision, and operator implementation can cause failures, including gradient overflow and loss divergence. Reproducing such failures directly on large models requires considerable time and computational resources. This paper systematically analyzes failures encountered during large-scale RL training on the Huawei Ascend platform, summarizes representative failure types, and identifies three model-side factors relevant to fault reproduction. Based on these factors, we propose a proxy-model construction method for low-cost fault investigation and auxiliary diagnosis. It employs structure-preserving, clustering-based expert pruning to select representative experts while retaining the model's backbone architecture, routing mechanism, and basic task capabilities. Our experimental results show that the proxy models reduce accelerator requirements by 50%-87.5% and achieve up to a 33.3x reduction in per-step NPU-hour cost, while preserving major training dynamics and reproducing fault responses consistent with the original models. Overall, the proxy models can serve as low-cost surrogates for fault reproduction, targeted validation, and auxiliary diagnosis in RL post-training.

Comment: Uses structure-preserving expert pruning to build cheaper MoE proxies that retain routing and reproduce large-model training failures.

Topic Match: The proxy construction directly manipulates MoE expert structure while preserving routing and failure dynamics for training diagnosis.

Relevance: 7 Novelty: 7


2. Policy-Masked Private Experts: Auditable and Reversible Capability Access Control in Sparse MoE Models

ArXiv ID: 2608.06690

Primary Topic: MoE Training

Also Matches: Architecture and Training Dynamics

Authors: Zhuoheng Huang, Mukesh Singh

Abstract: Most language-model access controls regulate behavior while leaving the same computation available to every request. We study a different systems question: can trusted authorization determine which newly trained parameters are reachable by the forward pass? Policy-Masked Private Experts freezes a pretrained sparse Mixture-of-Experts (MoE) model, trains a disjoint expert branch, and selects the public or private pool before top-k routing. The resulting claim is narrow but testable: under the declared trusted computing base (TCB), an unauthorized request executes no private expert. It does not imply that the public model lacks the same semantic capability. We test this separation between execution control and task utility in Qwen3-30B-A3B and DeepSeek-V2-Lite. Three Qwen BF16 seeds update all 32 private experts while the public fingerprint remains unchanged. Across 64 adversarial scenarios and 96 deny/fail-closed events, unauthorized private execution is zero; independent hooks exactly match 11,616 routed private rows and allow-deny-allow recovery is exact. On two prospectively frozen Qwen benchmarks, the private branch improves exact tool use by 5.0 percentage points (pp) (five versus zero discordances; one-sided Holm p = 0.03125, corresponding two-sided exact p = 0.0625) and 21.3 pp (percentile-bootstrap 95% CI [13.3, 29.3], Holm p = 0.000031). Three arm-blinded model evaluators retain a positive external effect of 18.7 pp (95% CI [9.3, 28.0]). A parameter-matched Lora has similar external utility, but a post-hoc request gate leaves 1,225 adapter calls under deny; the disjoint expert branch leaves none. DeepSeek reproduces the route invariant and gains 27.0 pp. A valid sealed evaluation is near-neutral. These results support auditable, reversible control over a trained parameter path, while showing that useful transfer remains distribution dependent.

Comment: Adds a policy-conditioned pre-router that selects disjoint public or private expert pools before standard top-k routing.

Topic Match: The strongest foundational contribution is a new expert-pool gating mechanism with auditable routing invariants.

Relevance: 7 Novelty: 7


3. A Theoretical Framework for Modular Learning of Robust Generative Models

ArXiv ID: 2602.17554

Primary Topic: MoE Training

Also Matches: Large-Scale Training Systems and Efficiency

Authors: Corinna Cortes, Mehryar Mohri, Yutao Zhong

Abstract: Training large-scale generative models is resource-intensive and relies heavily on heuristic dataset weighting. We address two fundamental questions: Can we train Large Language Models (LLMs) modularly, combining small, domain-specific experts to match monolithic performance, and can we do so robustly for any data mixture, eliminating heuristic tuning? We present a theoretical framework for modular generative modeling where a set of pre-trained experts are combined via a gating mechanism. We define the space of normalized gating functions $\mathcal{G}_{1}$ and formulate the problem as a minimax game to find a single robust gate that minimizes divergence to the worst-case data mixture. We prove the existence of such a robust gate using Kakutani's fixed-point theorem and show that modularity acts as a strong regularizer, with generalization bounds scaling with the lightweight gate's complexity. Furthermore, we prove that this modular approach can theoretically outperform models retrained on aggregate data, with the gap characterized by the Jensen-Shannon Divergence. Finally, we introduce a scalable Stochastic Primal-Dual algorithm and a Structural Distillation method for efficient inference. Empirical results on synthetic and real-world datasets confirm that our modular architecture effectively mitigates gradient conflict and can robustly outperform monolithic baselines.

Comment: Minimax-robust gating over pretrained domain experts with existence proof via Kakutani and generalization bounds scaling in gate complexity; a theory of expert combination and data-mixture robustness.

Topic Match: The contribution is a gating mechanism and its theory for combining experts, the core of MoE training, with a primal-dual algorithm for the mixture problem.

Relevance: 7 Novelty: 7


Large-Scale Training Systems and Efficiency (3)

1. SCOUT: Symmetric Consensus Outlier Detection for Failure Localization in LLM Pre-Training

ArXiv ID: 2608.11034

Primary Topic: Large-Scale Training Systems and Efficiency

Authors: Zhuang Wang

Abstract: In LLM pre-training, synchronization propagates rank-local stalls, slowdowns, and numerical errors into job-wide symptoms, obscuring their origin. Existing diagnosis often relies on in-process monitors that cannot report after the trainer blocks or terminates, or on post-mortem logs that preserve only synchronized symptoms; offline health tests lose the workload and operating conditions that triggered the failure. We present SCOUT, a unified runtime failure-localization framework built on one design principle: identify outliers through strict-majority consensus among equivalent replicas. SCOUT aligns replica progress, timing, and numerical evidence, then uses its Consensus Collective Communication (C3) abstraction to identify ranks whose compact signatures disagree with their peers. An out-of-band CPU observer remains responsive when training hangs, whereas in-situ replay exercises recurring stragglers and silent data corruption (SDC) beside the live job with its model state, kernels, allocations, communication path, and thermal and memory pressure present. Collective fingerprints expose rank-local protocol divergence. Clean replay coverage certifies checkpoint numerical integrity, preventing recovery from selecting state corrupted by SDC. SCOUT integrates with PyTorch, TorchTitan, Megatron-Core, and DeepSpeed without training-loop or framework-source modifications. SCOUT is open source at https://github.com/LMResiliency/lm-resiliency.

Comment: Localizes rank-level stalls and numerical faults during distributed pretraining through strict-majority replica consensus and live replay.

Topic Match: The work contributes a new reliability and failure-localization mechanism for large-scale distributed training jobs.

Relevance: 8 Novelty: 7


2. Scheduling Mixed RL Rollouts Beyond Prefix Locality

ArXiv ID: 2608.11152

Primary Topic: Large-Scale Training Systems and Efficiency

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Zetao Hong, Song Yuan, Yuanhao Ding, Yibo Zhu, Daxin Jiang, Zhibin Wang, Chen Tian

Abstract: Modern reinforcement learning (RL) post-training pipelines for large language models (LLMs) increasingly combine rollout workloads across multiple domains and feedback paradigms. Prefix-aware routing improves inference efficiency through cache reuse and load balancing, but it does not control how heterogeneous rollout sessions compete for KV-cache capacity. When reinforcement learning with verifiable rewards (RLVR), reinforcement learning from human feedback (RLHF), and agentic rollouts share an asynchronous inference service, their distinct sequence structures, interaction patterns, and KV-residency times create substantially different serving demands. Rollout scheduling must account for this heterogeneity without distorting the workload mixture specified by the trainer. We present MISA-T, a routing-layer admission policy for mixed rollout serving. MISA-T combines adaptive session admission, workload-aware KV-capacity allocation, and residency-time-aware KV accounting. In rollout-only ablations on Step3.7 and Qwen3.6-35B-A3B, MISA-T improves rollout throughput over a sweep-tuned cache-aware vLLM Router by 53.3% and 43.6%, respectively, while maintaining high prefix-cache hit rates. In a matched 50-iteration Step3.7 experiment, it increases rollout throughput by 35.6% and reduces mean iteration time by 22.8%, while keeping the consumed workload mixture close to the trainer target and achieving comparable task scores.

Comment: Adds workload-aware admission and KV-capacity allocation for heterogeneous asynchronous RL rollouts.

Topic Match: The central contribution is a scheduling algorithm that reduces rollout and end-to-end training iteration time.

Relevance: 7 Novelty: 7


3. Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks

ArXiv ID: 2608.10357

Primary Topic: Large-Scale Training Systems and Efficiency

Also Matches: Architecture and Training Dynamics, Efficiency, Compression, and Large-Scale Training

Authors: Zelei Cheng, Amritansh Mishra, Sambit Sahu, William Campbell

Abstract: Long-horizon tool-using agents must reason over user goals, domain policies, tool calls, simulator state, and delayed verifiable rewards. Reinforcement learning (RL) is a natural fit for this setting, but multi-turn on-policy rollouts create long contexts, while model-specific attention layers may require custom masks and learned sink normalization. We present SINKFLEX-RL, a modular training system for RL in dual-control tool-use environments. The system combines a Gymnasium-compatible environment wrapper, a VERL-style rollout dataflow, group-relative policy optimization without a separate value model, and a sink-aware FlexAttention path designed to preserve model-specific sink scaling under causal and sliding-window masks. In a preliminary Tau2Bench retail run, validation reward (mean@1) rises from 0.25 early in training to $0.44$ later in the observed training window, while training-score and trajectory-reward proxies also trend upward. In a fixed-configuration memory benchmark, the optimized attention path reduces peak VRAM from 28.06GB to 22.52GB at 4096 tokens, a $19.7\%$ reduction, and runs the measured 8192-token configuration using $25.53$~GB where the eager baseline runs out of memory. These results illustrate the value of integrating environment interfaces, RL dataflow, and attention-kernel design for memory-feasible long-horizon agent training.

Comment: Integrates asynchronous RL dataflow with sink-aware fused attention to make long-context agent training memory feasible.

Topic Match: The strongest contribution is a training system whose attention implementation materially reduces memory requirements.

Relevance: 7 Novelty: 6


Architecture and Training Dynamics (7)

1. Why Post-Norm Transformers Collapse: Attention Amplification and Gradient Repair Failure

ArXiv ID: 2608.09417

Primary Topic: Architecture and Training Dynamics

Also Matches: Large-Scale Training Systems and Efficiency

Authors: Xingjian Wang, Qingyu Han, Xiaodong Luo, Yin Zhang

Abstract: Deep decoder-only Transformers often replace the original Post-Norm architecture with Pre-Norm variants because Post-Norm training is highly sensitive to warmup and learning rate under conventional initialization schemes. Although prior work has identified rank collapse and gradient vanishing as related symptoms, it remains poorly understood how causal attention creates high-similarity representations and why training dynamics fail to repair them. We give a two-stage analysis of Post-Norm rank collapse using token similarity as a scalar state variable. First, at initialization, causal attention acts approximately as a prefix-averaging operator that increases token similarity across depth, while the SwiGLU branch contributes only a smaller damping effect. Second, once training enters a high-similarity regime, growth of pre-normalization residual norms makes the RMSNorm backward factor contractive; under mild conditions, gradients to earlier layers decay geometrically. As a complementary result, we characterize the properties of a collapsed network: its best predictor is frequency distribution with relatively high loss floor, and gradients in collapsed layers vanish at frequency distribution. Experiments on 48-layer decoder-only Transformers trained on C4 dataset match the predicted initialization-time similarity growth and collapse-time gradient contraction, and show that collapsed runs stay near the predicted frequency loss. Together, these results distinguish the forward similarity amplification and backward repair incapacity in Post-Norm collapse, while also characterizing the behavior of collapsed networks.

Comment: Two-stage account of Post-Norm rank collapse: causal attention acts as a prefix-averaging operator raising token similarity at init, then growing residual norms make the RMSNorm backward factor contractive so gradients decay geometrically; validated on 48-layer decoders trained on C4.

Topic Match: Squarely normalization and residual design plus optimisation dynamics explaining why deep transformers fail to train.

Relevance: 9 Novelty: 7


2. Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions

ArXiv ID: 2607.19378

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: David R. Wessels, Farhad Ramezanghorbani, Alireza Moradzadeh, David W. Romero, Olivia Viessmann, Maksim Zhdanov, John St. John, Ken Janik, David M Knigge, Yucheng Tang, Erik J Bekkers, Saee Gopal Paliwal

Abstract: Subquadratic alternatives to attention require compromises when applied to multi-dimensional data: standard convolutions lack global receptive fields and input dependency, while recurrent models require rasterizing data such as images, volumes, and partial differential equation (PDE) into an ad-hoc $1\rm D$ scan order that violates their spatial structure. We introduce \textit{HyenaND}, a subquadratic, global, input-dependent operator that acts directly on the native geometry of multidimensional data through convolutions with implicitly parametrized global, input-dependent multi-dimensional convolutional kernels. Our CUDA implementation, \texttt{nSubQ}, fuses the FFT-convolution path to turn HyenaND's $\mathcal{O}(L \log L)$ scaling into wall-clock speedups. Across long-context genomics, computer vision, medical imaging, and PDE modeling, pure HyenaND stacks match the accuracy of strong attention baselines, while hybrid configurations that interleave HyenaND and attention layers outperform both pure attention and strong recurrence-based hybrids.

Comment: Introduces an input-dependent multidimensional long-convolution operator and a fused FFT CUDA implementation with subquadratic scaling.

Topic Match: The primary advance is a new global sequence operator, complemented by a kernel that realizes its theoretical efficiency.

Relevance: 8 Novelty: 8


3. Spherical Flows for Sampling Categorical Data

ArXiv ID: 2605.05629

Primary Topic: Architecture and Training Dynamics

Authors: Jannis Chemseddine, Gregor Kornhardt, Gabriele Steidl

Abstract: We study the problem of learning generative models for discrete sequences in a continuous embedding space. Whereas prior approaches typically operate in Euclidean space or on the probability simplex, we instead work on the sphere $\mathbb S^{d-1}$. There the von Mises-Fisher (vMF) distribution induces a natural noise process and admits a closed-form conditional score. The conditional velocity is in general intractable. Exploiting the radial symmetry of the vMF density we reduce the continuity equation on $\mathbb S^{d-1}$ to a scalar ODE in the cosine similarity, whose unique bounded solution determines the velocity. The marginal velocity and marginal score on $(\mathbb S^{d-1})^L$ both decompose into posterior-weighted tangent sums that differ only by per-token scalar weights. This gives access to both ODE and predictor-corrector (PC) sampling. The posterior is the only learned object, trained by a cross-entropy loss. Experiments compare the vMF path against geodesic and Euclidean alternatives. The vMF path especially in combination with PC sampling significantly improves results on Sudoku, language modeling, and mathematical reasoning.

Comment: Formulates categorical sequence generation as score- and flow-based modeling on a sphere with tractable posterior-weighted dynamics.

Topic Match: The paper proposes a new foundational generative architecture and training formulation for discrete sequences.

Relevance: 7 Novelty: 8


4. Diffract: Spectral View of LLM Domain Adaptation

ArXiv ID: 2608.10850

Primary Topic: Architecture and Training Dynamics

Also Matches: Efficiency, Compression, and Large-Scale Training

Authors: Nikita Borodin, Maria Krylova, Artem Zabolotnyi, Dmitry Aspisov, Egor Shikov, Nikita Tyuplyaev, Oleg Travkin, Roman Alferov, Dmitry Vinichenko

Abstract: We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction, code, and natural text. Using singular value decomposition of weight matrices, we find that CPT leaves singular value spectra largely invariant, with adaptation driven mainly by changes in singular vectors. An analysis of attention-head projection matrices reveals strong, domain-dependent head heterogeneity, which we exploit to define a head importance criterion: up to 60% of head updates can be removed without measurable quality loss. Selectively rewinding low-importance heads to their pre-trained state improves benchmark accuracy by up to 4% versus the fully trained baseline. Finally, we identify domain connectivity - linear interpolation between CPT checkpoints yields smooth domain-quality interpolation without notable degradation on either domain - and release Diffract, an open-source toolkit for scalable spectral analysis of billion-parameter models.

Comment: Shows that continual pretraining primarily rotates singular vectors and reveals strongly domain-dependent attention-head update importance.

Topic Match: Its main value is mechanistic analysis of parameter and attention-head dynamics during continual pretraining.

Relevance: 7 Novelty: 7


5. Three Tokens Force Exponential Feature Rank in Nonnegative Kernel Attention

ArXiv ID: 2608.11427

Primary Topic: Architecture and Training Dynamics

Authors: Vicente Opazo

Abstract: Full attention exposes every token pair, whereas kernel attention compresses a sequence into a fixed-dimensional sketch. We show that this distinction becomes exponential at the first context length containing two competing candidates. On Min-IP over Boolean inputs, rank-one normalized kernel attention solves every sequence of length at most two exactly. In contrast, any single normalized nonnegative kernel-attention head that succeeds on all three-token sequences with error strictly below $1/2$ requires $2^{Ω(m)}$ features, even with arbitrary finite-dimensional tokenwise values and an arbitrary query-dependent affine readout. Dense softmax solves the same task with $m$-dimensional scores and constant temperature. The conclusion survives position-dependent token maps and a causal final query. As context length grows, the lower bound approaches the exact $2^m$-feature realization. Separately, for deterministic multihead, multilayer sketch models whose cross-token channels have finite alphabets, we prove a transcript lower bound linear in the number of independent answers and logarithmic in their alphabet size.

Comment: Proves an exponential feature-count separation between normalized nonnegative kernel attention and dense softmax on a three-token retrieval task, a mechanistic limit on linear/kernel attention variants.

Topic Match: Core contribution is an expressivity analysis of an attention mechanism variant, not an application.

Relevance: 7 Novelty: 7


6. Convergence of Sign-based Random Reshuffling Algorithms for Nonconvex Optimization

ArXiv ID: 2310.15976

Primary Topic: Architecture and Training Dynamics

Also Matches: Large-Scale Training Systems and Efficiency

Authors: Zhen Qin, Zhishuai Liu, Pan Xu

Abstract: signSGD is attractive in nonconvex optimization because it communicates sign-valued rather than full-precision gradients. Several standard analyses assume independent stochastic-gradient samples, whereas a common finite-sum implementation reshuffles the data and processes them sequentially. We study this variant, signSGD with random reshuffling (SignRR), and show that reshuffling does not in general repair the bias created by discarding gradient magnitudes. In particular, on a one-dimensional two-component strongly convex quadratic, the expected gradient norm at every SignRR inner iterate equals $1/2$. We complement this impossibility result with an alignment-explicit finite-time bound $O(\log(nT)/\sqrt{nT}+\varepsilon_{\mathrm{align}})$, where $\varepsilon_{\mathrm{align}}$ measures the averaged loss of descent caused by component-sign misalignment. A horizon-tuned constant stepsize improves the vanishing term to $O(1/\sqrt{nT})$, and a remaining-set alignment condition yields a residual-free $O(1/\sqrt{nT})$ guarantee. The alignment term is upper bounded by twice the averaged mean absolute gradient error and, in turn, by twice an averaged coordinatewise conditional root-mean-square error. As a variance-reduced alternative, we analyze SignRVR, which signs an SVRG estimator anchored at the beginning of every epoch. A pathwise argument gives a residual-free guarantee with an $O(\sqrt{d/T})$ averaged $\ell_1$-stationarity bound.

Comment: Characterizes signSGD under random reshuffling and derives alignment-dependent and variance-reduced convergence guarantees.

Topic Match: The work principally analyzes optimization dynamics, with secondary relevance to communication-efficient distributed optimization.

Relevance: 6 Novelty: 7


7. Convergence Guarantees of Gradient Descent for Neural Networks via Generalized Lipschitz Smoothness

ArXiv ID: 2608.11479

Primary Topic: Architecture and Training Dynamics

Authors: Siqiao Mu, Diego Klabjan

Abstract: We establish convergence guarantees of gradient descent for general feedforward neural networks of arbitrary width or depth, with no special requirements on the initialization or dataset. We only assume that the activation functions are Lipschitz smooth, Lipschitz continuous, and linearly bounded--- properties that hold for linear, tanh, softplus, and sigmoid activation functions. For the loss function, we require that it is Lipschitz smooth in the model outputs, which is true for mean-squared error. The key theoretical insight is that the Lipschitz properties of the activation functions are partially preserved even through repeated compositions, leading to a novel generalized Lipschitz smoothness condition where the change in gradient is upper bounded by the change in the parameter space, multiplied by polynomial terms of the parameter norms at both endpoints. This type of condition holds for both the model function and the loss function, enabling a descent lemma where the loss decreases as long as the learning rate is small enough with respect to the parameter norms. By ensuring that the parameter norms do not grow too quickly to infinity, we prove that the minimum squared gradient norm converges to zero in $T$ iterations at rate $O(1/T^{1/L})$ for an $L$-layer neural network.

Comment: Derives gradient-descent convergence guarantees for arbitrary-depth feedforward networks under generalized Lipschitz smoothness.

Topic Match: Its core contribution is optimization and convergence analysis for neural-network training dynamics.

Relevance: 6 Novelty: 7


Efficiency, Compression, and Large-Scale Training (13)

1. SQuaT: Self-Supervised Knowledge Distillation via Student-Aware Quantized Teacher Features

ArXiv ID: 2608.10709

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: HyeonJun Lee, Hyeonsik Jo, Jinwoo Chung, Jangho Kim

Abstract: Quantization-Aware Training (QAT) enables the deployment of quantized models with minimal accuracy degradation. However, in practical scenarios, training labels are often unavailable due to privacy, copyright, or cost constraints. Knowledge Distillation (KD) is a common approach to address this challenge, but we observe that prior work combining QAT with KD suffers from a fundamental limitation: during distillation, the range mismatch between the teacher and the quantized student model induces an unattainable residual, resulting in an irreducible lower bound on the distillation loss. Motivated by this observation, we propose SQuaT (Student-Aware Quantized Teacher Features), a label-free QAT framework with KD that theoretically eliminates this lower bound by applying the student's quantization parameters to quantize the teacher's features during distillation. Through comprehensive experiments across diverse settings, we demonstrate that SQuaT consistently outperforms strong baselines, with particularly pronounced gains in extreme low-bit (e.g., 1- and 2-bit) settings. Furthermore, extensive evaluations across various model design choices show that our approach does not rely on specific architectural assumptions, making it broadly applicable across diverse architectures and quantization settings. The source code is available at https://github.com/lcdbsa522/SQuaT.

Comment: Eliminates a distillation-loss floor in low-bit QAT by quantizing teacher features with the student's quantization parameters.

Topic Match: The work is directly centered on a new quantization-aware training mechanism, especially for extreme low-bit models.

Relevance: 8 Novelty: 7


2. Hybrid Token Compression for Vision-Language Models

ArXiv ID: 2512.08240

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Jusheng Zhang, Xiaoyang Guo, Tongyu Mo, Qinhan Lv, Wenhao Chai, Jian Wang, Keze Wang, Liang Lin

Abstract: Vision-language models (VLMs) rely on hundreds of visual tokens, leading to high computational and memory costs. Existing compression methods face a trade-off: continuous compression can weaken high-level semantics, while discrete quantization may lose fine-grained appearance details. We introduce HTC-VLM, a hybrid visual token compression framework that disentangles semantics and appearance through two complementary pathways. A continuous pathway preserves detailed ViT patch features, while a discrete pathway provides semantic anchors using MGVQ quantization represented by four tokens. The two pathways are fused into a 580-token hybrid sequence and compressed into a single token using a disentanglement attention mask and a bottleneck. Under the same one-token output budget, HTC-VLM achieves 87.2% average performance retention across seven benchmarks (GQA, VQAv2, MMBench, MME, POPE, SEED-Bench, and ScienceQA-Image), outperforming the leading continuous baseline at 81.0%. Attention analysis shows that the compressed token prioritizes discrete anchors, supporting their role as semantic guidance. We further study token-budget scaling, cross-architecture generalization, inference efficiency, robustness to codebook and masking variations, and the information-theoretic properties of the hybrid bottleneck. These results show that combining continuous appearance features with discrete semantic anchors enables effective extreme visual token compression for efficient VLMs.

Comment: Combines continuous appearance features and discrete semantic anchors in an extreme visual-token compression bottleneck.

Topic Match: The core contribution is a new token-compression mechanism that reduces VLM compute and memory requirements.

Relevance: 8 Novelty: 7


3. Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference

ArXiv ID: 2607.13205

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Soumil Mandal

Abstract: Attention-based KV cache eviction (H2O and its descendants) compresses the memory-constrained state of a long-context model by ranking tokens on accumulated attention mass, treated here as signal energy, and keeping the heaviest. On schema-dense input streams such as nested JSON, this score acts as a non-stationary filter that disproportionately retains noise: a non-content sink role (delimiters or whitespace) carries an order of magnitude more energy than any content role, and structural KEY tokens are over-retained at roughly 1.8x the rate of the answer-carrying VALUE tokens, collapsing exact-match accuracy from 88% to 0% at a 5% budget as the signal-to-noise ratio of the retained state degrades. A counterfactual experiment establishes that suppressing KEY tokens is the best deployable filter. Our retraining-free, role-conditional allocation over SnapKV's windowed score, governed by a single tuned hyperparameter, closes 63-98% of the H2O gap at sub-20% budgets and, at higher budgets, modestly matches or exceeds full-cache accuracy -- a small, seed-sensitive denoising effect (borderline significant at B=0.50; not distinguishable from zero at B=0.30 over four seeds). A 15 MB linear role probe supplies these labels at negligible inference cost, though matching parser-level downstream accuracy remains open.

Comment: Diagnoses structural-role bias in attention-based KV eviction and corrects it with role-conditional cache allocation.

Topic Match: This is directly a KV-cache compression method with a new mechanism for retaining useful long-context state.

Relevance: 8 Novelty: 7


4. ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization

ArXiv ID: 2608.11045

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: He-Yen Hsieh, H. T. Kung

Abstract: ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inherent in standard round-to-nearest (RTN) schemes when quantizing weights near the centers of quantization intervals. Starting from a pretrained LLM, ReRound trains a conditional diffusion model to produce continuous reconstructions of low-bit weights for the LLM. These reconstructed weights act as a guidance signal to disambiguate the rounding direction of weights located close to interval midpoints. To integrate this reconstruction-guided rounding with conventional RTN, ReRound introduces a tolerance metric measuring how far the quantized weight (not the final quantized integer) is away from the midpoint: quantized weights within a tolerance region around midpoints are quantized using diffusion-based reconstructions, whereas weights closer to quantization boundaries are quantized with RTN. By sweeping the tolerance parameter, ReRound generates multiple candidate quantized integer weight matrices and selects the de-quantized weight matrix candidate whose leading singular values most closely match those of the original full-precision weights. This selected candidate determines the tolerance parameter ReRound uses. ReRound is particularly effective for smaller LLMs. Across a range of such models, it consistently outperforms standard RTN for 3-bit and 4-bit weight quantization. ReRound achieves superior accuracy compared to an extensive set of calibration-free methods, remains competitive with calibration-dependent approaches, and operates entirely offline, introducing no additional overhead during low-bit inference. The ReRound strategy represents a new approach for low-bit quantization. The method applies to AI models beyond LLMs. This paper focuses on its applications to small LLMs.

Comment: Uses diffusion-reconstructed weights to resolve rounding ambiguity near quantization-interval midpoints without calibration data.

Topic Match: The paper is directly centered on a new calibration-free low-bit weight-quantization method.

Relevance: 8 Novelty: 7


5. Lifecycle-Optimal Tokenization: Vocabulary Size as a Deployment-Regime-Dependent Infrastructure Parameter

ArXiv ID: 2608.11361

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Large-Scale Training Systems and Efficiency

Authors: Rima Mittal, Ankit Gubrani, Satyanarayana Kakollu

Abstract: Tokenizer vocabulary size is a foundational design choice in large language model (LLM) infrastructure, yet it is typically fixed at training time based on convention rather than deployment analysis. We show that the cost-optimal vocabulary is not a constant but a function of the serving regime. We formalize total deployment cost as $C_{lifecycle}(V) = C_{train}(V) + λ\cdot C_{infer}(V, B)$, where $λ$ is inference volume and $B$ is the serving batch size. Through controlled experiments on two GPU families spanning the memory-bound to compute-bound regimes (A10G, ridge $\approx$ 117 FLOP/byte; A100, ridge $\approx$ 183 FLOP/byte), we demonstrate: (1) the inference-optimal vocabulary shifts 16x with serving batch, from 32k at $B=1$ to 524k at $B=64+$, driven by amortization of the $V \times d$ unembedding matrix read; (2) at 1.3-2.3B model scale, quality (bits per byte, BPB) is optimized at $V=65$k, confirming scale-dependent vocabulary preference; (3) the lifecycle-optimal vocabulary diverges from training-optimal by up to 16x for production deployments. Quality is approximately invariant across the optimal range ($<$2% BPB spread), making vocabulary a pure systems optimization with no quality penalty in the measured range. Our results provide actionable capacity planning guidance: on-device deployments ($B=1$) should use $V \approx 32$k; datacenter serving ($B \geq 64$, $λ\geq 10$) should use $V \approx 131$-262k.

Comment: Jointly models training and batched-inference costs to select vocabulary size for the full model lifecycle.

Topic Match: The work directly studies a model-design parameter through systems-aware training and inference cost scaling.

Relevance: 8 Novelty: 7


6. Weightless Fine-Tuning: Personalizing LLMs via Logit-Space Transport

ArXiv ID: 2608.11342

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Bohan Zhang, Anqi Ni, Yixin Wang, Paramveer S. Dhillon

Abstract: Supervised fine-tuning (SFT) is a standard approach for adapting LLMs to a target distribution, but in settings such as personalization, where each author requires separate weight access, optimization, storage, and retraining, its costs become prohibitive. We propose Weightless Fine-Tuning (WFT), a training-free decoding-time method that approximates the distributional effect of SFT without weight updates. WFT computes supervised residuals on an author's training sequence and transports them to the current prompt through a cross-prefix transport operator estimated from dropout-induced cross-covariance. The operator captures how a perturbation at one context propagates to predictions at another, replacing gradient-based parameter updates with logit-space corrections. On three LaMP personalization benchmarks, WFT achieves the best average performance across datasets, matches or exceeds SFT on individual tasks, and outperforms other lightweight baselines on average. In a budget-controlled comparison, WFT approaches SFT performance using less than 7% of the effective computation. Logit-level analysis shows a cosine similarity of 0.875 between the logit shifts induced by WFT and SFT over 95% of the next-token probability mass, suggesting that WFT captures the distributional effect of supervised adaptation without modifying model weights.

Comment: Approximates supervised fine-tuning through dropout-estimated transport of training residuals directly in logit space.

Topic Match: The strongest fit is a new adaptation mechanism that avoids weight updates and sharply reduces personalization compute.

Relevance: 7 Novelty: 8


7. HoloQ-VLA: Uniform W4A4 Quantization of Vision-Language-Action Models

ArXiv ID: 2605.28803

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Xinyu Wang, Mingze Li, Sicheng Lyu, Dongxiu Liu, Kaicheng Yang, Ziyu Zhao, Yufei Cui, Xiao-Wen Chang, Peng Lu

Abstract: Vision-Language-Action (VLA) models unify perception, reasoning, and control in a single policy, but their multi-billion-parameter backbones and diffusion-based action heads make on-device deployment prohibitively expensive. Low-bit post-training quantization (PTQ) is the natural remedy, yet the diffusion action head that emits continuous control signals is highly sensitive to it: a few weight and activation outliers are enough to destabilize the head, so prior work leaves it at full precision or falls back to mixed-precision schemes, and uniformly quantizing the whole model to low bit-width remains an open challenge. We present HoloQ-VLA, the first training-free PTQ framework that compresses both the language backbone and the entire diffusion action head to uniform W4A4 precision without mixed-precision allocation. Instead of trading weight quality against activation quality, HoloQ-VLA targets the two outlier sources with complementary transforms: a weight-adapted rotation composed with an activation-dispersing Hadamard transform, together with per-step scaling that absorbs the dynamic-range drift exhibited by the action head across denoising steps. On LIBERO, HoloQ-VLA compresses Pi-0.5 and GR00T-N1.5 to W4A4 with 98.0% and 87.8% task success rates, matching or exceeding their FP16 references of 97.1% and 87.0%, while reducing the static memory footprint by 74.2%. Real-world manipulation experiments further demonstrate that HoloQ-VLA maintains smooth and accurate control across diverse real-world scenarios.

Comment: Uniformly quantizes a VLA backbone and diffusion action head to W4A4 using complementary outlier-control transforms.

Topic Match: Its central methodological contribution is a low-bit post-training quantization scheme spanning heterogeneous model components.

Relevance: 7 Novelty: 7


8. Do LLMs Benefit From Their Own Words?

ArXiv ID: 2602.24287

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Jenny Y. Huang, Leshem Choshen, Wei Sun, Omar Khattab, Ramón Fernandez Astudillo, Mehul Damani, Tamara Broderick, Jacob Andreas

Abstract: In multi-turn conversations, large language models typically condition on the full conversation history: both past user prompts and assistant responses. We revisit this design choice by comparing full-context prompting to four alternative, substantially-reduced context configurations. Analyzing in-the-wild multi-turn conversations across three open reasoning and one state-of-the-art model, we find that response quality is largely preserved under aggressive context filtering: replacing all prior assistant turns with one-sentence summaries or keeping only the most recent user--assistant exchange often matches storing full context in performance while using roughly 8x less context. To understand this result, we observe that a substantial fraction of user turns (36.4%) in multi-turn conversations are self-contained and that many follow-up turns can be addressed by seeing only the immediately preceding user--assistant exchange. Furthermore, we find that when models condition on their own past responses, this can lead to context pollution, a phenomenon in which reasoning errors, hallucinations, or stylistic artifacts propagate across turns. Motivated by these findings, we design a context-filtering approach that selectively omits the assistant-side history. Taken together, these findings suggest moving away from storing full dialogue transcripts and instead retaining only what is relevant.

Comment: Shows that aggressively filtering prior assistant turns can preserve response quality while reducing conversational context roughly eightfold.

Topic Match: The main foundational contribution is context compression that materially reduces inference memory and compute.

Relevance: 7 Novelty: 7


9. Putting Registers to Work: Task Registers for Token Pruning in Vision Transformers

ArXiv ID: 2608.10989

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Hongsen Cao, Mona Jaber, Shanxin Yuan, Ahmed Sayed

Abstract: Token-pruning policies are usually designed for a single recognition pipeline, but pretrained Vision Transformers are reused across tasks with different spatial demands. We ask which parts of a pruning policy transfer across image classification, semantic segmentation, and object detection. For each pipeline, controlled probes freeze the no-pruning checkpoint and apply a series of parameter-free reduction criteria at one eligible layer at a time without retraining. The probes reveal three differences: segmentation and detection rank the criteria differently, classification is especially sensitive to attention-based pruning in the earliest layers, and the dense tasks prefer opposite recovery endpoints. These findings motivate Task-Adaptive Pruning (TAP). Existing register tokens serve as task-agnostic storage for feature artifacts. TAP instead introduces one task register per task and activates only the current one. Its evolving state ranks tokens, distributes an exact removal budget over depth, and sets the recovery scale for dense features. At a final keep rate of $ρ=0.5$, our jointly adapted model, TAP-J, reaches $47.0$ mIoU at $1.30\times$ encoder throughput on ADE20K and $53.7$ box AP at $1.32\times$ encoder throughput on COCO while remaining competitive on ImageNet-1K.

Comment: Introduces task registers that rank visual tokens and allocate an exact pruning budget across transformer depth.

Topic Match: Token pruning is the primary mechanism, with task-conditioned registers supplying the architectural control signal.

Relevance: 7 Novelty: 6


10. UPAIR: Diagnosing Reasoning States via Uncertainty-Progress Alignment for Selective Intervention

ArXiv ID: 2607.17188

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Cheng Yan, Zhijun Fan, Guangyang Ye, Fan Xu, Xiang Xia, Yawei Wang, Wuyang Zhang

Abstract: While test-time scaling improves the problem-solving ability of large reasoning models (LRMs) through additional inference-time computation, it can also exacerbate overthinking and underthinking, which we formulate as reasoning state--action mismatch. Resolving this mismatch requires reliable reasoning state diagnosis, yet single-signal monitors provide ambiguous evidence, while steering-based controllers often rely on outcome-labeled supervision or model-specific calibration. We introduce the Uncertainty--Progress Alignment Hypothesis, which posits that the relative transition timing of proxy answer uncertainty and latent reasoning progress distinguishes healthy, stagnant, and ready states that warrant different subsequent actions. Building on this insight, we propose UPAIR, a training-free framework that couples lightweight uncertainty monitoring with event-triggered joint diagnosis and maps the resulting state to native continuation, selective strategy switching, or verification-guided stopping. Across three LRMs and five cross-domain benchmarks, the stagnation diagnosis detects 64.3% of natural errors while flagging only 5.4% of correct samples, revealing a dynamic reasoning regularity shared across models and tasks. End to end, UPAIR improves accuracy by up to 16.67 percentage points and reduces generated tokens by up to 29.64%, demonstrating the effectiveness of its integrated diagnosis and intervention, while online diagnosis costs less than 1% of natural-generation time.

Comment: Dynamically switches, continues, or stops reasoning by jointly tracking uncertainty and latent progress.

Topic Match: Its strongest fit is adaptive inference computation that reduces generated tokens, supported by a dynamic reasoning-state mechanism.

Relevance: 6 Novelty: 7


11. GLAM: Efficient Continual Learning at Scale via Grouped LoRA Adapter Merging

ArXiv ID: 2509.13211

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Irene Testa, Luigi Quarantiello, Eric Nuertey Coleman, Samrat Mukherjee, Julio Hurtado, Vincenzo Lomonaco

Abstract: The ability to learn continuously over time remains a major challenge for modern machine learning systems, even in the era of Foundation Models. While the rich representations learned by large pre-trained models can partially mitigate catastrophic forgetting, they still struggle to adapt efficiently to evolving data distributions. A key challenge remains, how to continually add new knowledge to a large pretrained model in a way that is scalable and computationally efficient over long task sequences. In this work, we introduce GLAM, a simple and effective framework for class-incremental continual learning based on LoRA adapters merging. For each task, GLAM trains a lightweight low-rank adapter with an importance scalar, incurring minimal computational overhead. Adapters are then pruned, rescaled, and sequentially grouped to enable structured knowledge reuse and limit parameter growth. At inference, all groups are combined into a single module, ensuring constant computational cost regardless of the number of tasks. We evaluate GLAM on vision benchmarks with sequences of up to 50 tasks, significantly extending beyond standard protocols. GLAM achieves the highest accuracy across the evaluated benchmarks. Compared with the baseline attaining the highest average accuracy across benchmarks, it uses 16--19\% of the trainable parameters and reduces training time by approximately 63--74\%, demonstrating efficient and scalable continual learning. Our source code is publicly available at https://github.com/atlas-luiss/GLAM.

Comment: Prunes, rescales, and groups LoRA adapters to limit parameter growth and maintain constant inference cost in continual learning.

Topic Match: Its principal contribution is a low-rank adaptation and compression scheme that reduces training parameters and time.

Relevance: 6 Novelty: 6


12. Dynamic Context Adapters: Efficiently Infusing History into Vision-and-Language Models

ArXiv ID: 2608.10525

Primary Topic: Efficiency, Compression, and Large-Scale Training

Also Matches: Architecture and Training Dynamics

Authors: Yuhang Song, Bor-Jiun Lin, Jiaxu Liu, Te-Chuan Chiu, Anh Nguyen, Chun-Yi Lee

Abstract: Historical context integration presents a fundamental challenge for Vision-Language Models (VLMs) in sequential decision-making tasks. Current VLMs process visual inputs independently, which creates critical limitations for downstream applications that require temporal understanding. Direct incorporation of historical frames into Transformer inputs produces quadratic attention complexity and excessive memory consumption. Existing approaches suffer from significant drawbacks: computational inflation or substantial information loss through temporal compression. To address these challenges, we introduce Dynamic Context Adapter (DCA), a novel context injection approach for pretrained VLMs. Our method employs fixed-size, dynamically compressed memory to preserve historical semantics without frame concatenation. DCA bridges static VLMs and recurrent policies and enables memory capabilities in pretrained models while maintaining computational efficiency. DCA achieves over $25\%$ reduction in attention FLOPs and $13\%$ memory savings while improving performance on long-horizon tasks.

Comment: Injects fixed-size dynamically compressed history into pretrained VLMs without concatenating all historical frames.

Topic Match: The strongest fit is memory and attention-cost reduction through a new context-adapter architecture.

Relevance: 6 Novelty: 6


13. TACTICL: Task-Aware Compression of Tabular ICL Models

ArXiv ID: 2608.10837

Primary Topic: Efficiency, Compression, and Large-Scale Training

Authors: Mykhailo Koshil, Matthias Feurer, Katharina Eggensperger

Abstract: The strong performance of foundation models for tabular tasks comes at substantial inference costs. Distilling models into task-specific architectures reduces model size and computational demands but also sacrifices in-context adaptability. Here we introduce TACTICL, an automated task-aware compression framework for tabular in-context learning models that jointly prunes transformer layers and replaces them with lightweight adapters trained on downstream tasks, thus blending in-context with in-weight learning. We study TACTICL on 47 benchmark datasets and show that we can substitute up to 85% of layers without substantial performance drop on a given downstream task. We further show that TACTICL maintains robustness to data shifts, leaving its in-context ability intact. Overall, TACTICL provides a robust framework for exploiting the depth-wise redundancy of tabular foundation models by combining task-specific adaptation and structured compression. We provide the code at: https://github.com/Hebog/tfm_compression

Comment: Replaces pruned transformer layers with lightweight task adapters while retaining tabular in-context behavior.

Topic Match: Its core mechanism is structured depth pruning combined with parameter-efficient adaptation.

Relevance: 6 Novelty: 6


Paper Selection Prompt

System Prompt

You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.

User Prompt

Instructions

Respond in JSONL. Output exactly one JSON object per paper, one per line:

{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}

Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only: daily_hot, new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return []. - daily_hot means the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. - new_frontier means the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.

Scoring Criteria

Relevance and Novelty are independent axes. Score both from 1 to 10.

Relevance Scoring

  • 9-10: directly centered on the target foundational topics; highest when the core contribution is clearly within them.
  • 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
  • 5-6: touches the target topics, but the main contribution is elsewhere.
  • 3-4: largely outside the target topics, often application-focused or domain-specific.
  • 1-2: unrelated.

Important: Broad frontier relevance, major launch status, or daily buzz is not enough for a high Relevance score here. Those cases belong in the hotspot digest unless the paper strongly matches the specialized paper topics.

Novelty Scoring

  • 9-10: new paradigm, theory, or major methodological breakthrough.
  • 7-8: substantial methodological advance or strong new insight.
  • 5-6: meaningful but incremental extension or refinement.
  • 3-4: minor, narrow, or mostly engineering or domain-specific improvement.
  • 1-2: little originality; mainly standard application of existing methods.

Topic Registry

Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - moe_training: MoE Training - Mixture-of-Experts training: routing, load balancing, expert granularity and shared experts, training stability, MoE parallelism and communication, and MoE kernels. - training_systems: Large-Scale Training Systems and Efficiency - Large-scale training systems and efficiency: distributed training algorithms, optimizers, and communication. - architecture_training: Architecture and Training Dynamics - Core architectural or computational mechanisms, dynamic computation, and training-stability dynamics. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, memory or cache efficiency, and large-scale training systems that materially change cost or behavior.

Papers

[PAPER LIST HERE]

Relevant Topics

This is a training-side feed. The centre of gravity is Mixture-of-Experts training and the large-scale training systems it sits on; architecture and efficiency work is kept when it bears on how a large model is trained or what it costs to train.

Keep a paper when its CORE CONTRIBUTION falls in one of the four topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.

  1. MoE Training (primary) - Keep: expert routing and gating mechanisms; load balancing, with or without auxiliary losses; expert granularity, fine-grained and shared experts; MoE training stability (loss spikes, router collapse, z-loss and related fixes); expert and hybrid parallelism; all-to-all dispatch/combine and the communication schedules MoE needs; MoE kernels and grouped-GEMM style implementations; expert capacity, token dropping, and upcycling dense checkpoints into MoE. - Filter: papers that merely train on top of a MoE model without a new routing, balancing, stability, parallelism, or kernel contribution; "mixture of experts" in the classical ensemble or recommender sense.

  2. Large-Scale Training Systems - Keep: distributed training algorithms (data, tensor, pipeline, sequence, and expert parallelism, FSDP and ZeRO-style sharding); optimizers and preconditioners for large-scale pretraining; communication algorithms, overlap, and collective design; throughput and MFU work that materially changes what a training run costs or whether it converges; scaling-law work that informs how such runs are configured. - Filter: routine infrastructure tuning with no new algorithmic or systems idea.

  3. Architecture and Training Dynamics - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, dynamic or modular computation); optimisation and training-dynamics analysis that explains why large models train the way they do. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight.

  4. Efficiency and Compression - Keep: quantization, sparsity, pruning, low-rank adaptation, KV-cache and memory-efficient designs, and other work that materially changes the cost of running or training a large model, especially when the mechanism is new rather than a tuned variant. - Filter: straightforward application of standard efficiency methods, and deployment or serving engineering with no new idea.

Leave these to the hotspot digest unless the core contribution clearly falls in one of the four topics above: - major model or product releases - agent, tooling, and RAG launches - benchmark, leaderboard, or evaluation-only papers - representation-learning theory, interpretability, and feature-formation analysis - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning) - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains