This is a remedial run for missed papers from 08/15/2026 to 08/16/2026.
Results generated on 09/13/2026.
Personalized Daily ArXiv Papers 2026-08-17
| Model | Metric | Usage | Papers | ||||
|---|---|---|---|---|---|---|---|
| Prompt | Completion | Total | Total arXiv | Scanned | Relevant | ||
local agent CLI |
Tokens | not reported | not reported | not reported | 479 | 479 | 27 |
| Cost | not reported | not reported | not reported | ||||
Token counts are not reported for this run. 20 of 37 model calls succeeded, 16,825s of model wall clock.
Topic Coverage:
| Topic | Papers |
|---|---|
| Large-Scale Training Systems and Efficiency | 5 |
| Architecture and Training Dynamics | 7 |
| Efficiency, Compression, and Large-Scale Training | 15 |
Table of contents by topic:
Large-Scale Training Systems and Efficiency (5)
-
SubZero+: Efficient Zeroth-Order LLM Fine-Tuning via Large Learning Rates Authors: Ziming Yu, Shuyao Xiao, Xingyu Zhao, Sike Wang, Pan Zhou, Peiyu Zang, Xiangda Yan, Yongjie Yang, Jia Li
-
Scaling Laws for Dynamic Mini-Batch SGD in Sketched Linear Regression Authors: Ziyan Chen, Zhongzhu Zhou, Ding-Xuan Zhou
-
EMASAM: a Computationally Efficient Sharpness-Aware Minimization via EMA-Guided Perturbations Authors: Tanapat Ratchatorn, Masayuki Tanaka
-
Deterministic Adam-Inspired Methods with Accelerated Convergence Rate Authors: Yaxin Yu, Long Chen, Zeyi Xu
-
UniFed-VLM: Federated Instruction Tuning for Vision-Language Models with Multiple Heterogeneity Authors: Pengyu Wang, Baochen Xiong, Xiaoshan Yang, Yifan Xu, Zhang Qimeng, Haifeng Chen, Changsheng Xu
Architecture and Training Dynamics (7)
-
Directional Curvature from Armijo Backtracking: A Low-Cost Sharpness Probe and a Calibration-Free Learning-Rate Safeguard for Adam Authors: Ashmitha R, Jörg Frochte
-
Second-Moment Memory in Coordinatewise Adam Authors: Jeonseong Kim
-
Measuring Structured Predictability in Neural Training Dynamics: A Cross-Regime Study Authors: Fanqi Wang, Weisheng Tang, Hairong Qi
-
QuantumPhaseNet: A Gauge-Covariant Geometric and Quantum-Spectral Theory of Semantic Concept Hierarchies with Prototype Validation of a Classical Quantum-Inspired Model Authors: Kiyotaka Kasubuchi, Kazuo Fukiya
-
Layers Matter: Why Continual Learning Regularization Should Be Layer-Adaptive Authors: Brian B. Moser, Ahmed Anwar, Tobias Christian Nauen, Shishir Muralidhara, Federico Raue, René Schuster, Stanislav Frolov, Andreas Dengel
-
Near-Equilibrium Propagation training in nonlinear wave systems Authors: Karol Sajnok, MichaÅ Matuszewski
-
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Authors: Christian Schlarmann, Francesco Croce, Nicolas Flammarion, Matthias Hein
Efficiency, Compression, and Large-Scale Training (15)
-
Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't Authors: Ravi Satya Durga Prasad Yenugula
-
The Distributional View of Knowledge Distillation Authors: Gordei Verbii, Juho Lee
-
LACE-SVD: Loss-Aware SVD with Cumulative Error Correction for LLM Compression Authors: Zhuowen Liu, Longkun Hao, Shiyu Feng, Xiaowen Chang, Ruiqun Li, Changqun Li
-
SAPE: Sandwich Adapters for Parameter Efficiency in Large Language Model Fine-Tuning Authors: Mohammad Aref Jafari-Raddani, Morteza Mohajjel Kafshdooz
-
FedPA-LoRA: Product-Aligned Framework for Mitigating Aggregation and Initialization Errors in Heterogeneous Federated LoRA Authors: Juseok Jeon, Ramy E. Ali, Doyun Kwon, Myungbeom Her, Jinhwi Kim, Jinhyun So
-
DeltaLog: Deferred Materialization of Recurrent States for Linear Attention Decoding Authors: Junqing Lin, Jingwei Sun, Guangzhong Sun
-
Pricing the Risk of Runtime Compression: Anytime-Valid Admission and a Served-Output Law for Compressed Serving State Authors: Fanzhe Wei, Li Liu
-
KV-Rescue: Recovering Reasoning Language Model KV Eviction Loss via Stepwise Interleaving Authors: Minsoo Cheong, Woosang Lim, Vincent-Daniel Yun, Sungjoo Yoo
-
Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching Authors: Chang Zou, Shikang Zheng, Evelyn Zhang, Runlin Guo, Haohang Xu, Zhengyi Shi, Conghui He, Xuming Hu, Linfeng Zhang
-
Algorithm-Architecture Co-Design for Efficient VLA Inference via Speculative Inference and Verification Authors: Chunyu Qi, Zhuoran Song, Jian Weng, Haozhe Jiang, Xueyuan Liu, Naifeng Jing, Guanghui He, Xiaoyao Liang, Haibing Guan
-
Global Simulation-Guided Dynamic Operator Scheduling for Efficient Multi-Tenant Model Serving Authors: Weinan Liu, Zeyuan Ding, Dian Ding, Chengcheng Wan, Lu Tang, Guangtao Xue, Jiwu Shu, Yiming Zhang
-
Runtime Observability for Heterogeneous Attention Memory Authors: Fanzhe Wei, Li Liu, Ziyang Wang, Chenyu Wang
-
Funnel of Thoughts: Efficient Test-Time Scaling via Early Voting and Rollout Pruning Authors: Chanhee Park, Sungbin Han, Jeongho Yoon, Seongtae Hong, Heuiseok Lim
-
Whole-Piece Training for Symbolic Music Language Models via Full-Horizon Compressed Recurrence Authors: Yungang Yi, Weihua Li, Matthew Kuo, Catherine Shi, Quan Bai
-
PhaseLoRA: Control-Regime-Conditioned Low-Rank Adaptation for Continuous-Action Vision-Language-Action Policies Authors: Yufei Guo, Yinan Wu, Haoran Duan, Guiguang Ding, Jungong Han
Large-Scale Training Systems and Efficiency (5)
1. SubZero+: Efficient Zeroth-Order LLM Fine-Tuning via Large Learning Rates
ArXiv ID: 2608.15665
Primary Topic: Large-Scale Training Systems and Efficiency
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Ziming Yu, Shuyao Xiao, Xingyu Zhao, Sike Wang, Pan Zhou, Peiyu Zang, Xiangda Yan, Yongjie Yang, Jia Li
Abstract: Zeroth-order (ZO) optimization enables backpropagation-free fine-tuning of large language models, but existing ZO methods suffer from high-variance gradient estimators, making convergence unstable and highly sensitive to learning rates. We propose SubZero+, an improved SubZero framework that improves stability in three complementary ways: (i) multi-query gradient estimation within layer-specific low-rank subspaces to reduce variance without exhibiting the multi-query paradox; (ii) a subspace Adam optimizer that performs adaptive updates using in-subspace multi-query gradient statistics; and (iii) a sign correction for QR-based subspace construction to ensure Haar-distributed projection matrices, eliminating implementation-dependent orientation ambiguity. Experiments on models from 1.3B to 32B across SuperGLUE, under both full-parameter tuning and LoRA, show that SubZero+ consistently outperforms prior ZO baselines, enlarges the stable learning-rate range, and narrows the gap to first-order methods with minimal extra memory overhead.
Comment: Multi-query gradient estimation in low-rank subspaces stabilizes zeroth-order LLM optimization across larger learning rates.
Topic Match: Gradient-estimator and adaptive-optimizer design are central, with an additional match through memory-efficient, backpropagation-free LLM training.
Relevance: 9 Novelty: 6
2. Scaling Laws for Dynamic Mini-Batch SGD in Sketched Linear Regression
ArXiv ID: 2605.24316
Primary Topic: Large-Scale Training Systems and Efficiency
Also Matches: Architecture and Training Dynamics
Authors: Ziyan Chen, Zhongzhu Zhou, Ding-Xuan Zhou
Abstract: Mini-batching is central to large-scale optimization, yet its role in statistical scaling laws remains limited. We study one-pass and multi-pass batch SGD for sketched linear regression under power-law spectral and source conditions. Our analysis reveals a two-horizon phenomenon induced by warmup--stable--decay schedules: deterministic learning is governed by the full optimization trajectory, while stochastic error retains only a shorter terminal memory. For dynamic batch schedules, the individual batch sizes enter through influence-weighted summaries that measure how strongly each update affects the final risk. Consequently, batching leaves the approximation and optimization-bias laws unchanged at a fixed update horizon, but controls the one-pass variance and the multi-pass fluctuation around full-batch gradient descent. We obtain matching one-pass variance bounds and nearly matching multi-pass fluctuation bounds, recover static-batch and full-batch behavior as special cases, and derive an oracle square-root rule for allocating a fixed iteration budget. These results identify WSD horizon separation and final-risk influence as the mechanisms governing dynamic mini-batch scaling.
Comment: Derives dynamic-mini-batch scaling laws through a two-horizon account of optimization bias and stochastic variance.
Topic Match: Batch-schedule scaling laws most directly inform large-scale training configuration, while the horizon analysis also concerns training dynamics.
Relevance: 7 Novelty: 7
3. EMASAM: a Computationally Efficient Sharpness-Aware Minimization via EMA-Guided Perturbations
ArXiv ID: 2608.15105
Primary Topic: Large-Scale Training Systems and Efficiency
Authors: Tanapat Ratchatorn, Masayuki Tanaka
Abstract: Recent progress in optimization research has highlighted the sharpness of the loss landscape as a key factor in narrowing the generalization gap. Motivated by this insight, Sharpness-Aware Minimization (SAM) was proposed as a training strategy that enhances generalization. Despite the promising performance, SAM suffers from its twice computational cost due to its core algorithm requiring an extra gradient computation during the perturbation step. To overcome this limitation, we introduce Exponential Moving Average Sharpness-Aware Minimization (EMASAM), a computationally efficient variant of SAM. EMASAM does not require the loss gradient in the perturbation step. Instead, EMASAM defines the perturbation direction based on the discrepancy between the main model and the EMA shadow model. This perturbation travels away from the stable average position toward the less stable area, acting as a softer yet cheaper alternative to SAM's worst-case scenario perturbation. Moreover, since EMASAM's perturbation does not rely on noisy mini-batch gradients, it mitigates the gradient-induced instability inherent in SAM. Hence, EMASAM eliminates the need for an extra backpropagation while also preserving the generalization ability of the SAM-style training. Several experiments have been performed and confirm the efficiency and robustness of our method.
Comment: EMA-derived SAM perturbations eliminate the extra gradient computation required by conventional SAM.
Topic Match: The optimizer directly reduces training computation, although the abstract does not establish effectiveness at large-scale pretraining.
Relevance: 7 Novelty: 6
4. Deterministic Adam-Inspired Methods with Accelerated Convergence Rate
ArXiv ID: 2604.08742
Primary Topic: Large-Scale Training Systems and Efficiency
Authors: Yaxin Yu, Long Chen, Zeyi Xu
Abstract: Adam is widely used, but its convergence theory remains incomplete even in the deterministic full-batch setting because momentum and adaptive preconditioning are tightly coupled. For smooth convex objectives, we split the momentum variable through variable-and-operator splitting, which reveals the acceleration mechanism. We then combine a Hessian-driven correction with Adam-style feedback based on the gradient magnitude. The resulting Adam-HNAG (Hessian-driven Nesterov accelerated gradient with Adam-style adaptive preconditioning) flow admits a nonnegative energy that decays exponentially. Its discretization yields two methods, Adam-HNAG and the synchronous variant Adam-HNAG-s. Under the stated trajectory-bound and consistency conditions, both methods satisfy a discrete Lyapunov contraction. If the exact adaptive steps are accepted, this contraction gives an $O(k^{-2})$ objective-value bound. Numerical experiments illustrate their behavior. These results apply to the proposed methods, not to the original Adam recursion.
Comment: Momentum–preconditioner splitting yields conditional accelerated-convergence guarantees for new Adam-inspired optimizers.
Topic Match: Optimizer design is the nearest topic, but the guarantees concern new deterministic convex methods rather than original Adam or demonstrated large-scale pretraining.
Relevance: 6 Novelty: 7
5. UniFed-VLM: Federated Instruction Tuning for Vision-Language Models with Multiple Heterogeneity
ArXiv ID: 2608.15516
Primary Topic: Large-Scale Training Systems and Efficiency
Also Matches: Efficiency, Compression, and Large-Scale Training
Authors: Pengyu Wang, Baochen Xiong, Xiaoshan Yang, Yifan Xu, Zhang Qimeng, Haifeng Chen, Changsheng Xu
Abstract: Vision-Language Models (VLMs) have demonstrated strong performance in multimodal understanding and generation. However, fine-tuning of VLMs typically relies on centralized data, which raises privacy concerns in certain domains (e.g. healthcare). Federated Learning (FL) provides a natural solution by enabling model training without sharing raw data. However, applying FL to VLM instruction tuning is highly challenging. VLMs have substantial parameter scales, and in real-world scenarios, clients exhibit significant heterogeneity in tasks, modalities, and model architectures. Existing methods mainly focus on simplified settings and are unable to handle such multi-dimensional heterogeneous scenarios. In this work, we study federated instruction tuning under joint heterogeneity in tasks, modalities, and model architectures. We propose UniFed-VLM, a unified federated instruction tuning framework for VLMs that addresses multiple types of heterogeneity. It consists of two key components: 1) Federated Compensated Subspace Aggregation (FedCSA), which performs subspace-aligned aggregation of parameter-efficient adapters with dynamic weighting and compensation to mitigate heterogeneity-induced conflicts; 2) Two-stage Collaborative Distillation (TCoD), which enables effective knowledge transfer across heterogeneous models via a Mutual Distillation Adapter (MDA) and a mixture-of-experts-based distillation strategy. We conduct experiments on multiple benchmark datasets, and the results show that UniFed-VLM achieves stronger average performance across diverse tasks compared with existing FL methods. The source code is available at: https://github.com/wangpengyu2004/UniFed-VLM.
Comment: Performs subspace-aligned aggregation of parameter-efficient adapters across heterogeneous federated clients.
Topic Match: The strongest fit is its new distributed aggregation algorithm, with adapter efficiency as a secondary contribution.
Relevance: 6 Novelty: 6
Architecture and Training Dynamics (7)
1. Directional Curvature from Armijo Backtracking: A Low-Cost Sharpness Probe and a Calibration-Free Learning-Rate Safeguard for Adam
ArXiv ID: 2607.03998
Primary Topic: Architecture and Training Dynamics
Also Matches: Large-Scale Training Systems and Efficiency
Authors: Ashmitha R, Jörg Frochte
Abstract: The local sharpness of the loss, the top Hessian eigenvalue $λ_1$, determines the largest stable gradient step, but measuring it normally requires Lanczos or Hessian-vector products. A single Armijo backtracking line search already carries this information at the cost of a few forward passes: the accepted step $α$ brackets the directional curvature along the probed direction within the multiplicative band set by the backtracking factor: exactly the curvature averaged over the tested step, and empirically $q = g^\top H g/|g|^2$ to within that band. Across CIFAR-10, Fashion-MNIST and Imagenette, $\logα$ tracks $\logλ_1$ at Pearson $-0.91$ to $-0.95$, and the relation survives a per-run detrending check at $-0.60$ to $-0.70$, a low-cost online Edge-of-Stability reading of the slow sharpness component. Used as a safeguard rather than a faster optimiser, the reading caps a too-large initial learning rate. A single fixed protocol, probing along Adam's own update direction at initialisation and over the first fifty optimiser steps and capping the rate at twice the smallest reading, removes every divergence across learning-rate grids spanning $10^{-3}$ to $3.0$ and at GPT-2 pretraining scale, and all but one marginal case across the further architectures we test, at about $1\%$ overhead, and it leaves training bit-identical whenever the cap does not bind. No constant in the protocol is tuned per architecture; this is the sense in which the safeguard is calibration-free. The guarantee is divergence, not accuracy: where the productive range is narrow the capped run survives at strongly reduced accuracy (chance level on AG News at aggressive rates), and our measurements show why any cap frozen at initialisation must fail at pretraining scale: the loss surface sharpens within the first five optimiser steps, the gap warmup has always filled by convention.
Comment: Low-cost Armijo curvature probes safeguard Adam against early-training divergence through learning-rate caps.
Topic Match: The central contribution connects directional curvature to training stability and turns that insight into an optimizer safeguard tested during GPT-2 pretraining.
Relevance: 9 Novelty: 7
2. Second-Moment Memory in Coordinatewise Adam
ArXiv ID: 2608.15824
Primary Topic: Architecture and Training Dynamics
Authors: Jeonseong Kim
Abstract: Adam retains a moving average of past squared gradients in its denominator, but the optimization cost of this memory is not well understood. We show that second-moment memory can itself suppress progress toward the optimum even under finite-variance stochastic gradients. For a simple two-point oracle, the expected positive normalized update is $O(M_2^{-1/2})$ after an initialization transient, where $M_2=(1-β_2)^{-1}$ is the second-moment memory length. We convert this directional bound, under the stated memory and stepsize scaling, into an average-stationarity lower bound of the same order on a smooth convex problem with normalized gap, smoothness, and variance. Long second-moment memory can slow optimization even when the gradient noise has finite variance.
Comment: A lower bound shows how Adam's second-moment memory can suppress normalized updates and slow convergence.
Topic Match: The paper isolates a training-dynamics mechanism in Adam, with theoretical evidence from a controlled finite-variance convex setting.
Relevance: 8 Novelty: 7
3. Measuring Structured Predictability in Neural Training Dynamics: A Cross-Regime Study
ArXiv ID: 2608.15483
Primary Topic: Architecture and Training Dynamics
Authors: Fanqi Wang, Weisheng Tang, Hairong Qi
Abstract: Modern deep networks are trained through long update trajectories, yet their temporal organization remains less systematically characterized than architectures, losses, or optimizers. We study short-horizon predictability as a measure of temporal redundancy: where, when, and under which training conditions recent updates contain information about near-future parameter motion. We combine three complementary probe families, displacement-direction, subspace-residual, and predictor-based probes, with convention-aware, null-calibrated group-level readouts, and apply them to multi-pass vision training on CIFAR and public Pythia pretraining checkpoints. Across both regimes, vector-like tensors such as normalization parameters and biases (auxiliary parameters) exhibit simpler short-horizon dynamics than matrix-like feature-transforming weights (bulk parameters), whose predictable behavior concentrates in localized, time-varying pockets. Agreement within and across probe families, and with independent trajectory diagnostics, indicates that these measurements capture intrinsic trajectory structure, while probe differences distinguish complementary forms of temporal organization. Controlled CIFAR comparisons further show that architecture and training recipe systematically modulate the measured structure. A Pythia-70M case study further exposes a sequence of role-, depth-, and scale-dependent events, including bulk ESA falling below the random sign-agreement level and the emergence and redistribution of predictable qkv pockets across layers. These results position short-horizon predictability as a retrospective, parameter-resolved diagnostic of training dynamics.
Comment: Parameter-trajectory probes reveal more predictable short-horizon updates in normalization parameters and biases than in matrix weights.
Topic Match: The study directly characterizes optimization dynamics across parameter roles and training regimes, including Pythia pretraining.
Relevance: 8 Novelty: 6
4. QuantumPhaseNet: A Gauge-Covariant Geometric and Quantum-Spectral Theory of Semantic Concept Hierarchies with Prototype Validation of a Classical Quantum-Inspired Model
ArXiv ID: 2608.15820
Primary Topic: Architecture and Training Dynamics
Authors: Kiyotaka Kasubuchi, Kazuo Fukiya
Abstract: We present QuantumPhaseNet, a gauge-covariant geometric and quantum-spectral extension of Transformer representations. Context-dependent semantic states are modeled as complex amplitudes; a covariant phase rate induces a semantic wavelength used as a proxy for conceptual scale; and low-frequency graph modes define a document-level discourse direction. The theoretical part establishes local gauge invariance, unitarity of the quantum block, boundedness and conditional stability of WavePhase Attention, and a calibratable hallucination-risk formulation. We also implemented a fully offline Validation Studio for the classical quantum-inspired pipeline in Section 14.1 and evaluated the five research questions in Section 16.1 on its built-in synthetic setting (n=240, observation noise 0.22, circuit noise 0.08, five seeds). RQ1 yielded a wavelength-hierarchy Spearman correlation of 0.852 versus 0.707 for the baseline, 87.3% direction accuracy, and AUC 0.953. RQ2 achieved discourse alignment 0.933 versus 0.589 and 41.2 versus 16.2 paragraphs before drift. RQ3 achieved AUROC 0.881 versus cosine 0.765 and phase-shuffle 0.536. RQ4 achieved error-detection AUROC 0.854 versus entropy 0.634, with Brier 0.150 and ECE 0.098. RQ5 did not show quantum advantage: target probability and end-to-end cost efficiency were 25.5% and 0.107, compared with 70.7% and 0.707 for the Chebyshev classical approximation. These results provide initial synthetic evidence for the classical quantum-inspired components, but not external validity or unconditional quantum speedup.
Comment: Proposes gauge-covariant semantic states and WavePhase Attention as a quantum-spectral extension of transformer computation.
Topic Match: Its strongest foundational fit is the proposed attention and representation mechanism, despite validation being limited to synthetic data.
Relevance: 6 Novelty: 8
5. Layers Matter: Why Continual Learning Regularization Should Be Layer-Adaptive
ArXiv ID: 2608.15901
Primary Topic: Architecture and Training Dynamics
Authors: Brian B. Moser, Ahmed Anwar, Tobias Christian Nauen, Shishir Muralidhara, Federico Raue, René Schuster, Stanislav Frolov, Andreas Dengel
Abstract: Continual learning regularizers like EWC fight forgetting by penalizing changes from previous-task parameters with per-parameter importance, typically diagonal Fisher values. Per-parameter looks more flexible than per-layer, but each layer's diagonal Fisher is a weak summary of its actual curvature, missing the top-eigenvalue information that controls forgetting. Adversarial bit-flip attacks and Hessian-spectrum studies show that this missing per-layer sensitivity spans orders of magnitude in neural networks. Under a block-diagonal Hessian assumption, the layer-level analogue of EWC's existing diagonal assumption, we prove three things. Forgetting decomposes as a sum of per-layer terms weighted by each layer's top Hessian eigenvalue. Diagonal-Fisher weights cannot recover this eigenvalue. For instance, two layers with identical Fisher averages can have top eigenvalues differing by a factor as large as the layer width. For the same level of forgetting, uniform regularization loses new-task performance by an amount scaling with the layer condition number. Our theoretical analysis leads to a simple recipe: protect early layers strongly, let deeper layers move. We apply this recipe to EWC and SLCA and show clear improvements in average performance and forgetting metrics.
Comment: Layerwise Hessian sensitivity motivates adaptive regularization for the forgetting–plasticity tradeoff.
Topic Match: The core contribution explains how curvature affects regularization and forgetting, with scope concentrated on continual learning.
Relevance: 7 Novelty: 6
6. Near-Equilibrium Propagation training in nonlinear wave systems
ArXiv ID: 2510.16084
Primary Topic: Architecture and Training Dynamics
Authors: Karol Sajnok, MichaÅ Matuszewski
Abstract: Backpropagation learning algorithm, the workhorse of modern artificial intelligence, is notoriously difficult to implement in physical neural networks. Equilibrium Propagation (EP) is an alternative with comparable efficiency and strong potential for in-situ training. We extend EP learning to both discrete and continuous complex-valued wave systems. In contrast to previous EP implementations, our scheme is valid in the weakly dissipative regime, and readily applicable to a wide range of physical settings, even without well defined nodes, where trainable inter-node connections can be replaced by trainable local potential. We test the method in driven-dissipative exciton-polariton condensates governed by generalized Gross-Pitaevskii dynamics. Numerical studies on standard benchmarks, including a simple logical task and handwritten-digit recognition, demonstrate stable convergence, establishing a practical route to in-situ learning in physical systems in which system control is restricted to local parameters.
Comment: Extends equilibrium-propagation learning to weakly dissipative complex wave systems using trainable local potentials.
Topic Match: Alternative training dynamics is the closest fit, but the contribution concerns physical neural substrates with limited connection to large-model training.
Relevance: 6 Novelty: 7
7. FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
ArXiv ID: 2506.03096
Primary Topic: Architecture and Training Dynamics
Authors: Christian Schlarmann, Francesco Croce, Nicolas Flammarion, Matthias Hein
Abstract: Contrastive language-image pre-training aligns features of text-image pairs in a common latent space via distinct encoders for each modality. While this approach achieves impressive performance in several zero-shot tasks, it cannot natively handle multimodal inputs, i.e., encoding image and text into a single feature vector. As a remedy, it is common practice to use additional modules to merge the features extracted by unimodal encoders. In this work, we present FuseLIP, a new architecture for multimodal embedding. Leveraging recent progress in discrete image tokenizers, we propose to use a single transformer model operating on a unified vocabulary of text and image tokens. This early fusion approach allows the different modalities to interact at each depth of encoding and obtain richer representations compared to common late fusion. We collect new datasets for multimodal pre-training and evaluation, designing challenging tasks for multimodal encoders. We show that FuseLIP outperforms late fusion approaches in several multimodal and unimodal embedding tasks.
Comment: Fuses discrete image and text tokens from the first transformer layer through a single shared vocabulary and encoder.
Topic Match: The early-fusion transformer design is a genuine architectural mechanism, though developed for multimodal embedding tasks.
Relevance: 6 Novelty: 7
Efficiency, Compression, and Large-Scale Training (15)
1. Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't
ArXiv ID: 2608.02829
Primary Topic: Efficiency, Compression, and Large-Scale Training
Also Matches: Architecture and Training Dynamics
Authors: Ravi Satya Durga Prasad Yenugula
Abstract: Model families are typically trained size by size, each from scratch. Can a pretrained large model instead be converted into a smaller sibling? We characterize the 1.4B->410M conversion in Pythia end to end. Representations align strongly across sizes (ridge R^2=0.84) while parameters align weakly. Dense weight projection is functionally destructive, and a bit-exact control shows this is not an assembly artifact: basis mixing breaks rotary, per-head, GELU, and LayerNorm structure. After the best-fit linear operator, weight residuals are statistically indistinguishable from noise under shuffle controls. Conversion value therefore lives in initialization. In matched-budget continued pre-training we decompose conversion into two independent levers: least-squares compensation (function lever, best zero-shot) and variance-preserving rescale (dynamics lever, best endpoints). Compensation is a token-efficient, low-budget win rather than a universal one. At 30M tokens it beats the strongest subcloning variant on both a width-reduced pair (84.0 +/- 1.8 vs. 89.7 +/- 3.7, 3/3 seeds) and a held-out depth-reduced pair (109.3 vs. 117.9, 3/3 seeds), reaching a given quality with fewer tokens. At a 33x larger budget the two converge to parity (40.0 vs. 40.0), both far ahead of from-scratch, which transfer initialization always beats: by up to 18x at low budget, with the margin narrowing at convergence and at the largest scale. We also map the method's boundary. At about 5x the donor scale (6.9B->1.4B) stacking both levers over-corrects, consistent with ill-conditioning of the compensation solve at large width, which points to dimension-aware regularization as a fix. At matched budget our initialization also beats structured pruning with distillation, the standard pipeline, and improves further combined with it. Code, checkpoints, and the frozen evaluation corpus are released.
Comment: Least-squares compensation and variance rescaling improve the token efficiency of large-to-small checkpoint initialization.
Topic Match: Checkpoint reuse directly targets smaller-model pretraining cost, supported by analysis of structural constraints and initialization dynamics.
Relevance: 9 Novelty: 7
2. The Distributional View of Knowledge Distillation
ArXiv ID: 2608.15215
Primary Topic: Efficiency, Compression, and Large-Scale Training
Also Matches: Architecture and Training Dynamics
Authors: Gordei Verbii, Juho Lee
Abstract: Token-level knowledge distillation (KD) matches two conditional distributions per position, yet the standard objectives compare them pointwise: a Kullback-Leibler gradient is blind to which wrong token receives probability mass. We develop a distributional view in which the teacher is represented not by a single softened output but by a family of multi-temperature views - marginals of the annealing path of its logits - and the student is trained against a geometry-aware aggregate of these views under an embedding-based ground cost. We formalize the resulting design space (mixtures, log-linear pooling, entropic Wasserstein barycenters, and a debiased Sinkhorn-divergence flagship in hub and path forms), prove an exact collapse result showing log-linear pooling of tempered views is equivalent to a single temperature, and give a multi-marginal Schrodinger-bridge reading that yields falsifiable predictions. On instruction-tuned Pythia pairs, experiments yield three empirical laws: (i) dispersion law - the benefit of multi-temperature aggregation grows monotonically with the effective temperature dispersion of the views, not with their number; (ii) dispersed views unlock the aggregation operator - the barycenter separates from the arithmetic mixture exactly when transport-based aggregation starts to beat averaging; and (iii) two-regime picture governed by the ceiling gap $Î=\mathrm{PPL}{\mathrm{SFT}}-\mathrm{PPL}$: when the fine-tuned teacher barely beats a supervised student the gentle transport objective is the best KD loss but no KD beats supervised fine-tuning, whereas at a real ceiling the ranking inverts - and the sign of the fidelity-generalization correlation flips. We argue that "which distillation loss is the best" is not a fixed property of the loss but a function of $Î$.
Comment: Geometry-aware aggregation of multi-temperature teacher distributions introduces new knowledge-distillation objectives.
Topic Match: Compression-oriented distillation is primary; the analysis of teacher headroom also explains changes in distillation training behavior.
Relevance: 8 Novelty: 8
3. LACE-SVD: Loss-Aware SVD with Cumulative Error Correction for LLM Compression
ArXiv ID: 2607.03057
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Zhuowen Liu, Longkun Hao, Shiyu Feng, Xiaowen Chang, Ruiqun Li, Changqun Li
Abstract: The rapid growth in the parameter scale of large language models (LLMs) has created a strong demand for efficient compression techniques. As a hardware-agnostic and highly compatible approach, low-rank compression has been widely adopted to reduce both memory footprint and computational cost. However, existing SVD-based methods are still largely driven by local reconstruction objectives, overlooking two critical limitations: rank budgets are often allocated without explicitly considering layer-wise loss sensitivity, and local approximation errors can propagate and accumulate through the residual stream, leading to amplified global deviations from the original model. To address these issues, we propose LACE-SVD, a Loss-Aware SVD framework with Cumulative Error correction for LLM compression. LACE-SVD first estimates the calibration negative-log-likelihood increase induced by candidate layer-wise compression ratios and solves a budget-constrained allocation problem to assign rank budgets. It then refines the compressed model with closed-form local updates and introduces a propagation-aware correction for residual-stream output modules, reducing layer-output discrepancy as a proxy for cumulative error propagation. Experimental results demonstrate that at a high compression ratio (0.6), the WikiText-2 PPL of our method on LLaMA-7B (32.57) is significantly better than that of Dobi-SVD (46.18).
Comment: Loss-aware low-rank compression allocates rank budgets and corrects propagated approximation errors.
Topic Match: The core contribution directly improves LLM low-rank compression through sensitivity-aware budgeting and cumulative-error correction.
Relevance: 9 Novelty: 6
4. SAPE: Sandwich Adapters for Parameter Efficiency in Large Language Model Fine-Tuning
ArXiv ID: 2608.15360
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Mohammad Aref Jafari-Raddani, Morteza Mohajjel Kafshdooz
Abstract: While Parameter-Efficient Fine-Tuning (PEFT) has substantially reduced the hardware cost of adapting Large Language Models (LLMs) by decreasing the number of trainable parameters, recent studies have sought to further improve PEFT through parameter sharing. However, these approaches either employ uniform parameter sharing across layers, which can delay convergence, or rely on dynamic masking strategies, which add computational overhead. The potential of sharing patterns inspired by the inherent hierarchical structure of Transformer architectures remains unexplored in PEFT. To address this gap, we introduce SAPE (Sandwich Adapters for Parameter Efficiency), a PEFT framework based on a sandwich-style hard weight-sharing topology. SAPE routes intermediate Transformer layers through balanced shared group adapters while strictly isolating the input embedding and final projection boundary transformations to prevent gradient interference. This design significantly reduces memory consumption while eliminating the computational overhead associated with dynamic parameter-sharing methods. Extensive evaluations across encoder-only and causal decoder architectures demonstrate that SAPE achieves state-of-the-art performance in low-parameter regimes. On natural language understanding, SAPE outperforms proPETL on RoBERTa-large while utilizing only 10% of the baseline's parameter budget. On natural language generation and world knowledge reasoning with LLaMA-3.2 (3B) under a strict ~0.6M parameter constraint, SAPE outperforms AdaLoRA, yielding absolute improvements of +4.85% on GSM8K and +3.11% on CommonsenseQA. Furthermore, through comprehensive topological ablations, we formalize an inherent capacity trade-off: while hard parameter sharing strongly regularizes semantic generalization, it slightly smooths the sharp layer-wise transformations required for rigid multi-step arithmetic reasoning.
Comment: Boundary-isolated adapter sharing reduces trainable parameters and memory while limiting gradient interference.
Topic Match: The core contribution is a parameter-sharing topology that directly improves LLM adaptation efficiency.
Relevance: 9 Novelty: 6
5. FedPA-LoRA: Product-Aligned Framework for Mitigating Aggregation and Initialization Errors in Heterogeneous Federated LoRA
ArXiv ID: 2608.15381
Primary Topic: Efficiency, Compression, and Large-Scale Training
Also Matches: Large-Scale Training Systems and Efficiency
Authors: Juseok Jeon, Ramy E. Ali, Doyun Kwon, Myungbeom Her, Jinhwi Kim, Jinhyun So
Abstract: Low-Rank Adaptation (LoRA) enables efficient federated fine-tuning of large language models, but its factorized parameterization creates a tension between accurate aggregation of local updates and continuity of locally optimized factors. Factor-wise aggregation incurs aggregation mismatch but better preserves factor continuity, whereas product-space reconstruction reduces this mismatch at the cost of greater factor-level initialization mismatch from newly reconstructed factors. We propose FedPA-LoRA, a product-aligned federated LoRA framework that jointly addresses these limitations and provably converges under both homogeneous and heterogeneous client ranks. Each client preserves its local factors across communication rounds and aligns its product toward a rank-specific global reference, maintaining local optimization continuity while promoting global consistency under data heterogeneity. The server aggregates heterogeneous-rank updates in the common product space and efficiently reconstructs a rank-constrained global adapter without forming the dense aggregate. This design supports client-specific computation and communication budgets. Experiments on natural language understanding and generation tasks show that FedPA-LoRA consistently outperforms representative baselines across varying levels of data heterogeneity and homogeneous- and heterogeneous-rank settings, with up to a $6.82$ percentage-point improvement in average GLUE accuracy under heterogeneous client ranks.
Comment: Product-aligned aggregation supports heterogeneous-rank LoRA training without forming a dense aggregate.
Topic Match: Rank-constrained adapter optimization directly advances efficient LLM training; heterogeneous federated aggregation also contributes to distributed training algorithms.
Relevance: 8 Novelty: 7
6. DeltaLog: Deferred Materialization of Recurrent States for Linear Attention Decoding
ArXiv ID: 2608.15533
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Junqing Lin, Jingwei Sun, Guangzhong Sun
Abstract: Linear attention models eliminate the quadratic prefix computation and context-growing KV cache of softmax attention by replacing pairwise token interactions with recurrent state updates. However, existing decoding implementations often materialize and write back the full recurrent state after every generated token, making state maintenance a major source of memory traffic, especially for models with large states and many heads. This paper presents DeltaLog, a recurrent-state decoding scheme that reduces this overhead without changing the model semantics. Specifically, DeltaLog represents the recurrent state as a dense base state together with a bounded log of recent compact updates. Most decode steps append only compact update factors to this log, while periodic merge steps fold the accumulated updates back into the dense base state. Thus, the model observes the same dense state as in eager decoding, but most full-state write-backs are replaced by lightweight append operations. We implement DeltaLog for GDN, KDA, and RWKV6 and integrate it into a prototype serving stack. Across these models, DeltaLog accelerates the recurrent-state update kernel by up to $1.86\times$, reduces profiled recurrent-state write traffic by up to $7.83\times$, and achieves $1.05$--$1.20\times$ end-to-end serving speedups over dense recurrent baselines.
Comment: Deferred materialization of compact recurrent-state updates reduces linear-attention decoding memory traffic.
Topic Match: An exact state-representation and update scheme reduces memory writes, providing a substantive inference-efficiency mechanism across multiple recurrent architectures.
Relevance: 8 Novelty: 7
7. Pricing the Risk of Runtime Compression: Anytime-Valid Admission and a Served-Output Law for Compressed Serving State
ArXiv ID: 2608.15810
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Fanzhe Wei, Li Liu
Abstract: Runtime compression of serving state trades quality for capacity with no priced guarantee: systems adapt precision on load signals with no soundness statement, and certified approaches budget request-level risk by a union bound over a pre-declared event count. We show the union budget exhausts on every long request in a production serving stack (100% of requests), and replace it with an anytime-valid, physically accounted ledger whose bound holds at every one of 352,333 admission calls on live traffic and which, in a pre-registered held-out confirmatory round, halves the exact-fallback rate at matched risk (0.30 -> 0.14) -- coverage is bought at a price the account states. We then price the remaining distance from the certified witness to what a user experiences: a machine-checked design law (TV <= tanh(a_q w_thr)) turns the served-TV target into a threshold knob, and a three-layer audit of its instantiation -- an operator-norm query envelope measured 1.5x from tight, a measured-ellipsoid replacement for the Cauchy-Schwarz ball that buys nothing (0.89x, held-out sound), and the gate's operating point (~700x) -- localizes the entire 1064x gap to the operating point, a price the law now states rather than an unknown. A priced bound is worth nothing on a request one has not seen, so the third link is the quantifier: exchangeable extrapolation across 80 serving histories replaces binary conformal prediction's vacuous certificates with order-statistic bounds that discriminate (0.41 against 0.51 calibration risk). All probabilistic kernels are Lean 4-checked (228 exported theorems, no sorry); which object deserves this machinery at all is settled empirically in a companion paper that adjudicates -- and rejects -- the natural alternative of certifying routing. What ships is an account: risk you can spend, a gap you can read off a law, and a bound that survives the request you have not seen.
Comment: Anytime-valid compression-risk accounting reduces exact fallbacks at matched risk and links thresholds to output-distribution error.
Topic Match: Introduces risk-budgeting and admission mechanisms that improve the efficiency-quality trade-off of compressed serving state.
Relevance: 8 Novelty: 7
8. KV-Rescue: Recovering Reasoning Language Model KV Eviction Loss via Stepwise Interleaving
ArXiv ID: 2608.15797
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Minsoo Cheong, Woosang Lim, Vincent-Daniel Yun, Sungjoo Yoo
Abstract: KV-cache eviction caps the memory cost of long reasoning traces but is inherently lossy because the model decodes from a partial view of its history. Under aggressive budgets, this not only lowers accuracy but can also cause runaway degeneration, where the model produces incoherent or repetitive tokens until reaching the length limit. We characterize much of this loss as an information gapf caused by missing context, rather than a capability gap caused by limited model capacity. An evicted 7B model and a full-context 1.5B model make complementary errors, and an oracle choice between their answers recovers 79% of the accuracy gap to the full-KV 7B model. Based on this observation, we propose KV-Rescue, a training-free inference framework that bridges the information gap introduced by KV eviction using a lightweight full-context helper. KV-Rescue interleaves reasoning steps from the two models into a shared trajectory. An online detector uses entropy and compressibility to terminate the generation of incoherent or repetitive base-model candidates early. Across five math benchmarks with Qwen2.5-Math 7B and 72B, KV-Rescue recovers an average of 87% of the accuracy lost to eviction at eviction budget B=64. A decode-cost analysis further shows that preventing runaway degeneration cuts base-model token generation by 43% on average.
Comment: Stepwise interleaving with a small full-context helper recovers accuracy lost to aggressive KV-cache eviction.
Topic Match: The central mechanism addresses the accuracy and generation-cost consequences of memory-constrained KV caching.
Relevance: 8 Novelty: 7
9. Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching
ArXiv ID: 2412.18911
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Chang Zou, Shikang Zheng, Evelyn Zhang, Runlin Guo, Haohang Xu, Zhengyi Shi, Conghui He, Xuming Hu, Linfeng Zhang
Abstract: Diffusion Transformers (DiT) have become the dominant methods in image and video generation yet still suffer substantial computational costs. As an effective approach for DiT acceleration, feature caching methods are designed to cache the features of DiT in previous timesteps and reuse them in the next timesteps, allowing us to skip the computation in the next timesteps. Among them, token-wise feature caching has been introduced to perform different caching ratios for different tokens in DiTs, aiming to skip the computation for unimportant tokens while still computing the important ones. In this paper, we propose to carefully check the effectiveness in token-wise feature caching with the following two questions: (1) Is it really necessary to compute the so-called "important" tokens in each step? (2) Are so-called important tokens really important? Surprisingly, this paper gives some counter-intuition answers, demonstrating that consistently computing the selected important tokens'' in all steps is not necessary. The selection of the so-calledimportant tokens'' is often ineffective, and even sometimes shows inferior performance than random selection. Based on these observations, this paper introduces dual feature caching referred to as DuCa, which performs aggressive caching strategy and conservative caching strategy iteratively and selects the tokens for computing randomly. Extensive experimental results demonstrate the effectiveness of our method in DiT, PixArt, FLUX, and OpenSora, demonstrating significant improvements than the previous token-wise feature caching.
Comment: Alternating aggressive and conservative feature caching reduces diffusion-transformer computation.
Topic Match: The contribution changes intermediate-feature reuse and token-computation scheduling across large diffusion models, directly targeting inference efficiency.
Relevance: 8 Novelty: 6
10. Algorithm-Architecture Co-Design for Efficient VLA Inference via Speculative Inference and Verification
ArXiv ID: 2608.15636
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Chunyu Qi, Zhuoran Song, Jian Weng, Haozhe Jiang, Xueyuan Liu, Naifeng Jing, Guanghui He, Xiaoyao Liang, Haibing Guan
Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in the field of embodied AI, but their high computational cost and limited predicted action length hinder real-time deployment. Although Dadu-Corki, a dedicated accelerator for efficient embodied AI, has been introduced, it does not exploit the inherent interaction patterns between the robot and its environment, which results in a relatively short predicted action length. We observe that robotic environments naturally alternate between active states-where precise actions are crucial-and inactive states-where actions have limited impact on task success. This insight enables a new scheduling opportunity: long-action-length speculative prediction in inactive states, paired with selective verification in active states. We propose SpecVLA, an algorithm-system co-design framework that adaptively balances action length, inference latency, and task reliability. On the algorithm side, SpecVLA introduces a state-aware VLA inference execution paradigm and a hardware-friendly construction of a smaller verification model (sVLA) using differential residuals and block-wise mixed-precision quantization. On the system side, we develop a heterogeneous architecture consisting of a GPU and a robotic-specific hardware module, along with a speculative dataflow that decouples VLA and sVLA through parallel execution. Comprehensive evaluations on OpenVLA and RDT across LIBERO and ManiSkill benchmarks show that SpecVLA reduces end-to-end latency significantly while preserving task success rate. By enabling long-action-length speculative prediction with timely verification, SpecVLA achieves real-time robotic manipulation with both high efficiency and reliability.
Comment: Environment-state-aware speculation and selective verification reduce VLA inference cost.
Topic Match: The core contribution is an algorithm-hardware inference-efficiency mechanism, although its scheduling opportunities depend on robotic environment states.
Relevance: 7 Novelty: 7
11. Global Simulation-Guided Dynamic Operator Scheduling for Efficient Multi-Tenant Model Serving
ArXiv ID: 2608.15762
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Weinan Liu, Zeyuan Ding, Dian Ding, Chengcheng Wan, Lu Tang, Guangtao Xue, Jiwu Shu, Yiming Zhang
Abstract: Container-granularity scheduling leaves abundant short-lived idle slices within containers unexploited. Reallocating containers is too heavyweight to utilize such fine-grained opportunities under SLA constraints, and operator-level scheduling requires reasoning about dependencies, memory safety, and cluster-wide execution dynamics in real time. In this paper, we present SliceScheduler, a dynamic operator-level scheduling system for multi-tenant model serving. The key idea is to expose cluster-wide operator execution state and enable what-if reasoning over scheduling decisions. SliceScheduler consists of four key components. First, we introduce the Global Mapping Graph (GMG), a unified abstraction that captures operator dependencies, tensor shapes, resource mappings, and execution states, providing a real-time, cluster-wide view with explicit resource semantics. Second, we build a global simulator on top of GMG to predict operator-level execution and memory evolution under candidate placements. Third, we design an incremental, simulation-based scheduling module that selects placements to exploit fragmented idle slices while avoiding memory violations and preserving SLA. Finally, we develop an operator executor that materializes scheduling decisions on GPUs and coordinates computation and cross-accelerator transfers. We implement SliceScheduler as a PyTorch backend and evaluate it using production trace replay. Experimental results show that SliceScheduler improves token throughput by 1.10--2.29$\times$ compared to existing approaches, while maintaining SLA violations within 9\%. SliceScheduler demonstrates that operator-level scheduling is a practical and effective approach to improving GPU utilization for multi-tenant LLM serving.
Comment: Simulation-guided operator scheduling exploits fragmented GPU capacity while tracking dependencies and memory safety.
Topic Match: The operator-level scheduling mechanism materially improves LLM runtime efficiency, qualifying under the efficiency topic despite its serving focus.
Relevance: 7 Novelty: 7
12. Runtime Observability for Heterogeneous Attention Memory
ArXiv ID: 2608.05863
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Fanzhe Wei, Li Liu, Ziyang Wang, Chenyu Wang
Abstract: Modern models no longer keep a plain KV cache: latent caches, learned sparse selectors and recurrent states each carry the model's memory in a different form, and each fails differently under compression. We give a runtime observability contract that covers all four memory classes with three operators, instantiate it on six model configurations across five architecture families, and compose the per-stage bounds into an executable request-level risk ledger. Contracts carry their error metric as a type -- composition is only defined when metrics match, and this check rejected our own first composed chain; the repaired chain crosses metrics through two proved bridges, and whatever no formal system can certify is measured instead, dropping the composed tier to empirical automatically: every claim is certified, partially certified, or empirical, composition inherits the weakest tier, and the tier is decided by the machine. Replayed over $12.4$M entry reads and run under eight-way concurrency with per-request budgets and fail-closed identity attribution, the ledger quantifies the honest trade-off on today's witness and holds its risk budget with zero violations. A fused always-on probe observes a declared one-layer subset under CUDA graphs inside the serving noise floor. Applied to a served DeepSeek-V4 stack with a packed compressed-KV prototype, the same machinery localizes a silent corruption to a precise structural boundary -- exact in the eviction-free, identity-isolated regime, with every observed failure in an eviction or slot-reuse regime -- through a machine-adjudicated discrimination campaign whose calculus rejected two of our own confounded inferences along the way. All artifacts, guards, and the Lean development are released at https://github.com/metask-ai/witprobe-attention-memory; every number in this paper regenerates from the shipped artifacts by one command.
Comment: Typed error-bound composition tracks compression risk across KV, latent, sparse, and recurrent attention state.
Topic Match: New compression-risk contracts support memory-efficient inference, with the contribution concentrated on runtime guarantees and diagnostics.
Relevance: 7 Novelty: 7
13. Funnel of Thoughts: Efficient Test-Time Scaling via Early Voting and Rollout Pruning
ArXiv ID: 2608.15065
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Chanhee Park, Sungbin Han, Jeongho Yoon, Seongtae Hong, Heuiseok Lim
Abstract: Large Reasoning Models produce diverse, sometimes inconsistent answers across repeated queries on the same problem, so multi-sample inference is a prerequisite for reliable deployment. Majority voting at k rollouts is the standard solution and the de facto accuracy target for this regime, but it is prohibitively expensive at the scale LRMs require. We introduce Funnel of Thoughts (FoT), an inference-time method that preserves the full 32-trajectory voted accuracy while halving its attention FLOPs, a 28.8% reduction in full-model inference cost. Across 115K reasoning trajectories from six LRMs, we find that unproductive trajectories often reveal themselves through repeated hesitation markers such as "Wait", "Actually", and "perhaps." These trajectories are less likely to reach the correct answer and consume disproportionate attention FLOPs, degenerating into no-answer loops in the worst case. Built on this training-free lexical signal, FoT identifies the vocabulary that captures these pathological patterns and prunes affected trajectories before completion, reducing online generation attention FLOPs by 56.1% and wall time by 37.6% without any additional model inference; the same signal transfers without retuning across held-out architectures and out-of-domain tasks.
Comment: Early voting and lexical-signal rollout pruning reduce redundant reasoning-model inference computation.
Topic Match: Introduces a computation-saving inference mechanism, with relevance concentrated on test-time sampling efficiency.
Relevance: 7 Novelty: 6
14. Whole-Piece Training for Symbolic Music Language Models via Full-Horizon Compressed Recurrence
ArXiv ID: 2602.19816
Primary Topic: Efficiency, Compression, and Large-Scale Training
Also Matches: Architecture and Training Dynamics
Authors: Yungang Yi, Weihua Li, Matthew Kuo, Catherine Shi, Quan Bai
Abstract: For computational efficiency, modern language models are typically trained on independently sampled fixed-length sequences. Symbolic music language models largely inherit this paradigm, despite musical structure naturally unfolding over complete compositions rather than isolated excerpts. Fragmenting compositions into independent training instances therefore prevents continuous conditioning over the complete work. We present a practical framework for whole-piece training of symbolic music language models via Full-Horizon Compressed Recurrence (FHCR). FHCR preserves the full temporal horizon of recurrent memory while reducing the dimensionality of its key-value (KV) representation, making continuous whole-piece training practical under limited GPU memory. To directly assess functional long-range dependence, we introduce KV-Reset Context Utilization (KRCU), an evaluation-time diagnostic. On the MAESTRO symbolic piano dataset, KRCU shows that full-horizon models utilize context far beyond the local segment window, whereas reducing the temporal extent of recurrent memory substantially weakens this measurable long-range dependence. FHCR preserves long-range context utilization while substantially reducing recurrent memory cost. These findings show that preserving the temporal extent of recurrent history is important for efficient whole-piece modeling, and that memory cost can instead be reduced through KV representation compression. The project demos and generated music samples are available at https://wholemusic.github.io.
Comment: Compressed recurrent KV representations reduce training memory while retaining the full sequence horizon.
Topic Match: The substantive mechanism reduces recurrent-state memory during training, although validation is confined to symbolic music.
Relevance: 7 Novelty: 6
15. PhaseLoRA: Control-Regime-Conditioned Low-Rank Adaptation for Continuous-Action Vision-Language-Action Policies
ArXiv ID: 2608.15285
Primary Topic: Efficiency, Compression, and Large-Scale Training
Authors: Yufei Guo, Yinan Wu, Haoran Duan, Guiguang Ding, Jungong Han
Abstract: Parameter-efficient fine-tuning (PEFT) is a natural way to adapt pretrained vision-language-action (VLA) policies, but most adapter designs apply temporally static updates throughout a control rollout, overlooking the phase-dependent nature of continuous-action manipulation. Such policies traverse distinct regimes, including approach, contact transition, grasping, transport, and placement, each requiring different adaptation behaviors. We propose \textbf{PhaseLoRA}, a lightweight LoRA parameterization that conditions adaptation at each action-chunk prediction step using two weakly supervised descriptors: fine-control tendency and event/boundary intensity. PhaseLoRA modulates the LoRA left factor in the action expert, allowing the effective low-rank update direction to vary over time while keeping the backbone largely frozen. On LIBERO, PhaseLoRA improves average success rate by 12.2 points over a matched-parameter high-rank LoRA baseline and outperforms stronger LoRA variants. Ablations show that random temporal modulation and scalar gating do not reproduce the performance of the full model, while update-direction analyses reveal structured temporal variation associated with the predicted control descriptors. These results establish within-trajectory conditioning as an effective lightweight PEFT axis for continuous-action VLA policies.
Comment: Control-conditioned modulation changes LoRA update directions within a rollout under a fixed parameter budget.
Topic Match: The new low-rank parameterization is a substantive PEFT mechanism, with relevance narrowed by its control-specific conditioning and robotic evaluations.
Relevance: 7 Novelty: 6
Paper Selection Prompt
System Prompt
You are a helpful paper reading assistant whose job is to read daily posts from ArXiv and identify a few papers that your friend will enjoy reading. Your job is to carefully read the paper titles and abstracts below and find the ones that match the criteria below.
User Prompt
Instructions
Respond in JSONL. Output exactly one JSON object per paper, one per line:
{"ARXIVID":"...","COMMENT":"...","RELEVANCE":0,"NOVELTY":0,"PRIMARY_TOPIC_ID":"...","MATCHED_TOPIC_IDS":[],"TOPIC_MATCH_COMMENT":"...","HOTSPOT_PAPER_TAGS":[],"HOTSPOT_PAPER_COMMENT":"..."}
Rules: - ARXIVID: the arXiv ID. - COMMENT: identify the single strongest matching criterion. Be brief and specific. Do not rely on generic phrases like "language modeling" or "advancement". Do not mention non-matching criteria. - RELEVANCE: integer from 1 to 10. - NOVELTY: integer from 1 to 10. - PRIMARY_TOPIC_ID: exactly one stable topic ID from the allowed topic registry. - MATCHED_TOPIC_IDS: zero or more stable topic IDs from the same allowed set. Include PRIMARY_TOPIC_ID when there are multiple matches. - TOPIC_MATCH_COMMENT: briefly explain why the primary topic is the best fit. - HOTSPOT_PAPER_TAGS: zero or more tags from this exact set only:
daily_hot,new_frontier. - HOTSPOT_PAPER_COMMENT: briefly explain why the paper belongs in the daily hotspot paper feed when HOTSPOT_PAPER_TAGS is non-empty; otherwise use an empty string. - Use HOTSPOT_PAPER_TAGS sparingly. Most papers should return[]. -daily_hotmeans the paper feels broadly important to the day and belongs in the daily hotspot paper section even if it is not part of the personalized foundational reading list. -new_frontiermeans the paper appears to open a genuinely new direction, paradigm, or field, even if the work is still early. - Do not output markdown, code fences, or any extra text.Scoring Criteria
Relevance and Novelty are independent axes. Score both from 1 to 10.
Relevance Scoring
- 9-10: directly centered on the target foundational topics; highest when the core contribution is clearly within them.
- 7-8: substantially related, but partly peripheral or focused on a narrower aspect.
- 5-6: touches the target topics, but the main contribution is elsewhere.
- 3-4: largely outside the target topics, often application-focused or domain-specific.
- 1-2: unrelated.
Important: Broad frontier relevance, major launch status, or daily buzz is not enough for a high Relevance score here. Those cases belong in the hotspot digest unless the paper strongly matches the specialized paper topics.
Novelty Scoring
- 9-10: new paradigm, theory, or major methodological breakthrough.
- 7-8: substantial methodological advance or strong new insight.
- 5-6: meaningful but incremental extension or refinement.
- 3-4: minor, narrow, or mostly engineering or domain-specific improvement.
- 1-2: little originality; mainly standard application of existing methods.
Topic Registry
Use exactly one PRIMARY_TOPIC_ID chosen from the stable topic IDs below. - moe_training: MoE Training - Mixture-of-Experts training: routing, load balancing, expert granularity and shared experts, training stability, MoE parallelism and communication, and MoE kernels. - training_systems: Large-Scale Training Systems and Efficiency - Large-scale training systems and efficiency: distributed training algorithms, optimizers, and communication. - architecture_training: Architecture and Training Dynamics - Core architectural or computational mechanisms, dynamic computation, and training-stability dynamics. - efficiency_scaling: Efficiency, Compression, and Large-Scale Training - Compression, sparsity, memory or cache efficiency, and large-scale training systems that materially change cost or behavior.
Papers
[PAPER LIST HERE]
Relevant Topics
This is a training-side feed. The centre of gravity is Mixture-of-Experts training and the large-scale training systems it sits on; architecture and efficiency work is kept when it bears on how a large model is trained or what it costs to train.
Keep a paper when its CORE CONTRIBUTION falls in one of the four topics below. Filter it when the topic is merely the setting for a downstream application, a benchmark, or a product result.
MoE Training (primary) - Keep: expert routing and gating mechanisms; load balancing, with or without auxiliary losses; expert granularity, fine-grained and shared experts; MoE training stability (loss spikes, router collapse, z-loss and related fixes); expert and hybrid parallelism; all-to-all dispatch/combine and the communication schedules MoE needs; MoE kernels and grouped-GEMM style implementations; expert capacity, token dropping, and upcycling dense checkpoints into MoE. - Filter: papers that merely train on top of a MoE model without a new routing, balancing, stability, parallelism, or kernel contribution; "mixture of experts" in the classical ensemble or recommender sense.
Large-Scale Training Systems - Keep: distributed training algorithms (data, tensor, pipeline, sequence, and expert parallelism, FSDP and ZeRO-style sharding); optimizers and preconditioners for large-scale pretraining; communication algorithms, overlap, and collective design; throughput and MFU work that materially changes what a training run costs or whether it converges; scaling-law work that informs how such runs are configured. - Filter: routine infrastructure tuning with no new algorithmic or systems idea.
Architecture and Training Dynamics - Keep: new or analysed architectural mechanisms (attention variants, normalization and residual design, state-space and recurrent sequence modelling, dynamic or modular computation); optimisation and training-dynamics analysis that explains why large models train the way they do. - Filter: applying an existing architecture to a new task or benchmark with no mechanistic insight.
Efficiency and Compression - Keep: quantization, sparsity, pruning, low-rank adaptation, KV-cache and memory-efficient designs, and other work that materially changes the cost of running or training a large model, especially when the mechanism is new rather than a tuned variant. - Filter: straightforward application of standard efficiency methods, and deployment or serving engineering with no new idea.
Leave these to the hotspot digest unless the core contribution clearly falls in one of the four topics above: - major model or product releases - agent, tooling, and RAG launches - benchmark, leaderboard, or evaluation-only papers - representation-learning theory, interpretability, and feature-formation analysis - memory mechanisms and agent memory systems - world models, exploration, and reinforcement learning - alignment and post-training (RLHF, DPO, GRPO, RFT, instruction tuning) - downstream applications in medical imaging, segmentation, 3D vision, video understanding, information retrieval, summarization, recommendation, machine translation, speech recognition, time series, knowledge graphs, and similar domains