Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,117 records · Page 62Linked to original sources

CTE-Bench: Counterfactual Trace Evaluation for Stateful Software Simulators

Coding agents change running software: they patch a service's code or overwrite its stored state, and then act on their own expectation of how the service will respond afterwards. A wrong expectation may surface only several calls later. Function-level code-execution benchmarks omit persistent service state, and agent benchmarks score the actions an agent takes or the final state it reaches. We introduce CTE-Bench, which measures whether a model can predict how an intervention changes a stateful service's future behavior, without asking it to choose actions. Each scenario gives the model Python service code, the calls and responses observed before the intervention, the intervention itself (a source edit or a state overwrite), and 40 fixed future calls; the model predicts every future response, and predictions are checked by executing the service. Three memory protocols control whether the model sees the correct earlier responses, none of them, or its own earlier predictions. CTE-Bench-Core-v1 contains 255 scenarios over six deterministic Python services, giving 10,200 predictions per model. The main score is effect-step value match (VM): exact response equality on the 2,476 future calls whose response the intervention changes. With correct earlier responses revealed, four API-hosted models (DeepSeek V4-Flash, Kimi K2.5, Qwen3.6-35B-A3B, and Claude Sonnet 4.6) reach 54.3%-61.5% effect-step VM. Hiding those responses lowers effect-step VM to 23.2%-28.9%; conditioning on self-generated predictions gives 24.8%-33.2%, and at most 1.2% of scenarios are predicted exactly end to end. Current models thus track intervention effects mainly when correct feedback is supplied, and their errors compound over a rollout. We release CTE-Bench-Core-v1 with its executable oracle, evaluation scripts, and an evaluation card mapping each claim to its protocol.

cs.SE↗

VLM4Cluster: Benchmarking Deep Clustering In the Era of Vision-Language Pre-training

Vision-language pre-training has reshaped image clustering, giving rise to language-assisted image clustering (LaIC), which leverages textual semantics to complement visual representations. Despite the rapid proliferation of LaIC methods, it remains unclear how much LaIC has actually advanced image clustering, as existing studies generally suffer from major limitations, including inconsistent experimental settings, inadequate dataset selection, and limited evaluation dimensions. To address this gap, we introduce VLM4Cluster, a comprehensive benchmark for image clustering in the era of pre-trained vision-language models (VLMs). VLM4Cluster implements 17 representative methods spanning classical, deep, and language-assisted image clustering, and evaluates them on 20 datasets covering classical, challenging, fine-grained, large-scale, and out-of-distribution settings. Beyond effectiveness, VLM4Cluster systematically investigates image clustering along three complementary dimensions: robustness to adversarial perturbations, generalization under distribution shifts, and computational efficiency. Our study shows that LaIC substantially advances the clustering performance frontier on many semantically demanding benchmarks, generally exhibits stronger generalization under distribution shifts, and achieves a more favorable effectiveness-efficiency trade-off. However, its gains become less consistent on large-scale and fine-grained datasets, while language assistance does not systematically reduce sensitivity to adversarial perturbations. VLM4Cluster is released at https://github.com/YuanweiHuu/VLM4Cluster.

cs.CV↗

Elliptic flow in an expanding and rotating fireball

We investigate the role of initial orbital angular momentum in generating elliptic flow in off-central heavy-ion collisions. The conventional expanding fireball picture is extended for the first time to a simultaneously expanding and rotating fireball characterized by an angular velocity. Spherical and spheroidal emission geometries with two different rotating flow profiles, reminiscent of global and differential rotation, are considered for our investigation. While the transverse-momentum spectra at midrapidity are largely insensitive to the geometry and flow profile for realistic values of angular velocity, the elliptic flow is strongly affected by anisotropic expansion and rotation. A spherical fireball generates elliptic flow solely through the momentum anisotropy induced by rotation, but its magnitude remains below the observed values for realistic angular velocities. In contrast, an anisotropically expanding and rotating spheroidal fireball exhibits a $\sim 10\%$ enhancement of elliptic flow for $\sqrt{s_{\rm NN}}=7.7$ GeV, which reduces to a $2\%$ for $\sqrt{s_{\rm NN}}=39$ GeV relative to its non-rotating counterpart. Within our fireball framework, this indicates that vorticity can contribute at the $2-10\%$ level to the observed elliptic flow.

nucl-th↗

Threshold Tails of Black hole Differential Observables

We study how first-order differential observables affect the zero-frequency behavior of black-hole wave equations. Quasinormal modes and late-time tails do not always change in the same way. We give a simple condition for when an observable removes the leading threshold term of a Green function. For Schwarzschild Regge-Wheeler modes, the relevant operator is determined by the regular static solution. The transformation is nondegenerate at nonzero frequency but becomes globally degenerate at zero frequency. As a result, the nonzero quasinormal-mode problem is unchanged, while the leading fixed-radius late-time tail gains one extra inverse power of time.

gr-qc↗

FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution

Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI spends tokens on irrelevant content. We introduce FocusVTC, which breaks this trade-off through adaptive resolution while preserving general multimodal capabilities. It combines compressed low-DPI global views with selective region enhancement, integrating enhanced views into ongoing reasoning. We construct 29.4K high-quality Reasoning-Evidence Localization (REL) chain-of-thought examples (REL-CoT) that link reasoning traces to page indices and bounding boxes. Multi-resolution REL supervised fine-tuning (REL-SFT) teaches the model to localize relevant regions, and Group Relative Policy Optimization learns when to enhance resolution and how to use the resulting observations, without a separate continual-pretraining stage. At 72 DPI on RULER v1, FocusVTC scores 87.4 at $2.9\times$ input compression, including tool observations, versus 57.5 for Glyph at $3.0\times$ input compression. It surpasses its text-input backbone on LongBench (56.40 versus 55.86), improves the MRCR macro-average by 13.91 points, and achieves a 51.19 macro-average on VTCBench. The MRCR latency evaluation also shows a $2.79\times$ online end-to-end speedup over Text. General multimodal capabilities are preserved, with MMMU increasing from 65.12 to 66.73 and MME from 2424.02 to 2457.62.

cs.CV↗

RankBuffer: Efficient Ranking-Based Rewards for Open-Ended Generation

Open-ended generation lacks canonical answers, making pointwise rewards difficult to calibrate for group-based reinforcement learning. Directly ranking same-query rollouts provides a more suitable relative reward signal, but existing ranking-based reward methods can incur substantial judging cost. We introduce RankBuffer, which maintains an ordered, query-specific buffer of previously judged responses as a reusable quality scale. Each rollout is first inserted into an anchor interval through an independent coarse judgment, after which only rollouts assigned to the same interval undergo local fine ranking. The resulting complete order is converted into bounded rank rewards, while boundary expansion, local refinement, and inactive-anchor pruning adapt the buffer as the policy evolves. Across four open-ended benchmarks, RankBuffer consistently outperforms all pointwise baselines. It also achieves nearly on-par performance with the strongest ranking-based reward baseline while substantially reducing judging cost. Ablations demonstrate the importance of both local fine ranking and anchor response content, while buffer analyses show that rollout-derived anchors progressively extend and refine the covered quality scale. These results establish response reuse as an effective approach to efficient relative reward construction.

cs.AI↗

Scheduling Recursive Reasoning in Looped Transformers

Recurrent reasoning models have attracted growing attention for scaling test-time computation, typically by iteratively refining latent states with shared parameters. However, these models apply each learned update with a fixed unit scale, which can be conservative when updates make persistent progress and overly aggressive when they fluctuate, limiting the benefit of additional loops. To understand how the scale should vary along the trajectory, we first analyze the sensitivity of terminal loss to recurrent update scale. We show that its temporal average admits an exact decomposition into persistent-progress and centered-fluctuation contributions. Based on this, we introduce the Trajectory Adaptive Progress-Fluctuation Scheduler (TAPS), which tracks their balance across recurrent updates and adapts the step size online. Theoretically, we establish sufficient conditions under which TAPS reduces expected terminal loss and reaches a target quality in fewer recurrent loops. Empirically, we show that TAPS improves terminal accuracy across structured reasoning tasks without retraining. By further incorporating the progress-fluctuation principle into training, TAPS yields additional accuracy gains with up to 1.56 times wall-clock speedup at matched baseline accuracy. The broad applicability of TAPS is supported by its effectiveness across diverse recurrent architectures and inference strategies. Together, these results establish update scale as complementary control axis of recurrent inference alongside architecture and depth.

cs.LG↗

Replay the Curvature: Accurate and Scalable NVFP4 Quantization for Large Language Model Inference

Large language models make weight storage and memory traffic major inference costs, motivating low-precision formats that represent each weight with only a few bits. Such formats use a scale to map floating-point values into a small codebook; NVFP4 improves local range utilization by letting every 16 E2M1 weights share an E4M3 block scale. Choosing that scale is difficult in GPTQ because quantizing one column updates those that follow, so evaluating a block independently can misestimate its final reconstruction error. Large models pose a second challenge: full-precision weights, calibration activations, and second-order state cannot all remain on one accelerator, while assigning complete layers to devices leaves each time-consuming layer solve serial. We introduce \emph{Schur Replay}, a scale-selection algorithm that reproduces the GPTQ updates caused by each block scale and scores the resulting block error after accounting for compensation from unquantized columns. Separately, our execution infrastructure keeps only the active layer resident, tiers activations across device, host, and disk, retires full-precision layers after export, and distributes independent output rows across tensor-parallel ranks. Together, the algorithm and infrastructure attain $99.35\%$ and $100.84\%$ question-weighted recovery from BF16 across seven benchmarks on Qwen3.5-397B-A17B and Llama-3.3-70B-Instruct. On the 397B model, the infrastructure reduces measured per-layer time by $15.17\times$ over ModelOpt and $23.14\times$ over LLM Compressor, with lower memory used per GPU.

cs.LG↗

Not Every Correction Helps: Gain-Guided Continual Test-Time Adaptation

Continual test-time adaptation (CTTA) adapts a source model to an unlabeled test stream whose distribution may change over time. Existing TTA methods often assess prediction reliability using confidence or entropy, which primarily reflect the model's self-certainty for the current sample. In CTTA, accumulated target observations can provide complementary evidence for correcting the source prediction, but this history may become misaligned as the target distribution changes. The key question is therefore not how much the correction differs from the source prediction, but whether and how strongly it should be applied. This paper proposes Gain-Aware INtervention (GAIN), a backpropagation-free CTTA framework guided by a simple principle: history proposes, gain decides. GAIN maintains compact target statistics to form a correction proposal and a posterior-predictive evaluator that accounts for estimation uncertainty. The resulting source-relative gain estimates the proposal's benefit and determines a sample-specific intervention strength along a continuous path through efficient one-dimensional optimization. Gain-controlled predictions then update the target statistics online, limiting the propagation of unreliable corrections, all without backpropagation, sample storage, or replay. Across five benchmarks, our method achieves strong predictive performance, with favorable accuracy--calibration--efficiency trade-offs in continual adaptation. On ImageNet-C, for example, GAIN achieves 61.9% accuracy with near-source calibration. It remains stable under diverse and challenging continual shifts while running 15.9x faster than a representative optimization-based CTTA baseline.

cs.CV↗

An Ionized Superstructure at Cosmic Dawn Revealed by a Foundation Model for Astrophysical Research

The spatial structure of cosmic reionization remains poorly constrained. Using a foundation model trained for James Webb Space Telescope datasets, we identify a luminous galaxy at redshift $z=7.567$ showing strong Ly$α$ transmission at $z=6.80\pm0.07$. The transmission extends over $210.5^{+55.7}_{-53.5}$ comoving megaparsecs, indicating an ionized superstructure with hydrogen neutral fraction $x_{\rm HI}=(2.07^{+0.29}_{-0.30})\times10^{-6}$, in an epoch when the cosmic mean neutral fraction is $\sim0.4$. Cosmological simulations suggest a probability of $\approx10^{-5}$ for a random sightline to reproduce the signal. We find no significant galaxy overdensity associated with the transmission region, in tension with canonical inside-out reionization models. These observations severely challenge existing reionization models in the standard $Λ$CDM Universe, and provide the first example of deep learning discovering a previously unknown astrophysical phenomenon.

astro-ph.GA↗

Constitutional adapters: Inference-time interventions for misalignment and misuse

Training models to act in accordance with an explicitly defined set of principles, or "constitution," has shown promise as a robust and transparent mechanism for AI alignment. However, the generality and flexibility of such methods remain unclear. Here, we show that constitution-consistent behavior can be distilled from synthetic corpora into lightweight objects (low-rank adapters and steering vectors). Despite never seeing a harmful request or jailbreak during training, such objects increase jailbreak defense success and measured alignment -- particularly at long context lengths and against multi-turn attacks, where they outperform both prompted and steered baselines. Subtracting control-trained from constitution-trained objects further accentuates these effects, yielding defenses we call "constitutional adapters" (CAs). CAs can be trained on a base model, transferred zero-shot to its post-trained checkpoint, and scaled at inference time to predictably trade off defense for benign compliance. Taken together, these results recommend CAs as a lightweight, portable, and tunable lever for mitigating misalignment and misuse in API deployments.

cs.LG↗

Single-Photon Nonlinearity from a Nanobeam with a Quantum Dot

An efficient single-photon nonlinearity is a key resource for photonic quantum technologies. Single-photon nonlinearities have been demonstrated with quantum dots (QDs) across a range of cavity and waveguide geometries. However, none of these realizations simultaneously provide high-efficiency direct fiber coupling and compatibility with scalable on-chip integration. The nanobeam cavity addresses these limitations through its in-plane geometry, which enables highly efficient, direct fiber coupling alongside scalable homogeneous and heterogeneous on-chip integration. Here, we report a single-photon nonlinearity from a nanobeam cavity coupled to a single quantum dot. The nanobeam interfaces directly with a single-mode fiber with an efficiency of 60%. Driven by resonant picosecond pulses, the platform achieves a single-photon nonlinearity with a low threshold of 0.24 incident photons per pulse. The system operates in the strong-coupling regime, exhibiting a coupling rate of g/2π = 29.1 GHz and a cooperativity of C = 3.98. These results establish the nanobeam cavity-QD system as a compact, chip-integrable, and fiber-compatible building block for scalable photonic quantum technologies.

physics.optics↗

On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training

The strong generalization performance of on-policy post-training paradigms has motivated studies of their parameter update behaviors. However, these studies treat the observed behaviors only as byproducts in on-policy training, overlooking their potential to serve as optimization principles for improving the generalization of other paradigms such as supervised fine-tuning (SFT). To address this limitation, we investigate whether there exists a specific on-policy update behavior that can achieve such improvements. First, our analyses reveal that SFT updates parameters along consistent directions, while the on-policy paradigm continuously adjusts the direction during training. This difference inspires us to focus on the cumulative update direction of each parameter as a promising behavior. Then, we evaluate its effectiveness for improving generalization by proposing On-Policy direction-constrained Supervised Fine-Tuning (OPSFT), which constrains SFT updates to the direction identified by on-policy paradigms. The strong performance of OPSFT indicates that the generalization advantage of on-policy paradigms can be transferred to SFT through the parameter update direction. Once such a direction is identified, even SFT can generalize with its updates constrained to this direction. This finding offers two practical benefits by combining the strong generalization of on-policy paradigms with the advantages of SFT, including the high training efficiency and ability to leverage high-quality trajectories. For efficiency, we identify update directions that support strong generalization using a few on-policy training steps, and subsequently apply OPSFT to achieve high training efficiency. For leveraging high-quality trajectories, OPSFT can utilize these trajectories to continue improving a post-trained model along its update direction without disrupting the ability learned from on-policy training.

cs.LG↗

Byzantine-Robust Federated Representation Learning

We study federated learning (FL) with adversarial clients, where the goal is to minimize the average loss of the honest (non-adversarial) clients without knowing their identity. Under heterogeneity, a single shared model parameter is statistically inappropriate: it cannot capture the distinct data-generating processes across clients, incurring an irreducible model-heterogeneity bias and severely limiting robustness to adversarial clients (a.k.a. Byzantine-robustness). We address this problem through representation learning, where each client learns a personalized linear head, while collaboratively estimating a shared nonlinear representation through Byzantine-robust aggregation. We demonstrate that the heterogeneity among honest representation gradients is controlled by the representation error and statistical errors that decay either with the number of data samples per client ($τ$) or the number of iterations ($T$). In particular, our non-asymptotic parameter recovery error bound reveals three terms: (i) an initialization-dependent error that goes away with $T$, (ii) finite-sample noise terms that decreases with $τ$ and the number of honest clients, and (iii) a stochastic gradient variance term that also reduces with $T$. Importantly, with no irreducible model-heterogeneity bias in our bounds. We extend the regression analysis to multiclass classification, and empirically validate it on CIFAR-10, FEMNIST, and School Exam Score datasets.

cs.LG↗

You Only Reprogram Once: Rethinking Prolonged Training for Visual Reprogramming

Visual reprogramming is a parameter-efficient method for adapting pretrained models, yet its training can remain computationally expensive: even with a frozen backbone, visual prompts are often optimized through the full model for hundreds of epochs. Before changing what the pretrained model sees, we ask whether we are fully using what it already tells us. We find that modeling the full source response can already yield strong downstream predictions without prompt optimization. Motivated by this observation, we introduce You Only Reprogram Once (YORO), which constructs a downstream predictor from the frozen response space in a single forward-only traversal. Its Bayesian Discriminant Mapping (BDM) derives a covariance-aware affine mapping from streaming class statistics, requiring no backpropagation, optimizer updates, or repeated visits to the training set. When further input adaptation helps, YORO-FP optionally refines the visual prompt for 20 epochs. BDM also extends naturally to CLIP by treating attribute-prompt similarities as source responses. Across three full-data settings, YORO improves average accuracy over the strongest prior gradient-free mapping by 18.4--24.4\%. On 16-shot CLIP, it raises the four-backbone average from 71.4\% to 77.2\%. YORO-FP provides further gains on selected tasks, while validation often retains the one-pass predictor. These results suggest a different default for visual reprogramming: read out the frozen response first, and optimize the input only when needed.

cs.CV↗

AutoLoCo: Communication Efficient Distributed LLM Training via Adaptive Synchronization

The pre-training of Large Language Models (LLMs) is increasingly conducted across multiple data centers. As training scales to a larger number of accelerators, the fraction of time spent on computation decreases, while the fraction spent on communication increases. Therefore, frequent synchronization becomes a growing bottleneck. Local update methods reduce this cost by allowing workers to perform several optimizer steps between synchronizations. Most local update methods set the number of local optimizer steps between synchronizations before training and keep this interval fixed throughout the run. However, the best interval can change during the entire train process. If the interval and optimizer are adapted to the current training state, the communication frequency is reduced while maintaining the training performance. In this work, we introduce AutoLoCo, an adaptive training framework to reduce communication in LLM training. It adapts the local interval using scalar training statistics and corrects each outer update. Our method is motivated by two observations: 1) the appropriate local interval varies across training stages, and 2) changing the number of inner steps per interval creates a mismatch with an unchanged outer optimizer, requiring a correction to the outer update. We optimize this mismatch by correction of the outer optimizer for the momentum and the learning rate using the accumulated inner learning rate. Our experiments under communication constraints demonstrate that AutoLoCo reduces communication frequency by 27% relative to DiLoCo while maintaining training performance.

cs.LG↗

GenLimitLib: A Formal Library for Language Generation in the Limit and AI-Assisted Mathematical Research

We present GenLimitLib, a source-aligned Lean 4 library for language generation in the limit. Introduced by Kleinberg and Mullainathan at NeurIPS 2024, language generation in the limit studies a theoretical question motivated by LLMs: how to generate valid new strings from observed examples. This young and rapidly evolving field offers a natural testbed for studying large-scale formalization. GenLimitLib contains formal developments for 30 papers. It extracts shared definitions and reusable proof components while preserving paper-specific assumptions and statements, and records relationships across papers. In this way, GenLimitLib provides a concrete and structured view of the literature. We show through mathematical case studies and LLM experiments how our library can support both human mathematical research and AI-assisted research. Our Library: https://github.com/pengzhang91/generation-in-the-limit-lib.

cs.LG↗

Linear and nonlinear transport responses of topological nodal-line semimetals

Topological nodal-line semimetals are three-dimensional quantum materials characterized by band crossings that form closed loops in momentum space. In $\mathcal{PT}$-symmetric realizations, these nodal rings are stabilized in the absence of spin-orbit coupling, giving rise to drumhead surface states and unconventional transport responses. In this work, we study charge transport across a nodal-line semimetal containing a finite electrostatic barrier, with both leads described by the same equilibrium material. By solving the corresponding scattering problem, we show that the transmission across the barrier exhibits {Klein-tunneling behavior protected at normal incidence by the nodal topology}, despite the extended nodal-line dispersion, which can be traced back to Berry-curvature-induced momentum locking. Using the Landauer-Büttiker formalism, we derive general expressions for the linear and nonlinear conductances, including both longitudinal and Hall components, and evaluate them at zero and finite temperature. Our analytical and numerical results elucidate the dependence of the conductance on barrier height and width, as well as on a $\mathcal{PT}$-breaking mass term. We identify distinct transport regimes in which nonlinear contributions are strongly enhanced and transverse Hall currents emerge, providing clear transport signatures of nodal-line topology and suggesting potential routes toward device applications.

cond-mat.mes-hall↗