Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection

Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. Although recent language model-based log anomaly detectors achieve strong detection performance, their confidence estimates remain poorly calibrated. We show that these detectors frequently assign excessive confidence to incorrect predictions, particularly for anomalous logs under severe class imbalance. Moreover, confidence on erroneous predictions remains persistently high even when conventional calibration metrics indicate good calibration, creating a critical reliability gap for operational monitoring systems. To address this issue, we propose Log Reconstruction and Distance (LoRD), a lightweight post-hoc calibration framework for reliable log anomaly detection. LoRD learns prediction-route-specific reliability models from latent representations of correctly classified validation samples and estimates prediction reliability through route-wise reconstruction distances. Based on the estimated reliability, LoRD selectively recalibrates high-risk predictions to suppress overconfident errors while preserving reliable predictions. Extensive experiments on four large-scale log benchmark datasets and multiple language model-based detectors demonstrate that LoRD consistently improves confidence reliability and substantially reduces overconfident anomaly-related errors without sacrificing anomaly detection performance.

cs.LG↗

Beyond Forgetting: Diagnosing and Harnessing Shared Reasoning in Continual RLVR

Reinforcement learning with verifiable rewards (RLVR) commonly post-trains reasoning models on multiple tasks, while rerunning multitask RLVR (MTRL) as new tasks are added makes capability expansion costly. We therefore study continual RLVR, which updates the existing model as each task arrives. The central question is whether a model updated this way can perform as well as a jointly trained model. To answer this question, we introduce Continual Reasoning Gym, a continual-RLVR environment that organizes text and visual reasoning tasks into five task sequences. In this setting, we identify two key observations: Sequential RLVR exhibits modest forgetting, yet its final performance remains below that of MTRL. To understand the latter, we decompose final performance and show that forgetting accounts for only part of the gap. To explain the former, we identify shared reasoning: transferable reasoning structure allows training on one task to support others on average. We therefore introduce Continual Prompt Replay (CPR), which harnesses shared reasoning to improve learning on the arriving and future tasks by replaying previous-task prompts and regenerating their responses with the current policy. On average, only CPR reaches MTRL-level performance.

cs.LG↗

Metrics That Write Themselves: Evolving an Evaluator from Its Own Blind Spots

Agents improve quickly against a reliable automatic metric and stall without one, and the applications that need them most, report generation among them, are the ones nobody knows how to score. Can the metric write itself? Saying what makes an answer good is hard; pointing at something wrong with one is easier, so the metric we evolve is a pool of small Python operators that each flag a candidate for one named defect, or abstain, and vote. Asking a model for operators directly does not work: 183 candidates realise only 96 distinct behaviours, from one narrow region of an enormous space. EvalCEGAR instead borrows counterexample-guided abstraction refinement from program verification. It reads the pool as an abstraction and searches for a collision, two answers the operators score identically, one correct and one not. That pair, not a prompt, is the authoring request, and when a collision defeats every attempt the loop widens what an operator may read rather than resampling. On MBPP+ and HumanEval+, a sandbox whose hidden unit tests give exact ground truth, the loop writes a 55-line operator that closes 15.4% of the gap between flagging nothing and a perfect filter on 428 unseen tasks (+0.0065, p=0.0010) at a quarter of our best hand-written operator's flags. On the benchmark it never saw it matches that operator's effect exactly on a third of the flags. Six of eight runs admit such an operator and all six help out of sample; our 15 hand-written operators applied together as one filter lose accuracy. An LLM judge on the same information ties that delta on a nearly disjoint set of candidates, and charges a model call per candidate forever where the operator charges none.

cs.AI↗

Cluster Assignments in Soft Targets Shape Speech Representations: Evidence from S-JEPA

Cluster-based prediction is widely used in self-supervised speech learning. A soft target preserves a distribution over clusters rather than a single label. This distribution specifies both the probability values and which clusters receive them. Comparisons between soft targets and hard labels do not separate the contributions of these two aspects to the learned representation. We study this in S-JEPA, a recent high-performing self-supervised speech model trained with soft Gaussian mixture model (GMM) targets. We compare its original targets with counterfactual targets that preserve the most likely cluster and all probability values but change which remaining clusters receive the other probabilities. Across three training seeds, the original soft distribution is recovered more accurately from Encoders trained with the original than counterfactual targets. Because this could reflect target matching alone, we also test low-level acoustic and phonetic information. Both are more accessible from Encoders trained with the original targets. This suggests that cluster assignments affect acoustic and phonetic properties of the learned representation, not just recovery of the training target.

cs.LG↗

Information-Computation Inversion in Pseudo-Marginal MCMC

Observation refinement changes posterior uncertainty and the likelihood calculation in pseudo-marginal MCMC. We compare their combined effect through finite-run squared-error risk. A sufficient inversion condition relates information gain to accepted event flow and coarse-kernel contraction. A bootstrap construction realizes inversion at every fixed particle count. We then couple cross-event proposals and refresh same-event proposals independently. A swap identity establishes invariance; continuation identities describe subsequent risk. In the same finite model, a rational certificate proves inversion against optimized constant mixtures and repair by selective allocation over an initialization class at a common action-price budget. Reaction-network experiments measure CPU costs. Under finite-pool initialization, selective allocation reduces event mean-squared error by 40.5% against a tuned mixture at 25 post-initialization CPU seconds. Paired transcription observations show how increased particle effort can raise finite-budget error.

stat.ME↗

Supergroup Gauged Linear Sigma Models and their Physical Mathematics

We construct 2d $\mathcal{N}=(2,2)$ gauged linear sigma models with $\mathrm{U}(1|1)^N$ supergauge group possibly with superpotential. Despite being nonunitary, one can still study their space of supersymmetric states and explore their applications to mathematics. In particular, we find a relation between a nonlinear sigma model on a Calabi-Yau complete intersection of hypersurfaces in a super-Grassmannian and a supergauged Landau-Ginzburg orbifold, which can reduce to a regular Calabi-Yau/Landau-Ginzburg correspondence for complete intersections. This defines a super-Grassmannian/supergroup generalization of the correspondence proved by Clader [1] and Zhao [2]. Similarly, we find a relation between a nonlinear sigma model on a Calabi-Yau hypersurface in a product of super-Grassmannians and a hybrid NLSM/supergauged Landau-Ginzburg orbifold, which can reduce to a regular hybrid Calabi-Yau/Landau-Ginzburg correspondence for hypersurfaces in product space. This defines a super-Grassmannian/supergroup generalization of the correspondence proved by Fan-Jarvis-Ruan [3]. We also find that Calabi-Yau supervector bundles over a super-Grassmannian can undergo a physically related mild topology change which is reducible to a regular Atiyah-type flop transition. This defines a super-Grassmannian generalization of a birational equivalence of Calabi-Yau vector bundles in mathematics. Similarly, we find that a Calabi-Yau complete intersection of quadrics in a super-Grassmannian can also undergo a physically related topology change which is reducible to a regular conifold transition. This defines a super-Grassmannian generalization of a homological projective duality for Calabi-Yau quadrics by Kuznetsov-Perry [4] in mathematics.

hep-th↗

CLaST: Context-aware Contrastive VAE for Probabilistic Time Series Forecasting

Probabilistic forecasting models are widely used for time series forecasting in domains such as energy systems, finance, medicine, and transportation. In recent years, deep generative models have shown strong results on probabilistic forecasting, yet many conventional approaches struggle to capture internal temporal dependencies, leading to latent representations with limited expressive power. To address this limitation, we propose \textit{CLaST}, a VAE framework for probabilistic multivariate time series forecasting. Unlike existing generative models, CLaST learns embeddings that preserve contextual similarity between observations through our contrastive loss function. Experiments across nine widely adopted benchmarks demonstrate that CLaST consistently surpasses strong baseline methods. In short-term forecasting tasks, our approach achieves improvements of up to $16.4\%$ in CRPS and $14.4\%$ in NMAE over the second-best method. Furthermore, in long-term prediction CLaST attains superior overall performance, exceeding the second-best method by up to $48.6\%$ and $25.1\%$ in CRPS and NMAE, respectively.

cs.LG↗

DecoVAE: a Lightweight Interpretable Trend-Seasonal VAE Framework for Efficient Probabilistic Time Series Forecasting

Probabilistic time series forecasting remains challenging, largely because modeling distinct trend and seasonal dynamics requires specialized approaches. Existing methods often fail to capture the unique inner properties of these components, lack interpretability, or suffer from heavy memory and runtime overhead. To address these limitations, we propose DecoVAE, a lightweight interpretable trend-seasonal VAE framework that explicitly decomposes time series into trend and seasonal components by applying domain-specific inductive biases. The trend stream enforces structural smoothness using a differential regularizer on the latent trajectory, analogous to the Hodrick-Prescott filter. Concurrently, the seasonal stream operates in the frequency domain via a complex Gaussian VAE, natively capturing the amplitude and phase of periodic patterns. Extensive evaluations across seven real-world benchmarks show that DecoVAE consistently outperforms strong baselines. It achieves reductions of up to 14.96\% in CRPS and 23.30\% in NMAE for short-term forecasting, and up to 52.68\% and 26.51\% for long-term horizons. Crucially, DecoVAE yields these accuracy gains while remaining highly efficient, reducing model weight by up to 93\% and accelerating speed by up to 74\% compared to the second-best method.

cs.LG↗

Self-Normalizing Denominators in Rational Covariance Estimators

Many estimators are ratios of coprime polynomials in a sample covariance matrix, and their accuracy depends on the relative fluctuation of the sample denominator. Under Gaussian sampling in fixed dimension, we call a nonconstant polynomial denominator self-normalizing if the first-order variance of its relative error does not depend on the population covariance. We prove that these denominators are exactly the flag powers, nonzero constant multiples of products of positive integer powers of nested generalized variances. Equivalently, the denominator's sample-to-population ratio has a covariance-independent finite-sample law, which we determine explicitly. Sufficiency is classical; the new converse shows that a first-order variance condition forces an exact sampling law. We show that relative stability, meaning bounded first-order relative variance, characterizes uniform tightness of scaled relative errors over positive-definite covariances. It permits replacing the sample denominator by its population value in the limit theory of the ratio. Self-normalization is its rigid core. We locate these classes in applications, where regression on predecessors in a fixed order yields only constant or self-normalizing denominators, instrumental-variable formulas yield relatively unstable ones, and nonparametric identifiability does not guarantee relative stability.

math.ST↗

Beyond Attention Masks: Instruction Anchoring for Efficient In-Context Diffusion Generation

In-context diffusion transformers concatenate instruction, target, and reference tokens into a single sequence for joint attention. Reference-side computation must therefore be repeated at every denoising step, with the cost growing rapidly as more references are added. Decoupling reference tokens from the target enables exact key-value reuse across denoising steps, but prevents the references from attending to the instruction, degrading instruction following and reference fidelity. This trade-off cannot be resolved through attention-mask design alone. We introduce AnchorCache, a parameter-free token-layout and attention-mask co-design that inserts static text anchors. These anchors condition the reference representations on the instruction during cache construction, after which the resulting reference keys and values can be reused exactly across denoising steps. To recover the quality initially lost through this structural conversion, we apply teacher-forced velocity distillation followed by a short on-policy stage that queries the teacher at student-visited states. To our knowledge, this is the first use of on-policy distillation for architectural recovery in diffusion models. Across benchmarks spanning image, speech, and video generation, AnchorCache matches full-attention quality. Its efficiency gains increase with the reference-context size, reaching a 6.40x speedup in diffusion transformer inference.

cs.CV↗

Assessing the Impact of High-Resolution Imaging on Statistical Validation of TESS Planet Candidates

High-resolution imaging is widely used to constrain false-positive scenarios in exoplanet validation, but it is a finite follow-up resource that reaches only a subset of candidates, and its population-level impact on validation outcomes has not been quantified through controlled removal experiments. Using an automated pipeline built on TRICERATOPS, we compute the false-positive probability (FPP) of 443 TESS planet candidates. For the 264 planet candidates with high-resolution imaging observations, we compute FPP with and without the corresponding contrast curves, allowing us to quantify the impact of the additional data. We find that 72% of 68 contrast-curve bearing validated planets would fail validation without their adopted contrast curves. The fraction requiring imaging decreases with increasing planet size, from 100% below $1.7~R_\oplus$ to $33\%$ above $4~R_\oplus$: within our sample and TRICERATOPS-based analysis, the availability of high-resolution imaging directly limits the yield of small-planet validation and the supply of validated targets for atmospheric characterization. Our analysis statistically validates 64 new TESS planets with sizes spanning 0.94 to 7.83 $R_\oplus$ across hosts of spectral type M through F. Four of these are highly amenable to JWST observations based on the transmission and emission spectroscopy metrics, and each achieves validation only with its imaging constraint.

astro-ph.EP↗

Learning Generalizable Behaviors for Terminal Agents

Terminal agents are a compelling application of large language models (LLMs), with the potential to integrate deeply into users' daily workflows. Reinforcement learning (RL) is a key technique for improving their capabilities, making scalable training environments a central challenge. Since public real-user interaction data are scarce, synthetic environments provide a practical alternative, but often suffer from domain gaps and limited fidelity, leading to poor generalization. Existing work mainly scales the quantity and diversity of synthetic environments, while reward-signal quality and the mechanisms governing generalization remain under-explored. We study how RL improves terminal agents and propose the Agentic Compositional Generalization hypothesis: rather than teaching new domain-specific skills from scratch, RL primarily shapes high-level decision-making behaviors that compose and route low-level skills acquired during pre-training and supervised fine-tuning (SFT). This account is consistent with our empirical results and suggests that verifier quality, which determines which behaviors are reinforced, is more important than simply increasing environment quantity or diversity. Motivated by this insight, we propose River, a simple training recipe that improves reward quality by filtering low-quality environments and augmenting outcome rewards with process-level behavior regularization. Using this recipe, our RL-trained agent achieves the best performance among evaluated open-source RL-trained 8B models across four terminal-agent benchmarks. River also generalizes across model families, scales, agent harnesses, and RL objectives. Using fewer than 30% of the TMax training environments, River improves RL gains by 106% and 30% on average for models ranging from 2B to 27B on Terminal-Bench-Lite and Terminal-Bench-v2.1, respectively.

cs.LG↗

$Δ$ resonance contributions to QED radiative corrections in neutron and inverse beta decay

We incorporate the $Δ(1232)$ resonance into pion-induced QED radiative corrections to neutron decay and inverse beta decay (IBD). Within the framework of heavy-baryon chiral perturbation theory with explicit $Δ$ degrees of freedom, we compute additional contributions and study their impact on IBD cross sections and on the renormalization of the nucleon isovector vector and axial-vector charges. $Δ$ resonance does not renormalize the vector charge. For the axial-vector charge, including the $Δ$ resonance improves convergence and reduces the QED radiative correction to the experiment-over-lattice-QCD ratio $g_A/\left(g^\mathrm{QCD}_A g_V\right)$. $Δ$ resonance increases the pion-induced QED radiative corrections to IBD by a factor $1.2$-$1.3$.

hep-ph↗

Extreme-ultraviolet spectroscopy using quantum logic: a feasibility study for the 1S-2S transition in singly-ionized helium

Extreme-ultraviolet (XUV) spectroscopy represents an important new direction in precision physics, with potential applications ranging from the metrology of fundamental constants to tests of physics beyond the Standard Model. However, the application of quantum control methods for precision spectroscopy remains an open challenge in the XUV range. Here we present a novel quantum logic (QL) spectroscopy method for precision spectroscopy of weak XUV transitions, and numerically validate its feasibility for the $1S-2S$ transition at 40.81\,eV in singly-ionized helium (He$^{+}$). We propose a scheme based on a single He$^{+}$ ion co-trapped with a Be$^{+}$ ion in a Paul trap, and He$^{+}$ excitation with pairs of frequency-comb (FC) laser pulses upconverted to the XUV via High-Harmonic Generation (HHG). We investigate a nondestructive QL scheme to detect $1S-2S$ excitation, and compare its performance with a destructive readout based on state-selective ionization. Phase coherence of the XUV light is modelled and an optical cavity is used to filter the FC pulses prior to HHG. We model the motional excitation dynamics of trapped ions outside the Lamb-Dicke regime, and numerically validate a scheme we proposed in \cite{Grundeman} to cancel the first-order Doppler broadening and the recoil shift by synchronizing the ion's secular period with the time delay between the two excitation pulses. We show that precision spectroscopy of the $1S-2S$ transition in He$^{+}$ at the 10 kHz level is feasible, for improved tests of quantum electrodynamics (QED), a measurement of the Rydberg constant $R_{\infty}$ independent of hydrogen measurements, or an improved determination of the alpha particle and helion charge radii. The proposed method may also be applied to XUV spectroscopy of other ions outside the Lamb-Dicke regime.

physics.atom-ph↗

The Sharp Tail of Uniform Stability

Uniform stability controls how much one training example can change the loss at any test point. A new logarithmic-free upper bound shows that a $γ$-uniformly stable algorithm with loss in $[0,L]$ has generalization gap at most $O \left(γ\log(1/δ) +L\sqrt{\frac{\log(1/δ)}{n}}\right)$ with probability $1-δ$. Whether an actual bounded-loss learning algorithm can realize the linear dependence on $\log(1/δ)$ has remained open. The known construction realizes it only for auxiliary weakly dependent random variables whose pointwise range grows with $n$. The known learning lower bound holds only at constant probability. We close this gap. For every $n$, stability level $γ$, and loss bound $L$, we construct one deterministic $γ$-uniformly stable learning problem whose tail satisfies, simultaneously for $1\le p\le c n$, $\mathbb P \left( R(A_S)-R_S(A_S) \ge c'\min \left\{L,γp+L\sqrt{p/n}\right\} \right)\ge e^{-p}.$ The construction is ordinary bounded absolute-loss regression with constant labels. Its key is a multiscale collection of rare Rademacher features. A coordinatewise ramp is stable in sup norm, while an odd symmetrized maximum converts a unique extreme feature into a gap of order $γp$ without violating the loss bound. Geometrically spaced ramps put all confidence levels into the same problem. Together with the logarithmic-free upper bound, this determines the optimal high-probability and moment dependence of uniform stability up to universal constants.

cs.LG↗

Two Dimensions Govern Agnostic Multiclass Transductive Learning

In transductive classification, an adversary fixes a labeled population, one label is hidden uniformly, and the learner sees all remaining labels. For binary classes, agnostic transductive and PAC learning have the same minimax rate. Whether this extends to multiclass learning was open, especially for unbounded label spaces where uniform convergence can fail. We resolve the question up to logarithmic factors. For every multiclass class $\mathcal H$ with DS dimension $d_{DS}$ and Natarajan dimension $d_{\mathrm N}$, the optimal agnostic transductive excess error satisfies $\widetildeΘ\left(\frac{d_{DS}}{n}+\sqrt{\frac{d_{\mathrm N}}{n}}\right).$ The result holds for arbitrary label spaces. The two terms are both necessary. A DS pseudo-cube gives the realizable $d_{DS}/n$ obstruction, while a Natarajan cube with repeated points and fair labels gives the agnostic $\sqrt{d_{\mathrm N}/n}$ obstruction. The upper bound uses a random-reservation principle. The learner deliberately ignores a constant fraction of the visible labels, which makes the true test point uniform in a large unseen block. We combine realizable compression, a label-space reduction, and inside-menu agnostic compression across this finite-population split. A new without-replacement multiplicative-weights lemma preserves the fast $d_{DS}/n$ term. Consequently, agnostic multiclass PAC and transductive learning obey the same two-dimension law up to logarithmic factors.

cs.LG↗

J-Zero: Unified Challenger--Solver--Judge Self-Evolution from Zero Data

Self-evolving language models have recently emerged as a promising path toward superintelligence, with the advantage of reducing the cost of human supervision. While considerable progress has been made in verifiable domains, self-evolution in unverifiable domains remains less explored. We propose Judge co-adaptation from Zero data (J-Zero), a unified Challenger--Solver--Judge self-evolution framework that supports self-improvement across both domains. The Challenger and Solver co-evolve through an adversarial interaction: the Challenger generates increasingly difficult tasks, while the Solver learns to produce higher-quality responses to them. In parallel, the Judge co-adapts using preference pairs whose ordering is known in advance from how each response was produced, i.e., the Solver's answer over the Challenger's, and the Solver's decomposed-and-recombined answer over its one-shot answer, rather than from the Judge's own scores. J-Zero outperforms the baselines by an average of 4.2 points on verifiable and 8.0 points on unverifiable domains, and continues to improve through at least ten iterations, whereas the baselines degrade after two. Further analysis identifies Judge co-adaptation as the key driver of this sustained improvement.

cs.LG↗

On LCK Geometry of Gauduchon Connections

It is not a priori clear which of the Gauduchon connections is better suited to LCK and Vaisman manifolds. We thus investigate the geometry of these connections through Einstein problems and more general analytic and cohomological conditions on their Ricci tensors. Our results single out the Bismut connection as the privileged one for the non-Kähler lcK and Vaisman geometry. We therefore study the {second-Bismut--Einstein} equation and provide a characterization of these structures on Hopf manifolds.

math.DG↗