Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Convolutional non-parametric Gamma-Ray Signal and Background Separation

In very-high-energy gamma-ray astronomy, signals must be separated from the residual background which arises from misidentified cosmic-ray protons. Building on previous work, we introduce an ensemble of variational autoencoders that aims to perform automatic signal-background separation with minimal assumptions that include separability of the spatial and energy distributions in both the signal and background, with no prior specification of the number of components or the point-like/diffuse nature of the signal itself. In addition, we do not assume knowledge of which region of the coordinate space is background-dominated or where the signal is supposed to be located in the field of view. We test the model on an analytic point-source mixture scenario, a realistic simulation of dark matter annihilation in the Galactic centre, and real observations of the Crab nebula and MSH 15-52 from the public H.E.S.S data release. The model proves capable of completely reconstructing the signal and background at a pixel by pixel level in all scenarios, whilst also denoising the inputs. Stable performance and reasonable error estimates are obtained even for low signal-to-background ratios.

astro-ph.HE↗

Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations

Robust counterfactual explanations promise recourse that still works after the model behind it changes. Whether they keep that promise depends on what the change is. A small perturbation of the parameters, retraining on new data, and a new architecture are different events, and each existing method is evaluated against the one it was built for. Reported robustness scores, therefore, answer different questions and cannot be compared. We propose a unified cross-family evaluation protocol that holds factual instances and generated counterfactuals fixed while testing every method against the same eight types of model change. The benchmark compares six robust methods and two standard baselines on four tabular datasets. It characterizes every changed classifier through its outputs and reports empirical robustness together with coverage, base validity, and proximity. We find that relative performance and failure modes vary across change families. Bounded parameter perturbations change 0.95\% of test predictions on average, compared with 4.9\% for bootstrap retraining. Methods with guarantees for these perturbations do not necessarily transfer to other changes. RobX transfers most consistently in our experiments, although greater stability can require larger interventions. We argue that robust CFE methods should be evaluated through a common protocol that specifies the model changes, measures their realized behavioral magnitude, and keeps generation performance separate from robustness.

cs.LG↗

Upper and lower fluctuations in Shallit's law of leap years in Pierce expansions

Shallit's law of leap years in Pierce expansions gives, for Lebesgue-almost every $x\in[0,1]$, the upper and lower limits of a normalized difference between $Nx$ and the number of leap years up to the year $N$ determined by the Pierce expansion digits of $x$. In this paper, we show that, for every $x\in[0,1]$, both limits are determined by the lower limit of a normalized logarithm of the product of the first $n$ digits of $x$. In particular, the upper and lower limits always have the same absolute value. As applications, we characterize the sets of points with prescribed upper and lower limits, and show that each of these sets, if non-empty, is dense in $[0,1]$ and its intersection with any non-empty open subset of $[0,1]$ has full Hausdorff dimension. We also compare these sets with the sets defined by the growth rate of the digits, and show that, for each finite parameter, their difference has full Hausdorff dimension.

math.NT↗

Large-Data Global Well-Posedness for the Defocusing Cubic Schrödinger Equation in 3D Convex Domains

We establish the global well-posedness for the energy-subcritical defocusing cubic nonlinear Schrödinger equation (NLS) on the three-dimensional Friedlander model domain \( Ω=\{x\in\mathbb{R}:x\geq0\}\times\mathbb{R}^2_{y,z}, \) subject to homogeneous Dirichlet boundary conditions, for arbitrarily large initial data in the coercive energy space $H_0^1(Ω)$. The phase-space geometry of this configuration features a strictly convex boundary that admits a non-empty glancing set, trapping high-frequency wave packets within a boundary layer via generalized whispering-gallery caustics. These concentration phenomena induce an intrinsic derivative loss in the sharp linear Strichartz estimates, establishing a major microlocal obstruction to closing the nonlinear Duhamel iteration directly at the energy level via classical perturbative frameworks. To bridge the regularity deficit between the conservation laws and the linear theory, we construct a boundary-adapted family of continuous--discrete Bourgain--Strichartz restriction spaces $X^{s,b}$ built from the spectral decomposition of the Dirichlet realization of the model Friedlander operator \( Δ_g=\partial_x^2+(1+x)\partial_y^2+\partial_z^2. \) The normal variable is resolved through discrete Airy spectral modes, while the tangential variables are continuously Fourier analyzed. Within this functional framework, we implement a refined high--low frequency decomposition to formulate a perturbed nonlinear equation for the high-frequency remainder in the sub-energy space $H_0^s(Ω)$ for \( \frac{1}{2} λ}u_0\|_{H_0^s(Ω)} \label{eq:abs_decay} \lesssim λ^{s-1}\|u_0\|_{H_0^1(Ω)}, \) providing a small parameter to counteract the derivative loss.

math.AP↗

Counterexamples to the inhomogeneous Duffin-Schaeffer Conjecture

The Duffin-Schaeffer Conjecture, proven in 2020, gives a zero-full dichotomy for the approximation of irrational numbers with reduced fractions for arbitrary approximation functions. We construct approximation functions such that the inhomogeneous analogue of the Duffin-Schaeffer Conjecture fails. The counterexamples are obtained for all non-zero rationals and certain Liouville numbers.

math.NT↗

JevSoup: System-One Routing for Training-Free LoRA Composition

Building adaptable AI systems requires effective coordination of specialized capabilities across diverse tasks. Low-rank adaptation (LoRA) enables modular expertise, but existing routing approaches may require auxiliary data, additional training, or autoregressive decoding. We propose JevSoup, a training-free framework separating System One expert routing from System Two execution. Using only the input and expert descriptions, Jev selects two experts through structured probabilities. JevSoup retains the leading expert's update, projects the second onto the orthogonal complement of the first update's row space, and combines them with equal weights. Across 14 PorTAL tasks and three Qwen3 scales, JepSoup achieves absolute gains of up to 1.19\% in task-macro and 1.21\% in sample-micro accuracy over the strongest evaluated external baselines. Our code is available at https://github.com/Leowang980/JevSoup.

cs.AI↗

Improving Multi-Delay-ASL through specialized reconstruction

Purpose: Although image reconstruction has received relatively little attention in ASL research to date, it has the potential to address several challenges in ASL. As well as speeding up measurements by increasing undersampling and reducing measurement artifacts, it can improve the signal-to-noise ratio (SNR) of a given data set and the reproducibility of examinations. Methods: This work focuses on extending a dedicated ASL reconstruction approach (ASL-TGV) (Spann et al. (2020)) to multi-delay data and applying it to a high-resolution pCASL test-retest dataset. To show not only improvement in the Perfusion Weighted Images (PWIs), Cerebral Blood Flow and Arterial Transit Time was estimated and the test-retest reliability was estimated using the within subject Coefficient of Variance (wsCV), the Intraclass Correlation Coefficient (ICC) and Root-Mean-Squared-Error (RMSE). Results: The Perfusion-Weighted-Images reconstructed from highly undersampled single-shot data using ASL-TGV are clearly improved compared to a fully-sampled reference from the same data even after denoising, especially for long PLDs and the outermost slices. SNR calculated in grey and white matter ROIs shows an improvement of 62\% and 35\% respectively. The CBF maps produced from the ASL-TGV images have an improved test-retest reliability. Conclusion: ASL-TGV, which is now implemented in the Berkeley Advanced Reconstruction Toolbox (BART), can be used with any type of ASL labeling or data acquisition and any existing image post-processing pipeline and can improve SNR for the PWIs and reproducibility of the CBF maps.

physics.med-ph↗

Training-Free Pronunciation Transcription via Text-Constrained Acoustic Rescoring

Accurate and efficient pronunciation transcription is essential for preparing text-to-speech training data at scale. Existing approaches have different limitations: grapheme-to-pronunciation (G2P) and speech-to-pronunciation (S2P) methods each capture only partial information, using only text or only speech, while speech-and-text-to-pronunciation (ST2P) methods use both but require costly pronunciation-annotated data. To address this problem, we propose a training-free ST2P pipeline that integrates both lexical and acoustic information at inference time. Lexical resources and G2P tools generate text-constrained candidates, and a left-to-right greedy search selects the best one using whole-sequence negative log-likelihoods from frozen pretrained S2P models. On three Japanese corpora, our method reduces Character Error Rate (CER) from 0.60--1.40\% (text-only baseline) to 0.04--0.17\% with reference transcripts, and 0.64--1.58\% with ASR transcripts. It outperforms all baselines, including a trained ST2P model and commercial multimodal LLMs. Our greedy search method is 3--3.5$\times$ faster than beam search at similar CER, and the cascade is 2$\times$ faster than direct decoding ensuring the efficiency and accuracy. In Spanish, French, and preliminary English, it also surpasses four open multimodal LLMs and the best traditional methods.

cs.CL↗

How to break the Miranda signature scheme over matrix Gabidulin codes

The Miranda signature scheme relies on masking a matrix code which disposes of a masked underlying structure, and the knowledge of which allows for efficient error decoding. We consider a Gabidulin code which is expanded into a matrix code, that is only \Fq-linear. An additional masking is then applied to it. The attack proposed here shares similarities with that of [Le26, https://arxiv.org/abs/2608.03328] on the EGMC encryption scheme, which also follows the paradigm described above. It consists in recovering the Fqm-linear structure of a masked matrix Gabidulin code by reducing to a MinRank instance to be solved over the extension field Fqm, but where the matrices have coefficients in Fq. Such an instance can be efficiently solved. However, unlike the previous attack, it is possible to reduce in polynomial time to such a MinRank instance in the case of the Miranda signature scheme. This results in a particularly efficient key recovery attack against the parameters proposed for Miranda. For example, for the proposed parameter set with m=79, the complexity drops from 146 bits to 46 bits in this attack.

cs.CR↗

Minimal m-subharmonic functions with nonmaximal Hessian measures

For every $2\le m\le n$, we construct Hölder continuous functions on the closed unit ball that are minimal in the Cegrell class $\F_m$, although their Hessian measures are not maximal in the $m$-subharmonic ordering. This answers a question posed by the second author. We also obtain a minimality criterion based on partial pluricomplex energy and an explicit family of Monge--Ampère examples on a product domain. For $m=1$, minimality and maximality are equivalent on a ball.

math.CV↗

Bounds for Unions of Several Parts in Balanced Graph Partitions

Let $k\ge3$ and $1\le \ell\le k-1$. We study balanced $k$-partitions of a graph for which the union of any $\ell$ parts induces few edges. We show that every graph $G$ with $n$ vertices and $m$ edges admits a balanced partition $V_1,\ldots,V_k$ such that \begin{equation*} \max_{\substack{A\in\binom{[k]}{\ell}}}e_G\left(\bigcup_{i\in A}V_i\right)\le\frac{\ell^2}{k^2}m+\frac{\ell^2(k-\ell)}{k^2}(n-1)+\frac{\ell(k-\ell)}{k(k-1)}\sqrt{\left(\binom{k}{\ell}-1\right)m}. \end{equation*} In the case $\ell=2$, our result confirms a conjecture of Bollobás and Scott in a stronger form.

math.CO↗

UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound

Ultrasound is one of the most widely used medical imaging modalities, and recent large vision-language models(VLMs) have shown increasing capabilities in ultrasound image understanding. However, these models fail to provide pixel-level visual evidence aligned with their semantic predictions, and their fine-grained grounding capability in ultrasound remains largely unclear. We introduce UltraG-Bench, a large-scale multi-task benchmark for evaluating pixel-level evidence grounding in ultrasound. UltraG-Bench is built by annotating 40 public ultrasound segmentation datasets spanning 13 anatomical categories, and comprises three progressive tasks: instruction-guided segmentation, evidence-grounded VQA, and evidence-grounded report generation, with 331125, 666779, and 138832 annotations, respectively. Comprehensive evaluation of 14 state-of-the-art models reveals a substantial gap between semantic understanding and fine-grained pixel-level localization. We further propose UltraG-Agent, which combines the semantic reasoning capabilities of a VLM with the ultrasound-specific segmentation capability of UltraSAM3. Experiments show that UltraG-Agent substantially improves both semantic prediction and pixel-level visual grounding. Our dataset and code are available at https://github.com/zhuqh19/UltraG-Bench.

cs.CV↗

EPOC: Endpoint-Preserving Online Correction With Compressed Residual State for Multi-Horizon Time Series Forecasting

Completed multi-horizon forecasts provide residual feedback for a fixed forecaster, but retaining full residual blocks increases auxiliary state. We propose Endpoint-Preserving Online Correction (EPOC) with a compressed residual state. It stores low-order discrete cosine transform (DCT) coefficients and the final value of the preceding residual block. Within each channel, the endpoint is shared across component-wise online ridge regressions that also use current-forecast coefficients. The fitted DCT correction is blended with the base forecast. We evaluate eight multivariate series with DLinear and PatchTST, three seeds, and two training variants, yielding 96 matched fixed-base conditions at a 24-step horizon. EPOC achieves mean condition-wise reductions in mean squared error (MSE) and mean absolute error (MAE) of 15.40% and 9.35% from the uncorrected base, respectively, with a median of 6,352 B in retained auxiliary arrays. It has lower paired MSE than the $δ$-Adapter, COSA, FAC, and OMPB in a majority of conditions and uses less state than each. Full ELF achieves the largest mean MSE reduction, 19.29%, but its median retained state is 474,048 B ($\times$75 relative to EPOC). Equal-size summary controls favor the endpoint by 1.65--2.20% in paired MSE; a coefficient-reconstructed endpoint yields similar accuracy to the observed endpoint, highlighting its role as a shared input. Increasing the retained DCT component count from 4 to 8 adds 1.00 percentage point of MSE reduction for 5,728 B. On jointly trained bases, EPOC lowers MSE by 16.69--20.15% relative to globally blended TEFL-style adapters applied to the same base. The code and numerical records are available at https://github.com/keiotakmin/endpoint-preserving-residual-correction.

cs.LG↗

Tracking Dynamic Simplicial Complexes via Constrained State-Space Estimation

Simplicial complexes (SCs) extend graphs to represent higher-order interactions, but couple their simplex levels through the inclusion property. While static SC inference is an emerging research direction, tracking time-varying SCs remains largely unexplored. A central challenge in tracking SCs is to account for the inclusion property. To this end, we build a nonlinear state-space model tailored to SCs. For prediction, we introduce a closure-aware Markov generation model whose conditional mean preserves simplicial inclusion and whose covariance captures edge-triangle dependencies. For correction, we encode inclusion constraints as nonlinear pseudo-measurements, allowing the constraints to inform both the state and covariance updates. We investigate progressively richer treatments of these measurements through a standard extended Kalman filter, an iterated extended Kalman filter, and a Laplace approximation. Experiments then demonstrate the benefits of exploiting the proposed dynamics and constraint information.

eess.SP↗

Interpolating walk between discrete-time quantum walk and its intrinsic random walk on graph

We consider an interpolation with the parameter $p\in [0,1]$ between the discrete-time quantum ($p=0$) and its intrinsic random ($p=1$) walks on a finite connected graph. Its time evolution is defined by the convex combination of the Kraus (CPTP) maps of the quantum and random walks. The Kraus map of the random walk is represented by taking a kind of projection of that of the quantum walk. The time evolution of the interpolation is interpreted as the combination of the following two dynamics of $2$ walkers on the same graph correlated with each other: $2$ walkers move independently of each other (no-correlation) with probability $1-p$, while $2$ walkers always move into the same position (the maximal correlation) with probability $p$ at each time step. In this paper, we generalize the random walk to a new walk, namely, the correlated walk, which preserves the intrinsic property and relaxes the maximal correlation of the random walk. This correlation is determined by a partition of the arc sets in the underlying graph. We show that the eigenvalues of the interpolating walk between the correlated and quantum walks live in $\{ z\in \mathbb{C}\;|\;1-p\leq |z|\leq 1 \}$ and that the absorption state coincides with that of the correlated walk, which is characterized by the underlying graph structures.

math-ph↗

A Uniform Bound on Optimal Strategy Length in Water Transport Problem

We prove that every water transport problem on an $n$-vertex graph has an optimal strategy of length at most $n^{(2+o(1))n}$. More strongly, the convex hull of all strategy operators stabilizes within the same bound. We also give a five-vertex instance in which every optimal strategy repeats a nontrivial connected averaging set.

math.CO↗

Discrete orthogonal polynomials related to Hahn difference operator

This work studies general sequences of discrete orthogonal polynomials on linear lattices, including the so-called discrete Laguerre--Hahn orthogonal polynomials. We deduce the structure relations involving the orthogonal polynomials and their associated, as well as the explicit form of the fourth--order linear difference equation satisfied by these Laguerre--Hahn orthogonal polynomials. Furthermore, the particular cases corresponding to semiclassical and classical families of orthogonal polynomials, which lead to second--order difference equations, are also analyzed. Several examples illustrating the computation of the explicit coefficients of these difference equations for various families of orthogonal polynomials are also presented.

math.CA↗

ManiVid: Unified and Explainable Forensic Analysis of Manipulated Videos

Rapid advances in AI-generated video (AIGV) have increased the risks posed by deceptive video manipulation. Unlike fully synthetic videos, manipulated videos retain most source content and alter only localized regions, making forensic analysis particularly challenging. Existing video forgery research faces two limitations in both data and methodology: (1) High-quality datasets and benchmarks tailored for manipulated videos remain scarce. (2) Multimodal large language models (MLLMs) extend forgery analysis beyond binary classification but struggle to use low-level forensic cues and provide precise pixel-level grounding. Specifically, we introduce ManiVid, a unified forensic analysis task covering forgery detection, artifact grounding, and anomaly explanation for manipulated videos. We construct ManiVid-38K, the first dataset to combine paired, open-vocabulary localized manipulations of general videos with authenticity labels, forgery masks, and anomaly explanations. It comprises about 19K manually verified real-fake video pairs, mostly at 1080P resolution, generated under 2 paradigms with 15 powerful generation models. We sample 1K pairs for ManiVidBench, balanced across six manipulation types and generation models for fair evaluation. We further propose ManiVidLens, a unified framework for explainable video forgery analysis. Its Forensic Evidence Router supplies shared low-level forensic evidence for multimodal reasoning and video segmentation. Its Prompt Distill Module converts grounding states into semantic and geometric prompts and distills spatial priors for mask decoding and full-video propagation. ManiVidLens achieves relative gains over the strongest comparison methods in artifact grounding (+21.1% mIoU; +21.3% J&F) and anomaly explanation (+131.3% ROUGE-L; +9.9% CSS). Its forgery detection remains comparable to dedicated classifiers (0.914 Acc; 0.913 F1).

cs.CV↗