Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

MissClick: Execution-Aware Adversarial Attacks on Coordinate Generation in GUI Grounding Models

Recent GUI visual grounding models generate screen coordinates as digit-token sequences that are parsed into numerical values and mapped to executable clicks. This generation-to-execution interface creates an attack surface that existing objectives over visual representations or coordinate-token sequences do not explicitly model. Although each coordinate digit is predicted as a token, its spatial effect after parsing depends on decimal position: changing a hundreds-place digit by one shifts the coordinate by 100 units, whereas the same change at the ones place shifts it by one. This mismatch motivates attack objectives that account for both numerical coordinate structure and click execution. Moreover, untargeted and targeted attacks require different objectives because they aim to move the click outside the correct region and into an attacker-specified region, respectively. We propose MissClick, an execution-aware white-box attack that aligns optimization with click-level success conditions. MissClick-U maximizes soft-coordinate displacement for untargeted disruption, while MissClick-T minimizes a place-weighted target-digit loss for targeted redirection. On OS-Atlas and UGround across desktop, web, and mobile platforms, MissClick-U achieves untargeted success rates of 75.07% and 72.93% (+16.62 and +30.72 pp), while MissClick-T achieves targeted success rates of 44.86% and 62.67% (+31.73 and +47.06 pp). Among the evaluated objectives, soft-coordinate displacement performs best for untargeted attacks, whereas place-weighted target-digit optimization performs best for targeted attacks, supporting goal-specific execution-aware objective design.

cs.AI↗

OmniVR: Audio-Video Conditional Generation for Archival Footage Restoration

Archival footage often suffers from coupled visual and acoustic degradations, yet most restoration systems process the two modalities separately. To address this problem, we present OmniVR, the first systematic framework for joint audio-video restoration, covering data construction, model adaptation, efficient inference, and evaluation. We construct a high-quality audio-video corpus with detailed captions and use a joint degradation pipeline to produce aligned clean and degraded pairs. Using these pairs, we adapt a pretrained text-to-audio-video model (T2AV) by introducing degraded audio-video conditions (TAV2AV), then progressively replace sample captions with a fixed restoration prompt while retaining caption/null rehearsal. The resulting AV2AV model requires no user-provided text. Under a compatible residual-learning model, we prove that this condition-annealing schedule reduces gradient variance and expected restoration risk relative to direct fixed-prompt adaptation at the same training budget. For efficient deployment, OmniVR-Flash combines reduced-resolution video conditioning, MeanFlow-based one-step distillation, and Turbo VAE, achieving approximately 38 fps at 1K and 18 fps at 2K on a single B200 GPU. We further introduce OmniVRBench to evaluate four complementary dimensions: visual quality, audio quality, temporal consistency, and audio-visual synchrony. OmniVR achieves state-of-the-art results on public benchmarks and OmniVRBench. Data, code, and model weights will be released. Project Page: https://xin1u.github.io/OminiVR_PAGE/

cs.CV↗

Not Every Divergence Should Be Suppressed: Counterfactual Recoverability in On-Policy Distillation

On-policy distillation (OPD) supervises student-visited trajectories, yet divergence-based rules cannot determine whether an erroneous prefix remains correctable. We formulate this decision as counterfactual recoverability and replay each error state through budget-matched teacher-continuation and rollback branches. Based on their relative success, states are categorized as recoverable, irreversible-but-avoidable, or ambiguous, and these labels guide whether training retains, rolls back, or conventionally supervises the corresponding trajectory. On AIME branch diagnostics, the mean continuation-minus-rollback effect is 0.185 for recoverable states and -1.000 for irreversible-but-avoidable states, demonstrating opposite intervention preferences. A branch-derived recoverability proxy achieves an AUC of 1.000, substantially outperforming divergence alone at 0.392. Across frozen evaluations, recoverability-aware control achieves the strongest recorded performance, reaching 0.578 success on held-out AIME2025 compared with 0.517 for the best baseline. It also improves AIME2024-2025 average@32 from 0.2656 to 0.3125 and GPQA-Diamond average@32 from 0.2702 to 0.3070. Component ablations further show that retaining teacher-correctable prefixes provides the largest individual contribution. These findings establish recoverability as an outcome-grounded decision variable for selective supervision in OPD.

cs.LG↗

SSTQ:Privacy-Preserving Vector Quantization via Subsampled Stochastic TurboQuant

Achieving local differential privacy in distributed optimization while maintaining low communication cost remains challenging. Existing vector quantization methods, such as vqSGD, rely on high-dimensional geometric constructions but incur unfavorable dimension-dependent variance. In this work, we propose Subsampled Stochastic TurboQuant (SSTQ), a framework that combines a bounded Kashin representation, data-independent coordinate subsampling, and privacy-aware one-dimensional quantization. SSTQ includes two variants: (1) a Flat Randomized Response variant that is unbiased and, for a fixed codebook bit-width, frame redundancy, and dimension-independent Kashin level, achieves reconstruction MSE that scales linearly with the ambient dimension $d$, while using only $\lceil \log_2 N \rceil + b$ bits per message. Here, $N = Θ(d)$ denotes the frame size in the Kashin transform and $b$ is the codebook bit-width; and (2) a metric-aware truncated-Laplace variant that removes the exponential dependence on bit-width at the cost of a non-vanishing bias. We also derive a convex uniform-surrogate codebook objective whose worst-case codebook-dependent upper bound improves from $O(4^b)$ to $O(2^b)$. Experiments on synthetic regression, Fashion-MNIST, and CIFAR-10 compare the per-message privacy-utility and uplink-communication trade-offs of SSTQ with those of established baselines, demonstrating favorable utility and communication efficiency.

cs.LG↗

Full Justified Representation under Hare and Droop Quotas in Polynomial Time

I study Full Justified Representation (FJR) in approval-based multiwinner elections under both the Hare and Droop quota conventions. I introduce a descending-budget algorithm in which voters distribute their remaining budgets across their current representation gaps and candidates are purchased whenever the resulting offers cover a common price. With candidate price $λ_H=n/k$, the algorithm returns a Hare-FJR committee; with candidate price $λ_D=n/(k+1)$, it returns a committee satisfying the more demanding Droop-FJR axiom of Casey and Elkind. The two guarantees share a historical-payment invariant and a terminal row--column accounting argument, while the Droop proof requires a new residual-budget argument when all $k$ paid seats are filled. Both variants are deterministic once the voter and candidate orders are fixed and use $O(kmn)$ rational operations.

cs.GT↗

Which Decisions Low-Bit Quantization Breaks, and How to Predict Them

Quantization saves memory by storing model weights with fewer bits. It can also change model decisions, such as whether to call a tool or which option to choose from a finite set. We study these decision changes in 16 language models from 8 families at 4, 3 and 2 bits, across several post-training quantization settings. Our evaluation covers tool use, safety, general knowledge and social bias, using BFCL, XSTest, MMLU, BoolQ, BBQ and synthetic tasks. The decision margin is the score difference between two possible first tokens, measured before and after quantization. Writing the margin before quantization as $m$ and the margin after quantization as $m'$, we find an approximately linear relationship across decisions: $m' \approx c m + b$. The slope $c$ is usually below one and becomes smaller as precision falls, so quantization progressively shrinks decision margins. The offset $b$ is the same for every decision of one kind. Quantization therefore does not simply add random noise, and even a strong preference at full precision can flip. Quantization also affects different kinds of decisions to different degrees. Within tool use, whether to call a tool is often more sensitive than which tool to call: on 400 BFCL tasks, three of five models lose more completed calls than correct tool selections at 3-bit round-to-nearest. Under GPTQ and GGUF far fewer whether-to-call decisions flip than under plain rounding, so there is no single 3-bit failure point. The same relationship predicts how often decisions flip. Across 1,154 combinations of models, quantization settings, bit-widths and decision types drawn from our evaluation, we fit the slope, the offset and the spread around the fitted line on half of the decisions and predict the flip rate on the other half. The predicted flip rate differs from the observed flip rate by a median of 1.0 percentage point.

cs.LG↗

From Behavior to Mechanism: Tracing Divergent Response Modes in Frontier Language Models

Frontier language models are trained with distinct data, objectives, and safety pipelines, but whether those differences produce measurably different behavior under steering pressure has not been tested. We evaluate 6 frontier models from different labs on 300 paired base and steered items across 3 behavioral categories. All models also act as blind peer judges against fixed rubrics, and each response is labeled by consensus over 24,480 judgments, while leaving self-judgment out. Models differ both in how far steering moves them and in the kind of response they give. GPT-5 withholds its reasoning while still providing the answer on 99 of 100 steered items, against 0 in 500 for the others. Claude Opus 4.7 and GPT-5 resist explicit suppression instructions where the other four never do, and they resist differently. In Llama, the open-weight model, a linear probe reads the behavioral split from the residual stream before generation at 0.87 cross-validated accuracy. Injecting that direction drives the behavior from 0% to 86%, and ablating it cuts the natural rate by more than half, where a random direction of equal norm changes nothing. A second ablation on complementary items reproduces the effect more strongly, and its direction has cosine similarity 0.82 with the first.

cs.AI↗

MaskFlow: Precise, Consistent and Seamless Regional Image Editing

Regional image editing has attracted considerable attention for its spatial controllability. Although instruction-based and mask-based editing methods can achieve strong semantic alignment, reliable regional control remains challenging, where an edit must be accurately localized and naturally integrated with the preserved context. We propose MaskFlow, a training framework for precise localization, consistent background preservation, and seamless boundary transitions. MaskFlow incorporates the mask into the probability path and flow-matching objective, coordinating generation within the editable region with source preservation outside it. The proposed Soft-Poisson De-seaming module further refines the predicted vector field during both training and sampling to improve the smooth integration of the edited foreground with the preserved background. We also introduce MaskEdit-Benchmark for general scene and infographics editing, where prompts describe the desired edits without localization cues, leaving masks to specify the target regions. Experiments on natural scenes and infographic images demonstrate consistent improvements over competing methods in both quantitative and qualitative evaluations. Project page: https://reychiaro.github.io/MaskFlow

cs.CV↗

Vision Meets WiFi: Physics-Grounded Estimation of Volumetric Mechanical Properties

Estimating volumetric mechanical properties, including Young's modulus, Poisson's ratio, and density at each voxel, is intrinsically ambiguous from vision alone, as visually similar objects may have substantially different material compositions and physical behavior. Existing approaches predict these properties independently across voxels, overlooking the piecewise-constant material structure of real objects and producing noisy or inconsistent estimates for voxels that share the same material, while lacking an explicit mechanism to resolve visual ambiguity. We introduce ViWi (Vision Meets WiFi), an object-centric framework for volumetric mechanical-property estimation. ViWi represents each object using a compact set of material slots that aggregate evidence from voxels with a shared material identity and produce coherent slot-level property predictions. To complement visual appearance, ViWi incorporates a compact RF descriptor generated through WiFi-band electromagnetic simulation using permittivity and conductivity. The RF descriptor conditions the material slots with global composition cues that may be unavailable from images, while visual features preserve voxel-level spatial localization. On GVM, ViWi improves over the prior state of the art on four of six per-voxel metrics, while its vision-only variant improves all reported mass-estimation metrics on ABO-500. These results demonstrate that combining object-centric material structure with complementary RF evidence enables more accurate and physically coherent volumetric property estimation beyond what is possible from visual appearance alone.

cs.CV↗

Sharp functional quantization and empirical Wasserstein rates for Itô processes

We establish functional quantization and empirical Wasserstein rates for continuous Itô processes under the supremum norm. We assume that the initial condition and the drift and diffusion integrands are controlled by a time-uniform random upper bound with a finite $ρ$-moment for some $ρ>1$. Under this assumption, the $n$-point $L^q$ quantization error is at most $C(\log n)^{-1/2}$ for every $1\leq q<ρ$. The known Brownian lower bound shows that the exponent $1/2$ is optimal over this class. Our proof combines an adaptive dyadic time partition with localization according to the size of the integrands. A general transfer principle yields the sharp mean rate $(\log N)^{-1/2}$ for the $p$-Wasserstein distance between the empirical law of $N$ independent copies and their common path law, together with nonasymptotic deviation bounds, whenever $1\leq p<ρ$. Applications include empirical path-law estimates for path-dependent SDEs and a path-space limit-theory estimate for path-dependent McKean--Vlasov interacting particle systems with common noise.

math.PR↗

Memory--Batch Tradeoffs in Lipschitz Bandits

Lipschitz bandits admit near-optimal regret $\widetilde O_d(T^{(d+1)/(d+2)})$ with little memory under full adaptivity, or with few batches under unrestricted memory. We characterize the minimax expected pseudo-regret over $T$ rounds with $W$ bits of memory and at most $B$ batches in dimension $d$. For every memory budget $W$, we prove a lower bound of order $T^{\frac{d+2}{d+3}} (1+(B-1)W)^{-\frac{1}{d(d+3)}}$. When $W\gtrsim_d\log(eT)$, algorithms with fixed batch boundaries attain, up to logarithmic factors, the larger of this bound and the optimal $B$-batch regret with unrestricted memory. The analysis separates fine-scale comparisons from the regional information needed to allocate their samples at low regret. Attaining $\widetilde O_d(T^{(d+1)/(d+2)})$ regret requires both $B=Ω_d(\log\log T)$ and $(B-1)W=\widetildeΩ_d(T^{d/(d+2)})$. With one bit of memory, minimax regret is $\widetildeΘ_d(TB^{-1/(d+2)})$ under adaptive batch boundaries while $Θ_d(T)$ under fixed boundaries.

cs.LG↗

Evidence-RL: Towards Evidence-intensive Visual Reasoning

Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware post-training methods encourage image use through global perturbations or attention proxies, but they do not test whether a sampled answer causally depends on the local evidence that supports it. We propose Counterfactual Evidence Disentanglement (CED), a training-time evidence audit for VLM grounding. For each response, CED neutralizes an object-centric Evidence Region and compares the resulting support drop against matched non-evidence Regions. We combine this signal with answer correctness inside GRPO, rewarding correct answers that rely on the evidence path rather than shortcut or nuisance paths. CED uses weak object-level proposals, requires no question-specific evidence annotations, and adds no inference-time overhead. Across nine public benchmarks and four backbones, CED outperforms prior RL-based post-training methods, with targeted analyses verifying its object-centric signal.

cs.CV↗

SDDBMs: Soft Denoising Diffusion Bridge Models

Diffusion bridge models leverage Doob's \(h\)-transform to construct stochastic transports between arbitrary endpoint distributions, and have shown strong potential in image-to-image translation and restoration. However, most existing bridge models rely on hard endpoint conditioning, which forces the terminal state to match a prescribed target exactly. This hard constraint induces terminal-boundary singularities: the terminal law collapses to a Dirac measure, and the resulting drift coefficients become ill-conditioned near the endpoint. In this paper, we propose Soft Denoising Diffusion Bridge Models (SDDBMs), a generalized framework that regularizes diffusion bridges directly at the level of their terminal constraints. Instead of imposing an exact endpoint, SDDBMs prescribe a non-degenerate Gaussian terminal marginal under the transformed path measure, with a flexible terminal center and variance. Starting from this prescribed marginal, we develop a complete closed-form construction of the soft bridge, including the Gaussian terminal reweighting and soft \(h\)-function, the induced Gaussian forward marginals and \(\mathbf{x}_0\)-free dynamics. Theoretically, SDDBMs provide a unified probabilistic perspective that encompasses existing diffusion bridge models, including DDBMs, GOUB, and UniDB, as special cases under specific parameter choices. Extensive experiments on image restoration tasks demonstrate that SDDBMs achieve improved numerical stability and superior generation quality over existing bridge-based methods.

cs.AI↗

Preview-Based Relative-Motion Control of an Insertion Tool for Neural-Thread Placement in Pulsating Tissue

Flexible neural electrode threads must be placed at a prescribed depth while the cortical surface moves with cardiac and respiratory pulsation. A controller tracking a fixed point in the laboratory frame cannot distinguish commanded insertion from tissue motion; the error appears as both a depth offset and relative tip--tissue velocity during contact. This paper formulates thread insertion in tissue-relative coordinates: a harmonic observer predicts delayed cortical-surface motion over the control horizon, a constrained MPC regulates the tip relative to that prediction while limiting actuator effort and lateral relative velocity, and an augmented disturbance state removes the steady offset from persistent contact force and model mismatch. In a 1-DOF MuJoCo benchmark, the controller reaches RMS relative-placement errors of 12.0\um\ free-space and 1.9\um\ in contact, versus 18.3/176.8\um\ for delayed-feedback impedance and 286.1/275.5\um\ for laboratory-frame PD -- the lower contact offset costs more peak contact force (3.43 vs.\ 2.00~mN), since it drives to commanded depth rather than yielding to tissue. A 3-DOF extension reduces lateral shear velocity from 1.34 to 0.50~mm/s at 2.1\um\ lateral placement error, and a feasibility-restoring soft-slack formulation keeps the shear constraint solvable under degraded sensing where a matched hard-constraint controller fails. A two-vertex Lyapunov certificate for the finite-horizon gain holds over $-40\%/{+}50\%$ reflected-mass mismatch, and the 1-DOF QP solves in under 0.4~ms at the 95th percentile. These results are a simulation-based control benchmark, not a clinical safety claim: the modeled tip is a rigid contact point, and flexible-thread mechanics, a validated force constraint, biological damage thresholds, and hardware-realistic sensing and timing remain necessary before deployment.

eess.SY↗

HOPPER: Learnable Hop Extraction for Linearized Graph Sequence Models

Graph neural networks typically propagate information through repeated message-passing layers, coupling propagation distance with the number of nonlinear transformations applied. This coupling can make deep architectures difficult to optimize and lead to over-smoothing, over-squashing, and loss of long-range information. Linearized Graph Sequence Models (LGSMs) address this issue by separating propagation depth from processing depth and representing successive propagation states of each node as a sequence. However, existing LGSMs construct these sequences using fixed graph operators, limiting their ability to adapt propagation to the input graph, node features, and downstream task. We introduce HOPPER, an end-to-end learnable extension of LGSM that learns how hop sequences should be extracted before processing by a modern state-space model. HOPPER supports feature-conditioned, structure-aware, graph and hop-adaptive propagation while preserving permutation equivariance, with standard adjacency-based and non-backtracking LGSM sequences arising as special cases of the extractor family. HOPPER is state-of-the-art or competitive across ECHO-Synth and performs strongly on City-Networks. On the LRIM physics-based long-range dependency benchmark, varying the maximum neighborhood size used for message-backtracking cancellation, corresponding to the structural memory window, substantially affects performance. Ablations further isolate the contributions of the learnable extraction mechanism and its structural and feature-adaptive components, showing that adaptive hop-sequence construction provides gains beyond the downstream sequence model alone. Together, these results demonstrate that learnable sequence extraction is a flexible and effective framework for long-range graph representation learning across synthetic, physics-based and real-world graph benchmarks.

cs.LG↗

RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation

Efficient text-to-image generation requires both reinforcement-learning (RL)-based reward alignment and few-step distillation, yet these procedures are typically performed sequentially, increasing training cost and risking the loss of reward gains during compression. We take an RL-native perspective: diffusion RL already generates reward-scored finite-step trajectories, whose intermediate states provide distillation supervision. Based on this insight, we propose REST (Reward-Enhanced Scored-Trajectory Distillation), a single-stage co-training framework in which a decoupled student learns from the evolving RL teacher's trajectories without changing teacher optimization. Advantage-Modulated Distillation (AMD) transforms rollout advantages into signed weights, strengthening imitation of preferred trajectories and aligning distillation priorities with task value. The resulting framework is general and lightweight, requires no extra image rollouts, no separate distillation dataset, and no adversarial training. Experiments on compositional generation, visual text rendering, and human-preference alignment demonstrate competitive few-step, CFG-free generation with RAM or DiffusionNFT teachers. With only four sampling steps, REST-RAM achieves a DrawBench PickScore of 23.97, outperforming both the 40-step RAM teacher (23.95) and RTDMD (23.71).

cs.CV↗

From Objectives to What Models Learn: A Landau Theory of Invariant Learning

Invariant-learning objectives pursue similar goals yet produce qualitatively different regularization paths, leaving unclear when shortcuts can be suppressed without damaging stable modes. Our starting point is simple: in the desired shortcut-suppressed regime, shortcut loading is small, so the objective's low-order expansion governs local stability and residual amplitude. This brings the problem into the domain of Landau phase-transition theory. We establish a mathematical isomorphism between the near-critical normal form of predictive-mode learning and Landau theory, identifying the learning objective as an effective free energy and critical-mode amplitude as an order parameter. The resulting low-order objective signatures predict distinct regularization phenotypes: quadratic $R_2$ terms shift phase boundaries and enable finite-strength elimination, whereas quartic $R_4$ terms continuously attenuate acquired modes while leaving nonzero residual loading at finite strengths. Higher-order terms may further shape nonlinear tails. In a canonical bilinear model, we derive exact phase boundaries and equilibrium loadings, identifying conditions for a selective-retention window in which shortcut suppression preserves stable structure. Controlled bilinear and ReLU experiments, including an MNIST construction, support the predicted signature-phenotype relation, while coupled-feature experiments validate its extension to collective modes. The framework connects the mathematical structure of invariant-learning objectives to what models learn as regularization varies.

cs.LG↗

SpSYRK: Half the Work in Distributed Sparse Matrix Multiplication

The symmetric rank-$k$ update (SYRK), $C = AA^\top$, computes the dot product of each pair of rows of $A$, producing the Gram matrix $C$. Its sparse variant underpins similarity search in machine learning, graph analytics, and genomics, including Jaccard similarity on datasets too large for a single node. Despite the symmetry in its inputs and outputs, existing distributed sparse matrix multiplication algorithms such as Sparse SUMMA treat sparse SYRK as generic multiplication, computing the full output and materializing the explicit transpose even when the calling application uses only one triangle. Prior distributed $AA^\top$ computations in similarity search and genome assembly inherit this overhead from the underlying SpGEMM. This paper presents SpSYRK and CommSpSYRK, two distributed sparse SYRK algorithms that exploit symmetry. The first, SpSYRK, partitions the off-diagonal blocks of the output between the upper and lower triangular regions of the process grid and computes only the lower-triangular part of each diagonal block, halving per-process computation compared with state-of-the-art distributed SpGEMM. The second, CommSpSYRK, further reorders communication to avoid forming $A^\top$, which reduces per-process communication volume. On 32 nodes of the Perlmutter supercomputer, SpSYRK achieves a 2$\times$ speedup over an optimized Sparse SUMMA on matrices where local multiplication dominates the runtime; the advantage narrows on communication-bound inputs, a dependence that the cost model predicts from the arithmetic intensity. CommSpSYRK fixes this and consistently achieves superior scaling at high process counts. The approach is a drop-in replacement for any application computing $C = AA^\top$ via a distributed SpGEMM routine, and its triangular output can be consumed directly by subsequent operations, reducing both computation and memory footprint.

cs.DC↗