Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,333 records · Page 74Linked to original sources

Modular Functors with Singularities from Vertex Operator Algebras Beyond Rigidity and Finiteness

For a vertex operator algebra $V$ and a certain category of its modules, we propose a construction for spaces of conformal blocks organized into an open-closed modular functor with singularities. This is inspired by the idea of implementing directly from the start the principle of holomorphic factorization. More precisely, using the strategy of modular extension introduced by Costello and developed further in our previous work, we build for each surface $Σ$ with at least one boundary component per path component and specified boundary labels attached to marked intervals or boundary circles a representation $Ω_V(Σ;-)$ of the mapping class group of $Σ$. The construction can be described explicitly on generating Dehn twists. This approach is a priori independent from other constructions based on algebraic geometry or topological techniques involving e.g. surgery, but we include an overview over the available comparisons. In the special case in which the module category of $V$ is a not necessarily semisimple modular category $\mathcal{A}$, the spaces $Ω_V(Σ)$ are equivalent to the string-net spaces for $\mathcal{A}$ and hence to the modular functor for the Drinfeld center $Z(\mathcal{A})\simeq \bar{\mathcal{A}}\boxtimes\mathcal{A}$. However, the construction of $Ω_V$ in this paper has the advantage of being available beyond rationality, rigidity, self-contragredience and finiteness. Moreover, we prove that $Ω_V$ satisfies excision, is finite-dimensional in the $C_2$-cofinite case and produces in genus one a generalization of the elliptic double of Brochier-Jordan. We prove for the triplet $\mathcal{W}_{2,3}$ with non-exact fusion product that the boundary conditions introduced by Gaberdiel-Runkel-Wood produce, as expected by these authors, correlation functions, provided that one uses the notion of a modular functor with singularities that we develop.

math.QA↗

The reach of a verification tool decides its value: A controlled study of verification surface, artifact quality, and cost in AI coding agents

Modern artificial-intelligence coding agents can be equipped with tools for checking their own work e.g. a linter, a boot probe, a shell, a screenshot tool. We call this set the agent's verification surface. This study asks whether increasing only that surface, with everything else held fixed, produces a matching growth in the quality of the software the agent ships. We built a minimal coding agent whose tool list is the single controlled variable and used it to implement 1,116 web applications across six models and eight tool configurations. A condition-blind human graded every application against a frozen rubric, and automatic probes stress-tested the API-observable behaviors. Verification's cheapest benefit arrives first, which is to make sure that the application comes up. Without any tools, about one build in seven fails to launch at all and a single boot probe removes nearly all of these failures at roughly 35 percent of a full shell's token cost, while the full shell multiplies the no-tools cost by 2.35. Screenshots help most where mistakes are visible (e.g. element placement, interaction), though even there the gain over a shell is modest and does not survive correction for multiple statistical comparisons. In cases where failures can only be measured rather than seen, such as keeping scrolling smooth over a 100,000-row list, screenshots add nothing. A verification tool improves the output artifact only where its reach covers the way the application actually fails.

cs.SE↗

Flow-JEPA: Robust Latent Dynamics for JEPA World Models via Flow Matching

Joint-Embedding Predictive Architectures (JEPAs) provide a powerful framework for latent world modeling and planning in a reconstruction-free manner. Although numerous JEPA-based approaches have been proposed to mitigate representation collapse, our experiments on localized, out-of-distribution visual noise reveal that performance degradation remains pronounced and unresolved. We propose Flow-JEPA (F-JEPA), a flow-based latent dynamics model that jointly generates a sequence of future latent states conditioned on the current observation and actions. A Gaussian distribution serves as the flow source, exposing the vector field to perturbed latent trajectories as it learns to transport them toward clean future representations. This formulation retains the reconstruction-free JEPA framework while switching from pointwise transition regression to stochastic trajectory-level prediction. F-JEPA raises mean success from $86\%$ to $92\%$ under clean observations and from $67\%$ to $86\%$ under noisy conditions. Further evaluations over varying perturbation severity and inference settings show that the performance advantage persists across a broad range of conditions. These results suggest that conditional flow matching provides a promising alternative to deterministic autoregressive prediction as a dynamics formulation in JEPA world models.

cs.LG↗

Benevolent Bias in Multi-Turn Human-Agent Dialogue

Bias in human-agent interaction can manifest not only through hostile language but also as benevolent bias, whereby unequal treatment hides behind a warm, positive tone. To make it detectable, we operationalise benevolent bias along two dimensions, tone and treatment, yielding three classes: neutral support, overt bias, and benevolent bias. Building on these definitions, we construct BENEVDIAL, a class-balanced corpus of 362,880 multi-turn support dialogues spanning user and agent demographics, roles, and generators, to support controlled evaluation. We then test two detector families on it: off-the-shelf safety detectors and prompted large language model (LLM) judges. Our findings reveal a notable detection gap: off-the-shelf detectors reliably flag overt bias yet largely fail to identify benevolent bias. LLM judges improve sensitivity when guided by explicit detection criteria, but this comes at the cost of increased misclassification of neutral supportive statements as benevolent bias, a tendency that is further exacerbated by the presence of demographic context. These findings suggest that fair monitoring of human-agent dialogue must look beyond surface cues to whether the agent's treatment is disparate.

cs.AI↗

Review-Period Sensitivity in Multiclass Queue Scheduling

Optimal control of queueing networks is typically studied under continuous-time control. In many service settings, however, managers can adjust decisions only at discrete and potentially infrequent review epochs. We study a multiclass queue scheduling problem in which server assignments can be changed only at the beginning of discrete review periods, and examine how performance depends on the review-period length. We analyze a family of associated fluid control problems parameterized by the review-period length and characterize the first- and second-order sensitivity of the value function. For the two-class case, we derive explicit expressions for these derivatives and characterize their signs. We find that for short review periods, the value function need not be monotone. Once the review period is sufficiently large, the value function becomes monotone nondecreasing and may exhibit a convex or linear region before eventually becoming concave. These results provide insights on when the value of more frequent control, or \emph{flexibility} with respect to server assignments, is higher. We further numerically examine the robustness of these observations for the original stochastic scheduling problem and show that stochasticity smooths the nonmonotonicity at smaller scales, while the qualitative sensitivity patterns remain visible and become more pronounced as the scale of the system increases.

eess.SY↗

JPO: Juris Policy Optimization for Structured Legal Reasoning in Criminal Judgment Prediction

Criminal judgment prediction requires models to infer statutory articles, charges, and sentencing outcomes from case facts. Unlike standard classification tasks, it involves a structured reasoning process in which statutes should be matched with facts, charges should be justified by statutes, and sentencing outcomes should remain consistent with charges. Existing approaches optimize final labels, and while some have attempted to evaluate reasoning quality, their evaluations are indirect, often relying on LLM-generated rubrics that reflect model-internal preferences rather than the inherent logical structure of legal adjudication. We propose Juris Policy Optimization (JPO), a post-training framework for structured legal reasoning in Chinese criminal judgment prediction. JPO first uses teacher-generated rationales to supervise a standardized four-step reasoning process, and then applies reinforcement learning with a composite reward over legal prediction quality, reasoning structure completeness, and cross-step consistency. JPO further introduces token-level advantage reweighting and adaptive clipping for legally salient reasoning segments. Experiments on multiple open-source language models and three Chinese legal benchmarks show that JPO consistently improves both judgment prediction and reasoning quality over supervised fine-tuning and reinforcement learning baselines.

cs.CL↗

R-equivalence of quandle colorings and inner automorphisms

R-equivalence is an equivalence relation on the set of colorings of an oriented knot diagram by a quandle. In this paper, we show that a certain subgroup of the inner automorphism group of a quandle acts on the R-equivalence class of a given coloring by the quandle. We also determine the R-equivalence classes of colorings of a diagram of a $2$-bridge knot by a dihedral quandle completely, under a certain condition.

math.GT↗

Geometry-aware Latent Autoregressive Generative Model for PDEs in Complex Domains

Solving multiphysics partial differential equations (PDEs) remains a major challenge in scientific computing, especially for highly complex $μ$m-scale tortuous geometries critical to energy and chemical engineering. We address this challenge by proposing a Geometry-aware Latent Autoregressive generative Model for PDEs (GeoLAMP), which solves physics within highly irregular and tortuous structures by decoupling flow and transport physics. GeoLAMP introduces a dual-encoder architecture on graph representations to jointly capture global topology and fine-scale geometric features, enabling an effective transition from real-space fields to compact latent representations. In the latent space, we propose a causal self-attention transformer with flow matching to model temporal dynamics, allowing stable and scalable block-wise autoregressive prediction. In addition, we propose a grid-graph data fusion scheme that projects low-resolution grid-based approximate priors onto graph representations, improving prediction of flow in tortuous structures. We establish three multiphysics benchmark datasets in complex geometries, covering reactive flow, heat convection, and elasticity. GeoLAMP consistently achieves the most stable autoregression performance on these datasets. Our results provide a systematic study of geometry-aware learning for PDEs in $μ$m-scale complex geometries and offer new insights into block-wise time marching of latent autoregressive PDE modeling via a flow matching framework.

cs.LG↗

Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models

Speculative decoding accelerates generation without changing its output, but on vision-language models (VLMs) a self-reinforcing cycle holds it back. Because an autoregressive drafter pays a sequential pass for each drafted token, it must stay small and can ill afford to attend to the image at each pass. Prior work therefore compresses or hides the image, leaving the drafter weakest on the text the image determines. We present GLANCE, a one-pass block drafter that breaks this cycle on an unmodified VLM target. Its block-diffusion head drafts a whole block in one forward pass over the target's already fused vision-language states, reading the multimodal context once, however deep the draft. The target verifies a wide candidate tree in one pass and commits exactly its greedy output. In one production engine at a fixed round budget, GLANCE decodes up to 3.05 times faster than autoregressive decoding and outpaces the production EAGLE3-VL head on average and by about 11% on grounded tasks. An entropy law explains when drafting pays, predicting the longest accepted blocks on grounded tasks, where the target's next-token entropy is lowest. Our code is available at https://github.com/js-lee-AI/GLANCE.

cs.AI↗

Variable-Coefficient Parabolic Equations and Finite-Horizon Hamilton--Jacobi--Bellman Equations in Augmented Spectral Barron Spaces

We establish a whole-space solution framework in augmented spectral Barron spaces for uniformly elliptic parabolic equations with state-dependent principal coefficients. A frozen-symbol parametrix yields bounded Green and terminal operators and a one-derivative smoothing estimate on arbitrary finite horizons, without requiring the spatial variation of the principal coefficient to be perturbatively small. We then apply this linear framework to a finite-horizon Hamilton--Jacobi--Bellman equation. A semi-explicit gradient iteration converges on sufficiently short horizons at a Gamma-factorial rate in a whole-space augmented Barron norm, and its limit is identified with the stochastic-control value function. Finally, temporal Jackson approximation and spatial spectral Barron approximation give joint shallow cosine-network approximations in space and time for the value function and the optimal feedback, with quantitative error and neuron-count bounds. Taken together, these results connect variable-coefficient parabolic well-posedness and smoothing in augmented spectral Barron spaces with nonlinear HJB solvability and quantitative neural-network approximation.

math.OC↗

Quit While You're Ahead: Quit for Efficient Candidate Generation in Machine Translation Reranking

Reranking methods, such as Minimum Bayes Risk (MBR) decoding and Quality Estimation (QE) reranking, have been widely used in modern neural machine translation (NMT) to select an output from a set of candidate hypotheses. However, the performance gains come at the cost of high inference latency. Existing acceleration methods target MBR decoding and reduce only the reranking computation, leaving QE reranking unaddressed and candidate generation---which can be the larger computational bottleneck---largely untouched. In this work, we propose Quit (Quantifying Uncertainty for Incremental Termination), a novel early-stopping strategy for the entire generation--reranking pipeline. Quit treats candidate generation as a sequential decision-making process under uncertainty. It incrementally generates and reranks candidates, stopping when the best reranking score stabilizes. Comprehensive experiments with three NMT models across 19 language pairs show that Quit achieves end-to-end speedups of $1.47$--$2.66\times$ for MBR decoding and $3.43$--$4.12\times$ for QE reranking while preserving automatic metric scores.

cs.CL↗

A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss

An image may be worth a thousand words, but most captioning models describe it in only a few. Modern vision-language models produce fluent high-level captions, yet routinely miss the attributes, counts, textures, materials, and spatial relations that make an image visually specific. Recent multi-stage systems recover some of these details through generation, decomposition, verification, and rewriting, but they do so at the expense of substantially higher inference latency. We propose SimLoss, a reference-free embedding-space objective for single-pass fine-grained image captioning. SimLoss trains a vision-language model to align its projected hidden-state representation with a frozen image embedding through an InfoNCE contrastive loss, supplying a dense visual supervision signal before any text is decoded, and requiring neither human-written fine-grained captions nor pseudo-captions from a multi-stage pipeline. We instantiate it as SimLoss FFT, which backpropagates through a locally available embedding model, and SimLoss GRPO, which treats that model as a black-box reward. Compared with single-pass, multi-stage verification, reward-optimized, and perception-aware baselines, the fully differentiable fine-tuning variant, SimLoss FFT, achieves the highest precision while nearly matching the F1 score of the multi-stage method, all while retaining single-pass inference and running roughly 20 times faster than the multi-stage pipeline. The reward-based variant SimLoss GRPO attains the strongest recall. Together, these results show that embedding-space supervision can recover the quality of multi-stage verification at the latency of a single-pass captioner.

cs.CV↗

Coding for Multiple Reverse-Complement and Palindromic Duplications

Reverse-complement (RC) and palindromic (PAL) duplications copy a length-$k$ block, reverse the copy, and insert it immediately after the original block; an RC duplication also complements the copied symbols. We study $q$-ary codes correcting $t$ such operations performed sequentially, so a later operation may copy symbols created by an earlier one. For fixed $q\geq2$ and $k,t\geq1$, every length-$n$ code $C$ for either channel satisfies $n-\log_q|C|\geq t\log_q n-O_{q,k,t}(1)$; for fixed $q,k$ and $1\leq t=t(n)=o(n)$ the lower bound is $t\log_q(n/t)-O_{q,k}(t)$. For a single RC error over an even alphabet with a fixed-point-free complement, the previously known RC-specific construction applies at odd $k$ and does not cover even $k$. For every even $k$ and any involutive complement, we give a coordinate-wise bijection that turns each RC duplication into a PAL duplication. Applying this bijection to every codeword therefore converts any $t$-error-correcting RC code into a PAL code of the same size, and conversely; encoders and decoders transfer by adding linear-time coordinate passes. For every even $k$, we also determine the maximum number of distinct descendants produced by exactly two errors from one source word. Words alternating between any two distinct alphabet symbols attain this maximum for PAL, and the bijection gives the RC maximizers. For both PAL and RC at fixed even $k$, we prove the existence of two-error-correcting codes with redundancy $4\log_q n+O_{q,k}(1)$. In the binary two-error problem, the converse gives $2\log_2 n-O_k(1)$, leaving a factor-two gap in the best existence bounds.

cs.IT↗

Fourier--Mukai loci are open and base change

We develop a theory of relative (quasi-)perfect complexes for algebraic spaces. Our main result proves that for a pseudocoherent kernel relatively perfect over both factors, the loci where its fiberwise Fourier--Mukai transform is fully faithful or an equivalence are open. The kernel need not be perfect, and these loci commute with Noetherian base change. Moreover, we provide an explicit autoequivalence for elliptic fibrations, and establish that smoothness and Gorensteinness are derived invariants.

math.AG↗

MeshSplatBench: A Unified Benchmark for Triangle- and Mesh-Based Neural Rendering

Triangle- and mesh-based neural rendering aims to bridge neural scene representations and existing graphics engines (\textit{e.g.}, Unity and Blender) by leveraging triangle primitives compatible with standard rasterization hardware. However, existing methods are developed and evaluated under inconsistent settings, with limited comparison and little investigation into practical graphics engine deployment. This gap significantly hinders the understanding of their real-world usability. To address this issue, we introduce MeshSplatBench, the first benchmark for systematic evaluation of triangle- and mesh-based neural rendering from native rendering to graphics engine deployment. We propose a hierarchical deployment protocol with two options: (1) Standard deployment, using a conventional opaque mesh pipeline with vertex colors and hardware Z-buffering; and (2) Dedicated deployment, incorporating method-specific engine implementations to preserve appearance and compositing properties (e.g., alpha blending). For mesh splatting, we further introduce a structural audit to evaluate the topological and geometric integrity of exported surfaces for downstream graphics applications. Extensive evaluations reveal three key findings: (1) graphics engine deployment introduces noticeable quality degradation across methods, while mesh splatting approaches achieve relatively better robustness under standard deployment; (2) dedicated deployment can preserve most rendering fidelity at the cost of approximately 6-30$\times$ slowdown; and (3) explicit connectivity and shared vertex indexing in current mesh splatting methods remain insufficient to guarantee manifoldness or global connectivity. Our benchmark demonstrates that rasterizability alone does not imply graphics readiness and highlights the importance of evaluating practical engine compatibility. The benchmark and source code will be publicly released.

cs.GR↗

Optimal Fusion Strategies for Quantum Computation

Logical fusions are important for a number of tasks in quantum information, such as quantum error correction and quantum repeaters. In the photonic setting one must contend with the fact that physical fusions can fail, i.e. individual qubits are measured in a user-controlled product state basis with some probability, which can lead to a failure on the logical level. The choice of failure basis of each qubit is known as a fusion strategy, and finding good fusion strategies is important for optimizing performance of fusion-based quantum computation. Here we provide a complete characterization when $k=1$ qubits are encoded, and in particular characterize those codes and fusion strategies such that all but one physical fusion can fail, i.e.~\emph{perfect fusion strategies}. In doing so, we recover previously known perfect fusion strategies, and find perfect fusion strategies for quantum parity-check codes, answering an open question. We furthermore show that perfect fusion strategies are generic: random $[[n, 1, d]]$ graph codes admit a perfect fusion strategy with probability exponentially close to $1$. Additionally we motivate the study of a new graph parameter, namely the maximum degree of a graph at a given vertex taken over all LC-equivalent graphs, by giving a new operationally meaningful interpretation of it.

quant-ph↗

Instability Floors: Separating Bias from Noise in Fairness Audits of Clinical LLM Agents with FairMedAgent

Counterfactual fairness audits of clinical language-model agents report a flip rate: how often an action changes when only the patient's demographic descriptor changes. Part of that rate is not demographic. A stochastic agent also changes its own action when nothing changes, and a flip rate cannot be interpreted without knowing how often. We measured it. Re-running one condition ten times over sixteen synthetic vignettes at default sampling changed a clinical agent's action in 8.7 percent of replicate pairs, from 2.2 percent for intensive-care escalation to 17.9 percent for controlled-substance caution, an output given no operational criteria. Across six models from five vendors, pooled floors ranged from 2.5 to 23.7 percent; in this panel neither disclosed size, vendor, nor hosting ordered them. The floor depends on the decoding configuration: majority voting over five draws removed 39 percent of it (95 percent confidence interval (CI) 18 to 64); at temperature 0 three of four locally served models showed no disagreement, but a hosted model still did. We also show that, for a binary action, the flip rate expected under no demographic effect equals the floor and a real effect adds only its square, so a flip rate inside the floor is not evidence of fairness, and direction must be tested with a signed paired test. We give a four-step reporting procedure and release FairMedAgent, the harness, with its protocol, vignettes and analysis scripts, so any team can measure the floor for its own agent.

cs.CL↗

Is One Step Enough for Offline Policy Improvement?

Behavior regularization in offline reinforcement learning limits the exploitation of critic errors, but strong anchoring can also restrict policy improvement. We study how policy improvement is composed through multi-step proximal policy improvement (MPI), which re-centers each proximal objective on the preceding policy. We parameterize the procedure by a nominal total horizon $T$ and $K$ stages with local horizon $T/K$, distinguishing subdivision at a fixed total horizon from additional refinement at a common local horizon. Our analysis shows that sequential re-centering can reach endpoints unavailable to any single proximal step and characterizes how subdivision reduces the leading local discretization error of ideal updates under a fixed critic. We consider TD3+BC and IQL-based policy extraction to examine how improvement composition interacts with actor objectives and policy geometry. TD3+BC experiments on D4RL locomotion suggest that subdivision can broaden the range of useful total horizons, while adding refinement stages at a fixed small local horizon can improve return. The results identify improvement composition as a design choice alongside regularization strength, with distinct effects from horizon subdivision and additional policy extraction.

cs.LG↗