Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,549 records · Page 86Linked to original sources

Temporal-Attention Head Specialization During Video Diffusion Training

Video diffusion transformers depend on temporal attention to coordinate information across frames, yet nearly everything known about this mechanism comes from analyzing trained models, so when and where temporal-attention structure forms during training remains poorly characterized. Population averages can also hide it, since a few specializing heads and a diffusing majority cancel in the mean. We therefore conduct a checkpoint-resolved census of every temporal-attention head across nine Open-Sora STDiT training runs spanning three model scales (306M to 1.03B parameters), scoring each head with an entropy-normalized measure of cross-frame attention concentration (CFAC) under a preregistered change-point and effect-size selection rule. The census reveals the sparse picture that averages obscure. Aggregate CFAC is flat or decreasing in every run, while a small minority of heads, roughly 4--13% in full-grid runs, develops pronounced concentration. Across seeds, the reproducible signal is positional but block-level. Selected heads repeatedly arise in the first temporal block, whereas individual head coordinates do not reproduce once block membership is accounted for. Among the analyzed 760M selected heads, attention maps converge to a small repertoire of local frame-routing motifs, self-frame diagonals and adjacent-frame bands, even when the responsible coordinates differ across runs. Correlation and ablation analyses do not establish a causal link to generated video quality, and we bound our claims accordingly. Beyond this STDiT family, the study contributes a transferable methodology. Checkpoint-resolved, per-head analysis under fixed selection rules can expose sparse temporal organization in other factorized video diffusion transformers and, with adapted routing metrics, in joint spatio-temporal architectures.

cs.CV↗

The Price of a Familiar Perspective

Consumers learn how to interpret an information source through repeated exposure. We study how this understanding is priced and how access to historical reports changes competition. In a Gaussian model with inherited customer histories and terminal pricing, the familiar source charges a renewal premium. Better outcome feedback widens the premium when rival reports are inaccessible, but can narrow it once some rival records are available. For independent historical dates, we derive the exact threshold for this reversal. Within the covered interior market, archive opening narrows the premium and raises consumer surplus. Its effect on aggregate forecasting quality and total surplus is less direct: a partial opening can reduce both by reallocating consumers toward a source that remains less informative. Complete access nevertheless improves both outcomes relative to any incomplete archive. The results distinguish the quality of feedback from access to the evidence with which consumers combine it.

econ.TH↗

Video-to-Music Generation for Gameplay Videos

Video-to-music models have advanced considerably in the last few years, particularly in film and music video applications. In this paper, we investigate this problem in the video game domain, which introduces new challenges for these models: video frames are rendered graphics, music is mostly synthetic audio, and soundtracks loop across entire levels rather than following on-screen events. We introduce a new dataset of 217.6 hours of Super Nintendo (SNES) gameplay video paired with 485 hours of clean soundtracks, free of sound effects and voice-overs, matched to gameplay audio via audio fingerprinting. With this dataset, we train a simple encoder-decoder transformer that passes video features directly to a MusicGen decoder, comparing different encoding strategies: textual descriptions (T5), independent frames (ViT), or spatiotemporal patches (ViViT). Each encoder is tested both frozen and fine-tuned, while the decoder is always fine-tuned. Frozen encoders match or outperform their fine-tuned counterparts on every metric, and the frozen ViViT achieves the best overall results. We compare this model with state-of-the-art baselines using both objective metrics and a listening study (N = 96). Despite having up to 18% fewer parameters, our model outperforms all baselines on objective metrics, surpasses GVMGen in the listening study, and performs comparably to OSSL.

cs.SD↗

Read-Rezayi fractional Chern insulators in modulated Bernal graphene

Fibonacci anyons provide a universal platform for topological quantum computation, and emerge as low-energy excitations in the $\mathbb{Z}_3$ Read-Rezayi phase in the fractional quantum Hall effect. However, realistic microscopic realizations of this phase in the absence of a magnetic field have remained elusive. We study a model of periodically modulated Bernal bilayer graphene with gate-screened Coulomb interactions. Using the recently developed target-phase optimization method in conjunction with band-projected exact diagonalization, we identify at filling $ν=3/5$ a region of parameter space whose ground state is consistent with a Read-Rezayi fractional Chern insulator. The partially filled band from which it arises is a part of a two-band complex which mimics geometric aspects of the lowest and first Landau levels, with the ground state at $ν=1/2$ consistent with the Moore-Read state. Our results suggest that modulated Bernal graphene can realize delicate non-Abelian fractional quantum Hall states at zero magnetic field, while demonstrating target-phase optimization as a practical route to discovering such phases in realistic, high-dimensional microscopic models.

cond-mat.str-el↗

Averaged Mirror Descent and Dual Gradient Methods: Convergent Algorithms for Entropic Gromov-Wasserstein Problems

The Gromov-Wasserstein (GW) distance measures the discrepancy between metric measure (mm) spaces and identifies optimal alignments between them based solely on their intrinsic structure. Since it identifies isomorphic mm spaces, it provides a natural notion of distance for heterogeneous datasets which may admit isomorphic representations. In order to accelerate computation of GW distances, many practitioners employ entropic regularization to obtain an Entropic GW (EGW) problem. The most popular EGW solver is the Mirror Descent (MD) algorithm, which reduces EGW computations to an iterative process where an entropic optimal transport (EOT) problem is solved at each iteration. Despite its widespread use, the convergence of MD for this problem has only been established for restricted classes of costs. On the other hand, a recently proposed dual gradient method is available for general costs, but requires a choice of step size which depends on the regularization parameter. To address these two issues, we introduce Averaged Mirror Descent (AMD), which averages consecutive MD steps, and prove its convergence for arbitrary costs. Then, we establish that the dual gradient method with a fixed step size also converges for arbitrary costs at the cost of a more complicated iteration. In both cases, we also account for inexact iterations which are inescapable in practice. We compare the empirical performance of these methods across various settings and, in particular, show that AMD and the dual gradient method both converge on an example where classical MD fails.

cs.LG↗

LLM Judge Validation Under Sparse Overlap: From Inference to Design

Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this overlap sparsity is the first-order determinant of wrong deployment decisions: at 5% pairwise overlap, wrong-decision rates reach 25% and the probability of selecting the wrong best judge among ten candidates is 65%. The two actionable levers are overlap quantity and allocation. For quantity, we derive a minimum-overlap formula showing $ρ\geq 0.25$ suffices for non-borderline judges while borderline cases remain fundamentally hard. For allocation, a zero-cost stratified scheme halves false-rejection rates relative to random sampling when strata are informative. We validate on 10 LLM judges across four evaluation matrices spanning visual assessment, causal reasoning, and summarization.

cs.AI↗

Efficient Extraction for Effectful E-Graphs

Egraphs have enabled recent advances in program optimization, synthesis, and verification, yet remain difficult to apply to effectful programs whose memory and I/O operations must respect execution order. Existing effect-aware extraction algorithms rely on integer linear programming (ILP) and dominate total runtime. We introduce Statewalk DP, a new extraction algorithm that enforces effect ordering efficiently without external solvers. We prove that finding any effect-safe extraction is NP-complete, but show that Statewalk DP is tractable in statewalk width, a parameter that measures the complexity of dataflow interactions among effects. In practice, statewalk width generally remains small, enabling Statewalk DP to achieve order-of-magnitude speedups over ILP extraction while producing programs comparable to LLVM across our benchmarks. We implement the algorithm in eggcc, a prototype egraph-based compiler for imperative Bril programs and demonstrate that effect-aware extraction is no longer a bottleneck.

cs.PL↗

Joint Amplitude-Phase Optimization of Broadband Lasers Raises Absolute Two-Plasmon-Decay Thresholds Above Coherence-Time Predictions

Broadband lasers suppress the absolute two-plasmon-decay instability that preheats inertial-fusion targets. The usual scaling relates this suppression to the drive's coherence time, a power-spectrum statistic insensitive to spectral phase. By gradient descent through a differentiable enveloped wave solver, we show that at fixed power spectrum phase optimization raises the threshold from the random-phase median of $2.8\,I_\mathrm{mono}$ to $4.1\,I_\mathrm{mono}$, but produces a transform-limited (TL) pulse train with peak-to-average ratio $R=32$. Joint amplitude--phase optimization reaches $5.6\,I_\mathrm{mono}$ at $R=3.6$, above the best tested TL line-count result of $4.5\,I_\mathrm{mono}$ at $R=16$; relaxing the peak constraint raises the joint threshold to $6.4\,I_\mathrm{mono}$. An exact growth-rate budget evaluated inside the simulation separates the two mechanisms. TL suppresses coupling mainly by concentrating the pump field into short bursts, thereby lowering its time-averaged magnitude at fixed average intensity. The joint optima retain near-random values of this amplitude measure while suppressing both the available coupling and the fraction realized through pump--daughter alignment. The advantage of joint optimization persists from $1\%$ to $4\%$ bandwidth and from OMEGA-scale to ignition-scale conditions.

physics.plasm-ph↗

Discrete quality-factor control in a side-coupled photonic crystal microcavity: evanescent Bloch tunnelling and the finite-cell correction

A point-defect microcavity side-coupled to a W1 waveguide in a two-dimensional photonic crystal of silicon rods immersed in an aqueous analyte is studied by plane-wave expansion and finite-difference time-domain computation. The resonance moves continuously through the transverse-magnetic band gap with the square of the radius of the defect rod, as first-order perturbation theory predicts for the dielectric area restored to the lattice site. The quality factor is set instead by the separation between cavity and waveguide, counted in lattice rows: each added row multiplies it by 7.10 at the reference defect radius, whereas moving the defect rod by up to a tenth of a lattice period changes it by less than 7 per cent. The per-row factor follows, at two defect radii and with no adjustable parameter, from the decay of the slowest evanescent Bloch channel of the crystal at the wavevector of the guided mode. Two checks are needed before such values can be trusted in a finite cell: convergence in cladding thickness and removal of the reflections from the waveguide ends. Without them the per-row factor comes out too low and the quality factor wrong by up to a factor of two. Scaled to 1550 nm the design gives a sensitivity of 634 nm per refractive index unit, which perturbation theory reproduces, and a Fano fit to an independently normalised transmission spectrum confirms the linewidth. The absorption of water at this wavelength lowers the quality factor of the reference geometry by a third and caps it near 9200, which limits the benefit of adding further rows.

physics.optics↗

TriO: Tri-Modal Unsupervised Occupancy World Model for Anything Perception

We present TriO, a multi-modal unsupervised world model that predicts 4D occupancy, obstacle segmentation, flow and LiDAR. In contrast to prior work, TriO utilizes three distinct sensor modalities (camera, LiDAR, and RADAR) as both inputs and sources of self-supervision, eliminating the need for additional human annotations. Thanks to its novel supervision, the model is able to segment any occupancy from the drivable surface, overcoming the limitations of existing open-set methods in handling long-tail objects. TriO achieves state-of-the-art results in multiple 3D and 4D tasks, including occupancy, flow, and LiDAR prediction, as well as zero-shot road obstacle segmentation across multiple datasets such as Argoverse 2, and Spotting the Unexpected.

cs.CV↗

Residual Denoising Enables Sample-Efficient Multi-Agent Coordination on Demand

Pretrained robot policies offer strong manipulation skills but are typically limited to single-agent settings, where a robot acts in isolation. In this work, we study how to adapt pretrained single-agent diffusion policies to multi-agent settings using minimal collaborative data, co-optimizing for two key objectives: high coordination performance and single-agent skill retention. To this end, we introduce ALTER, an adaptation method for coordination on demand: the adapted policy coordinates with other robots when deployed in a team while remaining capable of acting independently when operating alone. Execution is decentralized: each robot acts only on its own visual observations, without explicit inter-agent communication. Our method trains a coordination head that predicts a residual denoiser to transform single-agent behavior into coordinated multi-agent behavior when necessary while also preserving single-agent capabilities. To preserve single-agent capabilities, we augment a small number of collaborative demonstrations with self-distilled data generated by the base policy during training of the residual denoiser. In simulation, ALTER achieves higher coordination success over our baselines while retaining much higher source-skill retention. In our hardware experiments, we find similar trends where ALTER better co-optimizes for coordination success and single-agent skill retention than the baselines.

cs.RO↗

CoMemBench: Benchmarking Collaborative Memory Boundaries across Multi-Agent Workflow Topologies

Multi-agent workflows require task-relevant information to be shared across agents, while irrelevant, stale, unverified, or incompatible information must remain isolated. We call this task-conditioned scope of information a collaborative memory boundary. Workflow topology determines which intermediate artifacts are applicable to which downstream workers and when they cease to be valid, thereby providing a structural stress dimension for sharing and isolation. Existing memory benchmarks primarily evaluate retention and retrieval, whereas multi-agent benchmarks emphasize coordination and end-to-end completion, leaving topology-conditioned memory boundaries largely unmeasured. We introduce CoMemBench, an execution-grounded benchmark for collaborative memory sharing and isolation across multi-agent workflow topologies. It constructs 800 composite workflows across four domains from source-grounded dependency graphs, with node-local specifications, verifiable artifact handoffs, native evaluators, and matched isolation challenges. CoMemBench measures workflow completion, verified node progress, required-handoff reliability, isolation robustness, and token cost. Experiments reveal a sharing-isolation trade-off: broader context improves information availability but can weaken isolation, while system rankings shift across topologies and artifact violations.

cs.AI↗

Instruct, Not Answer: Using Instruction Privileges in On-Policy Context Distillation

On-Policy Context Distillation (OPCD) has recently emerged as a powerful technique for transferring context to student models and for self-improvement. In OPCD, the teacher is conditioned on privileged information, and the goal is to minimize the Kullback-Leibler (KL) divergence between the privileged teacher and the student, evaluated on student-generated tokens. Many existing studies show that using instance-specific gold answers or gold demonstrations as the default privilege can hurt training performance, especially out-of-distribution (OOD). In this work, we instead design general instructions that target common student mistakes observed on the training samples, and show that such simple instructions can outperform gold as the OPCD privilege. In autoformalization tasks, using a matched formatting instruction as the privilege could outperform gold in OOD accuracy by a large margin. In 7 out of 8 experiments using ProverQA, ProofWriter, and ProntoQA as datasets, and Qwen3-Thinking and Olmo3-Thinking families as models, matched instruction privileges outperform gold in OOD by 4 to 17 points, while remaining on par with gold in-domain. Each instruction is only a few sentences (and thus contains much less information compared to all instance-specific gold) and is applied uniformly to every training sample. These results indicate that a general instruction, which applies equally to source and target domain examples, can be substantially more transferable than instance-specific gold in OPCD while maintaining in-domain performance.

cs.AI↗

Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs

Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing the sender's key-value (KV) cache avoids this redundancy, but prefill-free transfer across model families must handle differences in tokenization, model depth, and KV representations. To address these issues, we propose \textit{HeteroFold}, a prefill-free cross-family KV cache transfer method that keeps both the sender and receiver frozen. HeteroFold aligns model structures, maps the sender cache into the receiver space, and calibrates it to preserve receiver behavior. Across six transfer directions, HeteroFold achieves the best cache-transfer performance on all four long-context benchmarks and most short-context settings. It also matches text-based communication on the multi-agent benchmark. At 32K context length, Llama-3.1-8B$\rightarrow$Ministral-3-14B transfer is $10.7\times$ faster than Native Prefill and $1.18$--$1.47\times$ faster than the state-of-the-art prefill-free baselines, Dense Latent and KV Ridge. These results show that HeteroFold enables efficient cross-family KV reuse without receiver prefill.

cs.AI↗

ScentGen: Hierarchical Multimodal Olfactory Semantic Modeling for Molecular Odor Description Generation

In this paper, we introduce a molecular odor description generation task, which aims to generate natural language odor descriptions from molecular structures. Unlike conventional methods that describe molecular odor using discrete labels, this task generates expressive and human-interpretable sensory descriptions. To address this task, we propose a hierarchical multimodal olfactory semantic modeling framework, named ScentGen. ScentGen consists of three key components: an odor semantic planner, a semantic adapter, and a description generator. The odor semantic planner integrates complementary molecular information from 1D SMILES sequences, 2D molecular graphs, and 3D molecular conformations to learn discriminative and structured olfactory semantics. The semantic adapter further maps the learned olfactory representation into the hidden space of a large language model, transforming molecular odor semantics into language-compatible continuous prompts. Conditioned on these prompts, the description generator produces coherent odor descriptions that reflect plausible sensory characteristics of the input molecule. Considering the lack of molecular datasets with natural language odor descriptions, we further construct a molecular odor description dataset containing paired multimodal molecular representations and human-interpretable odor descriptions. Extensive experiments demonstrate that ScentGen generates coherent and expressive odor descriptions, providing a more flexible solution for molecular odor understanding beyond discrete odor label prediction.

cs.CE↗

Affordance-Conditioned Decision Making: Bridging the Semantic-Spatial Gap in Zero-Shot Cross-Floor Vision-and-Language Navigation

Vision-and-language navigation increasingly relies on general-purpose semantic planners, yet translating correct high-level intent into reliable physical execution remains difficult in spatially constrained transitions. Reaching a staircase, doorway, or narrow passage does not ensure traversal; the agent must identify an executable affordance pose and recover from accumulated action errors. We propose PACE (Preference-refined Affordance-Conditioned Execution), a supervised local execution module that augments frozen zero-shot semantic planners for reliable cross-floor navigation. PACE grounds transition-related semantics into a long-horizon, agent-centric traversable affordance pose and conditions short-horizon action generation on this spatial target, thereby aligning semantic goals with physical execution. We further post-train PACE through failure-aware preference refinement using rollout-derived pairs that contrast normal or recovery behaviors with deviation-amplifying behaviors, thereby improving closed-loop correction. We integrate PACE into six open-source zero-shot VLN navigators and demonstrate consistent improvements on the cross-floor subsets of R2R-CE and RxR-CE, increasing the average success rate from 16.35% to 27.65% and from 4.76% to 12.06%, respectively. Real-world experiments further demonstrate PACE's applicability in unseen environments, highlighting the potential of traversable affordances to bridge semantic intent and reliable embodied behavior.

cs.RO↗

EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding

Vision-language models (VLMs) have shown increasing potential for medical image understanding, yet their capabilities in ophthalmic imaging remain insufficiently characterized. Existing ophthalmic datasets are typically designed for individual diseases or specialized tasks, making it difficult to systematically evaluate whether VLMs can move beyond disease recognition toward comparative reasoning and fine-grained spatial grounding. We introduce EyeVQA, a unified visual question answering benchmark for comprehensive evaluation of ophthalmic VLMs. EyeVQA is constructed from 21 available ophthalmic datasets and contains 20,000 question-answer pairs spanning six disease groups and seven question types: Single-Choice, Multi-Select, Variable-Select, True-False, Ranking, Point Location, and Bounding Box. Gold answers are deterministically derived from source-provided diagnoses, severity grades, clinical findings, segmentation masks, bounding boxes, and anatomical landmarks, enabling reproducible evaluation without relying on model-generated annotations. Notably, 44.5% of the questions require reasoning across multiple images, extending evaluation beyond conventional single-image medical VQA. We benchmark fourteen representative general-purpose, scientific, and medically specialized VLMs under a unified zero-shot protocol. The best-performing model only achieves an overall score of 62.8, while substantial gaps remain in spatial grounding and cross-task generalization. These results highlight the limitations of current VLMs in comprehensive ophthalmic visual understanding and establish EyeVQA as a diagnostic benchmark for developing more reliable and spatially grounded ophthalmic multimodal models. The project page is available at https://github.com/PKUTHM/EyeVQA.

cs.CV↗

Matrix varieties over $S_7$: dimension-three stability and continuum-sized subvariety intervals

We study matrix semirings over the three-element flat additively idempotent semiring with elements one, a, and infinity, where the square of a is infinity. We determine the associated matrix-variety chain completely. The scalar, two-by-two, and three-by-three cases generate three distinct varieties, while every matrix dimension at least three generates the same variety. The proof shows that every failure of an identity in arbitrary dimension is already witnessed on three indices, and it yields a coordinate criterion for all identities in the stable variety. We also realize every graph semiring arising from a directed graph of in-degree and out-degree at most one as a divisor of a direct power of the two-by-two matrix semiring. Consequently, the variety generated by all three-nilpotent flat semirings is contained in the two-by-two matrix variety. Directed cycles and independent reversal identities then embed the power-set lattice of the odd primes into each interval between the base variety and a nontrivial matrix variety. Thus every such interval has continuum cardinality and contains continuum-sized chains and antichains. Finally, we determine the last three powers of the multiplicative subsemiring obtained by deleting the constant all-one matrix, and we develop general matrix operators on the lattice of additively idempotent semiring varieties, including stable closures, stable cores, and propagation of equality along matrix-dimension chains.

math.GR↗