SearcharxivSearch

arXiv subjects

Yue Cao

Publications and source records attributed to Yue Cao.

At least 19 recordsLinked to original sources

HERA: Historical Evidence Routing Adapter for Physical Prediction in Latent World Models

Predictive video models have emerged as promising world models by learning latent visual dynamics from large-scale video. Yet these models remain challenged by physical events under occlusion, where later predictions may depend on object evidence that is no longer available in the current view. Addressing this challenge requires historical evidence not only to be preserved but also to remain accessible when it becomes relevant to a subsequent prediction. Existing approaches mainly enlarge the temporal context, cache generic video features, or impose explicit object-centric states, thereby improving the capacity or structure of retained history. However, they do not directly address how relevant historical evidence can be selectively retrieved and integrated into a pretrained predictor without interfering with its native latent workspace. Accordingly, we introduce HERA (Historical Evidence Routing Adapter), a framework for routing retained historical evidence into a frozen latent predictor, and instantiate it with Register-Routed Patch Memory (RRPM), a lightweight adapter comprising a Structured Memory Bank, Memory Registers, and Workspace Registers. On the IntPhys2 Main split, HERA with RRPM improves the pairwise AvgSurprise accuracy of V-JEPA 2-G from 52.57% to 54.35%. Subgroup analysis shows particularly strong improvements on fixed-camera continuity, from 46.15% to 57.69%, and fixed-camera immutability, from 46.15% to 63.46%. These results support historical evidence routing as a practical adaptation strategy for physical prediction in latent world models.

cs.CV

Retrieve in Time, Correct in Frequency

Frozen vision-language-action (VLA) policies generate temporally extended action chunks, but long-horizon manipulation remains vulnerable to accumulated execution error and visual aliasing across task stages. Successful rollouts provide useful corrective evidence, yet current frame retrieval can return progress-misaligned actions,while direct replay or time-domain fusion can overwrite the reactive structure of the policy proposal. We introduce Retrieve in Time, Correct in Frequency (RTCF), a training-free test-time correction framework that improves frozen VLA performance with low model-side overhead.RTCF separates which experience to retrieve from which part of its action to transfer. Progressive Memory Alignment (PMA) causally aligns the growing visual execution history with complete successful trajectories through incrementally updated monotonic frontiers, jointly identifying a relevant memory and the current aligned memory position without stage labels. From the aligned action chunk,RTCF transfers a coefficient-wise-clipped low-frequency residual on motion channels. Higher-frequency components and gripper decisions remain inherited from the frozen policy. Across four LIBERO suites and 2,000 episodes per condition, RTCF raises aggregate success from 86.4% to 88.4% and improves LIBERO-Long from 61.6% to 68.6%.These gains require no parameter updates, repeated VLA inference, or additional GPU resources: correction can be performed on the client CPU after a single policy invocation, and the median latencies sum to only 10.99 ms per action chunk

cs.RO

Secure Long-Range Autonomous Valet Parking: A Reservation Scheme With Three-Factor Authentication and Key Agreement

Long-range autonomous valet parking (LAVP) is increasingly adopted to alleviate traffic congestion and parking difficulties. For large-scale parking demand, reservation can improve parking management. However, existing schemes mainly focus on parking request verification and parking check-in, and do not adequately protect identity legitimacy and communication security during passenger drop-off and pick-up. To address this problem, we propose SecLAVP, a provably secure three-factor authentication and key agreement protocol for LAVP reservation services. SecLAVP combines passwords, biometrics, and smart cards. With assistance from the drop-off/pick-up point (DP), the passenger and the autonomous vehicle (AV) achieve mutual authentication and establish a session key for secure communication. In the Real-Or-Random (ROR) model, we formally prove that SecLAVP provides session-key security. AVISPA simulations show that SecLAVP resists man-in-the-middle attacks, while informal analysis demonstrates that it satisfies 15 defined security goals. Finally, performance evaluation in terms of communication overhead, computational overhead, and scheduling shows that SecLAVP is feasible for practical deployment.

cs.CR

Electro-Optic Active Metasurfaces for High-Speed Photonic Applications

Metasurfaces are artificially engineered ultrathin nanostructured surfaces, capable of flexibly manipulating light-matter interactions on compact platforms, and thereby of great significance for a wide range of applications within modern optics and photonics, including communications, computing, sensing, and quantum technologies. However, the inherently static nature of conventional metasurfaces severely limits their functionalities and thus range of possible applications. Benefiting from integration of the metasurface platform for shaping optical wavefronts with ultrafast electro-optic (EO) materials, active EO metasurfaces have emerged as a frontier research direction targeting advanced photonic devices. This paper systematically reviews the latest progress in this field, featuring a comprehensive comparison of performances and application scenarios of mainstream EO materials such as lithium niobate, barium titanate and organic EO polymers. Modulation mechanisms based on the Pockels and Kerr effects along with the corresponding active metasurface implementations are summarized. Furthermore, improvements in modulation efficiency enabled by advantageously exploiting resonant structural designs and associated phenomena, including Fabry-Perot resonances, Mie resonances, surface plasmon polaritons, quasi-bound states in the continuum, surface lattice resonances, and guided-mode resonances, are presented and summerized in detail. Current challenges related to metasurface design, nanofabrication, performance and heterogeneous integration are also discussed. Finally, future research directions are outlined, highlighting interdisciplinary developments, novel material engineering, and AI-assisted design as key pathways to enable practical use of active EO metasurfaces in modern optics and photonics, including quantum information technologies.

physics.optics

Robust Monitoring of Arc Welding Processes: A Generalizable Framework with DVAE and Particle Filter

Arc welding processes are essential for continuous fabrication but prone to disturbances that impair weld quality, making real-time monitoring critical yet difficult due to complex visual patterns and nonlinear, time-varying dynamics. Deep learning shows promise but faces scalability limits because of its dependence on large labeled datasets and application-specific tuning. We explore whether a unified approach can characterize major arc welding processes across applications and improve scalability through consistent state monitoring. This paper introduces a robust and generalizable monitoring framework for arc welding. It combines unsupervised deep latent representation learning, which extracts compact features from weld pool images, with Bayesian filtering to handle persistent and fluctuating disturbances such as arc radiation and specular reflections. Specifically, a Dynamic Variational Autoencoder (DVAE), consisting of a CNN-based encoder-decoder and an LSTM-based transition model, jointly learns latent representations and their evolution under control inputs. For robust real-time inference, a specialized Particle Filter (PF) propagates the latent and LSTM hidden states, preserving process history while suppressing sensor noise. This design is well suited to welding's slow and inertial dynamics. Validation on GTAW and GMAW without process-specific tuning demonstrates the framework's generalizability and robustness.

eess.IV

Unified monogamy and polygamy relations for multipartite systems

For a bipartite entanglement measure $\mathcal{E}$ that satisfies the $\gamma$th-power monogamy inequality (Eq.~\eqref{e:chap1-ineq1}), and for its assisted counterpart $\mathcal{E}_a$ that obeys the $\delta$th-power polygamy inequality (Eq.~\eqref{e:chap1-ineq2}), we introduce a unified, tunable framework indexed by a parameter $m\geq1$. Within this framework, we derive two hierarchical families of refined inequalities: a tightened $\alpha$-power monogamy relation for $\mathcal{E}$, valid for all $\alpha \geq m\gamma$; a tightened $\beta$-power polygamy relation for $\mathcal{E}_a$, applicable for $(m-1)\delta < \beta \leq m\delta$. As $m$ increases, the bounds become progressively tighter, recovering known results at $m=1$. Notably, the optimal monogamy bound emerges as a piecewise function of $\alpha$, with additional correction terms activated as $\alpha$ crosses successive integer thresholds, thereby offering a sharper characterization of entanglement distribution. We demonstrate that our results generalize and strengthen existing monogamy and polygamy relations through analytical comparisons and numerical evaluations using concurrence and concurrence of assistance. This hierarchical, parameterized approach offers enhanced and flexible tools for applications in quantum communication, quantum networks, and multipartite quantum information processing.

quant-ph

Geometry-Optimized Complex-Domain error-diffusion encoding for Fourier Single-Pixel Imaging

This work proposes a geometry-optimized complex-domain error-diffusion encoding method for Fourier single-pixel imaging. Instead of independently binarizing multiple grayscale phase-shifting patterns, the proposed method directly represents each complex-valued Fourier basis pattern using K (K >= 3) weighted binary patterns while diffusing the residual error in the complex domain. A geometric interpretation is further established, revealing that the encoding process can be viewed as approximating the Fourier-basis unit circle by a regular polygon in the complex plane. Based on this geometric interpretation, practical optimization strategies are developed for K = 3, K = 4, and K = 7. Both numerical simulations and real-object experiments demonstrate consistently superior reconstruction quality compared with conventional phase-shifting dithering.

physics.optics

Near-real-time, meter-scale 3D urban wind modeling for low-altitude micrometeorology: numerical verification of a GPU-accelerated lattice Boltzmann framework

This study presents a near-real-time, meter-scale three-dimensional urban wind simulation framework for low-altitude flight events in complex urban meteorological environments. It reconstructs high-resolution wind fields by combining sparse observations with efficient microscale flow modeling. The framework integrates lattice Boltzmann method large-eddy simulation (LBM-LES), high-fidelity urban morphology reconstruction that explicitly resolves real building details, and observation-driven boundary assimilation into a rapid end-to-end pipeline for realistic urban domains. Multi-site Doppler lidar measurements from dense urban Guangzhou, China, are used for evaluation. The system reconstructs three-dimensional wind fields at 5 m resolution over kilometer-scale domains within minutes. Robustness and accuracy are tested through controlled observation reduction, independent validation against withheld lidar stations, and sensitivity analyses of grid resolution and precursor domain extent. Results show stable reproduction of vertical wind structures and key local flow features under complex morphology and limited observations, providing a scalable pathway for near-real-time urban wind reconstruction.

physics.flu-dyn

Equilibrium singular dividend control under ambiguity aggregation of heterogeneous discount rates

This paper studies a singular dividend control problem for a firm with heterogeneous shareholders whose discount rates follow a given distribution. The central planner aggregates expected discounted payoffs using an ambiguity aggregation function $phi$, which captures shareholder heterogeneity and ambiguity attitudes but also leads to time inconsistency. To address this issue, we seek a time-homogeneous equilibrium dividend law characterized by a partition of the state space into waiting and dividend-paying regions. We provide a rigorous mathematical characterization by proving a verification theorem and deriving necessary conditions for the equilibrium law. We then analyze barrier-type equilibria, showing non-existence for a class of aggregation functions that includes power-type and logarithmic aggregation functions, and establishing existence and uniqueness under linear and exponential aggregation. In the linear case, the bounded-rate equilibrium is shown to converge to the singular barrier-type equilibrium as the dividend rate bound tends to infinity. Numerical examples illustrate the effects of discount-rate heterogeneity and ambiguity aversion on the equilibrium barrier.

math.OC

DCSNet: Multiscale Feature Aggregation for Small Medical Object Segmentation with Detection-guided Hierarchical Cropping

Small object segmentation in medical imaging is primarily hindered by class imbalance and inherent boundary complexity. Consequently, conventional global networks frequently fail to detect sparse targets or suffer from severe edge degradation. To overcome these limitations, we propose the Detection-guided Cropping Segmentation Network (DCSNet), an end-to-end framework that transforms global dense prediction into a localized refinement process. This framework integrates two core components, namely Detection-guided Hierarchical Cropping (DGHC) and Multiscale Feature Aggregation (MSFA). The DGHC module leverages region proposals to dynamically extract object-centric features, effdataectively filtering out massive background interference to mitigate class imbalance. Subsequently, the MSFA module operates strictly within these purified regions, synergizing a Transformer encoder with a pixel-adaptive fusion strategy. This mechanism dynamically aggregates multiscale features to capture both semantic context and fine-grained details for sharp boundary delineation. Extensive experiments across three diverse medical datasets demonstrate that DCSNet significantly outperforms existing state-of-the-art methods, yielding substantial improvements in boundary precision and offering a highly robust solution for clinical micro-lesion segmentation.

cs.CV

Structural Assessment for Understanding and Guiding Dataset Distillation in Discrete Token Space

Dataset distillation (DD) has proven to reduce training cost while preserving accuracy. While promising, the factors that make one distilled dataset more effective than another remain poorly understood. In this work, we investigate this question through the lens of discrete visual tokenizers. Whereas many prior DD efforts emphasize matching global data distributions, we suggest that the effectiveness depends on which semantic concepts are captured and how they are composed. Discrete visual tokenizers provide a finite vocabulary that enables direct statistical analysis of such compositional structure. Through quantitative analysis of token-level statistics, we introduce the structural score to measure the adequacy of token compositions. We observe that distilled datasets with balanced token composition yield higher validation performance. On the other hand, divergence from the original data does not necessarily harm performance. We further show that samples with high structural scores in the discrete token space can effectively guide diffusion-based DD. Our findings highlight the importance of token composition in dataset effectiveness, offering a principled complement to distributional similarity considerations in DD.

cs.CV

Geometrically Constrained Stenosis Editing in Coronary Angiography via Entropic Optimal Transport

The scarcity of high-quality imaging data for coronary angiography (CAG) stenosis limits the clinical translation of automated stenosis detection. Synthetic stenosis data provides a practical avenue to augment training sets, improving data quality, diversity, and distributional coverage, and enhancing detection precision and generalization. However, diffusion-based editing commonly relies on soft guidance in a noise-initialized reverse process, offering limited pixel-level precision and structure preservation. We propose the OT-Bridge Editor, which reframes localized editing as a constrained entropic optimal transport (OT) problem and leverages geometric information to steer the generation path, enabling stronger geometric control. Extensive experiments show that our synthesized angiograms consistently improve downstream stenosis detection, yielding substantial relative gains of 27.8% on the public ARCADE benchmark and 23.0% on our multi-center dataset, supported by consistent qualitative results.

cs.CV

Physics-Grounded Understanding of Thermal Boundary Conductance between Ga$_2$O$_3$ and SiC from a Feedforward Neural Network Potential

Ga$_2$O$_3$/SiC heterointegration is attractive for ultra-wide-bandgap power electronics, but interfacial thermal boundary conductance (TBC) remains a major heat-removal bottleneck. Direct experimental access to intrinsic atomistic interfacial transport remains limited, particularly for ideally synthesized materials with defect-free interfacial contact. First-principles simulations are too expensive at relevant length and time scales, while empirical Molecular Dynamics (MD) potentials often lack transferability across oxide and carbide bonding environments. We develop a unified feedforward neural network potential and validate it against density-functional data, bulk phonon dispersions, and anisotropic thermal-conductivity trends in both $\beta$-Ga$_2$O$_3$ and SiC. Nonequilibrium simulations show that TBC decreases with transport length, increases with temperature, and is consistently higher for Ga$_2$O$_3$$(\bar{2}01)$/SiC(0001) than for Ga$_2$O$_3$(100)/SiC(0001). These trends are explained by attenuation of long-mean-free-path carriers, enhanced incoherent and anharmonic interfacial exchange within broadly unchanged spectral channels, and stronger bonding and vibrational coupling at the $(\bar{2}01)$ interface. The results show how a single transferable feedforward neural network potential can enable large-scale transport prediction and physics-grounded mechanistic understanding of thermal boundary conductance. Code for NEP training and simulation workflows is available at the project repository https://github.com/knowhow07/TBC_Ga2O3_SiC.git

cond-mat.mtrl-sci

Adaptor: Advancing Assistive Teleoperation with Few-Shot Learning and Cross-Operator Generalization

Assistive teleoperation enhances efficiency via shared control, yet inter-operator variability, stemming from diverse habits and expertise, induces highly heterogeneous trajectory distributions that undermine intent recognition stability. We present Adaptor, a few-shot framework for robust cross-operator intent recognition. The Adaptor bridges the domain gap through two stages: (i) preprocessing, which models intent uncertainty by synthesizing trajectory perturbations via noise injection and performs geometry-aware keyframe extraction; and (ii) policy learning, which encodes the processed trajectories with an Intention Expert and fuses them with the pre-trained vision-language model context to condition an Action Expert for action generation. Experiments on real-world and simulated benchmarks demonstrate that Adaptor achieves state-of-the-art performance, improving success rates and efficiency over baselines. Moreover, the method exhibits low variance across operators with varying expertise, demonstrating robust cross-operator generalization.

cs.RO

GIFT: Global Irreplaceability Frame Targeting for Efficient Video Understanding

Video Large Language Models (VLMs) have achieved remarkable success in video understanding, but the significant computational cost from processing dense frames severely limits their practical application. Existing methods alleviate this by selecting keyframes, but their greedy decision-making, combined with a decoupled evaluation of relevance and diversity, often falls into local optima and results in erroneously selecting irrelevant noise frames. To address these challenges, we propose GIFT: Global Irreplaceability Frame Targeting, a novel training-free framework that selects frames by assessing their intrinsic irreplaceability. Specifically, we first introduce Directed Diversity to quantify a frame's uniqueness conditioned on relevance, which allows us to formulate a unified irreplaceability score. Subsequently, our Budget-Aware Refinement strategy employs a adaptive iterative process that first secures a core set of frames with the highest irreplaceability, and then shifts its priority to building crucial temporal context around these selections as the budget expands. Extensive experiments demonstrate that GIFT achieves a maximum average improvement of 12.5% across long-form video benchmarks on LLaVA-Video-7B compared to uniform sampling.

cs.CV

Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model

We present daVinci-MagiHuman, an open-source audio-video generative foundation model for human-centric generation. daVinci-MagiHuman jointly generates synchronized video and audio using a single-stream Transformer that processes text, video, and audio within a unified token sequence via self-attention only. This single-stream design avoids the complexity of multi-stream or cross-attention architectures while remaining easy to optimize with standard training and inference infrastructure. The model is particularly strong in human-centric scenarios, producing expressive facial performance, natural speech-expression coordination, realistic body motion, and precise audio-video synchronization. It supports multilingual spoken generation across Chinese (Mandarin and Cantonese), English, Japanese, Korean, German, and French. For efficient inference, we combine the single-stream backbone with model distillation, latent-space super-resolution, and a Turbo VAE decoder, enabling generation of a 5-second 256p video in 2 seconds on a single H100 GPU. In automatic evaluation, daVinci-MagiHuman achieves the highest visual quality and text alignment among leading open models, along with the lowest word error rate (14.60%) for speech intelligibility. In pairwise human evaluation, it achieves win rates of 80.0% against Ovi 1.1 and 60.9% against LTX 2.3 over 2000 comparisons. We open-source the complete model stack, including the base model, the distilled model, the super-resolution model, and the inference codebase.

cs.CV

SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation

Recent advances in text-to-image (T2I) generation via reinforcement learning (RL) have benefited from reward models that assess semantic alignment and visual quality. However, most existing reward models pay limited attention to fine-grained spatial relationships, often producing images that appear plausible overall yet contain inaccuracies in object positioning. In this work, we present \textbf{SpatialReward}, a verifiable reward model explicitly designed to evaluate spatial layouts in generated images. SpatialReward adopts a multi-stage pipeline: a \emph{Prompt Decomposer} extracts entities, attributes, and spatial metadata from free-form prompts; expert detectors provide accurate visual grounding of object positions and attributes; and a vision-language model applies chain-of-thought reasoning over grounded observations to assess complex spatial relations that are challenging for rule-based methods. To more comprehensively evaluate spatial relationships in generated images, we introduce \textbf{SpatRelBench}, a benchmark covering object attributes, orientation, inter-object relations, and rendered text placement. Experiments on Stable Diffusion and FLUX show that incorporating SpatialReward into RL training consistently improves spatial consistency and overall generation quality, with results aligned more closely to human judgments. These findings indicate that verifiable reward models hold considerable potential for enabling more accurate and controllable optimization in text-to-image generation models.

cs.CV

AVION: Aerial Vision-Language Instruction from Offline Teacher to Prompt-Tuned Network

Adapting vision-language models to remote sensing imagery remains challenging due to two key factors: limited semantic coverage in textual representations and insufficient adaptability of visual features. These issues are particularly significant in aerial scenes, which involve various visual appearances and fine-grained object distinctions. We propose AVION, a knowledge distillation framework tailored for remote sensing adaptation of vision-language models. The teacher module constructs semantically rich textual prototypes by collecting descriptions from a large language model and verifying validity using remote sensing image features. The student module integrates lightweight and learnable prompts into both vision and language encoders, guided by the teacher to align embeddings and their cross-modal relationships. Once trained, the student operates independently during inference. Experiments on six optical remote sensing benchmarks show that AVION improves few-shot classification and base-class accuracy without degrading generalization to novel categories. It also enhances mean recall for cross-modal retrieval, with minimal additional trainable parameters.

cs.CV