SearcharxivSearch

arXiv subjects

Yue Yu

Publications and source records attributed to Yue Yu.

At least 19 recordsLinked to original sources

A multicenter benchmark and clinically structured metric for coronary CTA report generation

Reliable evaluation of automated coronary computed tomography angiography (CCTA) report generation requires standardized multicentre benchmarks and clinically structured metrics. We established a four-centre benchmark comprising 3,021 CCTA series from 818 patient-report pairs to evaluate seven open-source three-dimensional vision-language models. We developed CSM$_{\text{CCTA}}$, a clinically structured metric for CCTA report evaluation, with patient-, vessel-, and segment-level variables defined according to clinical guidelines. Report pairs are compared at the finest shared anatomical level, and the contributions of different clinical components are weighted based on expert assessments. We estimated these weights using 70 expert-scored cases and evaluated clinical alignment in a non-overlapping set of 30 cases. CSM$_{\text{CCTA}}$ showed a strong correlation with radiologist scores (Pearson's $r=0.97$, $p<0.001$), exceeding the next-best metric, FORTE ($r=0.70$), by 0.27, and agreed with expert preferences in 115 of 160 pairwise comparisons (71.9\%). Under controlled perturbations, CSM$_{\text{CCTA}}$ remained stable to clinically equivalent wording and decreased monotonically with progressive information omission. In the multicenter benchmark, the CCTA-trained C2RG model achieved the highest CSM$_{\text{CCTA}}$ scores across all four hospitals, although its performance remained far from optimal. In contrast, CCTA-irrelevant reports accounted for up to 98.7\% of the outputs from generalist models. Together, the benchmark provides a standardized setting for model comparison, while CSM$_{\text{CCTA}}$ enables clinically structured evaluation of finding agreement and anatomical specificity. These results support a more clinically aligned and anatomically resolved approach to evaluating CCTA report generation. Code is available at https://openi.pcl.ac.cn/OpenMedIA/CSM_CCTA.

cs.CV

GRAIN: Bridging Name and Narrative Shifts in Real-World Graph Reasoning through Invariance-Rewarded Agentic RL

Despite their potential in standardized graph tasks, Large Language Models (LLMs) remain brittle to real-world shifts in node identifiers and task formulation. While deterministic graph tools are invariant to such shifts, extracting topological structures from noisy text is highly fragile for LLMs, which often overfit to surface patterns. Moreover, mitigating these parsing failures via multi-agent systems incurs prohibitive latency. To address this, we propose GRAIN, a single-agent framework optimized via reinforcement learning. GRAIN models reasoning as a semantic parsing and tool-execution pipeline, guided by a Structure Invariance Reward. By validating extracted intermediate graphs against ground-truth topologies, this reward forces the LLM to learn robust text-to-structure mappings rather than memorizing linguistic artifacts. We also introduce GRIT, a benchmark evaluating sensitivity to such linguistic shifts. GRAIN outperforms multi-agent baselines by 16.45\% in accuracy with approximately 24\% lower latency. Furthermore, it demonstrates superior structural generalization, halving the out-of-distribution (OOD) gap of SFT models (from 15.77\% to 7.80\%) and maintaining robustness on large-scale graphs beyond the training distribution.

cs.AI

H-PAC Hand: Control-Oriented Modeling and Tendon-Elasticity Compensation for an Underactuated Robotic Hand

Underactuated tendon-driven hands offer compact actuation and passive compliance, but tendon elongation under restoring-spring loading introduces configuration-dependent joint deviations. This paper presents H-PAC, a modular 6-actuator, 15-DoF robotic hand with a control-oriented modeling and implementation framework. A sparse analytical actuator-joint model is derived from the tendon-routing geometry, and a mechanics-based compensation model is developed to account for tendon-elasticity-induced joint errors. The proposed method is implemented in a hierarchical architecture: a host computer performs workspace-constrained posture mapping and compensation, while an ESP32 generates synchronized commands for six position-controlled servos. The same control parameters and execution strategy are used across all tasks without task-specific retuning. Monotonic servo-sweep experiments show that the compensation substantially improves joint-angle prediction. The MAE of the index DIP joint decreases from 1.15 degrees to 0.18 degrees, and all nine evaluated joints achieve an MAE below 0.23 degrees. Representative postures and grasping configurations are further executed using the same control pipeline without external joint or force sensing in the control loop. The results demonstrate a practical approach to improving posture reproducibility in compact underactuated robotic end-effectors.

cs.RO

Few-cycle electro-optic light on thin-film lithium niobate

The twin fields of ultrafast optics and nonlinear photonics enable applications ranging from attosecond science [1, 2] and ultrafast electronics [3] to molecular spectroscopy [4, 5], nonlinear optics [6-8], quantum nanophotonics [9] and precision metrology [10]. However, bringing these capabilities-including ultrashort pulse generation, dispersion control, and strong nonlinear interactions-together within a scalable photonic integrated platform requires exceptional performance and cooperation between components while preserving sufficient optical power across the circuit. Here we demonstrate an integrated multi-functional nonlinear photonic system that transforms continuous-wave (CW) light into high-peak-power femtosecond pulses and harnesses them for pulse-driven nonlinear optics on thin-film lithium niobate (TFLN). Microwave-driven electro-optic (EO) broadening followed by integrated dispersive compression generates 230-fs Fourier-transform-limited pulses with energies up to 3.3 pJ at 30.7 GHz, representing orders of magnitude higher pulse energy than previous integrated pulse synthesis at comparable repetition rates [11]. In a 0.3-meter dispersion-engineered TFLN waveguide, soliton dynamics compress the pulses to 35 fs (6.7 optical cycles), accompanied by coherent spectral broadening exceeding 330 nm. In a fully-monolithic architecture, EO synthesis, dispersive compression and a high-Q nonlinear resonator are integrated on a single TFLN chip, enabling resonantly-enhanced pulse pumping and coherent spectral broadening at pulse energies as low as 400 fJ. By unifying microwave-controlled pulse synthesis and pulse-driven nonlinear interactions, our work establishes a direct path from CW excitation to few-cycle nonlinear optics on chip, with opportunities spanning microwave photonics [12], optical frequency synthesis and metrology, and mid-infrared and terahertz generation [13, 14].

physics.optics

A unifying framework for quantum algorithms for time-dependent non-unitary dynamics

Quantum algorithms for simulating linear differential equations have attracted growing interest, driven by applications ranging from Hamiltonian dynamics to general non-unitary dynamics. While time-independent cases are well studied, time-dependent non-unitary dynamics remains considerably less explored, and it is unclear how to systematically adapt existing solvers for time-independent systems to such problems. In this work, we address this gap by introducing an autonomization framework based on the clock-variable formulation, a technique originally developed for time-dependent Hamiltonian systems in~\cite{CJL23TimeSchr}. By lifting the original non-autonomous system to an autonomous transport-type equation on an extended space and applying the Fourier spectral discretization in the clock variable, we obtain an explicit time-independent linear system, together with a suitable initial state and a recovery map for the target solution. Crucially, this formulation decouples the treatment of time dependence from the choice of the quantum ODE solver, thereby enabling the direct application of existing solvers designed for time-independent systems to the resulting autonomous problem. We combine this framework with Schr\"odingerization and a Taylor-expansion-based quantum ODE solver. In the Schr\"odingerization-based combination, our complexity analysis shows that the precision dependence can scale as $\log^{5/4}(1/\varepsilon)$, improving upon the $\log^2(1/\varepsilon)$ scaling found in existing approaches. Numerical experiments validate the autonomization formulation and confirm the successful recovery of the target solution.

quant-ph

SphereVideo: Prototype-anchored Hyperspherical Boundary for Continual AI-generated Video Detection

AI-generated video (AIGV) detection aims to distinguish real videos from AI-generated ones. In practice, detectors trained on existing data often fail to generalize to newly emerging generative models, making this task challenging. Therefore, continual learning (CL) is essential for improving the adaptability. However, CL frameworks for this task remain underexplored. To this end, we propose SphereVideo, a novel CL framework for AIGV detection built on two key observations. First, real videos exhibit a compact feature distribution. Based on this, we encourage real video features to cluster around a real prototype on a hypersphere while repelling AI-generated samples, thereby establishing a decision boundary. This prototype serves as a stable anchor for CL, regulating boundary evolution and mitigating catastrophic forgetting. Second, existing methods tend to rely solely on spatial artifacts as shortcuts. To enhance temporal modeling, we introduce a strategy that models the temporal dynamics of real data at both frame and clip levels. By strengthening real data modeling, this strategy further facilitates learning a real prototype and forming a stable decision boundary. Moreover, we construct a comprehensive and challenging benchmark. Extensive experiments demonstrate that SphereVideo achieves an improved plasticity-stability trade-off, outperforming prior methods by 3.08% on seen data and 4.00% on unseen AI-generated data.

cs.CV

From Individual to Shared Ownership: A Coalitional Game Approach to Sustainable Co-investment

This paper proposes a cooperative game-theoretic framework for sustainable co-investment in shared infrastructure under regulatory incentives. Multiple heterogeneous operators co-invest in a common infrastructure whose production capability evolves over time and is subject to operational variability. A regulator supports the deployment through incentive mechanisms designed to align individual economic investment objectives with the coalitional one. We formulate the co-investment problem as a transferable-utility (TU) coalitional game in which the value generated by cooperation depends on heterogeneous operational profiles, dynamic resource availability, investment costs, and regulatory incentive level. We show that the proposed coalitional game can be reformulated as a linear production game (LPG), whose dual prices yield a constructive and stable allocation of the cooperative surplus. Finally, we illustrate the proposed framework through a case study on co-investment among data center operators in shared renewable energy infrastructure, supported by government subsidies promoting renewable energy consumption.

cs.GT

CameraAnything: Refilming Videos with Arbitrary Camera Control

We introduce CameraAnything, the first unified framework for camera controlled video editing that enables joint control of both intrinsic and extrinsic camera parameters. Existing approaches either rely on expensive 3D reconstruction to achieve full camera functionality or restrict editing to extrinsic parameter manipulation. Moreover, the coupled influence of intrinsic and extrinsic parameters on video appearance makes disentangled modeling particularly challenging. To address this, we adopt per-pixel Pl\"ucker ray injection alongside resolution-aware 3D RoPE in self-attention, building both camera conditioning and spatial positional encoding on the target latent to jointly control camera position, focal length, and native resolution editing without cropping or outpainting. To overcome the scarcity of paired training data, we further develop a scalable synthetic pipeline that constructs diverse dynamic scenes through structured multi-camera recording and generates synchronized videos with varied camera configurations. With a tailored orthogonal training strategy, CameraAnything enables expressive video reshooting with arbitrary viewpoint control, focal length adjustment, resolution adaptation, and multi-shot transitions within a single generation process, offering strong practical value for cinematic video editing and cross-platform content adaptation in video production.

cs.CV

Spin-Consistency Constraints in Noncollinear Tensor TDA

TDDFT for open-shell systems, whether spin-conserving or spin-flip, has long suffered from spin contamination. This problem arises because the single-excitation space built upon a single Kohn-Sham determinant is not spin-complete. Adopting spin tensor reference states therefore offers an elegant and promising route to resolving this issue. In this work, we revisit the tensor TDDFT equations within the Tamm-Dancoff approximation (TDA) from a noncollinear perspective. We show that, for S = 1/2 reference states, the internal consistency of the spin tensor formulation can, with the aid of the zero-excitation-energy theorem, be recast as a set of constraints that the exchange-correlation kernel must satisfy. Standard noncollinear functionals, however, generally fail to meet these constraints. To address this, we propose a kernel reconstruction scheme that is independent of the specific functional form and free of empirical parameters. This scheme enforces the required constraints, restoring internal consistency in the full spin tensor structure, with spin adaptation following as a natural consequence. Furthermore, when extended to tensor reference states with other values of S, such as S = 1 for the oxygen molecule, the scheme eliminates the so-called artifact states, namely solutions with severely underestimated excitation energies. In addition, the scheme allows the target states that ROKS reference states aim to describe to be expressed and computed, at the TDA level, within the same unified framework as other states, a capability that spin-adapted spin-conserving TDDFT has so far lacked.

physics.chem-ph

Overview and design optimization of a custom hybrid X-ray telescope for the International Axion Observatory (IAXO)

We present the design optimization for maximizing the effective area of a custom X-ray optic for the International Axion Observatory (IAXO) and BabyIAXO, including its novel hybrid configuration that enables full coverage of the 700-mm-diameter magnetic bore with minimal stress imposed on the mirrors; shell layout optimized for axion spectra and spatial distribution; and the coating recipes that enhance reflectivity in the energy range of interest. We evaluate how these design choices improve the observation signal-to-noise ratio (SNR) of BabyIAXO and IAXO by calculating the broad-band effective area and simulating the point spread function (PSF) and focal spot at the detector plane. The cost-effective and scalable optic offers an energy response from 0.03--15 keV, achieving an effective area that exceeds 2400 cm$^2$ near 1 keV - the peak of the ABC axion spectrum - and remains above 1700 cm$^2$ around 3 keV - the peak of the Primakoff axion spectrum. It yields a half-power diameter (HPD) of $\sim 46^{\prime\prime}$ for an on-axis point source at infinity, and a focal-spot HPD of $\sim 120^{\prime\prime}$ for the radial distribution expected for axion signals within the approximately $3^{\prime}$-radius solar core. A relatively generous fabrication-error budget is also summarized. The custom optic, accounting for fabrication errors, is anticipated to deliver a more than $55$-fold enhancement in the SNR.

physics.ins-det

Learning to Detect UI Principle Violations via Reinforcement Learning

Small language models and coding agents increasingly generate web front-end code, yet their outputs are typically evaluated primarily for functional correctness. A generated interface may compile, render, and pass unit tests while still violating established interface quality principles, including accessibility barriers, deceptive design patterns, poor visual hierarchy, and excessive decision complexity. Existing auditing approaches face a trade-off between cost, coverage, and scalability: expert human review provides rich judgment but is slow and expensive; frontier vision-language models offer broader reasoning capabilities but remain costly to deploy at scale; and rule-based tools such as axe-core and Lighthouse are inexpensive but primarily capture mechanically checkable accessibility issues. We investigate whether a lightweight vision-language model can serve as an effective critic for generated interfaces. We unify 19 interface-quality principles from three complementary sources of HCI knowledge: WCAG 2.2 accessibility standards, deceptive design taxonomies, and established theories of perception, cognition, and interaction. To train this critic, we construct a verified dataset of approximately 10,000 generated web pages by synthetically injecting known violations into clean, LLM-generated Tailwind pages. Continued reinforcement learning on a 4B vision-language model improves micro-F1 from 36\% to 84\%, with 13 of 19 principles exceeding 80\% F1. The resulting critic can audit generated interfaces, filter low-quality interface training data, and provide a reward signal for design-aware code generation. We release our data-generation recipe and injection/verification prompts to support reproducible evaluation and future work on scalable interface-quality assessment.

cs.CL

Fabrication status and expected performance of the inner-core X-ray optic for BabyIAXO

BabyIAXO, a pathfinder for the International Axion Observatory (IAXO), is designed to demonstrate all key technologies at scale while achieving an improvement in sensitivity over the recent CERN Axion Solar Telescope (CAST) experiment by approximately a factor of five. Such improvement is enabled by the X-ray optics, which allow for maintaining a high signal-to-noise ratio at the detector despite a cross-sectional area of the magnetic bore being over 250 times larger than that of CAST. The optic employs a hybrid design consisting of co-aligned inner core and outer corona optics that share a common optical axis and vacuum vessel but differ in focal length and manufacturing approach. Both are segmented glass optics, with the inner core fabricated from thermally slumped borosilicate glass and the outer corona from cold-slumped Corning Willow glass. To fabricate the inner-core optic, leveraging techniques developed for NuSTAR and HEFT optics, we reoptimized and streamlined the thermal-forming procedure. The quality of free-standing glass substrates was characterized by laser metrology, X-ray reflectometry, and atomic force microscopy. We developed a cutting technique that produces smooth edges at the micron scale. We used flat stacks of glass-epoxy-graphite layers to evaluate the performance of the epoxy bondline. The optic is expected to achieve an on-axis point spread function (PSF) with a half-power diameter (HPD) of < 90", enhancing the signal-to-noise ratio by more than 55 times.

physics.ins-det

Map as a Prompt: Learning Multi-Modal Spatial-Signal Foundation Models for Cross-scenario Wireless Localization

Accurate and robust wireless localization is a critical enabler for emerging 5G/6G applications, including autonomous driving, extended reality, and smart manufacturing. Despite its importance, achieving precise localization across diverse environments remains challenging due to the complex nature of wireless signals and their sensitivity to environmental changes. Existing data-driven approaches often suffer from limited generalization capability, requiring extensive labeled data and struggling to adapt to new scenarios. To address these limitations, we propose SigMap, a multimodal foundation model that introduces two key innovations: (1) A cycle-adaptive masking strategy that dynamically adjusts masking patterns based on channel periodicity characteristics to learn robust wireless representations; (2) A novel "map-as-prompt" framework that integrates 3D geographic information through lightweight soft prompts for effective cross-scenario adaptation. Extensive experiments demonstrate that our model achieves state-of-the-art performance across multiple localization tasks while exhibiting strong zero-shot generalization in unseen environments, significantly outperforming both supervised and self-supervised baselines by considerable margins.

eess.SP

Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable

The capability of a modern AI agent depends not only on its foundation model but also on its harness, which constructs prompts, manages state, invokes tools, and coordinates execution. As models, APIs, environments, and requirements evolve, the harness must be continually modified. Before such a change can be made, a developer or coding agent must identify all code locations that implement the target behavior. This is difficult because production harnesses are large, tightly coupled, and behaviorally distributed, while modification requests describe what the system should do and repositories are organized by files and modules. Code search, repository indexing, and long-context processing ease inspection, but still leave this behavior-to-code mapping to be recovered by hand. Behavior localization is therefore a central bottleneck in harness evolution. We introduce the Harness Handbook, a behavior-centric representation synthesized automatically from a harness codebase via static analysis and LLM-assisted structuring, linking each behavior to its corresponding source. We also introduce Behavior-Guided Progressive Disclosure (BGPD), which guides agents from high-level behaviors to relevant implementation details and verifies candidate locations against the current source. On diverse modification requests from two open-source harnesses, Handbook-Assisted planning improves behavior localization and edit-plan quality while using fewer planner tokens, with the largest gains on scattered sites, rarely executed paths, and cross-module interactions. Evolving complex agentic systems thus depends not only on generating edits, but also on determining where those edits should be made.

cs.AI

Precision masses of neutron-rich platinum and gold nuclei reveal enhanced $N=126$ shell strength below doubly-magic $^{208}$Pb

The heaviest stable nuclei in the universe owe their existence to quantum shell structure, the grouping of protons and neutrons into discrete energy levels separated by gaps. The largest known neutron shell gap in stable nuclei, at $N=126$, stabilizes doubly-magic $^{208}$Pb and is responsible for the characteristic abundance peak of heavy elements near gold and platinum produced by the rapid neutron-capture process (r-process). Whether this shell gap persists as protons are removed from lead is a question central to both nuclear structure and the modeling of heavy-element synthesis, yet it has remained unanswered due to the extraordinary difficulty of producing the relevant neutron-rich nuclei. Direct experimental knowledge in this region was essentially absent. Here we report the first precision mass measurements of $^{203,204}$Pt and $^{204,205,206}$Au, performed at GSI using a novel combination of Schottky and isochronous mass spectrometry in a heavy-ion storage ring. The $N=126$ isotones $^{204}$Pt and $^{205}$Au are more strongly bound than the extrapolated trend of the previously known mass surface by 403 and 464~keV, respectively, revealing an unexpectedly enhanced $N=126$ shell strength below doubly-magic $^{208}$Pb. Furthermore, the proton-neutron interaction strength exhibits a hitherto unobserved bifurcation at $N=126$ as protons are removed from $^{208}$Pb. Our results redefine the nuclear mass surface in the neutron-rich heavy-element region and provide direct experimental benchmarks for theoretical models whose extrapolations toward more exotic nuclei are essential for r-process nucleosynthesis calculations.

nucl-ex

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated only by their final outcome. This setup overlooks intermediate progress and partial solutions, yielding sparse reward signals and an incomplete picture of agent capability. We introduce Long-Horizon-Terminal-Bench, a terminal benchmark of 46 long-horizon tasks spanning nine categories, including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing. Each task follows a Terminal-Bench-style setup with a reference solution or simulation engine, but is further decomposed into fine-grained graded subtasks. This design enables dense intermediate rewards and partial credit, allowing evaluation to capture not only whether an agent reaches the final goal, but also how far it progresses on open-ended workflows. Tasks in Long-Horizon-Terminal-Bench typically require hundreds of episodes and minutes to hours of execution, stressing long-horizon planning, long-context management, and iterative debugging rather than one-shot problem solving. We evaluate 15 frontier models and find that agents consume on average 9.9M tokens per task, with roughly 231 episodes and 85.3 minutes of execution time per run, making Long-Horizon-Terminal-Bench more demanding than prior terminal-based benchmarks. Even the strongest tested model achieves 15.2% pass@1 at a partial-reward threshold of 0.95 and 10.9% at a perfect-reward threshold of 1.0, while the mean pass rate across models is 4.3% and 1.7% under the two thresholds, respectively. These results reveal headroom for improvement. We further analyze failure modes and error patterns, and release Long-Horizon-Terminal-Bench to support future progress on long-horizon terminal agents.

cs.AI

WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory

We present WorldDirector, a highly controllable video world model framework designed for persistent dynamic object memory and unrestricted viewpoint exploration. Unlike existing world models that entangle physical dynamics with pixel rendering and rely on continuous visual observation to sustain motion, our framework explicitly decouples semantic motion orchestration from visual generation. By leveraging an LLM to coordinate 3D trajectories with camera movements and subsequently employing these orchestrated trajectories as control signals for video generation, our approach ensures strict physical logic and appearance stability, successfully preserving the exact visual identities of dynamic entities even when they re-enter the scene after prolonged periods out of view. Experimental results demonstrate that our method supports the synthesis of complex and extended events with unprecedented controllability and persistent dynamic object memory. Project Page: https://worlddirector.github.io/

cs.CV

Schr\"odingerization based quantum algorithms for regularized Wasserstein proximal operators

We develop a quantum algorithm for the regularized Wasserstein proximal operator, which is a fundamental tool in optimal transport and mean-field games. The regularization introduces a small diffusive term into the continuity equation of the Benamou-Brenier formulation, which results in a forward-backward PDE system consisting of a Fokker-Planck equation and a viscous Hamilton-Jacobi equation with a quadratic Hamiltonian. Through the Cole-Hopf transformation, both equations are converted to forward heat equations, whose coupling requires a Hadamard division to prepare the initial data for the second heat equation and a Hadamard product to recover the terminal density. We solve these heat equations via the Schr\"odingerization method and implement the Hadamard division and product operations using simple matrix-vector multiplication representations. The complete quantum algorithm prepares an $\varepsilon$-approximation of the terminal density state with $\mathcal{O}(d N_x T \log^2(1/\varepsilon))$ query complexity, up to constants depending on the potential and initial density, where $d$ is the spatial dimension, $N_x$ is the number of grid points per spatial dimension and $T$ is the evolution time. The complexity depends only {\it linearly} on $d N_x$, yielding an {\it exponential} speedup over classical methods, whose cost scales as $N_x^d$ per time step. Numerical experiments validate the effectiveness of the proposed algorithm.

math.NA