SearcharxivSearch

arXiv subjects

Yue Yu

Publications and source records attributed to Yue Yu.

At least 37 records · Page 2Linked to original sources

CameraAnything: Refilming Videos with Arbitrary Camera Control

We introduce CameraAnything, the first unified framework for camera controlled video editing that enables joint control of both intrinsic and extrinsic camera parameters. Existing approaches either rely on expensive 3D reconstruction to achieve full camera functionality or restrict editing to extrinsic parameter manipulation. Moreover, the coupled influence of intrinsic and extrinsic parameters on video appearance makes disentangled modeling particularly challenging. To address this, we adopt per-pixel Plücker ray injection alongside resolution-aware 3D RoPE in self-attention, building both camera conditioning and spatial positional encoding on the target latent to jointly control camera position, focal length, and native resolution editing without cropping or outpainting. To overcome the scarcity of paired training data, we further develop a scalable synthetic pipeline that constructs diverse dynamic scenes through structured multi-camera recording and generates synchronized videos with varied camera configurations. With a tailored orthogonal training strategy, CameraAnything enables expressive video reshooting with arbitrary viewpoint control, focal length adjustment, resolution adaptation, and multi-shot transitions within a single generation process, offering strong practical value for cinematic video editing and cross-platform content adaptation in video production.

cs.CV

Spin-Consistency Constraints in Noncollinear Tensor TDA

TDDFT for open-shell systems, whether spin-conserving or spin-flip, has long suffered from spin contamination. This problem arises because the single-excitation space built upon a single Kohn-Sham determinant is not spin-complete. Adopting spin tensor reference states therefore offers an elegant and promising route to resolving this issue. In this work, we revisit the tensor TDDFT equations within the Tamm-Dancoff approximation (TDA) from a noncollinear perspective. We show that, for S = 1/2 reference states, the internal consistency of the spin tensor formulation can, with the aid of the zero-excitation-energy theorem, be recast as a set of constraints that the exchange-correlation kernel must satisfy. Standard noncollinear functionals, however, generally fail to meet these constraints. To address this, we propose a kernel reconstruction scheme that is independent of the specific functional form and free of empirical parameters. This scheme enforces the required constraints, restoring internal consistency in the full spin tensor structure, with spin adaptation following as a natural consequence. Furthermore, when extended to tensor reference states with other values of S, such as S = 1 for the oxygen molecule, the scheme eliminates the so-called artifact states, namely solutions with severely underestimated excitation energies. In addition, the scheme allows the target states that ROKS reference states aim to describe to be expressed and computed, at the TDA level, within the same unified framework as other states, a capability that spin-adapted spin-conserving TDDFT has so far lacked.

physics.chem-ph

Overview and design optimization of a custom hybrid X-ray telescope for the International Axion Observatory (IAXO)

We present the design optimization for maximizing the effective area of a custom X-ray optic for the International Axion Observatory (IAXO) and BabyIAXO, including its novel hybrid configuration that enables full coverage of the 700-mm-diameter magnetic bore with minimal stress imposed on the mirrors; shell layout optimized for axion spectra and spatial distribution; and the coating recipes that enhance reflectivity in the energy range of interest. We evaluate how these design choices improve the observation signal-to-noise ratio (SNR) of BabyIAXO and IAXO by calculating the broad-band effective area and simulating the point spread function (PSF) and focal spot at the detector plane. The cost-effective and scalable optic offers an energy response from 0.03--15 keV, achieving an effective area that exceeds 2400 cm$^2$ near 1 keV - the peak of the ABC axion spectrum - and remains above 1700 cm$^2$ around 3 keV - the peak of the Primakoff axion spectrum. It yields a half-power diameter (HPD) of $\sim 46^{\prime\prime}$ for an on-axis point source at infinity, and a focal-spot HPD of $\sim 120^{\prime\prime}$ for the radial distribution expected for axion signals within the approximately $3^{\prime}$-radius solar core. A relatively generous fabrication-error budget is also summarized. The custom optic, accounting for fabrication errors, is anticipated to deliver a more than $55$-fold enhancement in the SNR.

physics.ins-det

Fabrication status and expected performance of the inner-core X-ray optic for BabyIAXO

BabyIAXO, a pathfinder for the International Axion Observatory (IAXO), is designed to demonstrate all key technologies at scale while achieving an improvement in sensitivity over the recent CERN Axion Solar Telescope (CAST) experiment by approximately a factor of five. Such improvement is enabled by the X-ray optics, which allow for maintaining a high signal-to-noise ratio at the detector despite a cross-sectional area of the magnetic bore being over 250 times larger than that of CAST. The optic employs a hybrid design consisting of co-aligned inner core and outer corona optics that share a common optical axis and vacuum vessel but differ in focal length and manufacturing approach. Both are segmented glass optics, with the inner core fabricated from thermally slumped borosilicate glass and the outer corona from cold-slumped Corning Willow glass. To fabricate the inner-core optic, leveraging techniques developed for NuSTAR and HEFT optics, we reoptimized and streamlined the thermal-forming procedure. The quality of free-standing glass substrates was characterized by laser metrology, X-ray reflectometry, and atomic force microscopy. We developed a cutting technique that produces smooth edges at the micron scale. We used flat stacks of glass-epoxy-graphite layers to evaluate the performance of the epoxy bondline. The optic is expected to achieve an on-axis point spread function (PSF) with a half-power diameter (HPD) of < 90", enhancing the signal-to-noise ratio by more than 55 times.

physics.ins-det

Map as a Prompt: Learning Multi-Modal Spatial-Signal Foundation Models for Cross-scenario Wireless Localization

Accurate and robust wireless localization is a critical enabler for emerging 5G/6G applications, including autonomous driving, extended reality, and smart manufacturing. Despite its importance, achieving precise localization across diverse environments remains challenging due to the complex nature of wireless signals and their sensitivity to environmental changes. Existing data-driven approaches often suffer from limited generalization capability, requiring extensive labeled data and struggling to adapt to new scenarios. To address these limitations, we propose SigMap, a multimodal foundation model that introduces two key innovations: (1) A cycle-adaptive masking strategy that dynamically adjusts masking patterns based on channel periodicity characteristics to learn robust wireless representations; (2) A novel "map-as-prompt" framework that integrates 3D geographic information through lightweight soft prompts for effective cross-scenario adaptation. Extensive experiments demonstrate that our model achieves state-of-the-art performance across multiple localization tasks while exhibiting strong zero-shot generalization in unseen environments, significantly outperforming both supervised and self-supervised baselines by considerable margins.

eess.SP

Resonance-enhanced integrated acousto-optic beam steering

Optical beam steering is a key technology for free-space optical communication, sensing, and imaging. Mechanical beam steering systems suffer from limited scanning speed and bulky form factors, while existing solid-state solutions rely on pixelated synthetic aperture that requires complex fabrication and control architectures. Integrated acousto-optic beam steering (AOBS) is an emerging technology that enables continuous one-dimensional beam steering using integrated acoustic transducers and fixed-wavelength laser sources. Here, we integrate AOBS with an optical ring resonator on the same thin-film lithium niobate (TFLN) platform to significantly enhance beam steering efficiency and system functionality. The resulting device achieves a resonance-enhanced beam steering efficiency of up to $26\%$ and a field of view of $18^\circ$. Moreover, by leveraging integrated electro-optic control, we dynamically lock the ring-resonator's resonance to a chirped laser frequency, enabling frequency-modulated continuous-wave (FMCW) LiDAR operation. By combining lithium niobate's piezoelectric and electro-optic properties, this work establishes a compact, efficient, and scalable beam-steering platform with co-integrated acousto-optic modulation and electro-optic control for multifunctional applications.

physics.app-ph

Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable

The capability of a modern AI agent depends not only on its foundation model but also on its harness, which constructs prompts, manages state, invokes tools, and coordinates execution. As models, APIs, environments, and requirements evolve, the harness must be continually modified. Before such a change can be made, a developer or coding agent must identify all code locations that implement the target behavior. This is difficult because production harnesses are large, tightly coupled, and behaviorally distributed, while modification requests describe what the system should do and repositories are organized by files and modules. Code search, repository indexing, and long-context processing ease inspection, but still leave this behavior-to-code mapping to be recovered by hand. Behavior localization is therefore a central bottleneck in harness evolution. We introduce the Harness Handbook, a behavior-centric representation synthesized automatically from a harness codebase via static analysis and LLM-assisted structuring, linking each behavior to its corresponding source. We also introduce Behavior-Guided Progressive Disclosure (BGPD), which guides agents from high-level behaviors to relevant implementation details and verifies candidate locations against the current source. On diverse modification requests from two open-source harnesses, Handbook-Assisted planning improves behavior localization and edit-plan quality while using fewer planner tokens, with the largest gains on scattered sites, rarely executed paths, and cross-module interactions. Evolving complex agentic systems thus depends not only on generating edits, but also on determining where those edits should be made.

cs.AI

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated only by their final outcome. This setup overlooks intermediate progress and partial solutions, yielding sparse reward signals and an incomplete picture of agent capability. We introduce Long-Horizon-Terminal-Bench, a terminal benchmark of 46 long-horizon tasks spanning nine categories, including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing. Each task follows a Terminal-Bench-style setup with a reference solution or simulation engine, but is further decomposed into fine-grained graded subtasks. This design enables dense intermediate rewards and partial credit, allowing evaluation to capture not only whether an agent reaches the final goal, but also how far it progresses on open-ended workflows. Tasks in Long-Horizon-Terminal-Bench typically require hundreds of episodes and minutes to hours of execution, stressing long-horizon planning, long-context management, and iterative debugging rather than one-shot problem solving. We evaluate 15 frontier models and find that agents consume on average 9.9M tokens per task, with roughly 231 episodes and 85.3 minutes of execution time per run, making Long-Horizon-Terminal-Bench more demanding than prior terminal-based benchmarks. Even the strongest tested model achieves 15.2% pass@1 at a partial-reward threshold of 0.95 and 10.9% at a perfect-reward threshold of 1.0, while the mean pass rate across models is 4.3% and 1.7% under the two thresholds, respectively. These results reveal headroom for improvement. We further analyze failure modes and error patterns, and release Long-Horizon-Terminal-Bench to support future progress on long-horizon terminal agents.

cs.AI

Precision masses of neutron-rich platinum and gold nuclei reveal enhanced $N=126$ shell strength below doubly-magic $^{208}$Pb

The heaviest stable nuclei in the universe owe their existence to quantum shell structure, the grouping of protons and neutrons into discrete energy levels separated by gaps. The largest known neutron shell gap in stable nuclei, at $N=126$, stabilizes doubly-magic $^{208}$Pb and is responsible for the characteristic abundance peak of heavy elements near gold and platinum produced by the rapid neutron-capture process (r-process). Whether this shell gap persists as protons are removed from lead is a question central to both nuclear structure and the modeling of heavy-element synthesis, yet it has remained unanswered due to the extraordinary difficulty of producing the relevant neutron-rich nuclei. Direct experimental knowledge in this region was essentially absent. Here we report the first precision mass measurements of $^{203,204}$Pt and $^{204,205,206}$Au, performed at GSI using a novel combination of Schottky and isochronous mass spectrometry in a heavy-ion storage ring. The $N=126$ isotones $^{204}$Pt and $^{205}$Au are more strongly bound than the extrapolated trend of the previously known mass surface by 403 and 464~keV, respectively, revealing an unexpectedly enhanced $N=126$ shell strength below doubly-magic $^{208}$Pb. Furthermore, the proton-neutron interaction strength exhibits a hitherto unobserved bifurcation at $N=126$ as protons are removed from $^{208}$Pb. Our results redefine the nuclear mass surface in the neutron-rich heavy-element region and provide direct experimental benchmarks for theoretical models whose extrapolations toward more exotic nuclei are essential for r-process nucleosynthesis calculations.

nucl-ex

LectūraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching

Effective personalized AI-assisted learning demands systems that can not only generate accurate learner-specific educational materials, but also dynamically adapt their instruction to diverse learners. However, existing educational agents have primarily focused on lecture content automation and simulations, which often fall short of modelling multimodal and embodied instructional methods tailored for the individual learner. To this end, we propose LectūraAgents - a multi-agent framework that enables personalized learning through end-to-end adaptive embodied teaching. At its core, LectūraAgents mirrors a professor-student relationship, in which a ProfessorAgent leads a collaborative team of specialized subordinate agents through research, planning, review, and embodied delivery of lecture contents that adapt to a learner's needs. The framework offers three main contributions: (1) a hierarchical multi-agent architecture for end-to-end personalized learning; (2) an adaptive embodied teaching mechanism, wherein the ProfessorAgent executes visible and pedagogically motivated teaching actions (e.g., handwrite, highlight, underline, etc.) over contents in a teaching environment; and (3) a Teaching Action-Speech Alignment (TASA) algorithm that employs salience-based heuristics and temporal semantic segmentation to generate coherent teaching action sequences aligned with learner profiles. We evaluate LectūraAgents on diverse courses at high school, undergraduate, and graduate levels using sample-specific rubric-based analysis; with generated lecture materials and teaching actions assessed and validated by expert educators. Experimental results show consistent gains in lecture content quality, embodied teaching quality, assessment, and personalization over existing approaches, positioning LectūraAgents as a pedagogically well-grounded framework for personalized learning at scale.

cs.CL

WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory

We present WorldDirector, a highly controllable video world model framework designed for persistent dynamic object memory and unrestricted viewpoint exploration. Unlike existing world models that entangle physical dynamics with pixel rendering and rely on continuous visual observation to sustain motion, our framework explicitly decouples semantic motion orchestration from visual generation. By leveraging an LLM to coordinate 3D trajectories with camera movements and subsequently employing these orchestrated trajectories as control signals for video generation, our approach ensures strict physical logic and appearance stability, successfully preserving the exact visual identities of dynamic entities even when they re-enter the scene after prolonged periods out of view. Experimental results demonstrate that our method supports the synthesis of complex and extended events with unprecedented controllability and persistent dynamic object memory. Project Page: https://worlddirector.github.io/

cs.CV

Schrödingerization based quantum algorithms for regularized Wasserstein proximal operators

We develop a quantum algorithm for the regularized Wasserstein proximal operator, which is a fundamental tool in optimal transport and mean-field games. The regularization introduces a small diffusive term into the continuity equation of the Benamou-Brenier formulation, which results in a forward-backward PDE system consisting of a Fokker-Planck equation and a viscous Hamilton-Jacobi equation with a quadratic Hamiltonian. Through the Cole-Hopf transformation, both equations are converted to forward heat equations, whose coupling requires a Hadamard division to prepare the initial data for the second heat equation and a Hadamard product to recover the terminal density. We solve these heat equations via the Schrödingerization method and implement the Hadamard division and product operations using simple matrix-vector multiplication representations. The complete quantum algorithm prepares an $\varepsilon$-approximation of the terminal density state with $\mathcal{O}(d N_x T \log^2(1/\varepsilon))$ query complexity, up to constants depending on the potential and initial density, where $d$ is the spatial dimension, $N_x$ is the number of grid points per spatial dimension and $T$ is the evolution time. The complexity depends only {\it linearly} on $d N_x$, yielding an {\it exponential} speedup over classical methods, whose cost scales as $N_x^d$ per time step. Numerical experiments validate the effectiveness of the proposed algorithm.

math.NA

Mamba-FSCIL: Dynamic Adaptation with Selective State Space Model for Few-Shot Class-Incremental Learning

Few-shot class-incremental learning (FSCIL) aims to incrementally learn novel classes from limited examples while preserving knowledge of previously learned classes. Existing methods face a critical dilemma: static architectures rely on a constant parameter space to learn from data that arrive sequentially, making them prone to overfitting to the current session, while dynamic architectures continually expand the parameter space, leading to increased complexity. In this study, we explore the potential of Selective State Space Models (SSMs) for FSCIL. Mamba leverages its input-dependent parameters to dynamically adjust its processing patterns and generate content-aware scan patterns without session-wise projector expansion. This enables it to configure distinct processing for base and novel classes, helping preserve existing knowledge while adapting to new ones. To leverage Mamba's potential for FSCIL, we design two key modules: First, we propose a dual selective SSM projector that generates input-conditioned state-space parameters from intermediate features for dynamic adaptation. The dual design structurally decouples base and novel-class processing, employing a frozen base branch to maintain stable base-class features and a dynamic incremental branch that adaptively learns distinctive feature shifts for novel classes. Second, we develop a class-sensitive selective scan mechanism to guide dynamic adaptation of the incremental branch. It reduces the disruption to base-class representations caused by training on novel data, and meanwhile, encourages the selective scan to perform in distinct patterns between base and novel classes. Extensive experiments on miniImageNet, CIFAR-100, and CUB-200 demonstrate that Mamba-FSCIL achieves state-of-the-art performance.

cs.CV

HOLMES: Evaluating Higher-Order Logical Reasoning in LLMs

Logical reasoning is essential for reliable AI, yet existing benchmarks are largely first-order-logic-centric, focusing on object-level deduction over fixed predicates. This misses many realistic scenarios where models must reason over rules, predicates, functions, constraints, and decision procedures themselves. We introduce HOLMES (Higher-Order Logic Meets real-world Explainable Symbolic reasoning), the first real-world benchmark for higher-order symbolic reasoning in LLMs, containing 1379 instances. Built on higher-order logic, HOLMES pairs natural-language problems with HOL formalizations, ground-truth answers, verifiable reasoning traces, and fine-grained controllable reasoning factors across law and finance. Experiments show that current LLMs still struggle on HOLMES, with an average accuracy of only 50.64% and the best model reaching 59.54%. Our analyses further reveal that high final-answer accuracy can mask shortcut reasoning in conflict-resolution settings, while performance drops sharply under scope-conditioned and compositional reasoning. These findings identify higher-order symbolic reasoning as a key bottleneck for building reliable and verifiable LLMs. The project code and dataset are publicly available at https://github.com/wuyucheng2002/HOLMES.

cs.AI

Quantum preconditioning method for finite difference discretizations of the Poisson equation via Schrödingerization

We present a quantum preconditioning framework for solving linear systems arising from a finite difference discretization of the Poisson equation. It is based on the combination of the Schrödingerization technique \cite{JLY22b,JLYPRL24} and the BPX multilevel preconditioner in order to achieve near-optimal complexity. The Schrödingerization technique transforms linear partial and ordinary differential equations into Schrödinger-type systems with unitary evolution in one higher dimension, making them suitable for quantum simulation. A key contribution is a structure-aware construction of the block-encoding for the symmetrically preconditioned matrix $A_S = S^\top A S$, where $A$ is the stiffness matrix and $S$ encodes the BPX preconditioner in factored form. By establishing a novel commuting identity, we avoid the unfavorable normalization scaling that would otherwise arise from naive multiplication of block-encodings. This yields an exact block-encoding of $A_S$ with normalization $\mathcal{O}(d^2(L+1))$, where $d$ is the spatial dimension and $L$ is the number of levels. Combined with the Schrödingerization-based Hamiltonian simulation, the overall quantum algorithm achieves a query complexity of $\mathcal{O}\big(\mathrm{poly}(d)\varepsilon^{-1} \mathrm{polylog}(\varepsilon^{-1}) \big)$ for estimating linear functionals of the solution to a given tolerance $\varepsilon$.

math.NA

HandwritingAgent: Language-Driven Handwriting Synthesis in Scalable Vector Space

Teaching machines to emulate natural handwriting styles remains an open challenge, as it requires synthesizing stroke sequences that dynamically vary in shape, texture, pressure and script - not only across individuals, but also within a single person's handwriting. Attempts at this challenge have largely explored deep learning methods in both online and offline settings. However, these approaches are often constrained by style-specific architectural choices, heavy reliance on large datasets, high compute costs, and a lack of flexible control over writing styles through natural language. To this end, we introduce HandwritingAgent, a language-driven agent that can synthesize natural handwriting sequences directly in Scalable Vector Graphics (SVG) format with no need for style-specific training. The agent leverages a large reasoning model to geometrically analyse and autoregressively generate target handwritten glyphs as stroke sequences in a discrete grid canvas environment. Generation is conditioned on texts provided in either conversational or non-conversational mode, along with a reference handwriting-style image. Experiments on diverse handwriting tasks spanning imitation, recognition, multi-lingual handwriting synthesis, and generation of complex handwritten maths and science expressions indicate substantial improvement in performance, with HandwritingAgent matching or surpassing state-of-the-art generative handwriting models, while providing a more efficient, controllable, and generalizable synthesis method.

cs.CV

Field Demonstration of a Multi-User Continuous-Variable Quantum Access Network for Quantum-to-the-Home

Realizing scalable Quantum-to-the-Home (QTTH) faces a bottleneck: link asymmetry in broadcast continuous-variable quantum access networks (CV-QANs) hinders the selection of a globally optimal modulation variance. We demonstrate a downstream broadcast CV-QAN connecting a Quantum Line Terminal (QLT) to multiple Quantum Network Units (QNUs) over commercial fiber. Operating within a trusted local network domain, we establish a multi-user utility model to select the optimal shared variance, balancing network efficiency and user fairness. Supported by robust digital signal processing, our 1:16 field trial achieves Mbit/s-level asymptotic secure key rates, bridging theoretical protocols with Fiber-to-the-Home reality and guiding future scalable access architectures.

quant-ph

Rational Sparse Autoencoder

Sparse autoencoders (SAEs) are standard tools for mechanistic interpretability, but current SAE families are constrained by fixed encoder nonlinearities such as ReLU, JumpReLU, and TopK. This hard-codes a particular sparsity mechanism into the model and can distort the reconstruction-versus-sparsity trade-off. We introduce the Rational Sparse Autoencoder (RSAE), which replaces the fixed encoder activation with a trainable rational function. Rational activations are flexible enough to uniformly approximate the activation primitives used by existing SAE families on compact domains (for TopK, the thresholded gate obtained after a separating top-k threshold is supplied), while also providing a richer function class for adapting to the observed pre-activation geometry. We realise this idea through a two-stage pipeline: an initialisation procedure that copies the pre-trained baseline SAE weights, plugs in rational coefficients obtained by the relaxed Remez exchange on synthetic data, and calibrates the scale parameters along with the rational coefficients; followed by a fine-tuning step under the standard sparsity-regularised reconstruction objective. Empirically, on residual-stream activations of three open-weight language models and across all three baseline activation families, the RSAE strictly improves on it after the fine-tuning step, both on reconstruction-side metrics and on downstream-behaviour metrics, without sacrificing feature-level interpretability under sparse probing. These gains are consistent across host language models, across baseline activation families, and across the full range of baseline sparsity we tested, while the upgrade itself adds only a handful of scalar parameters per autoencoder and runs in minutes on a single consumer GPU.

cs.LG