Searcharxiv⌕ Search

arXiv subjects

Yan Li

Publications and source records attributed to Yan Li.

At least 37 records · Page 2Linked to original sources

OccAnyScene: Towards Unified Indoor-Outdoor 3D Occupancy Prediction

3D occupancy prediction is fundamental to scene understanding, yet existing 3D semantic occupancy methods are typically specialized to fixed scene types and occupancy protocols. We introduce Cross-Scene 3D Semantic Occupancy Prediction, a new task setting which requires a single model to handle heterogeneous indoor and outdoor scenes with varying cameras, spatial ranges, voxel specifications, and semantic taxonomies. This setting poses a fundamental challenge: achieving metric-consistent yet scene-adaptive image-to-3D lifting across varying camera configurations and scene scales. To address this challenge, we propose OccAnyScene, a pixel-frustum-centered Gaussian framework built upon a pretrained depth foundation model. Specifically, the framework employs Pixel-Aligned Frustum Feature Aggregation to construct a camera-aware frustum query for each feature pixel, and Frustum-Parameterized Gaussian Construction to decode each query into multiple Gaussians whose positions and sizes are constrained by the predicted pixel depth and corresponding frustum geometry. OccAnyScene sets new state-of-the-art results, achieving 59.92% mIoU on the indoor Occ-ScanNet and 23.06% mIoU on the outdoor SurroundOcc-nuScenes.

cs.CV↗

DisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and Editing

Recent image generation and editing models can produce visually appealing natural images, yet they remain unreliable when the target image is a knowledge-intensive diagram whose correctness depends on disciplinary concepts, symbolic structure, and precise spatial relations. We introduce DisciplineGen-1M, a million-scale multidisciplinary dataset that supports text-to-image generation and image editing. It contains 1.2M samples spanning mathematics, physics, chemistry, biology, geography, computer science, economics, history, music, and sports. To construct the dataset, we design a scalable framework that combines structured rendering, OCR-based editing, specialized programmatic synthesis, and large-scale text-to-image filtering. These pipelines produce captions, editing instructions, structured annotations, and paired images with controllable semantic differences. Building on DisciplineGen-1M, we further introduce a discipline-informed reasoning-generation model for both text-to-image generation and image editing. Experiments on discipline-related benchmarks, GenExam and GRADE, show substantial improvements over open-source baselines, while evaluations on general reasoning-informed benchmarks, WISE and RISE, further indicate broader transfer. The results suggest that large-scale structured academic visual data is a key ingredient for moving image generation from aesthetic plausibility toward verifiable knowledge-grounded visual creation. We will publicly release our dataset, model, and source code of the data curation pipeline to ensure reproducibility and benefit future research.

cs.CV↗

Probing Triton's Space Environment and Internal Structure: An Integrated Detection-and-Interpretation Framework

Triton, Neptune's largest moon, is a prime ocean-world target. Constraining ocean thickness, composition, and conductivity is essential for habitability assessment, but magnetic induction alone cannot resolve the thickness-conductivity degeneracy, and magnetic perturbations from Triton's space currents can obscure the internal induction signal. We present an integrated detection-and-interpretation concept linking four physically consistent calculations. Using `PlanetProfile', we construct a common radial interior structure (temperature, density, conductivity, seismic-wave speed). We then use `MoonMag' to compute the degree-one magnetic-induction response from that conductivity profile at the synodic, rotational, and orbital periods. We perform a multi-fluid `SWMF' simulation with the induced dipole as the inner-boundary condition and develop a Coulomb-gauge Poisson reconstruction to isolate space-current magnetic fields. Finally, we develop the `TritonSeis' workflow, three-dimensional seismic forward modeling plus hierarchical travel-time inversion, to constrain the ice-ocean and ocean-rock interface depths. We find that induction is substantially more sensitive to ocean conductivity than to layer thickness, and that space-current fields are comparable in amplitude to the internal induction signal. A five-station synthetic recovery test resolves both interfaces to first order, with errors of +8.4% for the ice shell and -12.5% for the ocean. Under a conservative noise assumption, the minimum detectable magnitudes are approximately 3.8-4.6 at epicentral distances of 100-1000 km. The Poisson reconstruction and end-to-end seismic recovery are, to our knowledge, the first such quantitative demonstrations for Triton. Coordinated magnetic, plasma, and seismic measurements are complementary and can break the conductivity-thickness degeneracy, providing a framework for future Triton exploration.

astro-ph.EP↗

Initialization Is Critical: Advancing Federated Short-Term Load Forecasting under Load Heterogeneity via Model Initialization

Short-term load forecasting (STLF) provides essential information for numerous applications in modern power systems. However, accurate STLF often relies on fine-grained smart-meter data from distributed users, raising increasing concerns about data privacy. Federated learning (FL) has therefore emerged as a promising privacy-preserving paradigm for STLF. Nevertheless, this paper reveals structured heterogeneity in clients' load data. Specifically, clients exhibit different responses to exogenous factors and distinct temporal load profiles, which can degrade forecasting performance in FL. To mitigate these issues, this paper studies the role of model initialization in federated STLF, and proposes two initialization strategies from global and local perspectives. For global model initialization, when auxiliary public load data are available, a pretrained initialization strategy is developed to initialize the global model before federated training, thereby reducing client drift during the training process. For local model initialization, we propose SLIAvg, a sequential local initialization strategy that promotes a more consistent training process by allowing participating clients to start from progressively adapted models within each communication round. Since the proposed strategies only modify the initialization process, they are compatible with most existing FL frameworks and privacy-enhancing techniques. Experiments on real smart-meter data with two representative forecasting architectures demonstrate that the proposed strategies effectively improve forecasting performance, as evidenced by reduced client drift, improved convergence behavior, and lower forecasting errors.

cs.LG↗

TrapVLA: Trapping Vision-Language-Action Models in Configured Failure Modes

This work introduces Configured Failure Trapping, a novel backdoor attack task against Vision-Language-Action (VLA) models, which aims to activate attacks through stealthy textual triggers and induce configured failure modes. Unlike prior backdoor attacks that treat any task failure as a successful attack, Configured Failure Trapping requires the attacker to control how the robot fails (e.g., causing the robot to grasp with a specified positional offset), making it substantially more challenging and hard to detect. To support the new task, we propose an effective data engine for synthesizing high-quality target trajectories and an automated suite for measuring configured-failure fidelity. Then, based on this foundation, we construct two new benchmarks, namely Trap-LIBERO and Trap-RoboTwin, that instantiate Configured Failure Trapping across four representative failure modes. To address this task, we identify sparse action deviation as a critical challenge and accordingly propose a novel method named TrapVLA, which explicitly learns trigger-induced action residuals to steer the policy toward the configured failure behavior. Extensive experiments across simulation benchmarks and real-world robotic settings show that TrapVLA effectively injects configured failure modes into VLA models while largely preserving performance on clean data. Project page: https://john-liua.github.io/TrapVLA/

cs.RO↗

Competing Auger-Like and Intraband Excitations Drive Anomalous Modulations in Indium Tin Oxide

Transparent conducting oxides near their epsilon-near-zero frequency exhibit near-unity ultrafast modulations of the refractive index which have enabled the field of time-varying metamaterials, yet the underlying carrier dynamics at high driving fluences remain poorly understood. Here, we report ultrafast modulations in the reflectivity and transmissivity of indium tin oxide, and a non-monotonic oscillatory behavior. This is especially evident in the time evolution of the complex Fresnel coefficients retrieved directly from pump-probe spectrograms using an optical gating technique, GRUMPY FROG. The dynamics of the retrieved plasma frequency and damping coefficient are well captured by an extended two-temperature model incorporating a competing nonlinear interband process: at high fluences, Auger-type scattering of hot conduction electrons promotes valence band carriers, increasing the plasma frequency while accelerating hot-electron cooling and raising the damping coefficient. These results clarify the origin of anomalous high-fluence dynamics in indium tin oxide and identify a fluence-tuneable modulation dynamic with direct implications for ultrafast refractive index engineering in time-varying photonic devices and optical switching.

physics.optics↗

Is Reasoning Always Useful? Rethinking Reasoning Utility in Universal Multimodal Embeddings

Reasoning-enhanced universal multimodal embeddings (UME) improve heterogeneous retrieval, but plausible rationales do not necessarily produce discriminative rankings. We study this gap by comparing the discriminative (DISC) and reasoning-driven generative (GEN) branches of UME-R1, a state-of-the-art reasoning UME method. We decompose reasoning utility into positive-target gain, hard-negative gain, and their margin difference. Positive similarity increases for 56.6%, but 15.7% are false-helpful cases where reasoning moves hard negatives closer even more. Local-neighborhood and token-attribution diagnostics suggest why: reasoning often de-condenses retrieved neighborhoods, but utility requires separator-aligned movement, while influential CoT tokens frequently encode evidence shared by positives and hard negatives. Motivated by these diagnostics, we propose SURE (Score-structure Utility Router for Embeddings), which improves UME-R1-7B by 1.5 points and yields consistent gains on two additional embedding models on MMEB-V2, without retraining, label-based policy selection, or extra VLM forward passes.

cs.AI↗

CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing

The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets mainly focus on single editing operations and fall short in supporting compositional instruction-guided video editing. In particular, multiple editing intents must be jointly understood and faithfully executed within the same video. To address this issue, we introduce CoinVE-200K, a large-scale, high-quality dataset for Compositional Instruction-Guided Video Editing. CoinVE-200K contains 1080p video-editing pairs of up to 201 frames, covering diverse compositional scenarios where each sample involves 2 to 5 atomic editing operations. The instructions target humans, objects, and backgrounds, and cover edit types such as addition, removal, modification, and stylization. All samples are built through a carefully designed generation and filtering pipeline to ensure instruction faithfulness, visual quality, temporal consistency, and compositional diversity. We also introduce CoinVE-Bench, a benchmark for compositional-instruction video editing across diverse subjects, operation types, and instruction complexities. Furthermore, we present CoinVE-Edit, a 22B compositional video editing model built upon Wan2.1-T2V-14B and Qwen3-VL-8B-Instruct. CoinVE-Edit disentangles region-aware attention for different editing instructions, enabling precise multi-region editing while preserving irrelevant content and temporal coherence. Experiments on CoinVE-Bench show that CoinVE-Edit achieves strong performance in instruction following, compositional editing accuracy, visual quality, and temporal consistency.

cs.CV↗

Multi-Partitioned Computing Quantum-Particle Approach: A Hybrid Quantum Framework for Fluid Flow

This study established a quantum-classical hybrid framework that integrates quantum computing paradigm with meshfree finite particle method. By harnessing quantum superposition and entanglement, it hybridized the critical computational kernels (termed as quantum finite particle method). A resource-efficient quantum computational strategy on multi-partitioned zones was proposed, which leverages a fixed small-scale quantum circuit as a fundamental processing unit to handle inner product for arbitrarily sized arrays. This approach employs iterative nesting of the quantum-core operation to accommodate varying input dimensions while maintaining hardware feasibility throughout. Motivated with developed quantum framework, the novel numerical discretization for hybrid quantum computational particle dynamics can be derived commonly and applied in fluid flows. Through a sequence of numerical experiments purposefully, the proposed numerical model was thoroughly validated and analyzed. Results demonstrate that integrating quantum computing to hybridize conventional linear combinations of particle dynamics serves as a novel computing paradigm. By further extending into the numerical investigation of viscoelastic, highly elastic, and purely elastic fluids under high Weissenberg number conditions, the applicability of simulation framework is broadened. Despite bottlenecks in quantum hardware and computational efficiency on this process, these advances offer critical insights for transitioning quantum-enhanced fluid simulation to practical engineering applications.

physics.flu-dyn↗

Closed-loop AI achieves certifiable engineering design

Agentic AI has automated parts of scientific discovery, including paper generation, expert-level coding, therapeutic proposal, and autonomous experimentation. Complex physical engineering design remains a gap, because candidates must satisfy simultaneous constraints in fluid dynamics, solid mechanics, and structural stability. We introduce The AI Engineer, an agentic framework that couples large language models (LLMs) to deterministic engineering backends in a closed loop: natural-language requirements are converted into design-domain geometry and mesh; topology is optimized with bi-directional evolutionary structural optimization (BESO) coupled to the CalculiX solver; and member sizes are refined with particle swarm optimization (PSO) coupled to Zwind under offshore aero-hydro-servo-elastic load cases. To explore many designs without per-candidate certification cost, an Automated Reviewer scores each candidate on five dimensions (capacity, steel intensity, unit cost, constructability, and fatigue life) using piecewise-linear functions calibrated on 11 real floating-wind projects. Search terminates only when a candidate reaches a composite score $S \ge 85$ (grade A) with no subscore below 60. We validated this gate by submitting the top-scoring design to the China Classification Society (CCS) for Approval in Principle (AIP), which it passed; AIP is thus an external check that the reviewer tracks professional judgment, not the daily objective. The certified design outperforms the human-optimized TuQiang baseline, reducing steel mass and unit capital cost by 8.1% each while meeting all AIP criteria. This verification-closed regime, in which every proposal is judged by deterministic physics and codified limit states, distinguishes The AI Engineer from open-ended generative systems. Remaining limits include detailed design and fabrication-hard constraints.

cs.AI↗

A Provable Oracle-Free Quantum Algorithm for Nonlinear Dynamics on Hybrid Oscillator-Qubit Processors

We develop a hybrid qubit--qumode algorithm for nonlinear ordinary differential equations of the form $\dot{\mathbf{x}}=\mathbf{f}(\mathbf{x})$ with drift of polynomial degree~$L$. Following the Fokker--Planck route of Tennie and Magri, the algorithm propagates the state density and returns the deterministic trajectory as the peak of that density in the small-noise limit. The discretised generator is carried into a parametrised family of Schrödinger equations by the warped-phase transformation of Jin, Liu, and Yu, and the Fourier-mode parameter of that family is placed on a single continuous-variable qumode. Our central structural result is that the Hermitian parts $H_{1}$ and $H_{2}$ of the discretised generator admit a bipartite Pauli decomposition that sorts the non-zero Pauli strings into $\mathcal{O}(\log N)$ mutually commuting families and factorises each family into a diagonal of degree at most $L$ tensored with a fixed rank-two bond operator. The factorisation renders each family exponential an exact product of $\mathcal{O}(n^{L})$ monomial-controlled momentum displacements, with no intra-family Trotter error. On a $d$-dimensional grid of $N=2^{n}$ points per axis the circuit costs $\mathcal{O}(d^{L+1}n^{L+2})$ gates per Trotter step. No sparse-access oracle and no block encoding is invoked: every gate is fixed in closed form by the polynomial coefficients of the drift. We also prove a bound on the numerical abscissa $λ_{\max}(H_{1})$ that fixes the recovery domain of the warped-phase transform and the post-selection cost. A classical simulation on two nonlinear benchmarks confirms the structural theorems, the shifted recovery, and the accuracy-per-resource advantage of the continuous-variable coupling over a discretised mode register.

quant-ph↗

SPLIT-Q: A Scalable Sequential Quantum Computing Framework for Coherent Controlled Islanding

Growing integration of distributed energy resources increases power-system variability and uncertainty. During disturbances, these effects can intensify generation-load imbalances and cascading failures. Controlled islanding limits their propagation by partitioning a compromised grid into connected, electrically sustainable islands. However, classical methods face rapidly growing computational costs as network size and island count increase. Quantum optimization offers an alternative for exploring this combinatorial partition space. Yet monolithic quantum formulations encode all assignment decisions in one circuit, causing qubit demand and circuit complexity to scale with network size. In this study, a qubit-bounded sequential distributed quantum approximate optimization algorithm (QAOA) framework is proposed to tackle coherent controlled islanding under limited quantum resources. It formulates the optimization as boundary-conditioned regional quadratic unconstrained binary optimization (QUBO) subproblems that are solved sequentially within a fixed qubit budget. Thus, circuit width remains independent of network size, with aggregate quantum workload scaling linearly on bounded-degree networks. Evaluation covers eleven IEEE systems from 9 to 300 buses using IBM quantum computing resources, with Gurobi and monolithic QAOA as references. Across all systems, the framework recovers feasible Gurobi-optimal partitions under noise, confirming the resilience of its solution quality. The results further show that the proposed method substantially reduces quantum-resource demand and circuit complexity relative to monolithic QAOA, allowing large islanding problems to be addressed within current hardware limits. The proposed framework provides a feasible and scalable pathway for quantum optimization in large-scale power systems.

quant-ph↗

Complete Rigidity at infinity and Existence of the Levinson Cavity

We present a potential theoretic approach reducing the analysis of the asymptotic shape of free surfaces to the analysis of a precise ordinary differential equation resulting from the reduction process. Although the approach relies mainly on the principal part of the PDE operator to allow for a representation formula and is thus not restricted to problems of elliptic type, we present it at the clean-cut example of three-dimensional axially symmetric steady incompressible cavity flows, which are Neumann-type Bernoulli free boundary problems and for which frequency formulas are unknown and, if they do exist, insufficient to yield the very precise asymptotic behavior we prove here. In 1946 Norman Levinson derived by a power-law ansatz with a slowly varying correction a precise formula for the asymptotic shape of such cavities. However his result requires very strong assumptions such that it has remained an open problem for 80 years whether the cavity solutions we know to exist by a result by Garabedian-Lewy-Schiffer [12] actually share this asymptotic behavior, or whether at least one solution possessing the Levinson asymptotics exists. Here we answer both questions affirmatively, and we obtain complete rigidity at infinity of the Levinson solution in the class of axially symmetric solutions, that is, any solution satisfying mild and natural assumptions at the fixed boundary and infinity converges asymptotically to the Levinson profile $(\log r)^{-1/4}\sqrt{r}$.

math.AP↗

Q-MAP: Multi-Platform Benchmarking of Distributed Quantum Computing for Coherent Controlled Islanding

The integration of distributed energy resources into power networks is accelerating. The resulting variability narrows operating margins, so a disturbance can cascade into a wide-area blackout. Controlled islanding arrests that propagation by splitting a compromised grid into self-sustaining islands that keep coherent generators together. Exact classical solutions become intractable as the bus count and island number grow. Gate-based quantum optimization provides a different route through this combinatorial space, although its reach is limited when one circuit carries every bus assignment, since qubit count and depth then follow grid size. In this study, a round-synchronous distributed quantum computing framework is developed for coherent controlled islanding under a fixed per-circuit qubit budget. Every round derives all regional subproblems from one frozen grid-wide snapshot, dispatches them to independent quantum backends at the same time, and merges the returned candidates classically into one globally evaluated update. Circuits executed in parallel therefore keep a constant size as the grid grows, and a round costs the slowest region rather than the sum of all of them. Benchmarking spans IEEE systems from 9 to 300 buses on simulation and on quantum processors of different architectures. The framework attains optimal and operationally feasible partitions on every platform under noise, even where compilation cost differs by nearly an order of magnitude. Bounding width in this way places grids beyond the reach of monolithic circuits within range of present devices and establishes a multi-backend baseline for quantum computing in large-scale power-system optimization.

quant-ph↗

Adaptable Fingerprinting with Nonlinear Shrinkage for Climate Change Detection and Attribution under Variance Heterogeneity

Detection and attribution of climate change relies on fingerprinting--a linear errors-in-variables regression framework in which both predictors and responses exhibit internal variability governed by a proportional covariance structure, subject to a variability inflation factor. Accurate estimation of the scaling factors (regression coefficients) depends on inferring the precision matrix of the regression errors from limited climate model control runs. In high-dimensional settings, existing approaches often overlook the variance inflation of the predictors and suffer from imprecise precision matrix estimates, yielding biased estimators, underestimated uncertainties, and confidence intervals with poor coverage. We propose a nonlinear, rotation-invariant shrinkage framework for estimating the precision matrix that restores the asymptotic optimality of the total least squares estimator in high-dimensional regimes. Our procedure jointly estimates the scaling factors and the variability inflation factor, thereby correcting estimation bias, and incorporates consistent variance estimators to enable valid uncertainty quantification. We also develop a residual consistency test to assess model adequacy. Numerical studies demonstrate precise estimation, improved confidence interval coverage, and higher efficiency. Applied to annual mean near-surface air temperature data from 1951--2020, our method produces narrower and more reliable confidence intervals, yielding refined attribution results.

stat.ME↗

Theoretical analysis towards accurate optomechanical detection of quantum gravity effects

Optomechanical systems offer a promising platform for observing dynamical signatures of quantum gravity through precision measurements of quantum harmonic oscillator dynamics. However, most existing analyses consider only the linear radiation-pressure interaction while neglecting higher-order optomechanical couplings and laser phase noise. These neglected contributions can be comparable in magnitude to the predicted quantum-gravity corrections and may therefore introduce spurious signals or mask the genuine physical effect. Here we reanalyze two experimentally realized platforms, a Fabry-Perot optomechanical system and a membrane-in-the-middle optomechanical system, by incorporating the complete nonlinear dynamics and realistic laser phase noise. Using measured device parameters, we derive revised protocols for generalized uncertainty principle tests and establish practical sensitivity bounds. Our results demonstrate that previous idealized estimates significantly overestimate the achievable resolution, underscoring the necessity of including higher-order interactions and implementing effective laser phase noise suppression in realistic assessments of optomechanical quantum gravity tests.

quant-ph↗

AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models

Agentic multimodal large language models (MLLMs) extend multimodal perception and reasoning with planning, tool use, and interaction in dynamic environments. Yet current models are specialized for particular tools or environments, complicating consolidation into a single generalist. We formulate Agentic MLLM Merging and identify two challenges: asymmetric capability preservation, whereby capabilities with different interaction complexity are retained unevenly, producing weak tasks after merging, and behavior-critical forgetting, whereby losing decisive actions can derail long-horizon execution. We propose AgentPatch, a training-free coarse-to-fine repair framework. It selects a stable merged backbone, restores diluted weak-task-specific signals through Weak-Task Unique Residual Recovery, and applies an Agent-Guided Behavior-Critical Patch that recovers decisive behaviors under explicit capability protection. AgentPatch produces a single static checkpoint without routing or ensembles. Experiments across six agentic and multimodal benchmarks show that AgentPatch improves diverse merged backbones, alleviates weak-task degradation, and better balances weak-task recovery with the preservation of complementary search and agentic visual processing capabilities. Code is available at https://github.com/ziboshao/AgentPatch.

cs.AI↗

Pulse-Duration Control of Subcycle Multiband Electron Dynamics Extends the High-Harmonic Cutoff in a Light-Driven Insulator

We demonstrate pathway-selective control of extreme-ultraviolet high-harmonic generation by jointly tuning laser pulse duration ($5$ - $29$ fs) and intensity ($0.8$ - $74$ TW/cm$^2$). Many-cycle pulses at moderate intensities, $\sim 6$ TW/cm$^2$, promote cumulative carrier transfer over successive optical cycles, progressively accessing higher conduction bands. In contrast, few-cycle, high-intensity, $\sim 22$ TW/cm$^2$, pulses drive subcycle multiband dynamics that reach $25$ - $50$ eV photon energies before decoherence can suppress coherent emission. These results reveal pulse duration and intensity as decisive control knobs for high-harmonic emission, opening a route to band-structure-guided pulse design for higher energy extreme-ultraviolet light sources.

physics.optics↗