Searcharxiv⌕ Search

arXiv subjects

Yan Li

Publications and source records attributed to Yan Li.

At least 55 records · Page 3Linked to original sources

Theoretical analysis towards accurate optomechanical detection of quantum gravity effects

Optomechanical systems offer a promising platform for observing dynamical signatures of quantum gravity through precision measurements of quantum harmonic oscillator dynamics. However, most existing analyses consider only the linear radiation-pressure interaction while neglecting higher-order optomechanical couplings and laser phase noise. These neglected contributions can be comparable in magnitude to the predicted quantum-gravity corrections and may therefore introduce spurious signals or mask the genuine physical effect. Here we reanalyze two experimentally realized platforms, a Fabry-Perot optomechanical system and a membrane-in-the-middle optomechanical system, by incorporating the complete nonlinear dynamics and realistic laser phase noise. Using measured device parameters, we derive revised protocols for generalized uncertainty principle tests and establish practical sensitivity bounds. Our results demonstrate that previous idealized estimates significantly overestimate the achievable resolution, underscoring the necessity of including higher-order interactions and implementing effective laser phase noise suppression in realistic assessments of optomechanical quantum gravity tests.

quant-ph↗

AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models

Agentic multimodal large language models (MLLMs) extend multimodal perception and reasoning with planning, tool use, and interaction in dynamic environments. Yet current models are specialized for particular tools or environments, complicating consolidation into a single generalist. We formulate Agentic MLLM Merging and identify two challenges: asymmetric capability preservation, whereby capabilities with different interaction complexity are retained unevenly, producing weak tasks after merging, and behavior-critical forgetting, whereby losing decisive actions can derail long-horizon execution. We propose AgentPatch, a training-free coarse-to-fine repair framework. It selects a stable merged backbone, restores diluted weak-task-specific signals through Weak-Task Unique Residual Recovery, and applies an Agent-Guided Behavior-Critical Patch that recovers decisive behaviors under explicit capability protection. AgentPatch produces a single static checkpoint without routing or ensembles. Experiments across six agentic and multimodal benchmarks show that AgentPatch improves diverse merged backbones, alleviates weak-task degradation, and better balances weak-task recovery with the preservation of complementary search and agentic visual processing capabilities. Code is available at https://github.com/ziboshao/AgentPatch.

cs.AI↗

Pulse-Duration Control of Subcycle Multiband Electron Dynamics Extends the High-Harmonic Cutoff in a Light-Driven Insulator

We demonstrate pathway-selective control of extreme-ultraviolet high-harmonic generation by jointly tuning laser pulse duration ($5$ - $29$ fs) and intensity ($0.8$ - $74$ TW/cm$^2$). Many-cycle pulses at moderate intensities, $\sim 6$ TW/cm$^2$, promote cumulative carrier transfer over successive optical cycles, progressively accessing higher conduction bands. In contrast, few-cycle, high-intensity, $\sim 22$ TW/cm$^2$, pulses drive subcycle multiband dynamics that reach $25$ - $50$ eV photon energies before decoherence can suppress coherent emission. These results reveal pulse duration and intensity as decisive control knobs for high-harmonic emission, opening a route to band-structure-guided pulse design for higher energy extreme-ultraviolet light sources.

physics.optics↗

Match One, Learn with Graph: One-to-Graph Query Collaboration with Backward Sharing for Object Detection

One-to-one (O2O) matching enables Detection Transformers (DETRs) to perform end-to-end set prediction by assigning each object to a single positive query. However, the strongest classification, center, scale, and overlap evidence for an object is often distributed across multiple queries. This mismatch leaves only the matched owner positively supervised for the object, while other evidence-bearing queries receive no box target for it. We term this query knowledge fragmentation. To exploit such complementary evidence without one-to-many supervision, we propose BS-O2G, a plug-in that builds a sparse prediction-aware graph from decoded features, boxes, and class distributions to organize query collaboration in feature and optimization spaces while preserving the original O2O matcher, positive labels, and objective. One-to-Graph (O2G) calibration propagates relative messages over this graph to consolidate query evidence in the forward pass, whereas Backward Sharing (BS) reuses its transposed detached adjacency to route gradients across persistent query basis vectors without changing the decoder input in the forward pass. Experiments across diverse DETR methods, backbones, COCO, and CrowdHuman show consistent gains and faster convergence with negligible parameter/FLOP growth and modest runtime overhead, supporting graph-based query collaboration as an alternative to expanding positive assignments.

cs.CV↗

Recompute or Reuse? Diagnosing and Mitigating Textual Shortcuts in VLM Self-Reflection

Vision-language models (VLMs) are expected to revise their reasoning when visual evidence changes. Failures to do so are often attributed to insufficient visual attention or contextual inertia, leaving unclear what models reuse instead of recomputing from the current image. We show that evidence-bearing reasoning in a prior chain of thought (CoT) can form a textual shortcut that competes behaviorally with visual recomputation. Across 16 VLMs, a matched counterfactual analysis identifies evidence-bearing content as the most robust carrier of prior-CoT influence. Removing this evidence-bearing content shifts answer preference more than removing length-matched non-evidence context or the final-answer span, with prior control weakening progressively as more stale evidence is removed. Reordering this evidence also weakens prior control, showing that its organization modulates shortcut strength. Beyond the immediate answer, the shortcut can retain residual influence after answer correction: weakening current-image support shifts preference back toward the prior answer, while repeated prior answers and reused premises arise mainly when the shortcut remains active. To limit this influence, we introduce Fresh-State Attention Firewall (FSAF), a training-free intervention that isolates fresh computation from the prior CoT. Across five VLMs, FSAF raises visual update rate from 35.28% to 53.61% and reduces prior-answer rate from 39.22% to 3.67%. Reliable VLM self-reflection therefore requires more than looking again: fresh visual recomputation must be protected from stale textual reuse.

cs.CV↗

Quasi-polar Decomposition of Quantum Neural Networks via Adaptive Non-local Observables

We use Diagonal Adaptive Non-local Observables (DANO) as a canonical decomposition for studying Variational Quantum Circuit model evolution. Separating each learned observable into a diagonal spectrum and a unitary basis gives a quasi-polar description: the spectral weights are viewed as radial coordinates, while the unitary circuit serves as angular coordinates through Lie group identifications. This turns the training process into a trajectory in spectral and Lie-algebra space. Experiments on two classification tasks show that DANO radial spectral expansion correlates with accuracy. DANO angle coordinates reveal a dominant accuracy-correlated component. The framework provides a different perspective to characterize quantum model behavior.

quant-ph↗

A universal framework for nonlinear frequency combs under electro-optic modulation

Nonlinear frequency combs, including electro-optic and Kerr combs, have become central platforms for chip-scale frequency synthesis. Recent breakthroughs in strong-coupling electro-optic modulation further expanded their accessible nonlinear dynamics, unlocking new phenomena and functionalities, but the underlying foundation remains largely unexplored. Here we establish a universal theoretical and experimental framework for nonlinear combs under arbitrary electro-optic modulation by introducing a general evolution equation (GEE) that transcends the mean-field Lugiato-Lefever equation. The GEE reduces to a discrete-time Integration Hamiltonian that provides a frequency-domain formalism unifying strong-coupling electro-optic modulation with photonic synthetic dimensions. Together with a band-wave correspondence linking modulation waveforms to synthetic band structures, the formalism enables programmable spectral control. We further show compatibility between Kerr nonlinearity and strong-coupling electro-optic modulation, highlighting their cooperative dynamics. Our work provides a foundational model for strong-coupling electro-optics in nonlinear combs, opening a route toward chip-integrated, microwave-programmable comb sources for metrology, spectroscopy, and emerging photonic technologies.

physics.optics↗

Spectral Analysis of the Schrödinger Operator for the Incommensurate System

Many novel and unique physical phenomena in incommensurate systems can be illustrated and predicted using their spectral structure and electronic state distributions. However, the absence of periodicity in these systems poses significant challenges for obtaining the associated information. In this paper, by embedding the system into higher dimensions together with introducing a regularization technique, we prove that the spectrum of the Schrödinger operator for the incommensurate system can be approximated by the spectra of a family of regularized Schrödinger operators, which are elliptic, retain periodicity, and enjoy favorable analytic and spectral properties. We also show the well-posedness of the probability density describing the electronic state distribution of the incommensurate system, which can be approximated by the ones generated by the Bloch solutions to the regularized model. Our analysis provides theoretical support for understanding and computing incommensurate systems.

math-ph↗

Spin and momentum fraction carried by partons in the nucleon

We determine the momentum fraction and angular momentum carried by quarks and gluons in the proton in lattice QCD. We use four ensembles simulated with up, down, strange and charm quarks with their masses tuned to their physical values. These ensembles have similar physical volume and different lattice spacings allowing us to take the continuum limit directly at the physical pion mass point. We extract the quark and gluon momentum fractions and total angular momentum in the continuum limit as well as the intrinsic quark spin and orbital angular momentum contributions to the proton spin. We find the total momentum fraction $\langle x_N \rangle= 0.995(60)(29)$ and the total spin $J_N = 0.507(43)(65)$, showing that both the momentum and spin sum rules are satisfied. We compare our results to those extracted from phenomenological analyses.

hep-lat↗

SlerpFlow: Spherical Trajectory Correction for Rectified Flow Inversion

Rectified-flow-based diffusion transformers, particularly FLUX, have demonstrated outstanding performance in high-quality image generation. However, achieving fast and accurate inversion--transforming images back to latent noise for faithful reconstruction and editing--remains a challenging bottleneck due to the discretization errors of linear solvers. This paper introduces SlerpFlow, a straightforward yet highly effective zero-shot approach that unlocks the full potential of FLUX for high-fidelity inversion and editing. Unlike existing approaches (e.g., RF-Solver) that rely on complex numerical approximations such as high-order Taylor expansions to correct trajectory errors, we present a geometric view based on the Manifold Hypothesis: the empirically observed trajectory curvature is not a numerical artifact, but rather serves as a necessary "centripetal force" that constrains the flow to remain on the data manifold. Guided by this insight, SlerpFlow integrates Spherical Linear Interpolation (Slerp) to rectify flow velocity directions on the hypersphere, strictly adhering to the intrinsic curvature of the latent space. Crucially, by caching the corrected velocity for subsequent steps, SlerpFlow achieves high-precision inversion while maintaining the computational efficiency of a first-order Euler solver. Extensive experiments on FLUX-based reconstruction and editing tasks demonstrate that SlerpFlow improves reconstruction fidelity and achieves stronger semantic alignment in editing without requiring additional training. Code is available at https://github.com/0answer0/SlerpFlow.

cs.CV↗

Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On

We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-purpose image editor, Oxygen-TryOn is fashion-native, built for try-on through a dedicated data engine and try-on-specific training. Given one or more reference items (clean product shots or in-the-wild worn-on photos) and a single target subject image, it synthesizes a photorealistic image of the subject wearing the items across virtually any fashion category. Prior systems handle a single garment category in a studio setting, and recent multi-reference methods remain garment-centric; in contrast, Oxygen-TryOn supports diverse items and scenarios, including full- and half-body views, a variable number of references, and free multi-item composition, while faithfully preserving both subject identity and item appearance. Instead of mask-based inpainting, we reformulate try-on as a multi-reference, understanding-driven generation task. We build a data engine that collects, manufactures, annotates, and filters high-quality try-on data at scale, and design a three-stage recipe of continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL). The RL stage uses a hybrid reward combining an in-house try-on reward model with a proprietary, rubric-guided general-purpose model, jointly supervising fine-grained consistency and instruction-level quality. It also follows general editing instructions (e.g., pose changes) in the same pass. Across public benchmarks and our in-house Oxygen-TryOn Bench, it achieves state-of-the-art consistency and realism on single-item try-on and leads on multi-item try-on, matching or surpassing both leading proprietary systems (Nano Banana Pro, GPT-Image-2, Seedream5 Lite) and open-source models (FLUX.2).

cs.CV↗

Efficient MPC-Based Energy Management System for Secure and Cost-Effective Microgrid Operations

Model predictive control (MPC)-based energy management systems (EMS) are essential for ensuring optimal, secure, and stable operation in microgrids with high penetrations of distributed energy resources. However, due to the high computational cost for the decision-making, the conventional MPC-based EMS typically adopts a simplified integrated-bus power balance model. While this simplification is effective for small networks, large-scale systems require a more detailed branch flow model to account for the increased impact of grid power losses and security constraints. This work proposes an efficient and reliable MPC-based EMS that incorporates power-loss effects and grid-security constraints. %, while adaptively shaping the battery power profile in response to online renewable inputs, achieving reduced operational costs. It enhances system reliability, reduces operational costs, and shows strong potential for online implementation due to its reduced computational effort. Specifically, a second-order cone program (SOCP) branch flow relaxation is integrated into the constraint set, yielding a convex formulation that guarantees globally optimal solutions with high computational efficiency. Owing to the radial topology of the microgrid, this relaxation is practically tight, ensuring equivalence to the original problem. Building on this foundation, an online demand response (DR) module is designed to further reduce the operation cost through peak shaving. To the best of our knowledge, no prior MPC-EMS framework has simultaneously modeled losses and security constraints while coordinating flexible loads within a unified architecture. The developed framework enables secure operation with effective peak shaving and reduced total cost. The effectiveness of the proposed method is validated on 10-bus, 18-bus, and 33-bus systems.

eess.SY↗

Nucleon unpolarized second Mellin moments using lattice QCD ensembles with physical quark masses and in the continuum limit

We compute the matrix elements of the energy-momentum tensor of the nucleon using four ensembles of twisted mass clover-improved fermions with the up, down, strange and charm quark masses tuned to approximately their physical values. The four ensembles have similar physical volume and lattice spacings $a=0.080$~fm, $0.068$~fm, $0.057$~fm, and $0.049$ fm, allowing us to take the continuum limit directly at the physical pion mass point. We compute both connected and disconnected quark contributions as well as gluon contributions. All renormalization functions, including the mixing of the quark singlet with the gluon, are determined non-perturbatively. We extract the gravitational form factors in the continuum limit at $Q^2=0$ and evaluate the contribution of quarks and gluons to the momentum and angular momentum of the proton. Using the values of the intrinsic quark spin computed using the same gauge ensembles we also determine the orbital angular momentum for each quark flavor.

hep-lat↗

Exploring the Small-scale Magnetic Fields in the Atmosphere of HD 49385 by Asteroseismic Analysis

Recent asteroseismic studies have shown convincing evidences that magnetic fields may exist in the interior of some pulsating red giants. Inspired by this breakthrough, we explored the effect of small-scale magnetic fields on the p-mode oscillations in an evolved star, HD 49385. {\bf We incorporate a modified Eddington $T$-$τ$ equation that phenomenologically mimics the effect of the magnetic fields in the atmosphere of HD 49385,} and calculate the frequencies of p-modes with $l=0$, 1, and 2. By comparing the calculated frequencies with the observed ones, we select two best-fit models with either GS98 or A09 chemical composition. Our best-fit models not only fit satisfactorily the observed frequencies, but also well reproduce some spectroscopically observed stellar parameters such as effective temperature and log\,$g$. Based on the two best-fit models, we have estimated that the small-scale magnetic fields possess a strength of approximately 80\,G and spread concentratively at approximately a height of 1850 km in the atmosphere. By selecting the best-fit models with special requirement on the avoided-crossing mode, we have confirmed that the frequency of the avoided-crossing mode is tightly related to the helium core of the star, and determined the size of the helium core as 0.117${\rm M}_\odot$ in mass and 0.078${\rm R}_\odot$ in radius. Based on the improvements of previous two sides, we can accurately determine the mass of HD 49385 to be $1.25\pm 0.02\,{\rm M}_\odot$ with an age of 4.1\,Gyr for GS98 composition and 4.5\,Gyr for A09 composition.

astro-ph.SR↗

Asteroseismic Analysis of a Red Giant KIC 9145955 by Including the Small-scale Magnetic Fields in the Atmosphere

Recent convincing evidence is found within asteroseismology that suggests the magnetic fields exist in three red giants. Research on small-scale magnetic fields in the Sun and HD 49385 has shown that they have a certain corrective effect on the systematic discrepancies between observed and theoretical frequencies. Here we apply a similar method applied for the Sun to a red giant, KIC 9145955, to explore the impact of small-scale magnetic fields in the photosphere on its frequencies. We find that the calculated frequencies of our best-fit model, which simulates the effect of the magnetic fields by artificially modifying the Eddington $T-τ$ relation, perfectly match those of the observed l = 0, 1, and 2 modes, indicating the existence of small-scale magnetic fields with an upper strength limit of 65 G and concentrating at a height 13,100 km in the photosphere. Based on the best-fit model, we revise the stellar parameters of KIC 9145955 as: $M = 1.23\pm0.04\,M_\odot$, $R = 5.57\pm0.06\,R_\odot$, $L = 19.85\pm0.5\,L_\odot$, $Age = 3.83\pm0.5$\,Gyr, $M_{\rm He} = 0.2108\pm0.0005M_\odot$, and $R_{\rm He} = 0.0306\pm0.0001R_\odot$.

astro-ph.SR↗

Giant third-order polarization rotation via wave-mixing-induced symmetry breaking in a Rydberg-EIT medium

We investigate how wave-mixing (WM)-induced symmetry breaking leads to giant third-order polarization rotation of a weak probe field in a Rydberg electromagnetically induced transparency (EIT) medium. A far-detuned counterpropagating WM field is adiabatically eliminated and retained solely as a Raman dressing of the lower Zeeman manifold. In this reduced description, the weak static magnetic field defines the two circular propagation channels, while the WM-induced Raman coherence breaks the symmetry between these channels, without acting as a gain channel or an independent nonlinear source. The weak-probe response is calculated using a reduced density-matrix expansion for van der Waals (vdW) correlations and self-consistent Maxwell-Bloch propagation, with the nonlinear rotation extracted by subtracting the linear propagation background. Including WM dressing increases the extracted third-order rotation from 1.06 degrees to 25.70 degrees, an enhancement of more than 24 times, for the parameters considered. The response is nonmonotonic in WM strength and can even reverse sign, revealing that the WM field controls the propagation channels through symmetry breaking rather than merely amplifying the probe. Eigenchannel diagnostics further indicate that this giant rotation requires coherent excitation of both WM-dressed propagation channels, which in turn depends on three factors: Raman-induced asymmetry, the EIT-supported Rydberg pathway, and vdW nonlocality. These results demonstrate a symmetry-breaking-controlled mechanism for Rydberg magneto-optics, with applications to weak-light polarimetry and all-optical polarization control.

quant-ph↗

Improving Code Understanding in Large Language Models through Concept-Aware Consistency Learning

Large language models (LLMs) have recently shown impressive results on diverse code-related tasks, benefiting from large-scale training and instruction tuning. However, studies reveal that their grasp of fundamental programming concepts, such as data flow and control flow, remains shallow, leading to fragile performance when code requires deeper reasoning. This limitation restricts the practical adoption of LLMs in real-world software development. To address this issue, this work introduces a counterfactual code augmentation framework combined with concept-aware tuning, designed to guide LLMs toward stronger conceptual understanding. Comprehensive evaluation across multiple models and benchmarks demonstrates the effectiveness of the proposed approach.

cs.SE↗

Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models

While Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in 2D medical image understanding, their extension to 3D volumetric imaging remains hindered by prohibitive annotation costs and dataset opacity. Current data formats, predominantly consisting of rigid Visual Question Answering (VQA) pairs or unstructured final clinical reports, typically fail to capture explicit clinical reasoning. To address this limitation, we introduce a large-scale structured reasoning dataset constructed via a novel slice-wise data synthesis paradigm. Inspired by the genuine diagnostic workflow of radiologists, this paradigm models visual cognition by decomposing the complex 3D reading process, translating global clinical priors into fine-grained, per-slice observations that are subsequently synthesized into an interpretable Chain-of-Thought (CoT). Crucially, this synthesized reasoning framework enforces essential clinical principles: sequential spatial tracking, multi-slice spatial awareness for artifact mitigation, and differential exclusion. To validate this approach, we instruction-tune a standard 2D-pretrained MLLM baseline using the synthesized data to enhance its volumetric comprehension. Comprehensive evaluations across multiple 3D medical benchmarks demonstrate that our method yields significant performance improvements over the 2D baseline. Furthermore, the resulting model exhibits robust spatial reasoning capabilities and rivals resource-intensive native 3D architectures, effectively bridging the performance gap. Ultimately, this data-centric strategy unlocks deep volumetric understanding and highly interpretable clinical logic without requiring computationally expensive 3D-specific pre-training. The complete repository, including datasets and training workflows, is publicly available at https://github.com/2020420145009/hounsfield.

cs.CV↗