SearcharxivSearch

arXiv subjects

Yuan Chen

Publications and source records attributed to Yuan Chen.

At least 19 recordsLinked to original sources

Data-driven Effective Modeling of Stochastic Chemical Reaction Networks

The Stochastic Simulation Algorithm (SSA), widely considered an exact algorithm for stochastic chemical reaction networks, suffers from high computational cost. In this work, we propose a data-driven effective model that operates on a user-defined coarse time step independent of the underlying microscopic reaction-event scale. This is accomplished by directly approximating the finite-time transition kernel of the continuous-time Markov chain induced by SSA, using a generative machine learning model trained on short bursts of SSA simulation data. The trained model constructs a stochastic propagator that recursively generates statistically consistent trajectories at the constant coarse time step, with significantly reduced computational cost. In this paper, we employ conditional normalizing flow as the stochastic propagator. A comprehensive set of numerical examples is presented to demonstrate the accuracy and efficiency of the proposed method.

math.NA

Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation

Recent advances in large language models (LLMs) have enabled their use as conversational recommender systems (CRS), demonstrating strong recommendation accuracy and natural dialogue. However, guiding multi-turn interactions to elicit user preferences effectively remains challenging. Existing approaches either use separate reinforcement learning agents with templated interactions or optimize for interactivity judged by another LLM, without measuring how much useful information is actually gained. We propose a new approach that quantifies the effectiveness of each interaction by the reduction in the assistant's uncertainty, measured via entropy over recommendations. We apply this entropy reduction as a reward---without relying on ground-truth recommendations, which are often unavailable in real-world scenarios---to fine-tune the LLM, enabling strategic interaction generation. Empirical results with supervised fine-tuning (SFT) and direct preference optimization (DPO) on the INSPIRED and ReDial datasets show that our method improves both recommendation quality and conversational efficiency.

cs.IR

Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?

Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-task success or single-frame grounding. Neither isolates whether a model can reconstruct the causal, task-relevant transition produced by an action- crucial for rejecting stale observations, verifying progress, and recovering from failure. This is difficult because inference, remote input, app rendering, and screenshot capture are asynchronous: the next observation may be delayed, occluded, transient, or unrelated, then misread as progress and carried into subsequent planning. We introduce Desktop-Delta Bench (DDB), an offline step-level benchmark with 2,013 human-verified instances from novel, multi-app Linux trajectories across ~15 applications and 50 task domains. DDB trajectories targets 3 failure dimensions- state verification, source tracking, and context-aware control- through 2 complementary tasks: 463 3-frame temporal-ordering instances, including 105 with a cross-trajectory decoy, and 1,550 before-after pairs labeled from 5 actions + its payload. We evaluate 8 closed and open-source model families across 32 ordering and 16 single-action settings, observing consistent gaps. Ordering remains unsaturated: best non-decoy and decoy exact-match rates are 65.1% and 65.7%. Task context improves decoy identification by 6.9 percentage points but reduces non-decoy exact match by 2.2 points; error analysis reveals systematic copying of the presented A-B-C order. Single-action results show that inferring the action family is harder than locating it: click F1 is 0.96 vs, 0.76 for drag, while recognized drags are generally localized well. DDB, thus, complements end-to-end benchmarks by filling the missing diagnostic layer between GUI grounding and final task success, enabling targeted improvements to desktop CUA verification, reliability, and recovery.

cs.AI

When HTTP 402 Meets the Blockchain: Risks on Emerging x402 Payments

x402 is an emerging payment protocol for Web APIs and autonomous AI agents. x402 extends HTTP 402 with a payment negotiation flow and delegates payment proof verification and on-chain settlement to third-party facilitators. As a result, facilitators serve as a shared payment infrastructure for many independent merchants. This centralizes trust and validation in one component, so a single flaw can affect many services. Despite rapid adoption by major vendors and economically meaningful mainnet activity, the security posture of real-world x402 deployments remains poorly characterized. We present the first systematic study of authorization correctness and execution safety in current facilitator-mediated x402 deployments in the wild, identifying eight security rules for facilitators as critical payment infrastructure. Based on our analysis of rule violations, we derive four new attack vectors, including Free Shopping, Asset Theft, Service Denial, and Gas Abuse. These attacks exploit weaknesses in the real-world facilitator and server implementations and cause severe harm, including direct financial loss to merchants, theft of facilitator-held assets, unbounded sponsor-paid gas/fees, and disruption of payment services. To assess the security of x402 deployments at scale, we propose a semi-automated black-box tool and apply it to 15 major x402 facilitators collectively used by over 60K sellers and 360K buyers. Alarmingly, we find violations in all evaluated facilitators. We responsibly disclosed our findings to the affected parties, who acknowledged the issues and adopted mitigations, including changes by Coinbase. Finally, we complement our controlled testing with an empirical measurement of over 119 million recent Base and Solana transactions, quantifying x402 adoption, facilitator centralization, and ecosystem-level risk indicators.

cs.CR

Criticality and reduced dynamical resilience in PM2.5 pollution systems

Concentration-based metrics underpin air-quality assessment, while dynamical persistence and recovery describe how rapidly high-PM2.5 episodes dissipate and how strongly they retain memory. Here we introduce a finite-memory multiplicative reversion (FMMR) process that links the lognormal concentration backbone of PM2.5 variability with event recurrence, temporal memory, variance amplification and local dynamical resilience. Across station observations and reanalysis data, elevated PM2.5 regimes show a coherent set of critical signatures: stronger memory, rising autocorrelation, broader upper tails, amplified variance, reduced resilience and more clustered exceedance events. Together, these co-occurring signals reveal dynamical criticality in PM2.5 pollution systems, with critical slowing down expressed as a loss of restoring capacity under high-pollution conditions. A gridded comparison across populated and emission-influenced regions further shows that areas with similar PM2.5 burden can differ in recovery capacity, while eastern China has shifted toward higher resilience during recent air-quality improvements and India and West Africa occupy lower-resilience states. By identifying where pollution burden and recovery capacity diverge, these findings establish dynamical persistence and resilience as complementary dimensions of PM2.5 risk and provide a quantitative basis for resilience-oriented air-quality assessment.

physics.soc-ph

R^3: Advertisement Compliance Rectification via Group-Relative Experience Extractor and Curriculum Reinforcement

Rigorous content moderation is crucial for online advertising but leads to millions of daily rejections. This scale renders manual rectification infeasible, particularly for video advertisements. However, existing safety-driven methods often suffer from aggressive over-editing, which compromises the advertiser's original semantic intent merely to satisfy compliance. In this work, we target the rectification of textual violations in video ads, covering both speech transcripts and on-screen text. We propose R^3, a novel framework designed to harmonize compliance with original semantic intent preservation. Our approach integrates three key innovations: (1) an experience-driven data synthesis framework that bootstraps high-quality supervision via a group-Relative compliance experience extractor; (2) a curriculum Reinforcement learning strategy with hierarchical rewards designed to enforce compliance while maximizing semantic consistency; and (3) a comprehensive video Rectification framework seamlessly integrating text recognition, rewriting, and re-rendering for industrial deployment. Extensive experiments on industrial datasets and online A/B testing demonstrate that R^3 significantly outperforms state-of-the-art baselines, achieving an optimal trade-off between violation rectification and intent preservation.

cs.CL

Comb-enabled spectral-domain image transport through perturbation-prone multimode fibers

Multimode fibers (MMFs) offer a compact platform for imaging, sensing, and information transport, but their practical deployment is hindered by sensitivity to fiber perturbations, which alter modal coupling and invalidate conventional speckle-based calibrations. Here, we demonstrate perturbation-resilient image transport through MMFs by combining image-to-spectrum encoding with dual-comb spectroscopy. Two-dimensional images are converted into comb-line-resolved spectral signatures before fiber transmission, allowing spatial information to be carried in the spectral domain rather than in the output speckle field. After propagation, dual-comb heterodyne detection maps the encoded spectrum into the radio-frequency domain, enabling massively parallel spectral readout with a single photodetector. Neural-network-assisted compressive reconstruction further enables high-fidelity imaging from sparse, noisy, and spectrally aliased measurements. Our approach achieves Pearson correlation coefficients exceeding 0.9 under strong fiber perturbations and supports frame rates up to 2.5 MHz, allowing the observation of transient switching dynamics in a digital micromirror device. These results establish a powerful tool for robust, real-time image transport through flexible MMFs, with potential applications in remote sensing and fiber-based optical instrumentation.

physics.optics

An ultralow-loss integrated photonic platform for discrete-variable quantum information processing

Photonic integrated circuits offer a scalable and robust route toward quantum information technologies by consolidating photon sources and linear optical networks onto compact, wafer-manufacturable chips. Although silicon photonics has enabled diverse discrete-variable quantum breakthroughs -- spanning multiphoton entanglement, quantum networking, and photonic qubit fusion for quantum computing -- scaling these platforms beyond proof-of-principle demonstrations remains severely constrained by a critical system-level bottleneck. Optical loss compounds rapidly across photon generation, routing, and state analysis, causing multiphoton generation probabilities to plummet exponentially as circuit depth and complexity grow. Here we overcome this rate-loss barrier by demonstrating a monolithic, ultralow-loss silicon nitride (Si$_3$N$_4$) integrated photonic platform engineered for high-performance discrete-variable quantum information processing. Our architecture seamlessly integrates narrowband photon-pair sources with low-loss qubit-fusion circuits and reconfigurable state-analysis interferometers. The on-chip sources prepare Einstein-Podolsky-Rosen (EPR) states with a fidelity of 0.9875(3) and exhibit near-unity photon indistinguishability, yielding a heralded Hong-Ou-Mandel interference visibility of 0.990(6). By executing on-chip fusion of two EPR states, we synthesize and characterize four-photon Greenberger-Horne-Zeilinger states with a record fidelity of 0.943(8) and a fourfold count rate of 27 Hz -- more than two orders of magnitude higher than previous silicon-photonic implementations. Combined with standard CMOS-compatible fabrication on 150-mm-diameter wafers, these results establish ultralow-loss Si$_3$N$_4$ integrated photonics as a definitive, manufacturable platform for deployable, large-scale quantum information processors.

quant-ph

Probing the ubiquity of complex ices in protostars with JWST: the first systematic quantification of weak ice bands between 6.8 and 7.9 micron

Complex organic molecules (COMs) are the key to understanding the chemical evolution from simple interstellar molecules to potential prebiotic material. Although COMs have been extensively studied in the gas phase toward protostars, their counterparts in ices, where they are thought to form at earlier stages, remain far less constrained. A number of diagnostic features of complex ices lie between 6.8 and 8.8 um, a region known as the "COM ice fingerprint range," but previous infrared facilities lacked the sensitivity and spectral resolution required to quantify the weak bands therein. With the unprecedented sensitivity and resolving power of JWST, these limitations can now be overcome. Here, we present the first large-sample quantitative study of the absorption features at 7.02, 7.24, 7.40, and 7.67 um, using MIRI-MRS spectra of 21 protostars. The CH4 band at 7.67 um is the strongest band and shows remarkably uniform peak positions (7.67-7.68 um) and FWHMs (0.06-0.08 um), suggesting CH4 ice as its dominant carrier. The 7.24 and 7.40 um bands exhibit larger source-to-source variations in peak positions and FWHMs, but their occurrence and intensities are strongly correlated with each other. Comparisons with existing and new laboratory spectra suggest HCOO- as the most likely carrier of these two bands, yet HCOO- cannot fully reproduce their intensity ratios, implying additional contributions from other species such as C2H5OH, CH3CHO, and CH3COCH3. Our results reveal, for the first time, the potential ubiquity of weak features of complex ices in protostars, which have remained largely undetected due to observational limitations.

astro-ph.GA

ASSCG: Just-Right Gating over Chattering for Fast-Slow LLM Planning in Autonomous Driving

Large language models (LLMs) can improve autonomous driving planning but are costly to query online, and existing fast-slow planners often rely on hand-designed triggering rules that either over-call the slow system or call it at the wrong times. We formulate slow-system invocation as a resource-aware sequential decision problem and propose the Adaptive Slow-System Control Gate (ASSCG), which makes frame-level Query/Cache/Drop decisions to refresh, reuse, or suppress slow guidance. ASSCG uses an RWKV backbone for efficient long-horizon gating and is trained with supervised fine-tuning followed by GRPO-style compute-aware reinforcement fine-tuning. We apply ASSCG to two different fast-slow architectures: (i) AsyncDriver on nuPlan Hard20 closed-loop evaluation, where ASSCG improves score to 67.28 (+2.28) while reducing average end-to-end inference latency by 60%; and (ii) a RecogDrive-based dual system that we build by replacing its original VLM-2B module with a lightweight ViT-based fast planner and adding an LLM slow planner, evaluated on NAVSIM, where ASSCG achieves 91.4 PDMS (+0.6) and increases average speed by 25%. The project page, including video visualizations and additional results, is available at https://williamxuanyu.github.io/asscg/.

cs.RO

Free-running single-cavity dual combs with Hz-level relative linewidth

Single-cavity dual-comb lasers provide a compact and efficient source for dual-comb spectroscopy in gas sensing applications; however, achieving sufficient free-running mutual coherence for comb-line-resolved, high-resolution measurements remains challenging. Here, we present a symmetry-engineered bidirectional single-cavity dual-comb laser based on an all-polarization-maintaining fiber architecture. The system exhibits exceptional free-running mutual coherence, achieving Hz-level relative linewidths without active feedback or phase correction. The time-averaged absolute jitter of the dual-comb repetition-rate difference reaches 4.7*10^-7 min-1, representing an improvement of nearly two orders of magnitude over previously reported free-running systems. As a spectroscopic demonstration, we resolve ~49,000 comb lines over a 5.4 THz optical bandwidth and measure the absorption spectrum of carbon monoxide (12CO), faithfully retrieving molecular line shapes with millisecond acquisition times. This architecture provides a compact and robust free-running platform for broadband molecular spectroscopy and millisecond-scale, line-shape-resolved gas sensing.

physics.optics

Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation

Self-distillation improves reasoning in large language models by using the model's own rollouts as training signal, typically through implicit logit-level alignment that minimizes KL divergence toward a privileged target distribution. However, because this supervision is generated via uncontrolled sampling, it provides no diagnostic insight into the model's specific errors or corrective guidance for its individual failure patterns. Consequently, the model learns to imitate a privileged distribution rather than receiving fine-grained corrections that pinpoint where and why its reasoning fails. In this paper, we propose Trajectory-Augmented Policy Optimization (TAPO), which advances self-distillation from implicit distributional alignment to explicit trajectory construction. During RL training, the model produces both correct and incorrect rollouts to the same query, and TAPO leverages this contrastive structure to construct micro-reflective corrections, new training trajectories that retain the model's erroneous reasoning up to the point of failure, then insert a natural-language diagnosis and corrected reasoning guided by a correct reference from the same sampling group. Since each trajectory is anchored in the learner's own prefix and solutions, the corrective signal preserves the model's on-policy distribution to a greater extent than the position-wise alignment imposed by KL-based methods. To integrate these trajectories, TAPO introduces difficulty-aware candidate selection at the model's capability boundary and decoupled advantage estimation to prevent gradient contamination. Experiments on AIME 2024, AIME 2025, and HMMT 2025 show that TAPO achieves consistent improvements over GRPO under the same number of training steps. Further analysis demonstrates that TAPO strengthens both first-pass reasoning and error-correction effectiveness.

cs.LG

FSS-Net: Frequency-Spatial Synergy Network with Wavelet Attention for Carotid Artery Ultrasound Segmentation

Accurate segmentation of carotid arteries in ultrasound imaging is critical for stroke risk assessment. However, speckle noise, low contrast, and blurred boundaries remain major challenges. In this paper, we propose a Frequency-Spatial Synergy Network (FSS-Net) to achieve noise-robust and high-precision carotid artery segmentation. The network integrates wavelet transform, multi-domain attention, and edge enhancement into a unified encoder-decoder architecture. Specifically, a Channel-Spatial-Wavelet Attention (CSWA) module is designed to suppress noise and purify semantic features in the frequency domain. A Wavelet-Enhanced Bottleneck (WEB) module is introduced to capture long-range global dependencies efficiently. Furthermore, a Laplacian-Guided Adaptive Edge Fusion (LAEF) module compensates high-frequency details and maintains boundary continuity. Extensive experiments on carotid ultrasound datasets show that FSS-Net achieves a Dice score (DSC) of 96.46% and strong robustness under low SNR conditions, outperforming several state-of-the-art methods. This method realizes accurate segmentation of carotid artery in ultrasonic imaging, effectively identifies carotid atherosclerotic plaque, and is verified by other task (such as segmentation of breast cancer), suggesting that it has good clinical application potential in identifying abnormal tissue masses in ultrasonic images.

cs.CV

Context-Aware Deep Learning for Defect Classification in Atomic-Resolution STEM

Artificial intelligence is rapidly advancing materials characterization, yet most applications in electron microscopy rely solely on image contrast, overlooking the chemical and experimental context that shapes image formation. This limitation makes defect classification inherently ambiguous, as similar contrasts can arise from different materials or imaging conditions. Here we develop a context-aware learning framework that integrates image-derived contrast with metadata describing composition, beam energy, and detector geometry. Using a systematically constructed dataset of ~55 million simulated patches spanning 576 cases across 96 doped monolayer transition-metal dichalcogenides, we show that conditioning on contextual variables transforms defect classification from an ill-posed image-only task into a well-posed, physically grounded problem. The framework achieves over 98% accuracy on simulations and near-human agreement on experimental data, with a 94% reduction in posterior entropy. By emphasizing contextual grounding over architectural complexity, this approach links experimental image contrast to the underlying chemical and imaging conditions, supporting physically grounded defect assignments and a general pathway toward multimodal AI models for autonomous materials characterization.

cond-mat.mtrl-sci

Robust Secure Beamforming for Movable Antenna Enhanced Integrated Sensing and Communications

In this letter, we investigate robust beamforming design for a movable antenna (MA)-enhanced secure integrated sensing and communications (ISAC) system with imperfect eaves?dropping channel state information (CSI). To improve radar sensing performance, we formulate a radar signal-to-interference?plus-noise ratio (SINR) maximization problem by jointly opti?mizing the transmit beamforming and antenna placement while ensuring communication data security. However, the resulting op?timization problem is inherently intractable due to the nonlinea mapping from antenna positions to channel coefficients, as well as the eavesdropper (Eve) channel uncertainty. To handle these challenges, we propose a block coordinate descent (BCD)-based algorithm incorporating successive convex approximation (SCA) and fractional programming (FP) techniques. Simulation results show that our proposed algorithm exhibits fast convergence and achieves a significant improvement in the radar SINR while guaranteeing communication security.

eess.SP

Ferroelectric brightening of spin forbidden dark excitons in a WSe2/hybrid perovskite heterostructure

Long-lived dark excitons in monolayer WSe2 present promising candidates for carrying spin and valley information, but their optical access and spin manipulation have conventionally required the use of strong external magnetic fields. Here, using a ferroelectric hybrid perovskite heterostructure, we leverage the ferroelectric proximity effect to break the WSe2's in-plane rotational symmetry and brighten the spin-forbidden dark excitons under zero magnetic field conditions. Furthermore, we show that the twist angle between the WSe2 and perovskite crystals controls the ferroelectric coupling strength and valley-contrasting polarization. Our proposed mechanism, supported by a four-band tight-binding model, suggests that the ferroelectric proximity effect induces an asymmetric intersublattice interaction, generating an effective in-plane spin-orbit coupling (SOC) field that rotates spin/valley polarization and brightens dark excitons. Our work establishes ferroelectric proximity coupling as an electrically reconfigurable, magnetic-field-free strategy for spin exciton control in two-dimensional semiconductors.

cond-mat.mtrl-sci

Free-Riding the Agentic Web: A Systematic Security Analysis of x402 Payments

The x402 protocol has crossed from prototype to infrastructure for the agentic web, driving 130 million all-time transactions and embedded in Google Cloud, Cloudflare, and Stripe. Yet bridging synchronous HTTP requests with asynchronous blockchain finality creates state-synchronization challenges, and x402's security has so far been examined only in piecemeal vendor disclosures. It is moreover not one artefact but a stack of an HTTP semantic, per-chain schemes, and a long tail of SDK and deployment choices whose required guarantees prior work has not established. We perform a systematic security analysis organized around five invariants grounded in specifications, literature, and vendor expectations, resolving every violation to the responsible layer. We identify four flaw classes: cross-resource substitution, duplicate-settlement race (independently corroborated by subsequent third-party reports), allowance overdraft, and denial of settlement. Against official SDKs and a production deployment, these reach resource-leakage ratios up to 100%. For pay-per-token scheme we prove a structural limit: no output-only pricing can be both fair to honest users and bounded against inflation of the hidden "thinking" tokens, the price of fairness being a $\sqrt{1+\Theta}$ manipulation gap. We propose per-flaw mitigations and a defense triple with provable guarantees, cutting per-call reasoning cost by 47% and inverting attacker leverage from 8.7$\times$ to 0.9$\times$ at only 2.8% overhead. All findings have been disclosed.

cs.CR

Lagged sea-surface-temperature precursors of the leading PM2.5 mode in China

Fine particulate matter(PM2.5) pollution in China is strongly modulated bymeteorological variability, yet its seasonal predictability from oceanic signals remains unclear. Here we identify the leading PM2.5 variability mode over China and show that it is preceded by coherent sea-surface-temperature anomaly clusters by more than one season. These oceanic precursors influence summer PM2.5 mainly by altering precipitation and lowlevel ventilation, and winter PM2.5 by modulating boundary-layer height and near-surface stagnation. Using the four largest precursor regions, a simple regression model achieves significant independent prediction skill for both summer and winter PM2.5 variability. Our results reveal a physical pathway linking sea-surface-temperature memory to regional aerosol pollution and provide a basis for seasonal air-quality risk assessment.

physics.ao-ph