SearcharxivSearch

arXiv subjects

Yutong Li

Publications and source records attributed to Yutong Li.

At least 19 recordsLinked to original sources

On overlap concentration in the Curie-Weiss Random Field model

We show that the overlaps of two independent replicas drawn from the Gibbs measure of the Curie--Weiss model with random field (CWRF) exhibit sub-Gaussian tails when the inverse temperature $ \beta < 1$ and the random field is drawn from a Gaussian distribution that makes the expected magnetization of the CWRF equal to its expected overlap in the thermodynamic limit. To show this, we prove finite-moment concentration via a rigorous Laplace approximation for a certain random rate function, and ``bootstrap'' it to a uniform bound for the moment-generating function of the squared overlap deviations using a sharp analysis of the associated rate function.

math.PR

Fundamental Limits of Quantum Metrology Beyond Fixed Causal Order

Quantum metrology with indefinite causal order (ICO) has attracted intense interest due to its potential to surpass the limitations of conventional fixed-order strategies. A key open question is whether ICO can fundamentally enhance asymptotic precision scaling. In this work, we bridge this gap for the estimation of a single parameter encoded in $N$ identical uses of a finite-dimensional quantum channel. We first establish a universal Heisenberg-scaling upper bound for the full general ICO process-matrix class and show that for unitary channels its optimal quantum Fisher information (QFI) coincides exactly with that of parallel strategies. For noisy channels, a structurally refined bound shows that channels restricted to the standard quantum limit (SQL) under parallel strategies remain SQL-limited under general ICO strategies. Most significantly, an asymptotically tight (AT) bound is derived to close the remaining possibility of an asymptotic ICO advantage by showing that general ICO and optimal parallel strategies have exactly the same leading QFI coefficient in both the SQL and Heisenberg regimes. For the operationally motivated class of quantum circuits with quantum control of causal order, we further obtain an iterative constraint on finite-query precision whose asymptotic limit agrees with that of the AT bound. Our results clarify the ultimate role of indefinite causality as a metrological resource for quantum channel estimation.

quant-ph

Type I Solar Radio Bursts Modulated by Solar Flares

Type I solar radio bursts (noise storms) are persistent meter-wave nonthermal emissions above active regions, with their occurrence and proper?ties closely related to the local magnetic configuration and nonthermal electron acceleration. This study examines a type I noise storm on 24 December 2023 and its relation to flare activity. The noise-storm source was co-spatial with active region AR 3529 and showed frequency-dependent spatial dispersion. The associated M2.9 flare strongly modulated the emission, with the storm intensity decreasing at flare onset, recovering afterward, and shifting to higher frequencies. Based on multiwavelength observations, we suggest that pre-flare small-scale reconnection supplied nonthermal electrons to overlying closed magnetic struc?tures and maintained the storm. During the flare, magnetic reconnection above the active region produced bidirectional plasma ejections and type III bursts with bidirectional frequency drifts; the gradually decreasing starting frequency of these bursts may indicate an upward-moving reconnection site. The resulting magnetic reconfiguration disrupted electron trapping and suppressed the storm, whereas post-flare magnetic recovery allowed the emission to resume. These results show that flares can modulate type I noise storms through magnetic restructuring and provide insight into the generation mechanism of noise storms.

astro-ph.SR

SG-UMP: Sequence-Guided Universal Multimodal Prioritization Calculation Framework

Multimodal sequential recommendation (MSR) improves recommendation by incorporating heterogeneous information such as text, images, and user interactions. However, existing MSR methods often fail to capture user-level preference heterogeneity and dataset-level modality bias, limiting their adaptability across users and datasets. To address this issue, we propose \textbf{S}equence-\textbf{G}uided \textbf{U}niversal \textbf{M}ultimodal \textbf{P}rioritization Calculation Framework (\textbf{SG-UMP}), a plug-and-play plugin for enhancing multimodal information processing in MSR. SG-UMP includes a Module Combiner for flexible multimodal processing and a Module Router for dynamic module ordering, enabling adaptation to both user preferences and dataset characteristics. Experiments on four real-world datasets show that SG-UMP consistently improves recommendation performance across different backbones and multimodal settings. The code is available at https://github.com/esemsc-xz524/SG-UMP .

cs.IR

DeMixPert: Decomposed Response Modeling with Gaussian Mixtures for OOD Single-Cell Perturbation Prediction

Predicting transcriptome-wide responses to unseen genetic perturbations remains a major computational challenge because accurate prediction requires recovering both perturbation-specific transcriptional shifts and heterogeneous cellular responses. Existing methods often entangle deterministic response structure with stochastic population-level variation, causing dominant shared patterns to mask weaker perturbation-specific signals and impair distributional modeling. To address these challenges, we propose \textbf{DeMixPert}, an approach for Decomposed response Modeling with Gaussian Mixtures for Out-Of-Distribution (OOD) single-cell Perturbation prediction. DeMixPert decomposes perturbation-induced changes into a basal-state-dependent systematic response, a perturbation-specific response, and population-level variation. The systematic component is derived from the basal state encoded from control-cell expression, whereas the perturbation-specific component is inferred from pretrained target embeddings for unseen-target generalization. DeMixPert models population-level variation using a Gaussian prototype Invertible Network and adaptively combines reusable Gaussian prototypes according to the basal state and perturbation condition. The resulting mixture is mapped to a condition-specific variation distribution. Sampled variations are integrated with the systematic and perturbation-specific components, followed by joint decoding with the basal state to reconstruct perturbed-cell gene expression. Experimental results show that DeMixPert effectively captures heterogeneous single-cell perturbation responses and achieves superior performance across unseen-perturbation settings. The source code is made publicly available upon publication.

cs.LG

From Metrics to Improvement: A Lifecycle-Aware LLM Feedback Framework for Research Software Quality

Research software is increasingly central to scientific workflows, yet it is often developed by researchers with limited software engineering expertise. This can lead to quality issues that hinder maintainability, reproducibility, reuse, and sustainability. Existing static analysis tools can identify such issues, but their outputs often require expert interpretation and provide limited support for translating quality assessments into actionable improvements. To address this gap, we propose a lifecycle-aware framework that integrates quantitative software quality assessment with Large Language Model (LLM)-based code refinement. The framework comprises two stages. First, a lifecycle-aware Quality Model is developed from established software quality standards and practitioner requirements. The model defines five quality dimensions and 25 candidate metrics, of which 14 are operationalized using existing analysis tools and custom measurements. Second, the resulting quality diagnostics are used as structured feedback within an iterative LLM-based refinement process, enabling generated improvements to be repeatedly reassessed against the Quality Model. We evaluate the framework on notebook-centric research software using multiple LLMs and compare iterative structured feedback with single-step feedback and unstructured prompting. The results show improvements in specific quality attributes, particularly code duplication and structural quality, while also revealing trade-offs among maintainability, code size, documentation, and complexity. These findings demonstrate the potential of metric-driven LLM feedback for research software quality improvement while highlighting its inherently multi-objective nature \footnote{The source code and experimental data are publicly available at https://github.com/QCDIS/Software_Quality_Control_LLM . }

cs.SE

Lagrangian Schur index and Bethe ansatz type formula

We propose a surprisingly elementary method to compute the Schur index in closed-form for general $\mathcal{N} = 2$ Lagrangian theories. The method is inspired by the Bethe ansatz type formula for $\mathcal{N} = 1$ superconformal index. We identify issues underlying the original derivation: the loss of periodicity property upon integration and the omitted poles outside of the annulus region. We circumvent the problems and transform integration into solving a simple difference equation. The final result is expressed as a finite sum of quasi-Jacobi forms involving twisted Eisenstein series. We test our method on different types of theories, including BCD-type $\mathcal{N} = 4$ theories and various $\mathcal{N} = 2$ quiver gauge theories.

hep-th

Enhancing Irregular Time Series Forecasting with Continuous-Time Modeling Framework

Irregular multivariate time series are widely encountered in applications such as healthcare monitoring, human activity recognition, and environmental sensing. Their core challenges stem from asynchronous observations, non-uniform sampling intervals, and the fact that temporal patterns themselves carry critical dynamic information. Existing approaches either rely on discretization-based preprocessing (e.g., interpolation, imputation, or aggregation), which disrupts the underlying continuous-time semantics, or adopt continuous-time modeling via ODE-based frameworks, which typically require specialized architectures and incur substantial computational overhead due to numerical solvers. To address these limitations, we propose WrapFlow, a continuous-time modeling framework for irregular time series forecasting. On the input side, WrapFlow introduces Continuous-Time Tokenization, which directly encodes raw observation events and explicitly models long unobserved intervals via gap-aware tokens. The resulting continuous-time tokens are then processed by a standard Transformer backbone to capture long-range temporal dependencies. On the output side, we develop a simulation-free training paradigm for Residual Flow Matching, which learns conditional residual vector fields around base predictions while avoiding numerical-solver simulation and backpropagation during training. This design enables high-quality continuous forecasting using only a small number of fixed rollout steps at inference. Extensive experiments on multiple real-world datasets demonstrate that WrapFlow achieves state-of-the-art performance.

cs.LG

$c$-axis strain tuning of superconductivity and symmetric elastoresistivity in CsV$_3$Sb$_5$

The kagome metal CsV$_{3}$Sb$_{5}$ hosts an intriguing interplay between charge-density-wave (CDW) order and superconductivity that is highly sensitive to lattice distortions. However, determining the specific roles of the in-plane ($A_{1g,1}$) and out-of-plane ($A_{1g,2}$) symmetric strain channels has been hindered by their intrinsic mixing in conventional piezo-based experiments. Here, we combine in-plane uniaxial strain with direct $c$-axis compression to independently access and disentangle these symmetry-resolved responses in CsV$_{3}$Sb$_{5}$. We reveal that $c$-axis compression drives a massive, linear enhancement of the superconducting transition temperature ($T_c$) alongside a suppression of $T_{\rm CDW}$. The tuning efficiency of this out-of-plane deformation acts with an opposite sign and far exceeds that of in-plane strain, demonstrating that $c$-axis lattice control dictates the phase competition. Furthermore, by isolating the pure elastoresistivity coefficients, we find that the out-of-plane cross-coupling coefficient ($m_{13}$) is comparable in magnitude but opposite in sign to the in-plane response ($m_{11}+m_{12}$). Unlike the sharply peaked in-plane response, $m_{13}$ exhibits a distinct, order-parameter-like onset across the CDW transition. Our results establish that out-of-plane lattice control plays a dominant role in tuning the intertwined states in CsV$_{3}$Sb$_{5}$ and provide a general pathway for resolving strain-coupled electronic responses in layered quantum materials.

cond-mat.supr-con

DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments

LLM-based agents have rapidly improved at operating individual digital environments such as mobile applications, desktop systems, and smart homes. However, real-world user goals often span multiple devices: information may come from a phone, be processed on a desktop, and the result may need to appear on another device. Most existing benchmarks center on a single dominant execution environment, making it difficult to evaluate whether agents can acquire and integrate information across heterogeneous devices and complete end-to-end tasks with cross-device dependencies. We introduce DevicesWorld, a large-scale executable benchmark for cross-device collaborative operation. DevicesWorld contains 6,140 tasks and integrates three classes of device environments -- mobile, desktop, and IoT -- into a unified cross-device interaction and evaluation framework. Each task defines a natural-language user goal, participating devices and initial states, executable actions, rule-based verifiers, and a cleanup procedure. A multi-stage construction and quality-control pipeline keeps tasks close to realistic user needs while allowing final outcomes to be automatically verified from device states and generated files. We evaluate five frontier LLM-agent systems on a fixed evaluation set. All methods achieve low success rates, with the best reaching only 12.5%. Among failed runs, about 28.7% satisfy at least one scoring condition yet still fail the full task. Trajectories show that agents become stuck acquiring information or manipulating interfaces, confuse source and output devices, or terminate before all conditions are jointly satisfied. DevicesWorld turns cross-device collaborative operation into an executable, reproducible, and diagnostically useful evaluation problem for research on reliable cross-device agents.

cs.CL

Rethinking Conditional Generation for Underwater Salient Object Detection

Salient Object Detection in underwater images remains challenging due to low contrast, uneven illumination, and color distortion caused by scattering and absorption effects, which limit the effectiveness of conventional SOD methods in underwater environments. To address these challenges, we propose a Degradation-aware Conditional Generation Network (DCGNet), specifically designed to construct reliable conditional features for underwater saliency generation. First, we design a Dynamic Multi-Granularity module (DMG) grounded in the human visual system to robustly detect salient objects of varying scales with blurred boundaries. Then, we develop an Underwater Physics-Prior module (UPP), which utilizes pseudo-depth guidance to estimate underwater light attenuation and backscatter, thereby restoring degradation-aware RGB features and mitigating color distortion and boundary ambiguity. Based on the physics-guided representation, we introduce an Underwater Spatial Gaussian module (USG), which constructs a spatial Gaussian saliency prior from the strongest guided response to enhance object-centered salient regions and suppress cluttered underwater backgrounds. In addition, a lightweight timestep-adaptive Diffusion Transformer (DiT) bottleneck is inserted into the denoising decoder to refine fused features at different diffusion timesteps. Comprehensive experiments on USOD10K, USOD, CSOD10K, MAS3K, and RMAS demonstrate that DCGNet significantly outperforms existing state-of-the-art methods, verifying its potential for complex underwater visual applications.

cs.CV

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

We introduce Embodied-R1.5, a unified Embodied Foundation Model (EFM) that integrates comprehensive embodied reasoning capabilities, spanning embodied cognition, task planning, correction, and pointing, within a single architecture toward general physical intelligence. Leveraging three automated data construction pipelines to significantly expand the data coverage of critical capabilities, we build a large-scale data system of over 15B tokens, and design a multi-task balanced RL recipe to alleviate heterogeneous task conflicts. We further introduce a Planner-Grounder-Corrector (PGC) closed-loop framework that enables a single model to autonomously execute and self-correct over long-horizon tasks. With only 8B parameters, Embodied-R1.5 achieves SOTA on 16 out of 24 embodied VLM benchmarks, surpassing leading models like Gemini-Robotics-ER-1.5 and GPT-5.4. Benefiting from the internalized embodied capabilities, Embodied-R1.5 can be fine-tuned into a VLA with only a small amount of data, outperforming leading VLA models like $\pi_{0.5}$ across 4 popular manipulation benchmark suites. We further conduct extensive zero-shot real-robot experiments, validating performance in instruction following, affordance grounding, articulated object manipulation, and long-horizon complex tasks, demonstrating strong generalization to the physical world. We open-source model weights, datasets, training code, and EmbodiedEvalKit, an evaluation framework tailored for embodied tasks, to facilitate future research in EFMs.

cs.RO

Teach Multimodal Recommendation Model to See via Personalized Visual Extraction and Adaptive Learning

Multimodal sequential recommendation (MSR) incorporates textual and visual information to improve recommendation quality. However, recent studies and our empirical analysis show that visual features are often underutilized, thereby contributing far less than textual signals. We attribute this issue to two factors: insufficient visual representation learning (pretrained encoders fail to capture preference-relevant cues) and unbalanced visual-text optimization (textual features dominate the learning process). To address these issues, we propose Teach Multimodal Recommendation Model to See via Personalized Visual Extraction and Adaptive Learning (REVEAL), a plug-and-play framework that enhances visual representation learning and cross-modal optimization without modifying the original recommendation backbone. REVEAL consists of Feedback-Guided Visual Extraction (FVE), which refines prompt-guided visual extraction through task-level feedback, and Adaptive Visual Learning (AVL), which dynamically reweights visual learning to alleviate modality imbalance. Experiments on multiple real-world datasets and MSR backbones demonstrate that REVEAL consistently improves recommendation performance. Further analysis shows that these gains arise from more effective attention to preference-relevant visual regions and better visual utilization during training. The code is available at https://github.com/YutongLi2024/REVEAL.

cs.IR

TROPHIES: Temporal Reconstruction of Places, Humans, and Cameras from Multi-view Videos

Reconstructing humans and their surrounding environments in a globally consistent 4D space is essential for comprehensive perception. However, prior works typically assume single-view inputs or decouple humans, scenes, and cameras, making them unable to recover coherent geometry, stable motion, and physically aligned trajectories. These limitations motivate us to introduce a new task: unified human-scene-camera reconstruction from multi-view videos, which aims to jointly estimate dynamic humans, static scenes, and camera poses in one global coordinate frame. We propose TROPHIES--Temporal Reconstruction of Places, Humans, and Cameras from Multi-view Videos-a unified framework tailored for this task. TROPHIES features a Human Branch that models humans through temporal and spatial reasoning, and a Scene Branch that reconstructs static geometry with human-aware attention. A global alignment and optimization module couples both branches by enforcing scale consistency, contact priors, and cross-view temporal coherence. Experiments on EgoHuman and EgoExo4D demonstrate that TROPHIES achieves globally aligned, physically plausible 4D reconstructions and consistently outperforms existing paradigms in both global fidelity and human-scene consistency.

cs.CV

EchoVQA: Enabling Conversational Assistance for Point-of-Care Cardiac Ultrasound

Point-of-care transthoracic echocardiography (TTE) enables cardiac assessment in virtually any clinical setting, yet its diagnostic utility remains constrained by the expertise required for image acquisition and interpretation. Visual question answering (VQA) offers a promising paradigm for bridging this expertise gap through interactive clinical assistance, but existing echocardiography VQA datasets are limited in scale, restricted to high-quality images, and only cover a few views. We introduce EchoVQA, the first large-scale VQA dataset for echocardiography, comprising 14,299 images and 74,819 question-answer pairs. The dataset integrates public sources (EchoNet-Dynamic, CAMUS) with our own point-of-care acquisitions from two handheld probes (Lumify, Clarius), spanning diverse views and including both high-quality and suboptimal images. Uniquely, EchoVQA includes acquisition guidance questions to help users optimize transducer positioning toward a diagnostic apical 4-chamber view for left ventricular ejection fraction estimation -- a challenging task for novice operators in point-of-care settings. We further develop a parameter-efficient method based on multimodal learnable prompts achieving state-of-the-art performance on most benchmarks, including EchoVQA, with significantly less trainable parameters than existing state-of-the-art approaches.

cs.CV

Coronal Diagnostics Via Modelling Periodic-Beaded Stripes of Solar Radio Bursts

Using high-resolution data from the Chashan Broadband Solar radio spectrometer at meter wavelengths (CBSm) of the Chinese Meridian Project-Phase II (CMP-II), Li et al. (2025) identified a novel fine spectral structure of solar radio bursts, termed periodic beaded stripes, and proposed a generation mechanism. Here we report additional events and develop a quantitative method to determine the physical conditions in the emission region. Periodic stripes tend to occur in the post-phase of flares and are associated with complex magnetic configurations. They repeat on sub-second timescales and show $\sim$0.1 s bead-like modulations, often accompanied by low-frequency absorptions. Modeling the chained stripes with linear kinetic theory of the double plasma resonance (DPR) instability constrains the source-region magnetic field to 0.2-1.7 G and the plasma density to (1-7) $\times 10^8$ cm $^{-3}$. The former follows the drift of individual stripes, and the latter tracks the overall trend. This study summarizes the key properties of periodic beaded stripes and establishes a quantitative DPR-based framework for coronal diagnostics.

astro-ph.SR

ORSIFlow: Saliency-Guided Rectified Flow for Optical Remote Sensing Salient Object Detection

Optical Remote Sensing Image Salient Object Detection (ORSI-SOD) remains challenging due to complex backgrounds, low contrast, irregular object shapes, and large variations in object scale. Existing discriminative methods directly regress saliency maps, while recent diffusion-based generative approaches suffer from stochastic sampling and high computational cost. In this paper, we propose ORSIFlow, a saliency-guided rectified flow framework that reformulates ORSI-SOD as a deterministic latent flow generation problem. ORSIFlow performs saliency mask generation in a compact latent space constructed by a frozen variational autoencoder, enabling efficient inference with only a few steps. To enhance saliency awareness, we design a Salient Feature Discriminator for global semantic discrimination and a Salient Feature Calibrator for precise boundary refinement. Extensive experiments on multiple public benchmarks show that ORSIFlow achieves state-of-the-art performance with significantly improved efficiency.

cs.CV

Scientific tools and Innovation: Big Science Facilities Yield More Novel and Interdisciplinary Knowledge

Scientific tools dictate the boundaries of human knowledge, serving as the foundation for perceptions and explorations. In the era of Big Science, science are increasingly dependent on advanced analytical technologies and experimental platforms. Over the past decades, national and supranational entities have invested massive financial resources, collaborative networks, and collective intelligence to construct Big Science Facilities (BSFs) aimed at generating cutting edge knowledge. However, empirical evaluations of these machines actual performance in driving scientific innovation remain scarce. To address this gap, we collected 310,086 publications from 88 global BSFs and constructed a matched control dataset of approximately 3 million publications sharing the same last authors. Our analysis reveals that the utilization of BSFs has expanded significantly since 1950s. Crucially, publications supported by these facilities exhibit higher recombinant novelty and interdisciplinary integration. Furthermore, this improvement is most pronounced in non physical sciences domains traditionally peripheral to BSFs core focus indicating the emergence of a powerful intra facility knowledge spillover effect. By enriching the Facilitymetrics framework, our findings provide empirical evidence that BSFs act as vital engines for scientific discovery, offering policymakers essential metrics to justify infrastructural investments, while prompting the science of science community to reassess the profound impact of scientific tools on knowledge production

cs.DL