SearcharxivSearch

arXiv subjects

Ying Li

Publications and source records attributed to Ying Li.

At least 19 recordsLinked to original sources

Low-cost algorithm-to-execution framework for surface-code quantum computing

The execution of useful quantum algorithms on fault-tolerant processors requires more than a mapping from logical gates to encoded operations: the spatial organization, non-Clifford resource supply, and execution schedule must also be determined while keeping physical overhead within practical limits. Although the theoretical hierarchy from logical circuits to fault-tolerant operations is well established, these implementation choices are often specified and optimized separately. Here we develop a low-cost algorithm-to-execution framework for surface-code quantum computing. From hierarchical algorithm descriptions, it constructs dependency-preserving logical schedules and an executable workload capturing logical interactions, operation parallelism, and time-resolved non-Clifford demand, thereby linking logical computation to surface-code organization, resource-state preparation, and fault-tolerant execution in a traceable workflow. We apply the framework to twenty benchmark circuits across seven algorithm families and a hierarchically composed application-scale elliptic-curve discrete-logarithm workload. Physical costs vary substantially even for circuits with similar logical resource counts. Under our direct-rotation calibration, non-Clifford implementation selection reduces space-time volume by up to 241.5 times versus an all-synthesis baseline for the QAOA amplitude-amplification workload. Circuit-specific surface-code layouts reduce routed-latency estimates for all twenty benchmarks; thirteen also reduce space-time volume because communication savings outweigh added spatial overhead. These results show that low-cost fault-tolerant execution depends on computation scheduling and organization, not aggregate logical resource counts alone.

quant-ph

Gradient estimates for the fractional $p$-Laplacian in the supercritical regime

We prove a pointwise gradient estimate for the fractional $p$-Laplace equation with measure data. Let $n\ge 2$, $p>1$ and $\max\{n/p,1/p'\}<s<1$, where $p'=p/(p-1)$. Suppose $u\in W^{s,p}(\mathbb R^n)$ is a weak solution of \[ (-\Delta_p)^s u=\mu \quad\text{in }\Omega, \] with $\mu\in\mathcal M_{\mathrm{loc}}(\Omega)$. Set $\gamma=s-(p-1)/p$. If $B_{2R}(x_0)\Subset\Omega$ and $\mathcal{W}_{\gamma,p}^{|\mu|}(x_0,2R)<\infty$, then $u$ is Fr\'echet differentiable at $x_0$, and \[|\nabla u(x_0)|\le C\bigl[\mathcal A(u;x_0,2R)+\mathcal{W}_{\gamma,p}^{|\mu|}(x_0,2R)\bigr], \] where $\mathcal{W}_{\gamma,p}^{|\mu|}$ denotes the truncated Wolff potential, and $\mathcal A$ depends on the local oscillation of $u$ and its nonlocal tail. The proof uses a comparison with fractional $p$-harmonic replacements and an affine decay estimate for homogeneous solutions, with constants independent of the affine slope. In the large slope regime, the affine decay estimate follows from a Schauder estimate for the linear nonlocal equation satisfied by the affine remainder.

math.AP

ES-VP : Energy-Shaped Dynamic Visual Prompting for Efficient Model Adaptation

Visual prompting (VP) has emerged as a parameter-efficient method for adapting pre-trained models to downstream tasks. However, existing approaches encounter a trade-off between flexibility and efficiency. Some methods apply a fixed prompt to all images, ignoring individual image characteristics, while others introduce auxiliary networks to generate diverse prompts. Although the latter can improve performance, it also significantly increases parameter usage and the potential for overfitting to specific datasets. Furthermore, the auxiliary networks, combined with inherent biases in pre-trained models, limit scalability and generalization. In this paper, we propose Energy-Shaped Visual Prompting (ES-VP), a novel approach that generates image-specific prompts using low-rank initialization and energy-guided dynamic adaptation, achieving superior performance with fewer parameters compared to single-prompt methods. ES-VP directly utilizes the pre-trained model for adaptive prompt generation, ensuring both parameter efficiency and improved generalization. Extensive experiments conducted on five architectures across fifteen datasets demonstrate that ES-VP consistently outperforms current state-of-the-art (SOTA) single and diverse VP methods. For instance, using the CLIP architecture across four datasets, ES-VP outperforms the SOTA method DAM-VP by an average of 2.6\% in accuracy while utilizing 590$\times$ fewer VP parameters, thereby establishing a new benchmark for efficient and generalizable model adaptation.

cs.CV

Quadrupole White-light Sources in an X1.2 Flare Observed by ASO-S/LST/WST and SDO/HMI

We present observations of an X1.2 white-light flare on 2023 January 6, which exhibits a rare quadrupolar white-light source configuration. This event was observed by the White-light Solar Telescope (WST; 3600 \AA) aboard the Advanced Space-based Solar Observatory and the Helioseismic and Magnetic Imager (HMI; 6173 \AA) aboard the Solar Dynamics Observatory. Four flare-related footpoints (labeled as FP1--FP4) were nearly simultaneously identified in both WST 3600 \AA\ and HMI 6173 \AA\ continua, associated with a quadrupolar magnetic configuration and a failed filament eruption. The inner sources of FP1 and FP2 showed a similar enhancement of $\sim$65%/10% in the WST/HMI continuum, while the outer sources of FP3 and FP4 exhibited weaker responses. The inner footpoints had earlier responses in UV and EUV bands and were spatially coincident with the hard X-ray (HXR) footpoint sources. The two southern footpoints (FP2 and FP4) showed stronger HXR and white-light emissions than their northern counterparts (FP1 and FP3), with FP4 uniquely exhibiting a distinct HXR emission above 60 keV, in contrast to the absence of such an emission at FP3. Notably, faint WST 3600 \AA\ enhancements at FP4 were observed during the gradual phase, temporally consistent with the fallback of filament material. This X1.2 flare presents a novel quadrupolar white-light structure, enriching our understanding of the generation and evolution of white-light flares.

astro-ph.SR

When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations

Large language models can write fluent stories, but open-ended storytelling requires more than local fluency. In evolving world simulations and AI-native games, models must preserve facts, relationships, causal dependencies, and character states as the world changes. We introduce WSE-bench, a process benchmark that separately evaluates sustained generation, canonical coherence, and meaningful development in dynamic LLM storytelling. Generation Coverage records the proportion of planned narrative steps produced; Consistency tracks when canon breaks; and Richness measures how meaningfully branching, player-shaped trajectories develop. Across frontier models, Consistency and Richness do not form a smooth trade-off: their empirical Pareto frontier is non-concave, with several non-dominated intermediate configurations that no positive linear weighting can select. Added structure can enrich trajectories, but it does not uniformly improve coherence and may shorten them. Model scale chiefly improves sustained generation, without producing reliable gains in canonical coherence or meaningful development. These results show that sustained generation, canonical coherence, and meaningful development are distinct and sometimes competing capacities. WSE-bench makes those dynamics visible by extending narrative evaluation from finished stories to the processes that create them.

cs.CL

Theoretical analysis towards accurate optomechanical detection of quantum gravity effects

Optomechanical systems offer a promising platform for observing dynamical signatures of quantum gravity through precision measurements of quantum harmonic oscillator dynamics. However, most existing analyses consider only the linear radiation-pressure interaction while neglecting higher-order optomechanical couplings and laser phase noise. These neglected contributions can be comparable in magnitude to the predicted quantum-gravity corrections and may therefore introduce spurious signals or mask the genuine physical effect. Here we reanalyze two experimentally realized platforms, a Fabry-Perot optomechanical system and a membrane-in-the-middle optomechanical system, by incorporating the complete nonlinear dynamics and realistic laser phase noise. Using measured device parameters, we derive revised protocols for generalized uncertainty principle tests and establish practical sensitivity bounds. Our results demonstrate that previous idealized estimates significantly overestimate the achievable resolution, underscoring the necessity of including higher-order interactions and implementing effective laser phase noise suppression in realistic assessments of optomechanical quantum gravity tests.

quant-ph

GenRec: An LLM-Backed Recommendation Ranker at Netflix

Large language models (LLMs) are reshaping recommender systems by enabling richer modeling of users, content, and context directly in natural language. At Netflix, we are exploring this direction through GenRec, an LLM-backed recommendation ranker built on top of an in-house foundational LLM. GenRec follows a two-phase framework: Phase 1 adapts an open-source LLM to Netflix data, developing deep understanding of the catalog and member behavior while balancing capabilities such as content understanding and instruction following. Phase 2 post-trains this foundation model with recommendation-ranking specific data, labels, and reward signals, aiming to align the ranker with business requirements and long-term member satisfaction. This paper focuses on Phase 2 and the transition from a traditional discriminative ranker with thousands of engineered features to an LLM-backed ranker driven by verbalized user histories and context. We describe our design for input verbalization and context engineering, post-training data construction, reward integration, model architecture, and a cost-constrained serving design based on a prefill-only inference approach. We report results from a large-scale A/B test comparing GenRec against the current production ranker model, where we show that a GenRec model trained with substantially fewer Phase-2 labeled training examples and input signals can achieve statistically significant gains in offline and online metrics. We discuss how LLM-backed recommenders could shift the recommendation paradigm: from feature engineering to context engineering, and from bespoke architectures to shared foundation backbones. We also outline practical lessons for serving such systems under real-world resource constraints.

cs.IR

ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion

Knowledge graph completion (KGC) aims to infer missing entities or relations from incomplete graph structures, and has evolved into multimodal knowledge graph completion (MMKGC), where entities are associated with multiple modalities such as text and images. Traditional representation learning approaches follow the embedding-based paradigm and may struggle when relation-specific evidence is limited. Meanwhile, LLM-based reasoning methods typically linearize graph structures into textual prompts, which obscures structural topology and neglects vital visual information. While vision-language models (VLMs) excel at multimodal reasoning, they cannot natively interpret structured graph topology, particularly when it comes to knowledge graphs where nodes and edges carry complex semantics. To bridge this gap, we propose ViSR-KGC, a visual subgraph reasoning approach for KGC. It integrates three complementary capabilities to capture semantic correlations: identifying global topology dependencies via representation learning, analyzing local multimodal evidence using VLMs, and providing necessary commonsense knowledge inherent in pre-trained models. Based on learned multimodal embeddings, our framework first extracts a compact and query-aware subgraph from the MMKG. Then, this subgraph is transformed into a visually interpretable image using a layout strategy selected through empirical comparison. Finally, the visualized subgraph, entity images, textual descriptions, and candidate answers are combined into a unified prompt, enabling the VLM to infer the missing entity.

cs.AI

BRiG-AFA: Bellman Risk-to-Go Learning for Non-Myopic Active Feature Acquisition

Active feature acquisition (AFA) asks which unobserved feature to measure next for each test instance under a budget. Greedy rules are easy to train but can overlook context features whose value is realized only through later acquisitions, while reinforcement-learning and generative approaches introduce difficult optimization or conditional-density estimation. We introduce \method, a deployable, supervised alternative that learns a separate candidate-conditioned risk-to-go function for every remaining budget. Starting from the one-step terminal classification risk, the functions are fitted backward with Bellman targets; inference greedily minimizes the learned terminal risk using only observed values, the mask, candidate identity, and remaining budget. A controlled non-myopic benchmark shows the expected mechanism: at budgets two and three, \method improves accuracy over its one-step ablation by $4.84\pm2.17$ and $4.39\pm1.10$ percentage points (mean $\pm$ standard error over five seeds). On Fashion-MNIST with 20 candidate pixels, it improves accuracy at every nontrivial reported budget on average, including $10.20\pm0.74$ points at four acquisitions; its mean paired gain across budgets $\{2,4,8,12,16\}$ is $3.50\pm0.37$ points. A three-seed MiniBooNE study is mixed at small budgets but positive at 8 and 16 acquisitions, identifying a current boundary rather than supporting a universal claim. These results establish a reproducible mechanism-level case for direct Bellman risk regression and delimit the experiments still needed for state-of-the-art comparison.

cs.LG

Modeling Unknown Nonlocal PDE Systems via Flow Map Learning

Nonlocal partial differential equations arise in many applications but are often difficult to model and learn because of the presence of nonlocal operators. We present a flow-map learning (FML) framework for modeling unknown nonlocal PDEs directly from solution data. Rather than learning or approximating the underlying nonlocal operators, the proposed approach learns the finite-time evolution operator in either modal or nodal space. Two complementary formulations are developed for spectral and grid-based solution representations. Numerical experiments on one- and two-dimensional fractional diffusion and wave equations demonstrate accurate and stable long-time prediction using only short observation windows. The proposed approach provides an effective data-driven framework for learning unknown nonlocal dynamics without explicit evaluation of nonlocal operators.

cs.LG

Role-Break in Attention Heads: Understanding and Detecting Hallucinations in VLMs

Despite remarkable progress in vision-language generation, Vision-Language Models (VLMs) remain prone to hallucinations, producing content that is inconsistent with or unsupported by the input image. Existing works largely design detection or mitigation methods around one specific hallucination pattern, such as visual-textual imbalance, but real VLM hallucinations arise from a mixture of multiple patterns, so signals bound to a single pattern struggle to remain stable across models and tasks. Under a unified head-level view, we find that hallucination-induced changes manifest as localized deviations from each head's faithful contextual behavior, a phenomenon we term Role-Break. Detailed analysis reveals that these deviations are systematically organized across attention heads, contextual sources, and deviation directions, and that the resulting signal is linearly readable once head identity is preserved. Based on these findings, we build a lightweight linear detector on top of Role-Break that requires no fine-tuning of the VLM, whose feature dimension stays below 5,000 and reaches an average AUROC of 93.23 across six VLMs and four benchmarks. A small-scale intervention experiment further shows that the detected tokens can be directly acted upon in the discriminative setting.

cs.CV

Qualitative properties of eigenfunctions in domains with small holes

In this paper we study qualitative properties of the eigenvalues and eigenfunctions of $-\Delta$ with Dirichlet boundary condition in a smooth bounded domain $\Omega$ with a small circular hole. In the literature, this is known as a "singular perturbation", in contrast with the "regular perturbation" case. Denoting by $\Omega_\epsilon:=\Omega\setminus B(P,\epsilon)$ where $B(P,\epsilon)$ is the ball centered at $P$ and radius $\epsilon$, for $P\in\Omega$ and $\epsilon$ small enough we investigate 1) quantitative estimates for the eigenfunctions of $-\Delta$ in $\Omega_\epsilon$; 2) the simplicity of the eigenvalues of $-\Delta$ in $\Omega_\epsilon$; 3) the behavior of nodal sets of the eigenfunctions of $-\Delta$ in $\Omega_\epsilon$. A key ingredient in our analysis consists of pointwise estimates on the so-called $u$-capacitary potential firstly introduced in \cite{afhl}.

math.AP

A Taxonomy of Performance Metrics for the Distributed Computing Continuum

Performance evaluation is essential for understanding, comparing, and improving computing systems, including Distributed Computing Continuum Systems (DCCS). In recent years, computational requirements have changed substantially with the growth of artificial intelligence and large-scale data-driven applications. These application tasks are increasingly distributed between resource-intensive data centers and resource-constrained edge environments. In this context, novel computing continuum architectures and algorithms are emerging, creating a need for transparent and consistent performance evaluation. However, existing evaluation practices often focus on isolated dimensions, such as computation, networking, energy efficiency, or application-level quality, and therefore provide only a partial view of cross-layer DCCS behavior. This paper presents a structured taxonomy of performance metrics for DCCS. The taxonomy organizes metrics into computing-level, network-level, and application/user-level categories, while also highlighting emerging dimensions such as sustainability, observability, adaptability, data locality, migration awareness, and continuum fragmentation. Further, we provide mathematical formulations and discuss their relevance to heterogeneous and dynamic continuum environments. We also summarize metric acquisition requirements in terms of acquisition scope, acquisition phase, and measurement method. These requirements help clarify whether a metric can be collected from a single node, multiple nodes, or the full system, and whether it is more suitable for operational monitoring or experimental evaluation.

cs.DC

Altermagnetism from a Cu-Fe Lieb Lattice in FeSe/Cuprate Heterostructures

Realizing altermagnetism in high-$T_c$ cuprate-based systems would provide a direct route for studying spin-split electronic bands in the absence of net magnetization and investigate their interplay with unconventional superconductivity. Here, we propose that FeSe/cuprate heterostructures offer such a platform, where a 45$^\circ$ twist of Cu and Fe layers creates an effective CuFe$_2$ Lieb lattice in which Fe magnetic order and Cu-Fe hybridization through the ligands induces altermagnetic $d$-wave spin splitting. A minimal tight-binding model shows that this mechanism is generic. Furthermore, a substrate-induced inequivalence of the two Se sites in FeSe provides a second route in which altermagnetism originates in the Fe layer and is transferred to the cuprate layer by proximity. Density functional theory calculations for FeSe/Bi$_2$Sr$_2$CuO$_6$ heterostructures confirm the viability of both mechanisms and reveal ways to enhance the spin splitting. These results establish superconducting cuprate/transition metal chalcogenide heterostructures as a promising setting for engineering altermagnetism and studying its coupling to unconventional superconductivity.

cond-mat.str-el

MEGA-CL: A Molecular Foundation Model for Generalizable ADMET Prediction through Graph External Attention and Contrastive Learning

Predicting the absorption, distribution, metabolism, excretion and toxicity (ADMET) properties of small molecules remains a major challenge in drug discovery. Here, we present MEGA-CL, a foundation graph neural network framework for universal molecular ADMET prediction. MEGA-CL integrates self-supervised contrastive learning with a multi-head external attention mechanism and an enhanced message-passing architecture, enabling simultaneous modeling of local chemical substructures and global inter-graph relationships while mitigating over-smoothing effects commonly observed in deep graph networks. Across 13 benchmark datasets and 21 downstream ADMET tasks, MEGA-CL consistently outperforms state-of-the-art baseline models. In particular, the framework demonstrates robust performance on challenging regression tasks, including clearance (CL) and steady-state volume of distribution (VDss), while maintaining strong generalization ability in independent external validation. Clinically relevant predictive accuracy was achieved, with more than 75% of predictions falling within a 3-fold error range. In an external evaluation on 18 novel compounds derived from recently approved FDA drugs, over 50% of human liver microsome clearance (HLMC) predictions were within a 2-fold error range. To further assess its practical applicability, MEGA-CL was prospectively evaluated on three preclinical drug candidates using in vitro hepatic microsomal metabolism assays and CYP450 inhibition assays guided by model predictions. The predicted HLMC values for all candidates were within 2.5-fold of the experimentally measured values, and 73.3% of CYP450 inhibition endpoints (11/15) were correctly classified. These results demonstrate the potential of MEGA-CL as a generalizable framework for accelerating in silico ADMET evaluation and early-stage drug candidate optimization.

cs.LG

Bifrost: Empowering Pretrained Language Model with Fallibility Representation for Log-Based Fault Diagnosis

Log-based fault diagnosis is crucial for runtime debugging and maintenance. Existing fault diagnosis methods use language models pre-trained on natural language (PLMs) for log representation. However, system faults are reflected in the multi-level structure of system logs. PLMs pre-trained on natural language struggle to comprehensively capture multi-level fault information, failing to meet the requirements of fault diagnosis. We refer to this information as fallibility representations. To address this problem, we propose a novel log representation learning method, Bifrost. It draws inspiration from the log analysis experience of Site Reliability Engineers and meticulously designs strategies based on self-supervised contrastive learning to learn the fallibility representations of logs. Across three public systems and one industrial ML-as-a-Service system, the log representations produced by Bifrost outperform existing PLMs by average margins of 9.83% in F1 for anomaly detection, 18.28% in HR@k for root cause localization, and 20.88% in Macro-F1 for fault identification.

cs.SE

Geo3R: Mitigating Spatial Reasoning Hallucination in Multimodal Large Language Models

Despite remarkable progress in visual understanding, Multimodal Large Language Models (MLLMs) remain prone to hallucinations when reasoning about spatial relationships, often producing judgments that contradict the true 3D structure of the scene. Though several existing works have proposed to mitigate hallucinations, our analysis indicates that they show limited effectiveness in spatial reasoning, as they fail to bridge the fundamental gap between 2D visual representations and 3D spatial reality. Based on this finding, we define hallucinations arising from insufficient spatial structure modeling as spatial reasoning hallucination, a subcategory of relation hallucination that existing mitigation methods fail to address. We further identify three typical scenarios where such hallucinations frequently occur: perspective effects, object orientation, and viewpoint changes. To this end, we propose Geo3R, a training-free, plug-and-play framework that incorporates geometric evidence and structured 3D reasoning to mitigate spatial reasoning hallucination. Experiments on three benchmarks, covering 18 tasks across all three scenarios, show that Geo3R substantially reduces spatial reasoning hallucination across diverse MLLMs without additional training, outperforming existing models and methods.

cs.CV

HalluScope: Fine-grained Hallucination Diagnosis for Multimodal Large Language Models

Although Multimodal Large Language Models have achieved strong performance across a wide range of vision-language tasks, they still suffer from hallucinations, where model outputs become inconsistent with the visual content, textual context, or commonsense knowledge. Existing studies primarily address this problem through coarse-grained detection. However, these approaches often provide insufficient diagnostic information for understanding hallucination types and supporting downstream hallucination mitigation. To bridge this gap, we propose fine-grained hallucination diagnosis for MLLMs, a new unified task that jointly performs hallucination detection, classification, and interpretable explanation generation. We develop an automated data generation pipeline and construct HalluScope-30K, a large-scale diagnostic dataset covering eight sources and five task categories. Based on this dataset, we design a multi-granular joint reward function and train two diagnosis models, HalluScope-4B and HalluScope-8B, which achieve state-of-the-art performance on both the MHALO benchmark and our fine-grained hallucination classification benchmark. Notably, detection and classification are mutually beneficial under joint optimization. Furthermore, diagnosis-driven feedback experiments show that the fine-grained diagnostic explanations produced by our model effectively guide target models to correct their hallucinations, with full diagnosis substantially outperforming all baselines on both Qwen3-VL-8B-Instruct and LLaVA-1.5-7B. Our code, data, and models are available at https://github.com/wkinglin/HalluScope.

cs.CV