SearcharxivSearch

arXiv subjects

Yuanlong Wang

Publications and source records attributed to Yuanlong Wang.

At least 19 recordsLinked to original sources

Unified formalism and adaptive algorithms for optimal quantum state, detector and process tomography

Quantum tomography is a standard technique for characterizing, benchmarking and verifying quantum systems/devices and plays a vital role in advancing quantum technology and understanding the foundations of quantum mechanics. Achieving the highest possible tomography accuracy remains a central challenge. Here we unify the infidelity metrics for quantum state, detector and process tomography in a single index $1-F(\hat S,S)$, where $S$ represents the true density matrix, POVM element, or process matrix, and $\hat S$ is its estimator. We establish a sufficient and necessary condition for any tomography protocol to attain the optimal scaling $1-F= O(1/N) $ where $N$ is the number of state copies consumed, in contrast to the $O(1/\sqrt{N})$ worst-case scaling of static methods. Guided by this result, we propose adaptive algorithms with provably optimal infidelity scalings for state, detector, and process tomography. Numerical simulations and quantum optical experiments validate the proposed methods, with our experiments reaching, for the first time, the optimal infidelity scaling in ancilla-assisted process tomography.

quant-ph

Estimation of a sparse multi-qubit Hamiltonian via compressed sensing

Hamiltonian estimation is an effective approach in studying the structure and dynamical evolution of quantum systems. The difficulty in estimating the Hamiltonian is that an $N$-qubit Hamiltonian has $4^N-1$ unknown parameters, requiring exponentially many equations for information extraction. In this paper we develop a method based on compressed sensing to estimate the Hamiltonian of a multi-qubit system. We identify a problem where as $N$ increases, the common sufficient condition (Restricted Isometry Property) for compressed sensing often fails, obstructing the application of compressed sensing in ($N\geq 3$)-qubit Hamiltonian estimation. To solve this problem, we propose a ``scale transformation" technique to restore RIP and ensure a compressive estimation of a $k$-sparse Hamiltonian using only $O(k\log(4^N/k))$ equations. In the numerical examples, we estimate the Hamiltonians of two 6- and 30-qubit systems, demonstrating the effectiveness of the method.

quant-ph

PBSBench: A Multi-Level Vision-Language Framework and Benchmark for Hematopathology Whole Slide Image Interpretation

Peripheral Blood Smear (PBS) is a critical microscopic examination in hematopathology that yields whole-slide imaging (WSI). Unlike solid tissue pathology, PBS interpretation focuses on individual cell morphologies rather than tissue architecture, making it distinct in both visual characteristics and diagnostic reasoning. However, current multimodal large language models (MLLMs) for pathology are primarily developed on solid-tissue WSIs and struggle to generalize to PBS. To bridge this gap, we construct PBSInstr, the first vision-language dataset for PBS interpretation, comprising 353 PBS WSIs paired with microscopic impression paragraphs and 29k cell-level image crops annotated with cell type labels and morphological descriptions. To facilitate instruction tuning, PBSInstr further includes 27k question-answer (QA) pairs for cell crops and 1,286 QA pairs for PBS slides. Building upon PBSInstr, we develop PBS-VL, a hematopathology-tailored vision-language model for multi-level PBS interpretation at both cell and slide levels. To comprehensively evaluate PBS understanding, we construct PBSBench, a visual question answering (VQA) benchmark featuring four question categories and six PBS interpretation tasks. Experiments show that PBS-VL outperforms existing general-purpose and pathology MLLMs, underscoring the value of PBS-specific data. We release our code, datasets, and model weights to facilitate future research. Our proposed framework lays the foundation for developing practical AI assistants supporting decision-making in hematopathology.

cs.CV

Achieving Fairness Without Harm via Selective Demographic Experts

As machine learning systems become increasingly integrated into human-centered domains such as healthcare, ensuring fairness while maintaining high predictive performance is critical. Existing bias mitigation techniques often impose a trade-off between fairness and accuracy, inadvertently degrading performance for certain demographic groups. In high-stakes domains like clinical diagnosis, such trade-offs are ethically and practically unacceptable. In this study, we propose a fairness-without-harm approach by learning distinct representations for different demographic groups and selectively applying demographic experts consisting of group-specific representations and personalized classifiers through a no-harm constrained selection. We evaluate our approach on three real-world medical datasets -- covering eye disease, skin cancer, and X-ray diagnosis -- as well as two face datasets. Extensive empirical results demonstrate the effectiveness of our approach in achieving fairness without harm.

cs.LG

Generalized collective quantum tomography: algorithm design, optimization, and validation

Quantum tomography is a fundamental technique for characterizing, benchmarking, and verifying quantum states and devices. It plays a crucial role in advancing quantum technologies and deepening our understanding of quantum mechanics. Collective quantum state tomography, which estimates an unknown state \r{ho} through joint measurements on multiple copies $ρ\otimes\cdots\otimesρ$ of the unknown state, offers superior information extraction efficiency. Here we extend this framework to a generalized setting where the target becomes $S_1\otimes\cdots\otimes S_n$, with each $S_i$ representing identical or distinct quantum states, detectors, or processes from the same category. We formulate these tasks as optimization problems and develop three algorithms for collective quantum state, detector and process tomography, respectively, each accompanied by an analytical characterization of the computational complexity and mean squared error (MSE) scaling. Furthermore, we develop optimal solutions of these optimization problems using sum of squares (SOS) techniques with semi-algebraic constraints. The effectiveness of our proposed methods is demonstrated through numerical examples. Additionally, we experimentally demonstrate the algorithms using two-copy collective measurements, where entangled measurements directly provide information about the state purity. Compared to existing methods, our algorithms achieve lower MSEs and approach the collective MSE bound by effectively leveraging purity information.

quant-ph

The Boundaries of Fair AI in Medical Image Prognosis: A Causal Perspective

As machine learning (ML) algorithms are increasingly used in medical image analysis, concerns have emerged about their potential biases against certain social groups. Although many approaches have been proposed to ensure the fairness of ML models, most existing works focus only on medical image diagnosis tasks, such as image classification and segmentation, and overlooked prognosis scenarios, which involve predicting the likely outcome or progression of a medical condition over time. To address this gap, we introduce FairTTE, the first comprehensive framework for assessing fairness in time-to-event (TTE) prediction in medical imaging. FairTTE encompasses a diverse range of imaging modalities and TTE outcomes, integrating cutting-edge TTE prediction and fairness algorithms to enable systematic and fine-grained analysis of fairness in medical image prognosis. Leveraging causal analysis techniques, FairTTE uncovers and quantifies distinct sources of bias embedded within medical imaging datasets. Our large-scale evaluation reveals that bias is pervasive across different imaging modalities and that current fairness methods offer limited mitigation. We further demonstrate a strong association between underlying bias sources and model disparities, emphasizing the need for holistic approaches that target all forms of bias. Notably, we find that fairness becomes increasingly difficult to maintain under distribution shifts, underscoring the limitations of existing solutions and the pressing need for more robust, equitable prognostic models.

cs.LG

Quantum Assemblage Tomography

A central requirement in asymmetric quantum nonlocality protocols, such as quantum steering, is the precise reconstruction of state assemblages -- statistical ensembles of quantum states correlated with remote classical signals. Here we introduce a generalized loss model for assemblage tomography that uses conical optimization techniques combined with maximum likelihood estimation. Using an evidence-based framework based on Akaike's Information Criterion, we demonstrate that our approach excels in the accuracy of reconstructions while accounting for model complexity. In comparison, standard tomographic methods fall short when applied to experimentally relevant data.

quant-ph

Memorization $\neq$ Understanding: Do Large Language Models Have the Ability of Scenario Cognition?

Driven by vast and diverse textual data, large language models (LLMs) have demonstrated impressive performance across numerous natural language processing (NLP) tasks. Yet, a critical question persists: does their generalization arise from mere memorization of training data or from deep semantic understanding? To investigate this, we propose a bi-perspective evaluation framework to assess LLMs' scenario cognition - the ability to link semantic scenario elements with their arguments in context. Specifically, we introduce a novel scenario-based dataset comprising diverse textual descriptions of fictional facts, annotated with scenario elements. LLMs are evaluated through their capacity to answer scenario-related questions (model output perspective) and via probing their internal representations for encoded scenario elements-argument associations (internal representation perspective). Our experiments reveal that current LLMs predominantly rely on superficial memorization, failing to achieve robust semantic scenario cognition, even in simple cases. These findings expose critical limitations in LLMs' semantic understanding and offer cognitive insights for advancing their capabilities.

cs.CL

Integrating Biological Knowledge for Robust Microscopy Image Profiling on De Novo Cell Lines

High-throughput screening techniques, such as microscopy imaging of cellular responses to genetic and chemical perturbations, play a crucial role in drug discovery and biomedical research. However, robust perturbation screening for \textit{de novo} cell lines remains challenging due to the significant morphological and biological heterogeneity across cell lines. To address this, we propose a novel framework that integrates external biological knowledge into existing pretraining strategies to enhance microscopy image profiling models. Our approach explicitly disentangles perturbation-specific and cell line-specific representations using external biological information. Specifically, we construct a knowledge graph leveraging protein interaction data from STRING and Hetionet databases to guide models toward perturbation-specific features during pretraining. Additionally, we incorporate transcriptomic features from single-cell foundation models to capture cell line-specific representations. By learning these disentangled features, our method improves the generalization of imaging models to \textit{de novo} cell lines. We evaluate our framework on the RxRx database through one-shot fine-tuning on an RxRx1 cell line and few-shot fine-tuning on cell lines from the RxRx19a dataset. Experimental results demonstrate that our method enhances microscopy image profiling for \textit{de novo} cell lines, highlighting its effectiveness in real-world phenotype-based drug discovery applications.

cs.CV

Efficient learning and optimizing non-Gaussian correlated noise in digitally controlled qubit systems

Precise qubit control in the presence of spatio-temporally correlated noise is pivotal for transitioning to fault-tolerant quantum computing. Generically, such noise can also have non-Gaussian statistics, which hampers existing non-Markovian noise spectroscopy protocols. By utilizing frame-based characterization and a novel symmetry analysis, we show how to achieve higher-order spectral estimation for noise-optimized circuit design. Remarkably, we find that the digitally driven qubit dynamics can be solely determined by the complexity of the applied control, rather than the non-perturbative nature of the non-Gaussian environment. This enables us to address certain non-perturbative qubit dynamics more simply. We delineate several complexity bounds for learning such high-complexity noise and demonstrate our single and two-qubit digital characterization and control using a series of numerical simulations. Our results not only provide insights into the exact solvability of (small-sized) open quantum dynamics but also highlight a resource-efficient approach for optimal control and possible error reduction techniques for current qubit devices.

quant-ph

SatHealth: A Multimodal Public Health Dataset with Satellite-based Environmental Factors

Living environments play a vital role in the prevalence and progression of diseases, and understanding their impact on patient's health status becomes increasingly crucial for developing AI models. However, due to the lack of long-term and fine-grained spatial and temporal data in public and population health studies, most existing studies fail to incorporate environmental data, limiting the models' performance and real-world application. To address this shortage, we developed SatHealth, a novel dataset combining multimodal spatiotemporal data, including environmental data, satellite images, all-disease prevalences estimated from medical claims, and social determinants of health (SDoH) indicators. We conducted experiments under two use cases with SatHealth: regional public health modeling and personal disease risk prediction. Experimental results show that living environmental information can significantly improve AI models' performance and temporal-spatial generalizability on various tasks. Finally, we deploy a web-based application to provide an exploration tool for SatHealth and one-click access to both our data and regional environmental embedding to facilitate plug-and-play utilization. SatHealth is now published with data in Ohio, and we will keep updating SatHealth to cover the other parts of the US. With the web application and published code pipeline, our work provides valuable angles and resources to include environmental data in healthcare research and establishes a foundational framework for future research in environmental health informatics.

cs.LG

Simultaneous estimations of quantum state and detector through multiple quantum processes

The estimation of all the parameters in an unknown quantum state or measurement device, commonly known as quantum state tomography (QST) and quantum detector tomography (QDT), is crucial for comprehensively characterizing and controlling quantum systems. In this paper, we introduce a framework, in two different bases, that utilizes multiple quantum processes to simultaneously identify a quantum state and a detector. We develop a closed-form algorithm for this purpose and prove that the mean squared error (MSE) scales as $O(1/N) $ for both QST and QDT, where $N $ denotes the total number of state copies. This scaling aligns with established patterns observed in previous works that addressed QST and QDT as independent tasks. Furthermore, we formulate the problem as a sum of squares (SOS) optimization problem with semialgebraic constraints, where the physical constraints of the state and detector are characterized by polynomial equalities and inequalities. The effectiveness of our proposed methods is validated through numerical examples.

quant-ph

Open-Set Heterogeneous Domain Adaptation: Theoretical Analysis and Algorithm

Domain adaptation (DA) tackles the issue of distribution shift by learning a model from a source domain that generalizes to a target domain. However, most existing DA methods are designed for scenarios where the source and target domain data lie within the same feature space, which limits their applicability in real-world situations. Recently, heterogeneous DA (HeDA) methods have been introduced to address the challenges posed by heterogeneous feature space between source and target domains. Despite their successes, current HeDA techniques fall short when there is a mismatch in both feature and label spaces. To address this, this paper explores a new DA scenario called open-set HeDA (OSHeDA). In OSHeDA, the model must not only handle heterogeneity in feature space but also identify samples belonging to novel classes. To tackle this challenge, we first develop a novel theoretical framework that constructs learning bounds for prediction error on target domain. Guided by this framework, we propose a new DA method called Representation Learning for OSHeDA (RL-OSHeDA). This method is designed to simultaneously transfer knowledge between heterogeneous data sources and identify novel classes. Experiments across text, image, and clinical data demonstrate the effectiveness of our algorithm. Model implementation is available at \url{https://github.com/pth1993/OSHeDA}.

cs.LG

Predictive Modeling with Temporal Graphical Representation on Electronic Health Records

Deep learning-based predictive models, leveraging Electronic Health Records (EHR), are receiving increasing attention in healthcare. An effective representation of a patient's EHR should hierarchically encompass both the temporal relationships between historical visits and medical events, and the inherent structural information within these elements. Existing patient representation methods can be roughly categorized into sequential representation and graphical representation. The sequential representation methods focus only on the temporal relationships among longitudinal visits. On the other hand, the graphical representation approaches, while adept at extracting the graph-structured relationships between various medical events, fall short in effectively integrate temporal information. To capture both types of information, we model a patient's EHR as a novel temporal heterogeneous graph. This graph includes historical visits nodes and medical events nodes. It propagates structured information from medical event nodes to visit nodes and utilizes time-aware visit nodes to capture changes in the patient's health status. Furthermore, we introduce a novel temporal graph transformer (TRANS) that integrates temporal edge features, global positional encoding, and local structural encoding into heterogeneous graph convolution, capturing both temporal and structural information. We validate the effectiveness of TRANS through extensive experiments on three real-world datasets. The results show that our proposed approach achieves state-of-the-art performance.

cs.LG

Broadband spectroscopy of quantum noise

Characterizing noise is key to the optimal control of the quantum system it affects. Using a single-qubit probe and appropriate sequences of $π$ and non-$π$ pulses, we show how one can characterize the noise a quantum bath generates across a wide range of frequencies -- including frequencies below the limit set by the probe's $\mathbb{T}_2$ time. To do so we leverage an exact expression for the dynamics of the probe in the presence of non-$π$ pulses, and a general inequality between the symmetric (classical) and anti-symmetric (quantum) components of the noise spectrum generated by a Gaussian bath. Simulation demonstrates the effectiveness of our method.

quant-ph

A two-stage solution to quantum process tomography: error analysis and optimal design

Quantum process tomography is a critical task for characterizing the dynamics of quantum systems and achieving precise quantum control. In this paper, we propose a two-stage solution for both trace-preserving and non-trace-preserving quantum process tomography. Utilizing a tensor structure, our algorithm exhibits a computational complexity of $O(MLd^2)$ where $d$ is the dimension of the quantum system and $ M $, $ L $ represent the numbers of different input states and measurement operators, respectively. We establish an analytical error upper bound and then design the optimal input states and the optimal measurement operators, which are both based on minimizing the error upper bound and maximizing the robustness characterized by the condition number. Numerical examples and testing on IBM quantum devices are presented to demonstrate the performance and efficiency of our algorithm.

quant-ph

Two-stage solution for ancilla-assisted quantum process tomography: error analysis and optimal design

Quantum process tomography (QPT) is a fundamental task to characterize the dynamics of quantum systems. In contrast to standard QPT, ancilla-assisted process tomography (AAPT) framework introduces an extra ancilla system such that a single input state is needed. In this paper, we extend the two-stage solution, a method originally designed for standard QPT, to perform AAPT. Our algorithm has $O(Md_A^2d_B^2)$ computational complexity where $ M $ is the type number of the measurement operators, $ d_A $ is the dimension of the quantum system of interest, and $d_B$ is the dimension of the ancilla system. Then we establish an error upper bound and further discuss the optimal design on the input state in AAPT. A numerical example on a phase damping process demonstrates the effectiveness of the optimal design and illustrates the theoretical error analysis.

quant-ph

Scalable multiparty steering based on a single pair of entangled qubits

The distribution and verification of quantum nonlocality across a network of users is essential for future quantum information science and technology applications. However, beyond simple point-to-point protocols, existing methods struggle with increasingly complex state preparation for a growing number of parties. Here, we show that, surprisingly, multiparty loophole-free quantum steering, where one party simultaneously steers arbitrarily many spatially separate parties, is achievable by constructing a quantum network from a set of qubits of which only one pair is entangled. Using these insights, we experimentally demonstrate this type of steering between three parties with the detection loophole closed. With its modest and fixed entanglement requirements, this work introduces a scalable approach to rigorously verify quantum nonlocality across multiple parties, thus providing a practical tool towards developing the future quantum internet.

quant-ph