SearcharxivSearch

arXiv subjects

Juhee Lee

Publications and source records attributed to Juhee Lee.

17 recordsLinked to original sources

Completing or Refusing Low-Dimensional Records of Structured Quantum Circuits: Measurement Loss, Compression Loss, and Hardware Drift

Structured quantum circuits are often summarized by low-dimensional records. Records lose control-dependent information when they merge outcomes whose probabilities respond differently to circuit parameters. We assess this loss without a parametric hardware-noise model or low-dimensional statistical family. Quantum Fisher information bounds premeasurement sensitivity; Fisher metrics of measurements and records describe device-attainable sensitivity. A Hellinger residual measures the response removed by a record, while its local quadratic term is the conditional covariance of the full-outcome score. Simultaneous confidence bounds support approval, refusal, or deferral. We prove a finite-library completion theorem: executable augmentations terminate with either a record preserving every declared response or proof that no library augmentation removes the loss. For an analytic control germ, integral closures characterize preservation along every analytic control arc, and finitely many Rees valuations detect failure. This state--measurement--record chain connects an all-arc criterion to executable completion or refusal, an attainable-Fisher local kernel criterion, and finite-sample decisions. A three-qubit calculation separates measurement loss from record loss. IBM experiments on Kingston and Marrakesh test decisions in fixed-particle-number and GHZ families, including negative controls. An end-to-end Kingston experiment reduces calibration shots by 33.3% while meeting prespecified noninferiority criteria on two held-out objectives. In two-epoch IQM Garnet data, median record-level drift is 0.0797 times full-distribution drift, but a significant residual remains in 35 of 36 settings. These experiments validate failure detection on the tested circuits; they do not establish universal compression performance or device quantum Fisher information.

quant-ph

Genetic association testing with multivariate survival phenotypes under interval censoring

Set-based genetic association tests provide a powerful framework for detecting genetic effects on complex traits by jointly analyzing multiple genetic variants. Although set-based methods have been developed for interval-censored survival outcomes, existing approaches primarily focus on a single survival phenotype and therefore do not fully use information from multiple correlated outcomes. In this paper, we develop two Weighted V Tests for Multivariate Interval-Censored Data (WV-M-IC), extending the weighted V-statistic framework (Wu et al., 2021) to the joint analysis of multiple correlated interval-censored survival outcomes. The performance of these methods is evaluated through simulation studies, showing that the proposed approaches can provide power gains compared with single-outcome analyses. We apply the proposed methods to the ZOE 2.0 study to investigate dental caries progression in children.

stat.ME

Exploring Academic Influence of Algorithms by Co-occurrence Network Based on Full-text of Academic Papers

Algorithms have become central to scientific research in the era of artificial intelligence (AI). Although algorithm mentions in papers are often used to indicate popularity and influence, existing studies usually evaluate individual algorithms in isolation and pay limited attention to the collective influence formed through their interconnections. This study constructs large-scale algorithm co-occurrence networks in natural language processing (NLP) based on the full text of academic papers and investigates algorithm influence from a network perspective. Using deep learning models, we extract algorithm entities and build overall, cumulative, and annual co-occurrence networks. We analyze their structural characteristics and apply multiple centrality measures to assess the group influence of algorithms across the whole field and over time. The results show that algorithm networks display typical features of complex networks, with increasingly dense connections developing over approximately two decades. Classic, high-performing algorithms and those located at the intersections of different research periods tend to have high popularity, control, centrality, and balanced influence. When the influence of an algorithm declines, it usually loses its core network position first, followed by weaker associations with other algorithms. This study is the first large-scale analysis of algorithm co-occurrence networks. Covering more than four decades of academic publications, it provides a temporal and structural view of algorithm influence and offers a foundation for future research on networks linking algorithms, scholars, and tasks.

cs.AI

Beyond average: heterogeneous first-passage dynamics in many-particle systems with resetting

We study how stochastic resetting affects first-passage processes in systems of many interacting particles. While resetting is well understood for single-particle dynamics, its consequences for collective behavior remain less clear. We consider a protocol in which all surviving particles are reset to the position of the most extreme one, motivated by problems in artificial selection and avoidance. Using stochastic simulations of particles diffusing in a confining potential with an absorbing boundary, we examine two notions of arrival: when the first particle reaches the boundary and the point at which half of the particles do. We find that resetting produces broad distributions of arrival times with heavy tails and extended plateaus that span several orders of magnitude. As the resetting rate increases, the mean arrival time grows and diverges beyond a threshold. Trajectory-level analysis also reveals strong heterogeneity, with very short and very long absorption times. These results show that collective resetting lacks a single characteristic time scale and that the definition of arrival is crucial for understanding and controlling such systems.

cond-mat.stat-mech

Bayesian Covariate-Varying Interaction Analysis for Multivariate Count Data: Application to Microbiome Studies

Understanding covariate-varying interdependencies among features is of great interest in various applications. Motivated by microbiome studies where microbial abundances and interactions vary with environmental factors, we develop a Bayesian covariate-varying factor model. This model flexibly estimates heteroscedasticity in the covariance matrix as a function of covariates. Specifically, our approach employs covariance regression through linear regression on a lower-dimensional factor loading matrix. This formulation, combined with joint sparsity induced by the Dirichlet--Horseshoe prior for the factor loadings, provides robust estimation of covariate-varying covariance in high-dimensional settings. The model simultaneously incorporates a regression structure for the mean abundance and jointly addresses the covariate-varying mean and covariance structure. Furthermore, the model tackles key statistical challenges such as discreteness, over-dispersion, compositionality, and high dimensionality, common in microbiome data analysis, using a flexible nonparametric Bayesian framework. We thoroughly investigate the properties of the model and conduct extensive simulation studies to examine its performance. Real microbiome data examples are provided for illustration.

stat.ME

RBF-Solver: A Multistep Sampler for Diffusion Probabilistic Models via Radial Basis Functions

Diffusion probabilistic models (DPMs) are widely adopted for their outstanding generative fidelity, yet their sampling is computationally demanding. Polynomial-based multistep samplers mitigate this cost by accelerating inference; however, despite their theoretical accuracy guarantees, they generate the sampling trajectory according to a predefined scheme, providing no flexibility for further optimization. To address this limitation, we propose RBF-Solver, a multistep diffusion sampler that interpolates model evaluations with Gaussian radial basis functions (RBFs). By leveraging learnable shape parameters in Gaussian RBFs, RBF-Solver explicitly follows optimal sampling trajectories. At first order, it reduces to the Euler method (DDIM). At second order or higher, as the shape parameters approach infinity, RBF-Solver converges to the Adams method, ensuring its compatibility with existing samplers. Owing to the locality of Gaussian RBFs, RBF-Solver maintains high image fidelity even at fourth order or higher, where previous samplers deteriorate. For unconditional generation, RBF-Solver consistently outperforms polynomial-based samplers in the high-NFE regime (NFE >= 15). On CIFAR-10 with the Score-SDE model, it achieves an FID of 2.87 with 15 function evaluations and further improves to 2.48 with 40 function evaluations. For conditional ImageNet 256 x 256 generation with the Guided Diffusion model at a guidance scale 8.0, substantial gains are achieved in the low-NFE range (5-10), yielding a 16.12-33.73% reduction in FID relative to polynomial-based samplers.

cs.LG

A General Bayesian Nonparametric Approach for Estimating Population-Level and Conditional Causal Effects

We propose a Bayesian nonparametric (BNP) approach to causal inference using observational data consisting of outcome, treatment, and a set of confounders. The conditional distribution of the outcome given treatment and confounders is modeled flexibly using a dependent nonparametric mixture model, in which both the atoms and the weights vary with the confounders. The proposed BNP model is well suited for causal inference problems, as it does not rely on parametric assumptions about how the conditional distribution depends on the confounders. In particular, the model effectively adjusts for confounding and improves the modeling of treatment effect heterogeneity, leading to more accurate estimation of both the average treatment effect (ATE) and heterogeneous treatment effects (HTE). Posterior inference under the proposed model is computationally efficient due to the use of data augmentation. Extensive evaluations demonstrate that the proposed model offers competitive or superior performance compared to a wide range of recent methods spanning various statistical approaches, including Bayesian additive regression tree (BART) models, which are well known for their strong empirical performance. More importantly, the model provides fully probabilistic inference on quantities of interest that other methods cannot easily provide, using their posterior distributions.

stat.ME

General Theory for Group Resetting with Application to Avoidance

We present a general theoretical framework for group resetting dynamics in a potential landscape. While traditional resetting models typically focus on a single particle, we consider a group of particles whose collective dynamics govern the resetting. We extend existing resetting theories to cover extreme-value group resetting. This has applications from bacterial evolution under antibiotic pressure to swarm-search optimization. Using renewal theory, we derive a Fokker-Planck equation for the spatial distribution of the group's center of mass, treated as an effective particle. This formalism yields analytical expressions for key observables such as the stationary mean position and variance. We also study a group avoidance problem, where the particles must avoid an undesirable region. Such problems have recently been studied in contexts such as preventing critically high water levels in dams and controlling excessive financial leverage. Our framework offers new insight into how resetting can optimize group-level search and avoidance strategies.

cond-mat.stat-mech

Natural Language Programming in Medicine: Administering Evidence Based Clinical Workflows with Autonomous Agents Powered by Generative Large Language Models

Generative Large Language Models (LLMs) hold significant promise in healthcare, demonstrating capabilities such as passing medical licensing exams and providing clinical knowledge. However, their current use as information retrieval tools is limited by challenges like data staleness, resource demands, and occasional generation of incorrect information. This study assessed the potential of LLMs to function as autonomous agents in a simulated tertiary care medical center, using real-world clinical cases across multiple specialties. Both proprietary and open-source LLMs were evaluated, with Retrieval Augmented Generation (RAG) enhancing contextual relevance. Proprietary models, particularly GPT-4, generally outperformed open-source models, showing improved guideline adherence and more accurate responses with RAG. The manual evaluation by expert clinicians was crucial in validating models' outputs, underscoring the importance of human oversight in LLM operation. Further, the study emphasizes Natural Language Programming (NLP) as the appropriate paradigm for modifying model behavior, allowing for precise adjustments through tailored prompts and real-world interactions. This approach highlights the potential of LLMs to significantly enhance and supplement clinical decision-making, while also emphasizing the value of continuous expert involvement and the flexibility of NLP to ensure their reliability and effectiveness in healthcare settings.

cs.AI

Bayesian Nonparametric Erlang Mixture Modeling for Survival Analysis

We develop a flexible Erlang mixture model for survival analysis. The model for the survival density is built from a structured mixture of Erlang densities, mixing on the integer shape parameter with a common scale parameter. The mixture weights are constructed through increments of a distribution function on the positive real line, which is assigned a Dirichlet process prior. The model has a relatively simple structure, balancing flexibility with efficient posterior computation. Moreover, it implies a mixture representation for the hazard function that involves time-dependent mixture weights, thus offering a general approach to hazard estimation. We extend the model to handle survival responses corresponding to multiple experimental groups, using a dependent Dirichlet process prior for the group-specific distributions that define the mixture weights. Model properties, prior specification, and posterior simulation are discussed, and the methodology is illustrated with synthetic and real data examples.

stat.ME

A Bayesian Feature Allocation Model for Identification of Cell Subpopulations Using Cytometry Data

A Bayesian feature allocation model (FAM) is presented for identifying cell subpopulations based on multiple samples of cell surface or intracellular marker expression level data obtained by cytometry by time of flight (CyTOF). Cell subpopulations are characterized by differences in expression patterns of makers, and individual cells are clustered into the subpopulations based on the patterns of their observed expression levels. A finite Indian buffet process is used to model subpopulations as latent features, and a model-based method based on these latent feature subpopulations is used to construct cell clusters within each sample. Non-ignorable missing data due to technical artifacts in mass cytometry instruments are accounted for by defining a static missing data mechanism. In contrast to conventional cell clustering methods based on observed marker expression levels that are applied separately to different samples, the FAM based method can be applied simultaneously to multiple samples, and can identify important cell subpopulations likely to be missed by conventional clustering. The proposed FAM based method is applied to jointly analyze three datasets, generated by CyTOF, to study natural killer (NK) cells. Because the subpopulations identified by the FAM may define novel NK cell subsets, this statistical analysis may provide useful information about the biology of NK cells and their potential role in cancer immunotherapy which may lead, in turn, to development of improved cellular therapies. Simulation studies of the proposed method's behavior under two cases of known subpopulations also are presented, followed by analysis of the CyTOF NK cell surface marker data.

stat.AP

Smartphone-based point-of-care lipid blood test performance evaluation compared with a clinical diagnostic laboratory method

Managing blood lipid levels is important for the treatment and prevention of diabetes, cardiovascular disease, and obesity. An easy-to-use, portable lipid blood test will accelerate more frequent testing by patients and at-risk populations. We used smartphone systems that are already familiar to many people. Because smartphone systems can be carried around everywhere, blood can be measured easily and frequently. We compared the results of lipid tests with those of existing clinical diagnostic laboratory methods. We found that smartphone-based point-of-care lipid blood tests are as accurate as hospital-grade laboratory tests. Our system will be useful for those who need to manage blood lipid levels to motivate them to track and control their behavior.

q-bio.QM

Induced interactions in the BCS-BEC crossover of two-dimensional Fermi gases with Rashba spin-orbit coupling

We investigate the Gorkov--Melik-Barkhudarov (GM) correction to superfluid transition temperature in two-dimensional Fermi gases with Rashba spin-orbit coupling (SOC) across the SOC-driven BCS-BEC crossover. In the calculation of the induced interaction, we find that the spin-component mixing due to SOC can induce both of the conventional screening and additional antiscreening contributions that interplay significantly in the strong SOC regime. While the GM correction generally lowers the estimate of transition temperature, it turns out that at a fixed weak interaction, the correction effect exhibits a crossover behavior where the ratio between the estimates without and with the correction first decreases with SOC and then becomes insensitive to SOC when it goes into the strong SOC regime. We demonstrate the applicability of the GM correction by comparing the zero-temperature condensate fraction with the recent quantum Monte Carlo results.

cond-mat.quant-gas

Comment on Influence of induced interactions on superfluid properties of quasi-two-dimensional dilute Fermi gases with spin-orbit coupling

In an article in 2013, Caldas et al. [Phys. Rev. A 88, 023615 (2013)] derived analytical expressions of the induced interaction within the scheme of Gorkov and Melik-Barkhudrov in quasi-two-dimensional Fermi gases with Rashba spin-orbit coupling (SOC). They claimed that the induced interaction is exactly the same as the one for the case without SOC when the SOC is weak, and in the region of strong SOC, it starts from a reduced value and then recovers the value for the zero SOC in the limit of large SOC. We point out that their calculations contain the critical errors and inconsistencies that significantly affect the basis of these claims.

cond-mat.quant-gas

A Bayesian feature allocation model for tumor heterogeneity

We develop a feature allocation model for inference on genetic tumor variation using next-generation sequencing data. Specifically, we record single nucleotide variants (SNVs) based on short reads mapped to human reference genome and characterize tumor heterogeneity by latent haplotypes defined as a scaffold of SNVs on the same homologous genome. For multiple samples from a single tumor, assuming that each sample is composed of some sample-specific proportions of these haplotypes, we then fit the observed variant allele fractions of SNVs for each sample and estimate the proportions of haplotypes. Varying proportions of haplotypes across samples is evidence of tumor heterogeneity since it implies varying composition of cell subpopulations. Taking a Bayesian perspective, we proceed with a prior probability model for all relevant unknown quantities, including, in particular, a prior probability model on the binary indicators that characterize the latent haplotypes. Such prior models are known as feature allocation models. Specifically, we define a simplified version of the Indian buffet process, one of the most traditional feature allocation models. The proposed model allows overlapping clustering of SNVs in defining latent haplotypes, which reflects the evolutionary process of subclonal expansion in tumor samples.

stat.AP

First-order phase transition and tricritical scaling behavior of the Blume-Capel model: a Wang-Landau sampling approach

We investigate the tricritical scaling behavior of the two-dimensional spin-$1$ Blume-Capel model using the Wang-Landau method measuring the joint density of states for lattice sizes up to $48\times 48$ sites. The first-order transition curve is systematically determined employing the method of field mixing in conjunction with finite-size scaling, showing a significant deviation from the previous data points. Deep in the first-order area of the phase diagram, we also find that the specific heat exhibits a double-peak structure of the Schottky-like anomaly appearing with the transition peak. At the tricritical point, we characterize the tricritical exponents through finite-size scaling analysis including the phenomenological finite-size scaling with thermodynamic variables. Our estimation of the tricritical eigenvalue exponents, $y_t = 1.804(5)$, $y_g = 0.80(1)$, and $y_h = 1.925(3)$, provides the first Wang-Landau verification of the conjectured exact values, demonstrating the effectiveness of the density-of-states-based approach in finite-size scaling study of multicritical phenomena.

cond-mat.stat-mech

Bayesian Inference for Tumor Subclones Accounting for Sequencing and Structural Variants

Tumor samples are heterogeneous. They consist of different subclones that are characterized by differences in DNA nucleotide sequences and copy numbers on multiple loci. Heterogeneity can be measured through the identification of the subclonal copy number and sequence at a selected set of loci. Understanding that the accurate identification of variant allele fractions greatly depends on a precise determination of copy numbers, we develop a Bayesian feature allocation model for jointly calling subclonal copy numbers and the corresponding allele sequences for the same loci. The proposed method utilizes three random matrices, L, Z and w to represent subclonal copy numbers (L), numbers of subclonal variant alleles (Z) and cellular fractions of subclones in samples (w), respectively. The unknown number of subclones implies a random number of columns for these matrices. We use next-generation sequencing data to estimate the subclonal structures through inference on these three matrices. Using simulation studies and a real data analysis, we demonstrate how posterior inference on the subclonal structure is enhanced with the joint modeling of both structure and sequencing variants on subclonal genomes. Software is available at http://compgenome.org/BayClone2.

stat.ME