SearcharxivSearch

arXiv subjects

Chihoon Lee

Publications and source records attributed to Chihoon Lee.

14 recordsLinked to original sources

Structure-Aware Variational State Preparation for Quantum Basket Option Pricing

Basket option pricing often relies on Monte Carlo estimation, for which quantum amplitude estimation (QAE) provides a quadratic speed-up. However, the practical benefit of QAE can be limited by the depth of the state-preparation circuit. We propose a structure-aware quantum state-preparation framework for QAE-based basket option pricing. The framework uses tensor-train (TT) rank information to design shallow variational state-preparation circuits. In the independent regime, TT ranks remove unnecessary entangling links from a hardware-efficient ansatz. In correlated basket settings, we instead prepare asset-wise marginals locally and train a compact latent block to match the basket cumulative distribution function. The Basket-CDF objective targets the basket pushforward distribution rather than the full joint state, directly aligning state preparation with basket-dependent payoffs. Numerical experiments show that the proposed circuits replace the exponential state-preparation depth scaling of exact amplitude loading with linear scaling, while maintaining low-percent basket-pricing errors. Additional sampling-based training experiments and an end-to-end QAE integration study support compatibility with sample-estimated training and standard QAE-based pricing workflows.

quant-ph

Predicting Current Outcomes From Historical Survey Data With Weighted Conformal Prediction

In large-scale complex surveys such as the National Health and Nutrition Examination Survey (NHANES), some outcomes are measured only in selected years, leaving incomplete records across survey waves. We develop a weighted conformal prediction framework that enables valid population-level prediction of unobserved outcomes using information from earlier surveys. The method accommodates covariate shift, where both continuous and categorical covariate distributions evolve over time while survey design affects representativeness. It integrates subgroup-specific density ratio and subgroup-proportion estimation to approximate likelihood ratios between the historical and target covariate distributions, and we establish coverage guarantees for the resulting prediction sets. Simulation studies and an application predicting low-density lipoprotein cholesterol (LDL-C) for the current U.S. population show that the proposed approach achieves coverage close to the nominal level and improved efficiency over existing methods, particularly when covariate distributions are complex or unknown.

stat.ME

Towards Fully-Automated Materials Discovery via Large-Scale Synthesis Dataset and Expert-Level LLM-as-a-Judge

Materials synthesis is vital for innovations such as energy storage, catalysis, electronics, and biomedical devices. Yet, the process relies heavily on empirical, trial-and-error methods guided by expert intuition. Our work aims to support the materials science community by providing a practical, data-driven resource. We have curated a comprehensive dataset of 17K expert-verified synthesis recipes from open-access literature, which forms the basis of our newly developed benchmark, AlchemyBench. AlchemyBench offers an end-to-end framework that supports research in large language models applied to synthesis prediction. It encompasses key tasks, including raw materials and equipment prediction, synthesis procedure generation, and characterization outcome forecasting. We propose an LLM-as-a-Judge framework that leverages large language models for automated evaluation, demonstrating strong statistical agreement with expert assessments. Overall, our contributions offer a supportive foundation for exploring the capabilities of LLMs in predicting and guiding materials synthesis, ultimately paving the way for more efficient experimental design and accelerated innovation in materials science.

cs.CL

Control in Stochastic Environment with Delays: A Model-based Reinforcement Learning Approach

In this paper we are introducing a new reinforcement learning method for control problems in environments with delayed feedback. Specifically, our method employs stochastic planning, versus previous methods that used deterministic planning. This allows us to embed risk preference in the policy optimization problem. We show that this formulation can recover the optimal policy for problems with deterministic transitions. We contrast our policy with two prior methods from literature. We apply the methodology to simple tasks to understand its features. Then, we compare the performance of the methods in controlling multiple Atari games.

cs.LG

A Method for Detecting Murmurous Heart Sounds based on Self-similar Properties

A heart murmur is an atypical sound produced by the flow of blood through the heart. It can be a sign of a serious heart condition, so detecting heart murmurs is critical for identifying and managing cardiovascular diseases. However, current methods for identifying murmurous heart sounds do not fully utilize the valuable insights that can be gained by exploring intrinsic properties of heart sound signals. To address this issue, this study proposes a new discriminatory set of multiscale features based on the self-similarity and complexity properties of heart sounds, as derived in the wavelet domain. Self-similarity is characterized by assessing fractal behaviors, while complexity is explored by calculating wavelet entropy. We evaluated the diagnostic performance of these proposed features for detecting murmurs using a set of standard classifiers. When applied to a publicly available heart sound dataset, our proposed wavelet-based multiscale features achieved comparable performance to existing methods with fewer features. This suggests that self-similarity and complexity properties in heart sounds could be potential biomarkers for improving the accuracy of murmur detection.

eess.SP

Stationary Distribution Convergence of the Offered Waiting Processes in Heavy Traffic under General Patience Time Scaling

We study a sequence of single server queues with customer abandonment (GI/GI/1+GI) under heavy traffic. The patience time distributions vary with the sequence, which allows for a wider scope of applications. It is known ([20, 18]) that the sequence of scaled offered waiting time processes converges weakly to a reflecting diffusion process with non-linear drift, as the traffic intensity approaches one. In this paper, we further show that the sequence of stationary distributions and moments of the offered waiting times, with diffusion scaling, converge to those of the limit diffusion process. This justifies the stationary performance of the diffusion limit as a valid approximation for the stationary performance of the GI/GI/1+GI queue. Consequently, we also derive the approximation for the abandonment probability for the GI/GI/1+GI queue in the stationary state.

math.PR

Entropy flow and De Bruijn's identity for a class of stochastic differential equations driven by fractional Brownian motion

Motivated by the classical De Bruijn's identity for the additive Gaussian noise channel, in this paper we consider a generalized setting where the channel is modelled via stochastic differential equations driven by fractional Brownian motion with Hurst parameter $H\in(0,1)$. We derive generalized De Bruijn's identity for Shannon entropy and Kullback-Leibler divergence by means of Itô's formula, and present two applications. In the first application we demonstrate its equivalence with Stein's identity for Gaussian distributions, while in the second application, we show that for $H \in (0,1/2]$, the entropy power is concave in time while for $H \in (1/2,1)$ it is convex in time when the initial distribution is Gaussian. Compared with the classical case of $H = 1/2$, the time parameter plays an interesting and significant role in the analysis of these quantities.

math.PR

Stationary Distribution Convergence of the Offered Waiting Processes for GI/GI/1+GI Queues in Heavy Traffic

A result of Ward and Glynn (2005) asserts that the sequence of scaled offered waiting time processes of the $GI/GI/1+GI$ queue converges weakly to a reflected Ornstein-Uhlenbeck process (ROU) in the positive real line, as the traffic intensity approaches one. As a consequence, the stationary distribution of a ROU process, which is a truncated normal, should approximate the scaled stationary distribution of the offered waiting time in a $GI/GI/1+GI$ queue; however, no such result has been proved. We prove the aforementioned convergence, and the convergence of the moments, in heavy traffic, thus resolving a question left open in Ward and Glynn (2005). In comparison to Kingman's classical result in Kingman (1961) showing that an exponential distribution approximates the scaled stationary offered waiting time distribution in a $GI/GI/1$ queue in heavy traffic, our result confirms that the addition of customer abandonment has a non-trivial effect on the queue stationary behavior.

math.PR

On "A General Framework for Pricing Asian Options Under Markov Processes"

Cai, Song and Kou (2015) [Cai, N., Y. Song, S. Kou (2015) A general framework for pricing Asian options under Markov processes. Oper. Res. 63(3): 540-554] made a breakthrough by proposing a general framework for pricing both discretely and continuously monitored Asian options under one-dimensional Markov processes. In this note, under the setting of continuous-time Markov chain (CTMC), we explicitly carry out the inverse Z-transform and the inverse Laplace transform respectively for the discretely and the continuously monitored cases. The resulting explicit single Laplace transforms improve their Theorem 2, p.543, and numerical studies demonstrate the gain in efficiency.

q-fin.PR

Parameter inference and model selection in deterministic and stochastic dynamical models via approximate Bayesian computation: modeling a wildlife epidemic

We consider the problem of selecting deterministic or stochastic models for a biological, ecological, or environmental dynamical process. In most cases, one prefers either deterministic or stochastic models as candidate models based on experience or subjective judgment. Due to the complex or intractable likelihood in most dynamical models, likelihood-based approaches for model selection are not suitable. We use approximate Bayesian computation for parameter estimation and model selection to gain further understanding of the dynamics of two epidemics of chronic wasting disease in mule deer. The main novel contribution of this work is that under a hierarchical model framework we compare three types of dynamical models: ordinary differential equation, continuous time Markov chain, and stochastic differential equation models. To our knowledge model selection between these types of models has not appeared previously. Since the practice of incorporating dynamical models into data models is becoming more common, the proposed approach may be very useful in a variety of applications.

stat.AP

A penalized simulated maximum likelihood approach in parameter estimation for stochastic differential equations

We consider the problem of estimating parameters of stochastic differential equations (SDEs) with discrete-time observations that are either completely or partially observed. The transition density between two observations is generally unknown. We propose an importance sampling approach with an auxiliary parameter when the transition density is unknown. We embed the auxiliary importance sampler in a penalized maximum likelihood framework which produces more accurate and computationally efficient parameter estimates. Simulation studies in three different models illustrate promising improvements of the new penalized simulated maximum likelihood method. The new procedure is designed for the challenging case when some state variables are unobserved and moreover, observed states are sparse over time, which commonly arises in ecological studies. We apply this new approach to two epidemics of chronic wasting disease in mule deer.

stat.ME

On drift parameter estimation for reflected fractional Ornstein-Uhlenbeck processes

We consider a reflected Ornstein-Uhlenbeck process $X$ driven by a fractional Brownian motion with Hurst parameter $H\in (0, \frac12) \cup (\frac12, 1)$. Our goal is to estimate an unknown drift parameter $α\in (-\infty,\infty)$ on the basis of continuous observation of the state process. We establish Girsanov theorem for the process $X$, derive the standard maximum likelihood estimator of the drift parameter $α$, and prove its strong consistency and asymptotic normality. As an improved estimator, we obtain the explicit formulas for the sequential maximum likelihood estimator and its mean squared error by assuming the process is observed until a certain information reaches a specified precision level. The estimator is shown to be unbiased, uniformly normally distributed, and efficient in the mean square error sense.

math.ST

Direction-Projection-Permutation for High Dimensional Hypothesis Tests

Motivated by the prevalence of high dimensional low sample size datasets in modern statistical applications, we propose a general nonparametric framework, Direction-Projection-Permutation (DiProPerm), for testing high dimensional hypotheses. The method is aimed at rigorous testing of whether lower dimensional visual differences are statistically significant. Theoretical analysis under the non-classical asymptotic regime of dimension going to infinity for fixed sample size reveals that certain natural variations of DiProPerm can have very different behaviors. An empirical power study both confirms the theoretical results and suggests DiProPerm is a powerful test in many settings. Finally DiProPerm is applied to a high dimensional gene expression dataset.

stat.ME

Non-Markovian state dependent networks in critical loading

We establish heavy traffic limit theorems for queue-length processes in critically loaded single class queueing networks with state dependent arrival and service rates. A distinguishing feature of our model is non-Markovian state dependence. The limit stochastic process is a continuous-path reflected process on the nonnegative orthant. We give an application to generalised Jackson networks with state-dependent rates.

math.PR