SearcharxivSearch

arXiv subjects

Catherine Lee

Publications and source records attributed to Catherine Lee.

15 recordsLinked to original sources

A Statistical Framework for Understanding Causal Effects that Vary by Treatment Initiation Time in EHR-based Studies

Standard practice in electronic health record (EHR)-based studies evaluating the comparative effectiveness of bariatric surgery relative to no surgery is to estimate and report a constant treatment effect across calendar time. However, real-world treatment strategies can evolve, particularly when comparators include standard of care or surgical procedures where techniques may improve, making it clinically important to ascertain whether efficacy of bariatric surgery has changed over time. Efforts to determine whether treatment efficacy itself is evolving are complicated by changing patient populations, with potential covariate shift in key effect modifiers. Through a comprehensive analysis of EHR data from Kaiser Permanente following two bariatric surgical procedures compared to standard of care, we develop a statistical framework to estimate calendar time-specific average treatment effects and describe both how and why effects vary across treatment initiation time in EHR-based studies. Our approach projects doubly robust, time-specific treatment effect estimates onto candidate marginal structural models and uses a model selection procedure to best describe how effects vary by treatment initiation time. We further introduce a novel summary metric, based on standardization analysis, to quantify the role of covariate shift in explaining observed effect changes and disentangle changes in treatment effects from changes in the patient population receiving treatment.

stat.ME

Robust Causal Inference for EHR-based Studies of Point Exposures with Missingness in Eligibility Criteria

Missingness in variables that define study eligibility criteria is a seldom addressed challenge in electronic health record (EHR)-based settings. It is typically the case that patients with incomplete eligibility information are excluded from analysis without consideration of (implicit) assumptions that are being made, leaving study conclusions subject to potential selection bias. In an effort to ascertain eligibility for more patients, researchers may look back further in time prior to study baseline, and in using outdated values of eligibility-defining covariates may inappropriately be including individuals who, unbeknownst to the researcher, fail to meet eligibility at baseline. To the best of our knowledge, however, very little work has been done to mitigate these concerns. We propose a robust and efficient estimator of the causal average treatment effect on the treated, defined in the study eligible population, in cohort studies where eligibility-defining covariates are missing at random. The approach facilitates the use of flexible machine-learning strategies for component nuisance functions while maintaining appropriate convergence rates for valid asymptotic inference. This method is directly motivated by, and applied throughout to EHR data from Kaiser Permanente to analyze differences between two common bariatric surgical interventions for long-term weight and glycemic outcomes among a cohort of severely obese patients with type II diabetes mellitus.

stat.ME

Enhancing Phishing Email Identification with Large Language Models

Phishing has long been a common tactic used by cybercriminals and continues to pose a significant threat in today's digital world. When phishing attacks become more advanced and sophisticated, there is an increasing need for effective methods to detect and prevent them. To address the challenging problem of detecting phishing emails, researchers have developed numerous solutions, in particular those based on machine learning (ML) algorithms. In this work, we take steps to study the efficacy of large language models (LLMs) in detecting phishing emails. The experiments show that the LLM achieves a high accuracy rate at high precision; importantly, it also provides interpretable evidence for the decisions.

cs.CR

Adjusting for Selection Bias Due to Missing Eligibility Criteria in Emulated Target Trials

Target trial emulation (TTE) is a popular framework for observational studies based on electronic health records (EHR). A key component of this framework is determining the patient population eligible for inclusion in both a target trial of interest and its observational emulation. Missingness in variables that define eligibility criteria, however, presents a major challenge towards determining the eligible population when emulating a target trial with an observational study. In practice, patients with incomplete data are almost always excluded from analysis despite the possibility of selection bias, which can arise when subjects with observed eligibility data are fundamentally different than excluded subjects. Despite this, to the best of our knowledge, very little work has been done to mitigate this concern. In this paper, we propose a novel conceptual framework to address selection bias in TTE studies, tailored towards time-to-event endpoints, and describe estimation and inferential procedures via inverse probability weighting (IPW). Under an EHR-based simulation infrastructure, developed to reflect the complexity of EHR data, we characterize common settings under which missing eligibility data poses the threat of selection bias and investigate the ability of the proposed methods to address it. Finally, using EHR databases from Kaiser Permanente, we demonstrate the use of our method to evaluate the effect of bariatric surgery on microvascular outcomes among a cohort of severely obese patients with Type II diabetes mellitus (T2DM).

stat.ME

Causal Quantile Treatment Effects with missing data by double-sampling

Causal weighted quantile treatment effects (WQTE) are a useful complement to standard causal contrasts that focus on the mean when interest lies at the tails of the counterfactual distribution. To-date, however, methods for estimation and inference regarding causal WQTEs have assumed complete data on all relevant factors. In most practical settings, however, data will be missing or incomplete data, particularly when the data are not collected for research purposes, as is the case for electronic health records and disease registries. Furthermore, such data sources may be particularly susceptible to the outcome data being missing-not-at-random (MNAR). In this paper, we consider the use of double-sampling, through which the otherwise missing data are ascertained on a sub-sample of study units, as a strategy to mitigate bias due to MNAR data in the estimation of causal WQTEs. With the additional data in-hand, we present identifying conditions that do not require assumptions regarding missingness in the original data. We then propose a novel inverse-probability weighted estimator and derive its asymptotic properties, both pointwise at specific quantiles and uniformly across a range of quantiles over some compact subset of (0,1), allowing the propensity score and double-sampling probabilities to be estimated. For practical inference, we develop a bootstrap method that can be used for both pointwise and uniform inference. A simulation study is conducted to examine the finite sample performance of the proposed estimators. The proposed method is illustrated with data from an EHR-based study examining the relative effects of two bariatric surgery procedures on BMI loss at 3 years post-surgery.

stat.ME

Development of a Boston-area 50-km fiber quantum network testbed

Distributing quantum information between remote systems will necessitate the integration of emerging quantum components with existing communication infrastructure. This requires understanding the channel-induced degradations of the transmitted quantum signals, beyond the typical characterization methods for classical communication systems. Here we report on a comprehensive characterization of a Boston-Area Quantum Network (BARQNET) telecom fiber testbed, measuring the time-of-flight, polarization, and phase noise imparted on transmitted signals. We further design and demonstrate a compensation system that is both resilient to these noise sources and compatible with integration of emerging quantum memory components on the deployed link. These results have utility for future work on the BARQNET as well as other quantum network testbeds in development, enabling near-term quantum networking demonstrations and informing what areas of technology development will be most impactful in advancing future system capabilities.

quant-ph

Minimum coprime graph labelings

A coprime labeling of a graph $G$ is a labeling of the vertices of $G$ with distinct integers from $1$ to $k$ such that adjacent vertices have coprime labels. The minimum coprime number of $G$ is the least $k$ for which such a labeling exists. In this paper, we determine the minimum coprime number for several well-studied classes of graphs, including the coronas of complete graphs with empty graphs and the joins of two paths. In particular, we resolve a conjecture of Seoud, El Sonbaty, and Mahran and two conjectures of Asplund and Fox. We also provide an asymptotic for the minimum coprime number of the Erdős-Rényi random graph.

math.CO

Expected Chromatic Number of Random Subgraphs

Given a graph $G$ and $p \in [0,1]$, let $G_p$ denote the random subgraph of $G$ obtained by keeping each edge independently with probability $p$. Alon, Krivelevich, and Sudokov proved $\mathbb{E} [χ(G_p)] \geq C_p \frac{χ(G)}{\log |V(G)|}$, and Bukh conjectured an improvement of $\mathbb{E}[χ(G_p)] \geq C_p \frac{χ(G)}{\log χ(G)}$. We prove a new spectral lower bound on $\mathbb{E}[χ(G_p)]$, as progress towards Bukh's conjecture. We also propose the stronger conjecture that for any fixed $p \leq 1/2$, among all graphs of fixed chromatic number, $\mathbb{E}[χ(G_p)]$ is minimized by the complete graph. We prove this stronger conjecture when $G$ is planar or $χ(G) < 4$. We also consider weaker lower bounds on $\mathbb{E}[χ(G_p)]$ proposed in a recent paper by Shinkar; we answer two open questions of Shinkar negatively and propose a possible refinement of one of them.

math.CO

Incidence geometry and universality in the tropical plane

We examine the incidence geometry of lines in the tropical plane. We prove tropical analogs of the Sylvester-Gallai and Motzkin-Rabin theorems in classical incidence geometry. This study leads naturally to a discussion of the realizability of incidence data of tropical lines. Drawing inspiration from the von Staudt constructions and Mnëv's universality theorem, we prove that determining whether a given tropical linear incidence datum is realizable by a tropical line arrangement requires solving an arbitrary linear programming problem over the integers.

math.CO

SemiCompRisks: An R Package for Independent and Cluster-Correlated Analyses of Semi-Competing Risks Data

Semi-competing risks refer to the setting where primary scientific interest lies in estimation and inference with respect to a non-terminal event, the occurrence of which is subject to a terminal event. In this paper, we present the R package SemiCompRisks that provides functions to perform the analysis of independent/clustered semi-competing risks data under the illness-death multi-state model. The package allows the user to choose the specification for model components from a range of options giving users substantial flexibility, including: accelerated failure time or proportional hazards regression models; parametric or non-parametric specifications for baseline survival functions; parametric or non-parametric specifications for random effects distributions when the data are cluster-correlated; and, a Markov or semi-Markov specification for terminal event following non-terminal event. While estimation is mainly performed within the Bayesian paradigm, the package also provides the maximum likelihood estimation for select parametric models. The package also includes functions for univariate survival analysis as complementary analysis tools.

stat.CO

Metropolitan quantum key distribution with silicon photonics

Photonic integrated circuits (PICs) provide a compact and stable platform for quantum photonics. Here we demonstrate a silicon photonics quantum key distribution (QKD) transmitter in the first high-speed polarization-based QKD field tests. The systems reach composable secret key rates of 950 kbps in a local test (on a 103.6-m fiber with a total emulated loss of 9.2 dB) and 106 kbps in an intercity metropolitan test (on a 43-km fiber with 16.4 dB loss). Our results represent the highest secret key generation rate for polarization-based QKD experiments at a standard telecom wavelength and demonstrate PICs as a promising, scalable resource for future formation of metropolitan quantum-secure communications networks.

quant-ph

High-rate field demonstration of large-alphabet quantum key distribution

Quantum key distribution (QKD) exploits the quantum nature of light to share provably secure keys, allowing secure communication in the presence of an eavesdropper. The first QKD schemes used photons encoded in two states, such as polarization. Recently, much effort has turned to large-alphabet QKD schemes, which encode photons in high-dimensional basis states. Compared to binary-encoded QKD, large-alphabet schemes can encode more secure information per detected photon, boosting secure communication rates, and also provide increased resilience to noise and loss. High-dimensional encoding may also improve the efficiency of other quantum information processing tasks, such as performing Bell tests and implementing quantum gates. Here, we demonstrate a large-alphabet QKD protocol based on high-dimensional temporal encoding. We achieve record secret-key rates and perform the first field demonstration of large-alphabet QKD. This demonstrates a new, practical way to optimize secret-key rates and marks an important step towards transmission of high-dimensional quantum states in deployed networks.

quant-ph

On-Chip Detection of Entangled Photons by Scalable Integration of Single-Photon Detectors

Photonic integrated circuits (PICs) have emerged as a scalable platform for complex quantum technologies using photonic and atomic systems. A central goal has been to integrate photon-resolving detectors to reduce optical losses, latency, and wiring complexity associated with off-chip detectors. Superconducting nanowire single-photon detectors (SNSPDs) are particularly attractive because of high detection efficiency, sub-50-ps timing jitter, nanosecond-scale reset time, and sensitivity from the visible to the mid-infrared spectrum. However, while single SNSPDs have been incorporated into individual waveguides, the system efficiency of multiple SNSPDs in one photonic circuit has been limited below 0.2% due to low device yield. Here we introduce a micrometer-scale flip-chip process that enables scalable integration of SNSPDs on a range of PICs. Ten low-jitter detectors were integrated on one PIC with 100% device yield. With an average system efficiency beyond 10% for multiple SNSPDs on one PIC, we demonstrate high-fidelity on-chip photon correlation measurements of non-classical light.

physics.optics

Finite-key analysis of high-dimensional time-energy entanglement-based quantum key distribution

We present a security analysis against collective attacks for the recently proposed time-energy entanglement-based quantum key distribution protocol, given the practical constraints of single photon detector efficiency, channel loss, and finite-key considerations. We find a positive secure-key capacity when the key length increases beyond $10^4$ for eight-dimensional systems. The minimum key length required is reduced by the ability to post-select on coincident single-photon detection events. Including finite-key effects, we show the ability to establish a shared secret key over a 200 km fiber link.

quant-ph

High-dimensional quantum key distribution using dispersive optics

We propose a high-dimensional quantum key distribution (QKD) protocol that employs temporal correlations of entangled photons. The security of the protocol relies on measurements by Alice and Bob in one of two conjugate bases, implemented using dispersive optics. We show that this dispersion-based approach is secure against general coherent attacks. The protocol is additionally compatible with standard fiber telecommunications channels and wavelength division multiplexers. We offer multiple implementations to enhance the transmission rate and describe a heralded qudit source that is easy to implement and enables secret-key generation at up to 100 Mbps at over 2 bits per photon.

quant-ph