SearcharxivSearch

arXiv subjects

Jeongjin Lee

Publications and source records attributed to Jeongjin Lee.

16 recordsLinked to original sources

Kalman Filtering and Smoothing for Improving Precision in Horvitz--Thompson Estimation of Infectious Disease Prevalence

Horvitz--Thompson (HT) estimators can provide unbiased daily estimates of infectious disease prevalence under repeated surveillance by correcting for nonrandom testing induced by scheduled, symptom-based, and contact-tracing components. However, because each HT estimate is based on the testing data available for that day and may involve highly variable inverse probability weights, it can be noisy, have precision that varies over time, and become unavailable during temporary interruptions in testing. The daily HT estimator is modeled as a noisy observation of an underlying prevalence process, with day-specific observation variances estimated using a delete-a-group jackknife. Our primary specification is a joint local linear trend state-space model that extends the standard level-only random walk by adding a latent slope. The Kalman filter improves precision by borrowing information from past estimates. At each time $t$, the process variances are estimated or carried forward using only observations available through time $t$, so the resulting filtered estimate is available in real time. We also describe the corresponding Kalman smoother as a retrospective extension based on the full observed series. When daily HT estimates are missing, the Kalman filter proceeds through prediction-only updates, whereas the corresponding smoother retrospectively reconstructs those periods using later observations. In simulations, the joint Kalman filter substantially improves precision relative to the raw daily HT estimator while preserving the main temporal pattern, and the smoother provides a more stable retrospective summary. In The Ohio State University's fall 2020 SARS-CoV-2 surveillance data, the filter and smoother provide estimates on no-testing days, when the HT estimator provides neither point nor interval estimates, and yield narrower confidence intervals than HT intervals on observed days.

stat.ME

A Counterfactual Framework for Estimating Infectious Disease Prevalence under Repeated Testing with Symptomatic and Contact-Tracing Components

This paper addresses the problem of estimating infectious disease prevalence under longitudinal testing programs that include scheduled, symptomatic, and contact-tracing testing. Our study is motivated by data from The Ohio State University, where a mandatory once-per-week COVID-19 testing and isolation program was implemented during the Fall 2020 semester, supplemented by additional testing for symptomatic individuals and identified contacts. In this setting, the probability of being tested depends on symptoms or contact-tracing status, creating a complex observation process. We develop a counterfactual framework that links the observation process to a hypothetical process in which infection is prevented. This formulation enables unbiased estimation of disease prevalence by modeling the testing process, possibly nonparametrically, without requiring explicit modeling of transmission dynamics, even though the testing and infection processes are jointly dependent.

stat.ME

Are Video Models Zero-Shot Learners and Reasoners in Education? EduVideoBench, A Knowledge-Skills-Attitude Benchmark for Educational Video Generation

Video generation models (VGMs) are rapidly entering classrooms, yet existing benchmarks evaluate only perceptual quality, intrinsic faithfulness, generic safety, or video as a reasoning medium, and none assesses whether the outputs are educationally valid. In this work, we present EduVideoBench, the first balanced benchmark in the education domain, grounded in the Knowledge-Skills-Attitude (KSA) framework so that pedagogical adequacy and educational safety are evaluated jointly rather than as ad-hoc quality dimensions. Across five frontier VGMs, our results show substantial room for improvement across knowledge, skills, and attitude before they are classroom-ready. We complement this with a qualitative analysis of expert comments, finding that educational validity is multi-component, where a single misaligned element such as pacing, legibility, or notation can invalidate an otherwise correct video. We hope EduVideoBench will guide the development of VGMs that are pedagogically grounded and safe for the classroom.

cs.CL

Analysing Extreme Rainfall via a Geometric Framework

Motivated by the EVA 2025 Data Challenge, we address the problem of predicting extreme rainfall in the eastern United States using data from a large ensemble of climate model runs. The challenge focuses on three quantities of interest related to the spatial extent and/or temporal duration of extreme rainfall, each requiring extrapolation. To tackle these questions, we adopt the recently developed geometric framework for extreme-value analysis, offering substantial flexibility for capturing complex extremal dependence structures and enabling extrapolation across the entire multivariate tail. In this work, we focus on the spatial geometric framework for analysing the spatial extent and consider a sampling procedure that retains the temporal information in the data, thereby enabling estimation of the duration of extreme rainfall events. We also account for the non-stationary behaviour, arising from topographical and seasonal effects, that commonly characterises extreme weather events in both space and time. Using diagnostic metrics, we demonstrate that the proposed model is appropriate for inferring extreme events on this dataset and apply it to estimate target quantities of interest.

stat.ME

Counterfactual Survival Q-learning via Buckley-James Boosting, with Applications to ACTG 175 and CALGB 8923

We propose a Buckley James (BJ) Boost Q learning framework for estimating optimal dynamic treatment regimes from right censored survival outcomes in longitudinal randomized clinical trials, motivated by the clinical need to support patient specific treatment decisions when follow up is incomplete and covariate effects may be nonlinear. The method combines accelerated failure time modeling with iterative boosting using flexible base learners, including componentwise least squares and regression trees, within a counterfactual Q learning framework. By modeling conditional survival time directly, BJ Boost Q learning avoids the proportional hazards assumption, yields clinically interpretable time scale contrasts, and enables estimation of stage specific Q functions and individualized decision rules under standard potential outcomes assumptions. In contrast to Cox based Q learning, which relies on hazard modeling and can be sensitive to nonproportional hazards and model misspecification, our approach provides a robust and flexible alternative for regime learning. Simulation studies and analyses of the ACTG175 HIV trial and the CALGB 8923 two stage leukemia trial show that BJ Boost Q learning improves treatment decision accuracy and produces more stable within participant counterfactual contrasts, particularly in multistage settings where estimation error and bias can compound across stages.

stat.ML

Transformed Linear Prediction for Extremes

We address the problem of prediction for extreme observations by proposing an extremal linear prediction method. We construct an inner product space of nonnegative random variables derived from transformed-linear combinations of independent regularly varying random variables. Under a reasonable modeling assumption, the matrix of inner products corresponds to the tail pairwise dependence matrix, which can be easily estimated. We derive the optimal transformed-linear predictor via the projection theorem, which yields a predictor with the same form as the best linear unbiased predictor in non-extreme settings. We quantify uncertainty for prediction errors by constructing prediction intervals based on the geometry of regular variation. We demonstrate the effectiveness of our method through a simulation study and its applications to predicting high pollution levels, and extreme precipitation.

stat.ME

Hypothesis testing for partial tail correlation in multivariate extremes

Statistical modeling of high dimensional extremes remains challenging and has generally been limited to moderate dimensions. Understanding structural relationships among variables at their extreme levels is crucial both for constructing simplified models and for identifying sparsity in extremal dependence. In this paper, we introduce the notion of partial tail correlation to characterize structural relationships between pairs of variables in their tails. To this end, we propose a tail regression approach for nonnegative regularly varying random vectors and define partial tail correlation based on the regression residuals. Using an extreme analogue of the covariance matrix, we show that the resulting regression coefficients and partial tail correlations take the same form as in classical non-extreme settings. For inference, we develop a hypothesis test to explore sparsity in extremal dependence structures, and demonstrate its effectiveness through simulations and an application to the Danube river network.

stat.ME

Geometric criteria for identifying extremal dependence and flexible modeling via additive mixtures

The framework of geometric extremes is based on the convergence of scaled sample clouds onto a limit set, characterized by a gauge function, with the shape of the limit set determining extremal dependence structures. While it is known that a blunt limit set implies asymptotic independence, the absence of bluntness can be linked to both asymptotic dependence and independence. Focusing on the bivariate case, under a truncated gamma modeling assumption with bounded angular density, we show that a ``pointy'' limit set implies asymptotic dependence, thus offering practical geometric criteria for identifying extremal dependence classes. Suitable models for the gauge function offer the ability to capture asymptotically independent or dependent data structures, without requiring prior knowledge of the true extremal dependence structure. The geometric approach thus offers a simple alternative to various parametric copula models that have been developed for this purpose in recent years. We consider two types of additively mixed gauge functions that provide a smooth interpolation between asymptotic dependence and asymptotic independence. We derive their explicit forms, explore their properties, and establish connections to the developed geometric criteria. Through a simulation study, we evaluate the effectiveness of the geometric approach with additively mixed gauge functions, comparing its performance to existing methodologies that account for both asymptotic dependence and asymptotic independence. The methodology is computationally efficient and yields reliable performance across various extremal dependence scenarios.

stat.ME

OASIS: Object-based Analytics Storage for Intelligent SQL Query Offloading in Scientific Tabular Workloads

Computation-Enabled Object Storage (COS) systems, such as MinIO and Ceph, have recently emerged as promising storage solutions for post hoc, SQL-based analysis on large-scale datasets in High-Performance Computing (HPC) environments. By supporting object-granular layouts, COS facilitates column-oriented access and supports in-storage execution of data reduction operators, such as filters, close to where the data resides. Despite growing interest and adoption, existing COS systems exhibit several fundamental limitations that hinder their effectiveness. First, they impose rigid constraints on output data formats, limiting flexibility and interoperability. Second, they support offloading for only a narrow set of operators and expressions, restricting their applicability to more complex analytical tasks. Third--and perhaps most critically--they fail to incorporate design strategies that enable compute offloading optimized for the characteristics of deep storage hierarchies. To address these challenges, this paper proposes OASIS, a novel COS system that features: (i) flexible and interoperable output delivery through diverse formats, including columnar layouts such as Arrow; (ii) broad support for complex operators (e.g., aggregate, sort) and array-aware expressions, including element-wise predicates over array structures; and (iii) dynamic selection of optimal execution paths across internal storage layers, guided by operator characteristics and data movement costs. We implemented a prototype of OASIS and integrated it into the Spark analytics framework. Through extensive evaluation using real-world scientific queries from HPC workflows, OASIS achieves up to a 32.7% performance improvement over Spark configured with existing COS-based storage systems.

cs.DB

Counterfactual Q Learning via the Linear Buckley James Method for Longitudinal Survival Data

Treatment strategies are critical in healthcare, particularly when outcomes are subject to censoring. This study introduces the Counterfactual Buckley-James Q-Learning framework, which integrates the Buckley-James method with reinforcement learning to address challenges posed by censored survival data. The Buckley-James method imputes censored survival times via conditional expectations based on observed data, offering a robust mechanism for handling incomplete outcomes. By incorporating these imputed values into a counterfactual Q-learning framework, the proposed method enables the estimation and comparison of potential outcomes under different treatment strategies. This facilitates the identification of optimal dynamic treatment regimes that maximize expected survival time. Through extensive simulation studies, the method demonstrates robust performance across various sample sizes and censoring scenarios, including right censoring and missing at random (MAR). Application to real-world clinical trial data further highlights the utility of this approach in informing personalized treatment decisions, providing an interpretable and reliable tool for optimizing survival outcomes in complex clinical settings.

stat.ME

X-Vine Models for Multivariate Extremes

Regular vine sequences permit the organisation of variables in a random vector along a sequence of trees. Regular vine models have become greatly popular in dependence modelling as a way to combine arbitrary bivariate copulas into higher-dimensional ones, offering flexibility, parsimony, and tractability. In this project, we use regular vine structures to decompose and construct the exponent measure density of a multivariate extreme value distribution, or, equivalently, the tail copula density. Although these densities pose theoretical challenges due to their infinite mass, their homogeneity property offers simplifications. The theory sheds new light on existing parametric families and facilitates the construction of new ones, called X-vines. Computations proceed via recursive formulas in terms of bivariate model components. We develop simulation algorithms for X-vine multivariate Pareto distributions as well as methods for parameter estimation and model selection on the basis of threshold exceedances. The methods are illustrated by Monte Carlo experiments and a case study on US flight delay data.

stat.ME

Individual Tooth Detection and Identification from Dental Panoramic X-Ray Images via Point-wise Localization and Distance Regularization

Dental panoramic X-ray imaging is a popular diagnostic method owing to its very small dose of radiation. For an automated computer-aided diagnosis system in dental clinics, automatic detection and identification of individual teeth from panoramic X-ray images are critical prerequisites. In this study, we propose a point-wise tooth localization neural network by introducing a spatial distance regularization loss. The proposed network initially performs center point regression for all the anatomical teeth (i.e., 32 points), which automatically identifies each tooth. A novel distance regularization penalty is employed on the 32 points by considering $L_2$ regularization loss of Laplacian on spatial distances. Subsequently, teeth boxes are individually localized using a cascaded neural network on a patch basis. A multitask offset training is employed on the final output to improve the localization accuracy. Our method successfully localizes not only the existing teeth but also missing teeth; consequently, highly accurate detection and identification are achieved. The experimental results demonstrate that the proposed algorithm outperforms state-of-the-art approaches by increasing the average precision of teeth detection by 15.71% compared to the best performing method. The accuracy of identification achieved a precision of 0.997 and recall value of 0.972. Moreover, the proposed network does not require any additional identification algorithm owing to the preceding regression of the fixed 32 points regardless of the existence of the teeth.

cs.CV

Liver Segmentation in Abdominal CT Images via Auto-Context Neural Network and Self-Supervised Contour Attention

Accurate image segmentation of the liver is a challenging problem owing to its large shape variability and unclear boundaries. Although the applications of fully convolutional neural networks (CNNs) have shown groundbreaking results, limited studies have focused on the performance of generalization. In this study, we introduce a CNN for liver segmentation on abdominal computed tomography (CT) images that shows high generalization performance and accuracy. To improve the generalization performance, we initially propose an auto-context algorithm in a single CNN. The proposed auto-context neural network exploits an effective high-level residual estimation to obtain the shape prior. Identical dual paths are effectively trained to represent mutual complementary features for an accurate posterior analysis of a liver. Further, we extend our network by employing a self-supervised contour scheme. We trained sparse contour features by penalizing the ground-truth contour to focus more contour attentions on the failures. The experimental results show that the proposed network results in better accuracy when compared to the state-of-the-art networks by reducing 10.31% of the Hausdorff distance. We used 180 abdominal CT images for training and validation. Two-fold cross-validation is presented for a comparison with the state-of-the-art neural networks. Novel multiple N-fold cross-validations are conducted to verify the performance of generalization. The proposed network showed the best generalization performance among the networks. Additionally, we present a series of ablation experiments that comprehensively support the importance of the underlying concepts.

cs.CV

Pose-Aware Instance Segmentation Framework from Cone Beam CT Images for Tooth Segmentation

Individual tooth segmentation from cone beam computed tomography (CBCT) images is an essential prerequisite for an anatomical understanding of orthodontic structures in several applications, such as tooth reformation planning and implant guide simulations. However, the presence of severe metal artifacts in CBCT images hinders the accurate segmentation of each individual tooth. In this study, we propose a neural network for pixel-wise labeling to exploit an instance segmentation framework that is robust to metal artifacts. Our method comprises of three steps: 1) image cropping and realignment by pose regressions, 2) metal-robust individual tooth detection, and 3) segmentation. We first extract the alignment information of the patient by pose regression neural networks to attain a volume-of-interest (VOI) region and realign the input image, which reduces the inter-overlapping area between tooth bounding boxes. Then, individual tooth regions are localized within a VOI realigned image using a convolutional detector. We improved the accuracy of the detector by employing non-maximum suppression and multiclass classification metrics in the region proposal network. Finally, we apply a convolutional neural network (CNN) to perform individual tooth segmentation by converting the pixel-wise labeling task to a distance regression task. Metal-intensive image augmentation is also employed for a robust segmentation of metal artifacts. The result shows that our proposed method outperforms other state-of-the-art methods, especially for teeth with metal artifacts. The primary significance of the proposed method is two-fold: 1) an introduction of pose-aware VOI realignment followed by a robust tooth detection and 2) a metal-robust CNN framework for accurate tooth segmentation.

cs.CV

Deeply Self-Supervised Contour Embedded Neural Network Applied to Liver Segmentation

Objective: Herein, a neural network-based liver segmentation algorithm is proposed, and its performance was evaluated using abdominal computed tomography (CT) images. Methods: A fully convolutional network was developed to overcome the volumetric image segmentation problem. To guide a neural network to accurately delineate a target liver object, the network was deeply supervised by applying the adaptive self-supervision scheme to derive the essential contour, which acted as a complement with the global shape. The discriminative contour, shape, and deep features were internally merged for the segmentation results. Results and Conclusion: 160 abdominal CT images were used for training and validation. The quantitative evaluation of the proposed network was performed through an eight-fold cross-validation. The result showed that the method, which uses the contour feature, segmented the liver more accurately than the state-of-the-art with a 2.13% improvement in the dice score. Significance: In this study, a new framework was introduced to guide a neural network and learn complementary contour features. The proposed neural network demonstrates that the guided contour features can significantly improve the performance of the segmentation task.

cs.CV

Automatic Registration between Cone-Beam CT and Scanned Surface via Deep-Pose Regression Neural Networks and Clustered Similarities

Computerized registration between maxillofacial cone-beam computed tomography (CT) images and a scanned dental model is an essential prerequisite in surgical planning for dental implants or orthognathic surgery. We propose a novel method that performs fully automatic registration between a cone-beam CT image and an optically scanned model. To build a robust and automatic initial registration method, our method applies deep-pose regression neural networks in a reduced domain (i.e., 2-dimensional image). Subsequently, fine registration is performed via optimal clusters. Majority voting system achieves globally optimal transformations while each cluster attempts to optimize local transformation parameters. The coherency of clusters determines their candidacy for the optimal cluster set. The outlying regions in the iso-surface are effectively removed based on the consensus among the optimal clusters. The accuracy of registration was evaluated by the Euclidean distance of 10 landmarks on a scanned model which were annotated by the experts in the field. The experiments show that the proposed method's registration accuracy, measured in landmark distance, outperforms other existing methods by 30.77% to 70%. In addition to achieving high accuracy, our proposed method requires neither human-interactions nor priors (e.g., iso-surface extraction). The main significance of our study is twofold: 1) the employment of light-weighted neural networks which indicates the applicability of neural network in extracting pose cues that can be easily obtained and 2) the introduction of an optimal cluster-based registration method that can avoid metal artifacts during the matching procedures.

cs.CV