SearcharxivSearch

arXiv subjects

Chae Young Lim

Publications and source records attributed to Chae Young Lim.

9 recordsLinked to original sources

MEDLAYXPLAIN: Benchmarking the Expert-Lay Gap in Medical Vision-Language Models

Medical Vision-Language Models (Med-VLMs) achieve strong expert-level performance, yet their ability to generate patient-accessible descriptions remains underexplored. With the 21st Century Cures Act now mandating immediate patient access to diagnostic imaging results, evaluating whether Med-VLMs can bridge this Expert-Lay Gap is both urgent and clinically consequential for patient education and shared decision-making. To this end, we introduce MedLayXPlain, the first large-scale multimodal benchmark and evaluation framework for Medical Lay Language Generation (MLLG). MedLayXPlain-122K provides 122,789 region-grounded samples across 8 imaging modalities from 12 publicly available source datasets, each comprising a medical image with paired expert and lay captions anchored in a three-level Unified Medical Language System (UMLS) ontology hierarchy spanning 7 semantic groups, 43 semantic types, and 2,411 medical concepts. Lay captions are constructed via Hierarchical Ontology-Verified Refinement (HOVER), a three-step pipeline combining patient-centric vocabulary mapping, LLM-based constrained rewriting, and cross-model visual verification to enforce semantic equivalence while preventing hallucination. We further introduce MedLayEval, a lightweight 3B evaluator distilled from a 27B verifier that scores expert-lay alignment across five clinically grounded attributes, addressing the poor correlation between standard NLG metrics and clinical judgment. Benchmarking 33 VLMs on MedLayXPlain-122K reveals a systematic Expert-Lay Gap: medical VLMs achieve strong expert captioning but suffer significant lay-register degradation, while general-purpose VLMs produce more accessible language yet lack clinical precision, confirming that neither current paradigm adequately serves patient-facing communication.

cs.CV

Cluster-Aware Conformal Calibration for Spatio-Temporal Distributional Prediction

DeepKriging-style models, such as Spatio-Temporal DeepKriging, improve scalability through basis-function embeddings and stochastic gradient learning; however, fixed regular-grid spatial bases remain inefficient under highly non-uniform sampling patterns, often over-allocating capacity to sparse regions while under-resolving dense clusters. To address this limitation, we propose a practical extension of DeepKriging for reliable spatio-temporal distributional forecasting, incorporating cluster-adaptive spatial bases - whose centers and scales are initialized from {the spatial sampling density} - to better capture heterogeneous spatial sampling, together with cluster-aware conformal calibration that determines prediction-interval widths within spatial clusters (with a global fallback when calibration samples are insufficient). The resulting calibration pipeline explicitly targets spatial heterogeneity and local miscalibration, and experiments, including simulation studies and PM$_{2.5}$ data analysis, demonstrate substantially improved coverage accuracy and tail reliability under clustered observation patterns compared with a global conformal baseline.

stat.ME

Differentially private synthesis of Spatial Point Processes

This paper proposes a method to generate synthetic data for spatial point patterns within the differential privacy (DP) framework. Specifically, we define a differentially private Poisson point synthesizer (PPS) and Cox point synthesizer (CPS) to generate synthetic point patterns with the concept of the $α$-neighborhood that relaxes the original definition of DP. We present three example models to construct a differentially private PPS and CPS, providing sufficient conditions on their parameters to ensure the DP given a specified privacy budget. In addition, we demonstrate that the synthesizers can be applied to point patterns on the linear network. Simulation experiments demonstrate that the proposed approaches effectively maintain the privacy and utility of synthetic data.

stat.AP

Regularized Nonlinear Regression with Dependent Errors and its Application to a Biomechanical Model

A biomechanical model often requires parameter estimation and selection in a known but complicated nonlinear function. Motivated by observing that data from a head-neck position tracking system, one of biomechanical models, show multiplicative time dependent errors, we develop a modified penalized weighted least squares estimator. The proposed method can be also applied to a model with non-zero mean time dependent additive errors. Asymptotic properties of the proposed estimator are investigated under mild conditions on a weight matrix and the error process. A simulation study demonstrates that the proposed estimation works well in both parameter estimation and selection with time dependent error. The analysis and comparison with an existing method for head-neck position tracking data show better performance of the proposed method in terms of the variance accounted for (VAF).

stat.ME

Spatial Regression With Multiplicative Errors, and Its Application With Lidar Measurements

Multiplicative errors in addition to spatially referenced observations often arise in geodetic applications, particularly in surface estimation with light detection and ranging (LiDAR) measurements. However, spatial regression involving multiplicative errors remains relatively unexplored in such applications. In this regard, we present a penalized modified least squares estimator to handle the complexities of a multiplicative error structure while identifying significant variables in spatially dependent observations for surface estimation. The proposed estimator can be also applied to classical additive error spatial regression. By establishing asymptotic properties of the proposed estimator under increasing domain asymptotics with stochastic sampling design, we provide a rigorous foundation for its effectiveness. A comprehensive simulation study confirms the superior performance of our proposed estimator in accurately estimating and selecting parameters, outperforming existing approaches. To demonstrate its real-world applicability, we employ our proposed method, along with other alternative techniques, to estimate a rotational landslide surface using LiDAR measurements. The results highlight the efficacy and potential of our approach in tackling complex spatial regression problems involving multiplicative errors.

stat.ME

Bayesian estimation of the autocovariance of a model error in time series

Autocovariance of the error term in a time series model plays a key role in the estimation and inference for the model that it belongs to. Typically, some arbitrary parametric structure is assumed upon the error to simplify the estimation, which inevitably introduces potential model-misspecification. We thus conduct nonparametric estimation of it. To avoid the difficult bandwidth selection issue under the traditional nonparametric truncation approach, this paper conducts the Bayesian estimation of its spectral density in a frequency domain. To this end, we consider two cases: fixed error variance and time-varying one. Each approach is taken to estimate the spectral density of the autocovariance and the model parameters. The methodology is applied to exchange rate forecasting and proves to compete favorably against some benchmark models, including the random walk without drift.

stat.ME

Regularized Nonlinear Regression for Simultaneously Selecting and Estimating Key Model Parameters

In system identification, estimating parameters of a model using limited observations results in poor identifiability. To cope with this issue, we propose a new method to simultaneously select and estimate sensitive parameters as key model parameters and fix the remaining parameters to a set of typical values. Our method is formulated as a nonlinear least squares estimator with L1-regularization on the deviation of parameters from a set of typical values. First, we provide consistency and oracle properties of the proposed estimator as a theoretical foundation. Second, we provide a novel approach based on Levenberg-Marquardt optimization to numerically find the solution to the formulated problem. Third, to show the effectiveness, we present an application identifying a biomechanical parametric model of a head position tracking task for 10 human subjects from limited data. In a simulation study, the variances of estimated parameters are decreased by 96.1% as compared to that of the estimated parameters without L1-regularization. In an experimental study, our method improves the model interpretation by reducing the number of parameters to be estimated while maintaining variance accounted for (VAF) at above 82.5%. Moreover, the variances of estimated parameters are reduced by 71.1% as compared to that of the estimated parameters without L1-regularization. Our method is 54 times faster than the standard simplex-based optimization to solve the regularized nonlinear regression.

stat.ME

Spatial Clustering of Curves with Functional Covariates: A Bayesian Partitioning Model with Application to Spectra Radiance in Climate Study

In climate change study, the infrared spectral signatures of climate change have recently been conceptually adopted, and widely applied to identifying and attributing atmospheric composition change. We propose a Bayesian hierarchical model for spatial clustering of the high-dimensional functional data based on the effects of functional covariates and local features. We couple the functional mixed-effects model with a generalized spatial partitioning method for: (1) producing spatially contiguous clusters for the high-dimensional spatio-functional data; (2) improving the computational efficiency via parallel computing over subregions or multi-level partitions; and (3) capturing the near-boundary ambiguity and uncertainty for data-driven partitions. We propose a generalized partitioning method which puts less constraints on the shape of spatial clusters. Dimension reduction in the parameter space is also achieved via Bayesian wavelets to alleviate the increasing model complexity introduced by clusters. The model well captures the regional effects of the atmospheric and cloud properties on the spectral radiance measurements. The results elaborate the importance of exploiting spatially contiguous partitions for identifying regional effects and small-scale variability.

stat.AP

A generalized mixed model framework for assessing fingerprint individuality in presence of varying image quality

Fingerprint individuality refers to the extent of uniqueness of fingerprints and is the main criteria for deciding between a match versus nonmatch in forensic testimony. Often, prints are subject to varying levels of noise, for example, the image quality may be low when a print is lifted from a crime scene. A poor image quality causes human experts as well as automatic systems to make more errors in feature detection by either missing true features or detecting spurious ones. This error lowers the extent to which one can claim individualization of fingerprints that are being matched. The aim of this paper is to quantify the decrease in individualization as image quality degrades based on fingerprint images in real databases. This, in turn, can be used by forensic experts along with their testimony in a court of law. An important practical concern is that the databases used typically consist of a large number of fingerprint images so computational algorithms such as the Gibbs sampler can be extremely slow. We develop algorithms based on the Laplace approximation of the likelihood and infer the unknown parameters based on this approximate likelihood. Two publicly available databases, namely, FVC2002 and FVC2006, are analyzed from which estimates of individuality are obtained. From a statistical perspective, the contribution can be treated as an innovative application of Generalized Linear Mixed Models (GLMMs) to the field of fingerprint-based authentication.

stat.AP