SearcharxivSearch

arXiv subjects

JoonHo Lee

Publications and source records attributed to JoonHo Lee.

At least 19 recordsLinked to original sources

How Reliable Are Psychological Measurements? The Distribution of Marginal Reliability Across 889 Item-Response Datasets

Nearly every quantitative study in psychology reports a reliability coefficient, so the field knows a great deal about the reliability that authors choose to publish. It knows much less about the reliability of the data psychology actually produces, because published coefficients pass through decisions about what to compute and what to report. We therefore measure reliability directly, applying the same estimators under the same rules to 889 datasets from the Item Response Warehouse, a public collection of item-response data that spans cognitive tests, clinical screeners, personality inventories, and attitude scales. Three findings emerge. Low reliability is common: even under a lenient definition of reliability, 30% of datasets fall below the conventional .80 threshold. The variation across datasets is real rather than statistical, since estimation noise accounts for only about one percent of it. Finally, the answer depends on the definition itself: under a strict definition the share below .80 rises to 52%, a difference large enough to change what one concludes about the field. We conclude that a reliability report should say which definition it uses, attach a measure of uncertainty, and give a strict coefficient alongside a lenient one.

stat.AP

How Common Are Estimated Latent-Distribution Departures From Normality? Evidence From 504 Item-Response Data Sets

Item response models usually assume a normal trait distribution, yet little is known about how often fitted distributions in real studies differ substantially from normality or which reported results are most affected. We analyzed 504 itemresponse data sets from 273 studies in the Item Response Warehouse, fitting each with the normal assumption and with a flexible distribution estimated from the responses, and we compared reliability, item estimates, predicted test responses, and person scores between the two calibrations. In more than half of the data sets the two fitted distributions differed by at least 10 percentage points of cumulative probability at some point on the trait scale, with a median maximum difference of about 11 points; about one-third reached 15 points, and nearly one-fifth reached 20 points. The estimated shapes included skewness, heavy tails, flat regions, and occasional multimodality. Differences were larger in several attitudinal, affective, and behavioral domains, although these contrasts weakened when data sets were compared within item-model families. Reliability usually changed little when item estimates were held fixed, whereas refitting the full model produced larger changes in some data sets, and item estimates, predicted test responses, and person scores did not change in parallel. Alternative flexible methods generally identified the same data sets as most unusual but disagreed somewhat about magnitude. Applied analyses should state the distributional assumption, examine a flexible alternative, and report sensitivity separately for each result used in interpretation or decision making.

stat.AP

Design Effect Ratios for Bayesian Survey Models: A Diagnostic Framework for Identifying Survey-Sensitive Parameters

Bayesian hierarchical models are increasingly fitted to complex survey data through weighted pseudo-posteriors, with a post-processing step that rescales the posterior to match a design-based sandwich covariance. Applied to all parameters at once, this correction can be harmful as well as protective: parameters stabilized by hierarchical shrinkage or identified by between-group variation have their credible intervals needlessly widened or spuriously narrowed. We propose the Design Effect Ratio (DER), the ratio of a parameter's design-based sandwich variance, computed for a declared variance target, to its model-based posterior variance, as a per-parameter diagnostic for applying the correction selectively. For hierarchical Gaussian models we derive exact finite-sample expressions under explicit balance and common-design-effect hypotheses: fixed-effect DERs scale with the design effect attenuated by the between-group share of identifying variation, and random-effect DERs factor into design effect, shrinkage, and a group-count term, with a conservation identity linking the two levels. A general matrix formula covers arbitrary parameter blocks and non-Gaussian likelihoods. A simulation study with a genuine informative two-stage sampling mechanism and a re-analysis of the 2019 National Survey of Early Care and Education, reported under both the design-PSU and model-group variance targets, show the diagnostic separating survey-sensitive from shrinkage-protected parameters and the selective correction avoiding the damage of blanket rescaling. The R package svyder implements the workflow.

stat.ME

A Bayesian Hierarchical Hurdle Beta-Binomial Model for Survey-Weighted Bounded Counts and Its Application to Childcare Enrollment

Bounded discrete proportions -- counts out of known totals -- present modeling challenges when data exhibit structural zeros, overdispersion, and hierarchical clustering. We develop a Bayesian hierarchical hurdle beta-binomial model with state-varying coefficients that addresses all four features. The framework makes three methodological contributions: (i) it studies cross-margin dependence via a cross-block covariance component and clarifies when and how this parameter is identified through the hierarchical layer rather than the conditional likelihood; (ii) it proposes a Cholesky-based sandwich variance calibration for pseudo-posterior inference under survey weights, guided by a parameter-specific design effect ratio diagnostic; and (iii) it introduces a log-scale marginal effect decomposition for hurdle models that translates regression coefficients into policy-relevant quantities. Applied to 6,785 childcare providers across 51 states from the 2019 National Survey of Early Care and Education, the model reveals a "poverty reversal": poverty reduces enrollment participation yet increases intensity among participants, with the extensive margin accounting for two-thirds of the total effect. Design-calibrated simulation shows that sandwich-corrected intervals substantially improve coverage, reaching 82--88.5% at the 90% nominal level for fixed effects. The R package hurdlebb implements all methods.

stat.ME

Design-Conditional Prior Elicitation for Dirichlet Process Mixtures: A Unified Framework for Cluster Counts and Weight Control

Dirichlet process mixtures describe variation in effects or latent traits across sites, studies, or examinees. The concentration parameter controls the number and relative sizes of the clusters, but its hyperprior is often chosen by default. We describe a method for choosing this hyperprior when the number of units is fixed by design. Analysts specify an expected number of clusters and their uncertainty about that count. We translate these judgments into a hyperprior and examine what it implies about the sizes of the clusters. When the count and size judgments conflict, the Dual-Anchor procedure chooses and reports the trade-off. Simulations show that the default hyperprior can favor a single cluster when estimates for individual units are noisy. Applications to a multisite trial, a meta-analysis, and a vocabulary test show sensitivity in heterogeneity summaries despite relatively small changes in individual estimates. The DPprior R package implements the calibration and diagnostics.

stat.ME

Beyond the Null Effect: Unmasking the True Impact of Teacher-Child Interaction Quality on Child Outcomes in Early Head Start

In Early Head Start (EHS), teacher-child interactions are widely believed to shape infant-toddler outcomes, yet large-scale studies often find only modest or null associations. This study addresses four methodological sources of attenuation -- item-level measurement error, center-level confounding, teacher- and classroom-level covariate imbalance, and overlooked nonlinearities -- to clarify classroom process quality's true influence on child development. Using data from the 2018 wave of the Early Head Start Family and Child Experiences Survey (Baby FACES), we applied a three-level generalized additive latent and mixed model (GALAMM) to distinguish genuine classroom-level variability in process quality, as measured by the Classroom Assessment Scoring System (CLASS) and Quality of Caregiver-Child Interactions for Infants and Toddlers (QCIT), from item-level noise and center-level effects. We then estimated dose-response relationships with children's language and socioemotional outcomes, employing covariate balancing weights and generalized additive models. Results show that nearly half of each item's variance reflects classroom-level processes, with the remainder tied to measurement error or center-wide influences, masking true classroom effects. After correcting for these biases, domain-focused dose-response analyses reveal robust linear associations between cognitive/language supports and children's English communicative skills, while emotional-behavioral supports better predict social-emotional competence. Some domains display plateaus when pushed to extremes, underscoring potential nonlinearities. These findings challenge the "null effect" narrative, demonstrating that rigorous methodology can uncover the critical, domain-specific impacts of teacher-child interaction quality, offering clearer guidance for targeted professional development and policy in EHS.

stat.AP

Reliability-Targeted Simulation of Item Response Data: Solving the Inverse Design Problem

Monte Carlo simulations are the primary methodology for evaluating Item Response Theory (IRT) methods, yet marginal reliability - the fundamental metric of data informativeness - is rarely treated as an explicit design factor. Unlike in multilevel modeling where the intraclass correlation (ICC) is routinely manipulated, IRT studies typically treat reliability as an incidental outcome, creating a "reliability omission" that obscures the signal-to-noise ratio of generated data. To address this gap, we introduce a principled framework for reliability-targeted simulation, transforming reliability from an implicit by-product into a precise input parameter. We formalize the inverse design problem, solving for a global discrimination scaling factor that uniquely achieves a pre-specified target reliability. Two complementary algorithms are proposed: Empirical Quadrature Calibration (EQC) for rapid, deterministic precision, and Stochastic Approximation Calibration (SAC) for rigorous stochastic estimation. A comprehensive validation study across 960 conditions demonstrates that EQC achieves essentially exact calibration, while SAC remains unbiased across non-normal latent distributions and empirical item pools. Furthermore, we clarify the theoretical distinction between average-information and error-variance-based reliability metrics, showing they require different calibration scales due to Jensen's inequality. An accompanying open-source R package, IRTsimrel, enables researchers to standardize reliability as a controlled experimental input.

stat.ME

SGuard-v1: Safety Guardrail for Large Language Models

We present SGuard-v1, a lightweight safety guardrail for Large Language Models (LLMs), which comprises two specialized models to detect harmful content and screen adversarial prompts in human-AI conversational settings. The first component, ContentFilter, is trained to identify safety risks in LLM prompts and responses in accordance with the MLCommons hazard taxonomy, a comprehensive framework for trust and safety assessment of AI. The second component, JailbreakFilter, is trained with a carefully designed curriculum over integrated datasets and findings from prior work on adversarial prompting, covering 60 major attack types while mitigating false-unsafe classification. SGuard-v1 is built on the 2B-parameter Granite-3.3-2B-Instruct model that supports 12 languages. We curate approximately 1.4 million training instances from both collected and synthesized data and perform instruction tuning on the base model, distributing the curated data across the two component according to their designated functions. Through extensive evaluation on public and proprietary safety benchmarks, SGuard-v1 achieves state-of-the-art safety performance while remaining lightweight, thereby reducing deployment overhead. SGuard-v1 also improves interpretability for downstream use by providing multi-class safety predictions and their binary confidence scores. We release the SGuard-v1 under the Apache-2.0 License to enable further research and practical deployment in AI safety.

cs.CL

Exploring OCR-augmented Generation for Bilingual VQA

We investigate OCR-augmented generation with Vision Language Models (VLMs), exploring tasks in Korean and English toward multilingualism. To support research in this domain, we train and release KLOCR, a strong bilingual OCR baseline trained on 100M instances to augment VLMs with OCR ability. To complement existing VQA benchmarks, we curate KOCRBench for Korean VQA, and analyze different prompting methods. Extensive experiments show that OCR-extracted text significantly boosts performance across open source and commercial models. Our work offers new insights into OCR-augmented generation for bilingual VQA. Model, code, and data are available at https://github.com/JHLee0513/KLOCR.

cs.CV

Preference Consistency Matters: Enhancing Preference Learning in Language Models with Automated Self-Curation of Training Corpora

Inconsistent annotations in training corpora, particularly within preference learning datasets, pose challenges in developing advanced language models. These inconsistencies often arise from variability among annotators and inherent multi-dimensional nature of the preferences. To address these issues, we introduce a self-curation method that preprocesses annotated datasets by leveraging proxy models trained directly on them. Our method enhances preference learning by automatically detecting and selecting consistent annotations. We validate the proposed approach through extensive instruction-following tasks, demonstrating performance improvements of up to 33\% across various learning algorithms and proxy capabilities. This work offers a straightforward and reliable solution to address preference inconsistencies without relying on heuristics, serving as an initial step toward the development of more advanced preference learning methodologies. Code is available at https://github.com/Self-Curation/ .

cs.CL

Valid standard errors for Bayesian quantile regression with clustered and independent data

In Bayesian quantile regression, the most commonly used likelihood is the asymmetric Laplace (AL) likelihood. The reason for this choice is not that it is a plausible data-generating model but that the corresponding maximum likelihood estimator is identical to the classical estimator by Koenker and Bassett (1978), and in that sense, the AL likelihood can be thought of as a working likelihood. AL-based quantile regression has been shown to produce good finite-sample Bayesian point estimates and to be consistent. However, if the AL distribution does not correspond to the data-generating distribution, credible intervals based on posterior standard deviations can have poor coverage. Yang, Wang, and He (2016) proposed an adjustment to the posterior covariance matrix that produces asymptotically valid intervals. However, we show that this adjustment is sensitive to the choice of scale parameter for the AL likelihood and can lead to poor coverage when the sample size is small to moderate. We therefore propose using Infinitesimal Jackknife (IJ) standard errors (Giordano & Broderick, 2023). These standard errors do not require resampling but can be obtained from a single MCMC run. We also propose a version of IJ standard errors for clustered data. Simulations and applications to real data show that the IJ standard errors have good frequentist properties, both for independent and clustered data. We provide an R-package, IJSE, that computes IJ standard errors for clustered or independent data after estimation with the brms wrapper in R for Stan.

stat.ME

Improving Instruction Following in Language Models through Proxy-Based Uncertainty Estimation

Assessing response quality to instructions in language models is vital but challenging due to the complexity of human language across different contexts. This complexity often results in ambiguous or inconsistent interpretations, making accurate assessment difficult. To address this issue, we propose a novel Uncertainty-aware Reward Model (URM) that introduces a robust uncertainty estimation for the quality of paired responses based on Bayesian approximation. Trained with preference datasets, our uncertainty-enabled proxy not only scores rewards for responses but also evaluates their inherent uncertainty. Empirical results demonstrate significant benefits of incorporating the proposed proxy into language model training. Our method boosts the instruction following capability of language models by refining data curation for training and improving policy optimization objectives, thereby surpassing existing methods by a large margin on benchmarks such as Vicuna and MT-bench. These findings highlight that our proposed approach substantially advances language model training and paves a new way of harnessing uncertainty within language models.

cs.CL

V-STRONG: Visual Self-Supervised Traversability Learning for Off-road Navigation

Reliable estimation of terrain traversability is critical for the successful deployment of autonomous systems in wild, outdoor environments. Given the lack of large-scale annotated datasets for off-road navigation, strictly-supervised learning approaches remain limited in their generalization ability. To this end, we introduce a novel, image-based self-supervised learning method for traversability prediction, leveraging a state-of-the-art vision foundation model for improved out-of-distribution performance. Our method employs contrastive representation learning using both human driving data and instance-based segmentation masks during training. We show that this simple, yet effective, technique drastically outperforms recent methods in predicting traversability for both on- and off-trail driving scenarios. We compare our method with recent baselines on both a common benchmark as well as our own datasets, covering a diverse range of outdoor environments and varied terrain types. We also demonstrate the compatibility of resulting costmap predictions with a model-predictive controller. Finally, we evaluate our approach on zero- and few-shot tasks, demonstrating unprecedented performance for generalization to new environments. Videos and additional material can be found here: https://sites.google.com/view/visual-traversability-learning.

cs.RO

LiDAR-UDA: Self-ensembling Through Time for Unsupervised LiDAR Domain Adaptation

We introduce LiDAR-UDA, a novel two-stage self-training-based Unsupervised Domain Adaptation (UDA) method for LiDAR segmentation. Existing self-training methods use a model trained on labeled source data to generate pseudo labels for target data and refine the predictions via fine-tuning the network on the pseudo labels. These methods suffer from domain shifts caused by different LiDAR sensor configurations in the source and target domains. We propose two techniques to reduce sensor discrepancy and improve pseudo label quality: 1) LiDAR beam subsampling, which simulates different LiDAR scanning patterns by randomly dropping beams; 2) cross-frame ensembling, which exploits temporal consistency of consecutive frames to generate more reliable pseudo labels. Our method is simple, generalizable, and does not incur any extra inference cost. We evaluate our method on several public LiDAR datasets and show that it outperforms the state-of-the-art methods by more than $3.9\%$ mIoU on average for all scenarios. Code will be available at https://github.com/JHLee0513/LiDARUDA.

cs.CV

Improving the Estimation of Site-Specific Effects and their Distribution in Multisite Trials

In multisite trials, researchers are often interested in several inferential goals: estimating treatment effects for each site, ranking these effects, and studying their distribution. This study seeks to identify optimal methods for estimating these targets. Through a comprehensive simulation study, we assess two strategies and their combined effects: semiparametric modeling of the prior distribution, and alternative posterior summary methods tailored to minimize specific loss functions. Our findings highlight that the success of different estimation strategies depends largely on the amount of within-site and between-site information available from the data. We discuss how our results can guide balancing the trade-offs associated with shrinkage in limited data environments.

stat.ME

Unsupervised Accuracy Estimation of Deep Visual Models using Domain-Adaptive Adversarial Perturbation without Source Samples

Deploying deep visual models can lead to performance drops due to the discrepancies between source and target distributions. Several approaches leverage labeled source data to estimate target domain accuracy, but accessing labeled source data is often prohibitively difficult due to data confidentiality or resource limitations on serving devices. Our work proposes a new framework to estimate model accuracy on unlabeled target data without access to source data. We investigate the feasibility of using pseudo-labels for accuracy estimation and evolve this idea into adopting recent advances in source-free domain adaptation algorithms. Our approach measures the disagreement rate between the source hypothesis and the target pseudo-labeling function, adapted from the source hypothesis. We mitigate the impact of erroneous pseudo-labels that may arise due to a high ideal joint hypothesis risk by employing adaptive adversarial perturbation on the input of the target model. Our proposed source-free framework effectively addresses the challenging distribution shift scenarios and outperforms existing methods requiring source data and labels for training.

cs.CV

TerrainNet: Visual Modeling of Complex Terrain for High-speed, Off-road Navigation

Effective use of camera-based vision systems is essential for robust performance in autonomous off-road driving, particularly in the high-speed regime. Despite success in structured, on-road settings, current end-to-end approaches for scene prediction have yet to be successfully adapted for complex outdoor terrain. To this end, we present TerrainNet, a vision-based terrain perception system for semantic and geometric terrain prediction for aggressive, off-road navigation. The approach relies on several key insights and practical considerations for achieving reliable terrain modeling. The network includes a multi-headed output representation to capture fine- and coarse-grained terrain features necessary for estimating traversability. Accurate depth estimation is achieved using self-supervised depth completion with multi-view RGB and stereo inputs. Requirements for real-time performance and fast inference speeds are met using efficient, learned image feature projections. Furthermore, the model is trained on a large-scale, real-world off-road dataset collected across a variety of diverse outdoor environments. We show how TerrainNet can also be used for costmap prediction and provide a detailed framework for integration into a planning module. We demonstrate the performance of TerrainNet through extensive comparison to current state-of-the-art baselines for camera-only scene prediction. Finally, we showcase the effectiveness of integrating TerrainNet within a complete autonomous-driving stack by conducting a real-world vehicle test in a challenging off-road scenario.

cs.RO

Feature Alignment by Uncertainty and Self-Training for Source-Free Unsupervised Domain Adaptation

Most unsupervised domain adaptation (UDA) methods assume that labeled source images are available during model adaptation. However, this assumption is often infeasible owing to confidentiality issues or memory constraints on mobile devices. Some recently developed approaches do not require source images during adaptation, but they show limited performance on perturbed images. To address these problems, we propose a novel source-free UDA method that uses only a pre-trained source model and unlabeled target images. Our method captures the aleatoric uncertainty by incorporating data augmentation and trains the feature generator with two consistency objectives. The feature generator is encouraged to learn consistent visual features away from the decision boundaries of the head classifier. Thus, the adapted model becomes more robust to image perturbations. Inspired by self-supervised learning, our method promotes inter-space alignment between the prediction space and the feature space while incorporating intra-space consistency within the feature space to reduce the domain gap between the source and target domains. We also consider epistemic uncertainty to boost the model adaptation performance. Extensive experiments on popular UDA benchmark datasets demonstrate that the proposed source-free method is comparable or even superior to vanilla UDA methods. Moreover, the adapted models show more robust results when input images are perturbed.

cs.CV