SearcharxivSearch

arXiv subjects

Sunbeom Kwon

Publications and source records attributed to Sunbeom Kwon.

2 recordsLinked to original sources

Can We Trust Item Response Theory for AI Evaluation?

AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality. However, AI benchmark data often departs from the data regime of human testing, for which standard IRT estimation tools were originally developed: benchmarks typically involve fewer evaluated models, far more items, and capability distributions that may be skewed, clustered, or multimodal. We examine how these regime mismatches challenge the reliability of IRT modeling for AI evaluation. Using item parameters and capability distributions derived from six widely used LLM benchmarks, we simulate response matrices under three common IRT models and compare four estimation tools used in recent benchmark studies: marginal maximum likelihood, Markov chain Monte Carlo, variational inference, and a neural pseudo-Siamese estimator. Across 18,000 simulation conditions, we systematically evaluate computational feasibility, scalability, and the reliability of IRT inferences about model rankings, predicted performance, and item characteristics. Results show that classical estimators can become infeasible in large benchmark settings, whereas scalable estimators can produce unreliable item-level and ranking inferences with small or non-normally distributed model sets. This study identifies when latent trait models reliably support or risk distorting AI benchmarking claims, and what sample sizes and diagnostics are needed for trustworthy use.

cs.AI

PR-CARA: Proactive V2X Resource Allocation with Extended 1-Stage SCI and Deep Learning-based Sensing Matrix Estimator

Distributed resource allocation algorithms differ from centralized methods by relying on locally collected information for resource selection, leading to a low vehicle-to-everything (V2X) communication quality of service (QoS) in high-traffic congestion. To overcome these challenges, this study proposes a proactive received signal strength indicator (RSSI)-based collision avoidance resource allocation (PR-CARA) algorithm. This algorithm features an extended 1-stage SCI system, which is a critical component that enables resource monitoring of adjacent vehicle user equipment (VUE). Monitored resources were then processed through a deep learning-based proactive RSSI estimator. The estimated proactive RSSI helps avoid resource selection, which leads to packet collisions, thereby significantly reducing the occurrence of this issue during resource allocation. The proposed algorithm is tested in a cooperative adaptive cruise control (CACC)-based platoon driving scenario that requires ultra-reliable and low-latency communication (URLLC) performance. Simulation results demonstrate that the proposed deep-learning-based proactive resource allocation algorithm, with the extended 1-stage SCI system, reduces packet collisions and improves the transmission signal-to-interference-plus-noise ratio (SINR), thereby significantly enhancing communication reliability compared to the benchmark resource allocation algorithm.

eess.SY