SearcharxivSearch

arXiv subjects

Kensuke Okada

Publications and source records attributed to Kensuke Okada.

8 recordsLinked to original sources

Optimality of the Half-Order Exponent in the Turing-Good Identities for Bayes Factors

Bayes factors are widely computed by Monte Carlo, yet heavy-tailed sampling distributions can make numerical validation unreliable. The Turing--Good identities provide exact moment equalities for powers of a Bayes factor (a density ratio). When these identities are used as Good-check diagnostics, the power choice becomes a statistical design parameter. We develop a nonasymptotic variance theory for Monte Carlo evaluation of the identities and show that the half-order (square-root) power is uniquely minimax-stable: it equalizes variability across the two model orientations and is the only choice that guarantees finite second moments in a distribution-free worst-case sense over all mutually absolutely continuous model pairs. This yields a balanced two-sample half-order diagnostic that is symmetric in model labeling and has a uniform variance bound at fixed computational budget; in small-overlap regimes it is guaranteed to be no less efficient than the standard one-sided Turing check. Simulations for binomial Bayes factor workflows illustrate stable finite-sample behavior and sensitivity to simulator--evaluator mismatches. We further connect the half-order overlap viewpoint to stable primitives for normalizing-constant ratios and importance-sampling degeneracy summaries.

stat.ME

Quantifying and Mitigating Socially Desirable Responding in LLMs: A Desirability-Matched Graded Forced-Choice Psychometric Study

Human self-report questionnaires are increasingly used in NLP to benchmark and audit large language models (LLMs), from persona consistency to safety and bias assessments. Yet these instruments presume honest responding; in evaluative contexts, LLMs can instead gravitate toward socially preferred answers-a form of socially desirable responding (SDR)-biasing questionnaire-derived scores and downstream conclusions. We propose a psychometric framework to quantify and mitigate SDR in questionnaire-based evaluation of LLMs. To quantify SDR, the same inventory is administered under HONEST versus FAKE-GOOD instructions, and SDR is computed as a direction-corrected standardized effect size from item response theory (IRT)-estimated latent scores. This enables comparisons across constructs and response formats, as well as against human instructed-faking benchmarks. For mitigation, we construct a graded forced-choice (GFC) Big Five inventory by selecting 30 cross-domain pairs from an item pool via constrained optimization to match desirability. Across nine instruction-following LLMs evaluated on synthetic personas with known target profiles, Likert-style questionnaires show consistently large SDR, whereas desirability-matched GFC substantially attenuates SDR while largely preserving the recovery of the intended persona profiles. These results highlight a model-dependent SDR-recovery trade-off and motivate SDR-aware reporting practices for questionnaire-based benchmarking and auditing of LLMs.

cs.CL

Sparse Bayesian joint modal estimation for exploratory item factor analysis

This study presents a scalable Bayesian estimation algorithm for sparse estimation in exploratory item factor analysis based on a classical Bayesian estimation method, namely Bayesian joint modal estimation (BJME). BJME estimates the model parameters and factor scores that maximize the complete-data joint posterior density. The algorithm's scalability is achieved through an alternating optimization scheme that iteratively updates model parameters and latent variables. Simulation studies show that the proposed algorithm has high computational efficiency and accuracy in variable selection over latent factors and the recovery of the model parameters. Moreover, we conducted a real data analysis using large-scale data from a psychological assessment that targeted the Big Five personality traits. This result indicates that the proposed algorithm achieves computationally efficient parameter estimation and extracts the interpretable factor loading structure.

stat.ME

Commissioning of a compact multibend achromat lattice: A new 3 GeV synchrotron radiation facility

NanoTerasu, a new 3 GeV synchrotron light source in Japan, began user operation in April 2024. It provides high-brilliance soft to tender X-rays and covers a wide spectral range from ultraviolet to tender X-rays. Its compact storage ring with a circumference of 349 m is based on a four-bend achromat lattice to provide two straight sections in each cell for insertion devices with a natural horizontal emittance of 1.14 nm rad, which is small enough for soft X-rays users. The NanoTerasu accelerator incorporates several innovative technologies, including a full-energy injector C-band linear accelerator with a length of 110 m, an in-vacuum off-axis injection system, a four-bend achromat with B-Q combined bending magnets, and a TM020 mode accelerating cavity with built-in higher-order-mode dampers in the storage ring. This paper presents the accelerator machine commissioning over a half-year period and our model-consistent ring optics correction. The first user operation with a stored beam current of 160 mA is also reported. We summarize the storage ring parameters obtained from the commissioning. This is helpful for estimating the effective optical properties of synchrotron radiation at NanoTerasu.

physics.acc-ph

Assessing the Performance of Diagnostic Classification Models in Small Sample Contexts with Different Estimation Methods

Fueled by the call for formative assessments, diagnostic classification models (DCMs) have recently gained popularity in psychometrics. Despite their potential for providing diagnostic information that aids in classroom instruction and students' learning, empirical applications of DCMs to classroom assessments have been highly limited. This is partly because how DCMs with different estimation methods perform in small sample contexts is not yet well-explored. Hence, this study aims to investigate the performance of respondent classification and item parameter estimation with a comprehensive simulation design that resembles classroom assessments using different estimation methods. The key findings are the following: (1) although the marked difference in respondent classification accuracy was not observed among the maximum likelihood (ML), Bayesian, and nonparametric methods, the Bayesian method provided slightly more accurate respondent classification in parsimonious DCMs than the ML method, and in complex DCMs, the ML method yielded the slightly better result than the Bayesian method; (2) while item parameter recovery was poor in both Bayesian and ML methods, the Bayesian method exhibited unstable slip values owing to the multimodality of their posteriors under complex DCMs, and the ML method produced irregular estimates that appear to be well-estimated due to a boundary problem under parsimonious DCMs.

stat.CO

Variational Bayesian Inference for a Polytomous-Attribute Saturated Diagnostic Classification Model with Parallel Computing

As a statistical tool to assist formative assessments in educational settings, diagnostic classification models (DCMs) have been increasingly used to provide diagnostic information regarding examinees' attributes. DCMs often adopt a dichotomous division such as the mastery and non-mastery of attributes to express the mastery states of attributes. However, many practical settings involve different levels of mastery states rather than a simple dichotomy in a single attribute. Although this practical demand can be addressed by polytomous-attribute DCMs, their computational cost in a Markov chain Monte Carlo estimation impedes their large-scale application due to the larger number of polytomous-attribute mastery patterns than that of binary-attribute ones. This study considers a scalable Bayesian estimation method for polytomous-attribute DCMs and developed a variational Bayesian (VB) algorithm for a polytomous-attribute saturated DCM -- a generalization of polytomous-attribute DCMs -- by building on the existing literature on polytomous-attribute DCMs and VB for binary-attribute DCMs. Furthermore, we proposed the configuration of parallel computing for the proposed VB algorithm to achieve better computational efficiency. Monte Carlo simulations revealed that our method exhibited the high performance in parameter recovery under a wide range of conditions. An empirical example is used to demonstrate the utility of our method.

stat.CO

Scalable Bayesian Approach for the DINA Q-matrix Estimation Combining Stochastic Optimization and Variational Inference

Diagnostic classification models (DCMs) offer statistical tools to inspect the fined-grained attribute of respondents' strengths and weaknesses. However, the diagnosis accuracy deteriorates when misspecification occurs in the predefined item-attribute relationship, which is encoded into a Q-matrix. To prevent such misspecification, methodologists have recently developed several Bayesian Q-matrix estimation methods for greater estimation flexibility. However, these methods become infeasible in the case of large-scale assessments with a large number of attributes and items. In this study, we focused on the deterministic inputs, noisy "and" gate (DINA) model and proposed a new framework for the Q-matrix estimation to find the Q-matrix with the maximum marginal likelihood. Based on this framework, we developed a scalable estimation algorithm for the DINA Q-matrix by constructing an iteration algorithm that utilizes stochastic optimization and variational inference. The simulation and empirical studies reveal that the proposed method achieves high-speed computation, good accuracy, and robustness to potential misspecifications, such as initial value's choices and hyperparameter settings. Thus, the proposed method can be a useful tool for estimating a Q-matrix in large-scale settings.

stat.CO

Observation of W decay in 500 GeV p+p collisions at RHIC

W boson production is observed in $\sqrt{s}=500$ GeV proton proton collisions at RHIC-PHENIX experiment. The single longitudinal spin asymmetry $A_L(\overrightarrow{p}p\rightarrow W^+X)$ is measured via decay positrons in the mid rapidity region. The asymmetry $-0.83\pm0.31\pm(11%$ scale uncertainty) is consistent with calculations from various polarized parton distribution functions.

hep-ex