SearcharxivSearch

arXiv subjects

Jaeyong Lee

Publications and source records attributed to Jaeyong Lee.

At least 19 recordsLinked to original sources

Bayesian Estimation of the Eigenstructure in High-Dimensional Approximate Factor Models

High-dimensional economic datasets often display strong co-movement driven by a small number of latent factors, which are typically modeled using approximate factor models. When the number of variables is large relative to the sample size, the eigenvalues and eigenvectors of the sample covariance matrix are severely distorted, which in turn makes principal component based estimators of the factor structure unstable. To address the high-dimensional problem, we propose a Bayesian model for approximate factor structures. We show that the posterior convergence rate is of the same order as benchmark results for high-dimensional spiked covariance models. Simulation studies show that the proposed method more accurately recovers the factor structure in approximate factor models than existing methods. Real data analyses on macro--financial datasets illustrate that the proposed method provides interpretable estimates of latent factor structure and performs competitively in forecasting exercises.

stat.ME

First Measurement of the $K^-$ Escape Cross Section in the ${}^{12}{\rm C}(K^{-},p)$ Reaction

We investigated the $\bar{K}$-nucleus interaction through the simultaneous measurement of the inclusive $^{12}{\rm C}(K^-, p)$ and exclusive $K^-$-escape $^{12}{\rm C}(K^-, p K^-_{esc})$ reactions at $1.8$ GeV/$c$ at J-PARC. The present measurement explicitly focuses on the $K^-$ escape process for the first time, successfully accomplishing a direct experimental determination of the imaginary part of the $K^-$ optical potential. The differential cross section for the $K^-$-escape reaction was determined to be $436 \pm 6\:(\text{stat.}) \pm 44\:(\text{syst.})~\mu\text{b/sr}$. A simultaneous likelihood fit yielded real and imaginary potential strengths of $V_0 = -72\:^{+3}_{-5}\:(\text{stat.})\:^{+0}_{-8}\:(\text{syst.})~\text{MeV}$ and $W_0 = -100\:^{+7}_{-1}\:(\text{stat.})\:^{+0}_{-16}\:(\text{syst.})~\text{MeV}$ at the nuclear center, respectively. The derived $W_0$ is significantly stronger than that predicted by theoretical models based on one-nucleon processes, suggesting possible contribution of multi-nucleon involving processes.

nucl-ex

Posterior Contraction of Lévy Adaptive B-spline Regression in Besov Spaces

We investigate the asymptotic properties of the Lévy Adaptive B-spline (LABS) regression model, a Bayesian nonparametric method that incorporates B-spline kernels into the Lévy Adaptive Regression Kernel (LARK) model. LABS applies splines of varying degrees with independently defined knots, yielding a flexible model class capable of adapting to irregular and locally structured features of the true function. Within the nonparametric regression framework with univariate random design and Gaussian errors, we establish that the LABS posterior contracts around the true function in Besov classes at nearly minimax-optimal rates, up to a logarithmic factor, while adapting automatically to unknown smoothness. This study contributes to filling a gap in the literature, where theoretical results on posterior contraction of the LARK model in Besov spaces remain scarce. Simulation experiments on standard test functions in Besov spaces, including Blocks, Bumps, HeaviSine, and Doppler, complement the theoretical results and demonstrate the practical utility of LABS.

stat.ML

Posterior Contraction Rates for Sparse Kolmogorov-Arnold Networks in Anisotropic Besov Spaces

We study posterior contraction rates for sparse Bayesian Kolmogorov-Arnold networks (KANs) over anisotropic Besov spaces, providing a statistical foundation of KANs from a Bayesian point of view. We show that sparse Bayesian KANs equipped with spike-and-slab-type sparsity priors attain the near-minimax posterior contraction. In particular, the contraction rate depends on the intrinsic anisotropic smoothness of the underlying function. Moreover, by placing a hyperprior on a single model-size parameter, the resulting posterior adapts to unknown anisotropic smoothness and still achieves the corresponding near-minimax rate. A distinctive feature of our results, compared with those for standard sparse MLP-based models, is that the KAN depth can be kept fixed: owing to the flexibility of learnable spline edge functions, the required approximation complexity is controlled through the network width, spline-grid range and size, and parameter sparsity. Our analysis develops theoretical tools tailored to sparse spline-edge architectures, including approximation and complexity bounds for Bayesian KANs. We then extend to compositional Besov spaces and show that the contraction rates depend on layerwise smoothness and effective dimension of the underlying compositional structure, thereby effectively avoiding the curse of dimensionality. Together, the developed tools and findings advance the theoretical understanding of Bayesian neural networks and provide rigorous statistical foundations for KANs.

stat.ML

Transferability of Token Usage Rights: A Design Space Analysis of Generative AI Services

With the rapid spread of generative AI services, the token has gained value not only as a technical unit of language processing but also as an economic currency for accessing AI services. Major AI model providers have adopted token-based billing as their default service model, requiring users to purchase platform-bound, fixed token usage rights. However, the fixedness of these usage rights is grounded in the billing-policy decisions of service providers rather than in any technical necessity. This study defines the Transferability of token usage rights as a design property that allows users to flexibly reallocate purchased data resources free from the constraints of time, account, and service. Drawing on the Design Space Analysis framework of MacLean et al. (1991), we identify five design axes (Target, Direction, Unit, Control, Reversibility) and five concrete Transferability types (carry-over, co-management, transfer, conversion, and trade) by analyzing the billing policies and terms of service of four major LLM services (ChatGPT, Claude, Gemini, Grok). Our analysis reframes the token from a purely economic-technical primitive into a core element of user-centered system design that expands user choice and autonomy.

cs.HC

Conformal mapping based Physics-informed neural networks for designing neutral inclusions

We address the neutral inclusion problem with imperfect boundary conditions, focusing on designing interface functions for inclusions of arbitrary shapes. Traditional Physics-Informed Neural Networks (PINNs) struggle with this inverse problem, leading to the development of Conformal Mapping Coordinates Physics-Informed Neural Networks (CoCo-PINNs), which integrate geometric function theory with PINNs. CoCo-PINNs effectively solve forward-inverse problems by modeling the interface function through neural network training, which yields a neutral inclusion effect. This approach enhances the performance of PINNs in terms of credibility, consistency, and stability.

cs.LG

STAR: Improving Lifetime and Performance of High-Capacity Modern SSDs Using State-Aware Randomizer

Although NAND flash memory has achieved continuous capacity improvements via advanced 3D stacking and multi-level cell technologies, these innovations introduce new reliability challenges, particularly lateral charge spreading (LCS), absent in low-capacity 2D flash memory. Since LCS significantly increases retention errors over time, addressing this problem is essential to ensure the lifetime of modern SSDs employing high-capacity 3D flash memory. In this paper, we propose a novel data randomizer, STate-Aware Randomizer (STAR), which proactively eliminates the majority of weak data patterns responsible for retention errors caused by LCS. Unlike existing techniques that target only specific worst-case patterns, STAR effectively removes a broad spectrum of weak patterns, significantly enhancing reliability against LCS. By employing several optimization schemes, STAR can be efficiently integrated into the existing I/O datapath of an SSD controller with negligible timing overhead. To evaluate the proposed STAR scheme, we developed a STAR-aware SSD emulator based on characterization results from 160 real 3D NAND flash chips. Experimental results demonstrate that STAR improves SSD lifetime by up to 2.3x and reduces read latency by an average of 50% on real-world traces compared to conventional SSDs

cs.AR

Eigenstructure inference for high-dimensional covariance with generalized shrinkage inverse-Wishart prior

In multivariate statistics, estimating the covariance matrix is essential for understanding the interdependence among variables. In high-dimensional settings, where the number of covariates increases with the sample size, it is well known that the eigenstructure of the sample covariance matrix is inconsistent. The inverse-Wishart prior, a standard choice for covariance estimation in Bayesian inference, also suffers from posterior inconsistency. To address the issue of eigenvalue dispersion in high-dimensional settings, the shrinkage inverse-Wishart (SIW) prior has recently been proposed. Despite its conceptual appeal and empirical success, the asymptotic justification for the SIW prior has remained limited. In this paper, we propose a generalized shrinkage inverse-Wishart (gSIW) prior for high-dimensional covariance modeling. By extending the SIW framework, the gSIW prior accommodates a broader class of prior distributions and facilitates the derivation of theoretical properties under specific parameter choices. In particular, under the spiked covariance assumption, we establish the asymptotic behavior of the posterior distribution for both eigenvalues and eigenvectors by directly evaluating the posterior expectations for two sets of parameter choices. This direct evaluation provides insights into the large-sample behavior of the posterior that cannot be obtained through general posterior asymptotic theorems. Finally, simulation studies illustrate that the proposed prior provides accurate estimation of the eigenstructure, particularly for spiked eigenvalues, achieving narrower credible intervals and higher coverage probabilities compared to existing methods. For spiked eigenvectors, the performance is generally comparable to that of competing approaches, including the sample covariance.

math.ST

ScholarBench: A Bilingual Benchmark for Abstraction, Comprehension, and Reasoning Evaluation in Academic Contexts

Prior benchmarks for evaluating the domain-specific knowledge of large language models (LLMs) lack the scalability to handle complex academic tasks. To address this, we introduce \texttt{ScholarBench}, a benchmark centered on deep expert knowledge and complex academic problem-solving, which evaluates the academic reasoning ability of LLMs and is constructed through a three-step process. \texttt{ScholarBench} targets more specialized and logically complex contexts derived from academic literature, encompassing five distinct problem types. Unlike prior benchmarks, \texttt{ScholarBench} evaluates the abstraction, comprehension, and reasoning capabilities of LLMs across eight distinct research domains. To ensure high-quality evaluation data, we define category-specific example attributes and design questions that are aligned with the characteristic research methodologies and discourse structures of each domain. Additionally, this benchmark operates as an English-Korean bilingual dataset, facilitating simultaneous evaluation for linguistic capabilities of LLMs in both languages. The benchmark comprises 5,031 examples in Korean and 5,309 in English, with even state-of-the-art models like o3-mini achieving an average evaluation score of only 0.543, demonstrating the challenging nature of this benchmark.

cs.CL

Bayesian Analysis of Spiked Covariance Models: Correcting Eigenvalue Bias and Determining the Number of Spikes

We study Bayesian inference in the spiked covariance model, where a small number of spiked eigenvalues dominate the spectrum. Our goal is to infer the spiked eigenvalues, their corresponding eigenvectors, and the number of spikes, providing a Bayesian solution to principal component analysis with uncertainty quantification. We place an inverse-Wishart prior on the covariance matrix to derive posterior distributions for the spiked eigenvalues and eigenvectors. Although posterior sampling is computationally efficient due to conjugacy, a bias may exist in the posterior eigenvalue estimates under high-dimensional settings. To address this, we propose two bias correction strategies: (i) a hyperparameter adjustment method, and (ii) a post-hoc multiplicative correction. For inferring the number of spikes, we develop a BIC-type approximation to the marginal likelihood and prove posterior consistency in the high-dimensional regime $p>n$. Furthermore, we establish concentration inequalities and posterior contraction rates for the leading eigenstructure, demonstrating minimax optimality for the spiked eigenvector in the single-spike case. Simulation studies and a real data application show that our method performs better than existing approaches in providing accurate quantification of uncertainty for both eigenstructure estimation and estimation of the number of spikes.

math.ST

Bayesian Bootstrap based Gaussian Copula Model for Mixed Data with High Missing Rates

Missing data is a common issue in various fields such as medicine, social sciences, and natural sciences, and it poses significant challenges for accurate statistical analysis. Although numerous imputation methods have been proposed to address this issue, many of them fail to adequately capture the complex dependency structure among variables. To overcome this limitation, models based on the Gaussian copula framework have been introduced. However, most existing copula-based approaches do not account for the uncertainty in the marginal distributions, which can lead to biased marginal estimates and degraded performance, especially under high missingness rates. In this study, we propose a Bayesian bootstrap-based Gaussian Copula model (BBGC) that explicitly incorporates uncertainty in the marginal distributions of each variable. The proposed BBGC combines the flexible dependency modeling capability of the Gaussian copula with the Bayesian uncertainty quantification of marginal cumulative distribution functions (CDFs) via the Bayesian bootstrap. Furthermore, it is extended to handle mixed data types by incorporating methods for ordinal variable modeling. Through simulation studies and experiments on real-world datasets from the UCI repository, we demonstrate that the proposed BBGC outperforms existing imputation methods across various missing rates and mechanisms (MCAR, MAR). Additionally, the proposed model shows superior performance on real semiconductor manufacturing process data compared to conventional imputation approaches.

stat.ME

Conditional Dirichlet Processes and Functional Condition Models

In this paper, we study the conditional Dirichlet process (cDP) when a functional of a random distribution is specified. Specifically, we apply the cDP to the functional condition model, a nonparametric model in which a finite-dimensional parameter of interest is defined as the solution to a functional equation of the distribution. We derive both the posterior distribution of the parameter of interest and the posterior distribution of the underlying distribution itself. We establish two general limiting theorems for the posterior: one as the total mass of the Dirichlet process parameter tends to zero, and another as the sample size tends to infinity. We consider two specific models, the quantile model and the moment model, and propose algorithms for posterior computation, accompanied by illustrative data analysis examples. As a byproduct, we show that the Jeffreys substitute likelihood emerges as the limit of the marginal posterior in the functional condition model with a cDP prior, thereby providing a theoretical justification that has so far been lacking.

math.ST

Pedagogy-R1: Pedagogically-Aligned Reasoning Model with Balanced Educational Benchmark

Recent advances in large reasoning models (LRMs) show strong performance in structured domains such as mathematics and programming; however, they often lack pedagogical coherence and realistic teaching behaviors. To bridge this gap, we introduce Pedagogy-R1, a framework that adapts LRMs for classroom use through three innovations: (1) a distillation-based pipeline that filters and refines model outputs for instruction-tuning, (2) the Well-balanced Educational Benchmark (WBEB), which evaluates performance across subject knowledge, pedagogical knowledge, tracing, essay scoring, and teacher decision-making, and (3) a Chain-of-Pedagogy (CoP) prompting strategy for generating and eliciting teacher-style reasoning. Our mixed-method evaluation combines quantitative metrics with qualitative analysis, providing the first systematic assessment of LRMs' pedagogical strengths and limitations.

cs.AI

Cross section Measurements for $^{12}$C$(K^-, K^+Ξ^-)$ and $^{12}$C$(K^-, K^+ΛΛ)$ Reactions at 1.8 GeV$/c$

We present a measurement of the production of $Ξ^-$ and $ΛΛ$ in the $^{12}$C$(K^-, K^+)$ reaction at an incident beam momentum of 1.8 GeV/$\mathit{c}$, based on high-statistics data from J-PARC E42. The cross section for the $^{12}$C$(K^-, K^+Ξ^-)$ reaction, compared to the inclusive $^{12}$C$(K^-, K^+)$ reaction cross section, indicates that the $Ξ^-$ escaping probability peaks at 70\% in the energy region of $E_Ξ=$100 to 150 MeV above the $Ξ^-$ emission threshold. A classical approach using eikonal approximation shows that the total cross sections for $Ξ^-$ inelastic scattering ranges between 42 mb and 23 mb in the $Ξ^-$ momentum range from 0.4 to 0.6 GeV/c. Furthermore, based on the relative cross section for the $^{12}$C$(K^-, K^+ΛΛ)$ reaction, the total cross section for $Ξ^-p\toΛΛ$ is estimated in the same approach to vary between 2.2 mb and 1.0 mb in the momentum range of 0.40 to 0.65 GeV/c. Specifically, a cross section of 1.0 mb in the momentum range of 0.5 to 0.6 GeV/c imposes a constraint on the upper bound of the decay width of the $Ξ^-$ particle in infinite nuclear matter, revealing $Γ_Ξ< \sim 0.6$ MeV.

nucl-ex

STRAW: A Stress-Aware WL-Based Read Reclaim Technique for High-Density NAND Flash-Based SSDs

Although read disturbance has emerged as a major reliability concern, managing read disturbance in modern NAND flash memory has not been thoroughly investigated yet. From a device characterization study using real modern NAND flash memory, we observe that reading a page incurs heterogeneous reliability impacts on each WL, which makes the existing block-level read reclaim extremely inefficient. We propose a new WL-level read-reclaim technique, called STRAW, which keeps track of the accumulated read-disturbance effect on each WL and reclaims only heavily-disturbed WLs. By avoiding unnecessary read-reclaim operations, STRAW reduces read-reclaim-induced page writes by 83.6\% with negligible storage overhead.

cs.AR

Error analysis for finite element operator learning methods for solving parametric second-order elliptic PDEs

In this paper, we provide a theoretical analysis of a type of operator learning method without data reliance based on the classical finite element approximation, which is called the finite element operator network (FEONet). We first establish the convergence of this method for general second-order linear elliptic PDEs with respect to the parameters for neural network approximation. In this regard, we address the role of the condition number of the finite element matrix in the convergence of the method. Secondly, we derive an explicit error estimate for the self-adjoint case. For this, we investigate some regularity properties of the solution in certain function classes for a neural network approximation, verifying the sufficient condition for the solution to have the desired regularity. Finally, we will also conduct some numerical experiments that support the theoretical findings, confirming the role of the condition number of the finite element matrix in the overall convergence.

math.NA

Scalable and optimal Bayesian inference for sparse covariance matrices via screened beta-mixture prior

In this paper, we propose a scalable Bayesian method for sparse covariance matrix estimation by incorporating a continuous shrinkage prior with a screening procedure. In the first step of the procedure, the off-diagonal elements with small correlations are screened based on their sample correlations. In the second step, the posterior of the covariance with the screened elements fixed at $0$ is computed with the beta-mixture prior. The screened elements of the covariance significantly increase the efficiency of the posterior computation. The simulation studies and real data applications show that the proposed method can be used for the high-dimensional problem with the `large p, small n'. In some examples in this paper, the proposed method can be computed in a reasonable amount of time, while no other existing Bayesian methods can be. The proposed method has also sound theoretical properties. The screening procedure has the sure screening property and the selection consistency, and the posterior has the optimal minimax or nearly minimax convergence rate under the Frobeninus norm.

stat.ME

Asymptotic Properties for Bayesian Neural Network in Besov Space

Neural networks have shown great predictive power when dealing with various unstructured data such as images and natural languages. The Bayesian neural network captures the uncertainty of prediction by putting a prior distribution for the parameter of the model and computing the posterior distribution. In this paper, we show that the Bayesian neural network using spike-and-slab prior has consistency with nearly minimax convergence rate when the true regression function is in the Besov space. Even when the smoothness of the regression function is unknown the same posterior convergence rate holds and thus the spike-and-slab prior is adaptive to the smoothness of the regression function. We also consider the shrinkage prior, which is more feasible than other priors, and show that it has the same convergence rate. In other words, we propose a practical Bayesian neural network with guaranteed asymptotic properties.

stat.ML