SearcharxivSearch

arXiv subjects

Kexin Xie

Publications and source records attributed to Kexin Xie.

6 recordsLinked to original sources

TCARD: Nearly Balanced Two-Level Designs with Treatment Cardinality Constraints with an Application to LLM Prompt Engineering

Modern experimental designs often face the so-called treatment cardinality constraint, which is the constraint on the number of included factors in each treatment. Experiments with such constraints are commonly encountered in engineering simulation, AI system tuning, and large-scale system verification. This calls for the development of adequate designs to enable statistical efficiency for modeling and analysis within feasible constraints. In this work, we study two-level designs under this $k$-treatment cardinality constraint (TCARD), where the design matrix $\mathbf{X} \in \{0,1\}^{n \times p}$ has constant row sums equal to $k$. Although TCARDs are closely related to balanced incomplete block designs (BIBDs), exact BIBD structure is unavailable for many practical $(n,p,k)$ combinations. This leads to the notion of nearly balanced TCARDs, which we prove minimize the first two components of the generalized word-length pattern. We also show that good projection behavior in this setting is governed by two count-based regularities: balanced factor replications and uniform pairwise concurrences. Motivated by this characterization, we then propose the Balanced Concurrence Deviation ($Φ_{\mathrm{BCD}}$), a model-free objective that jointly penalizes replication imbalance and concurrence dispersion. We further show that this criterion is closely connected to classical optimality principles, including $(M,S)$-optimality, centered $\mathrm{UE}(s^2)$ criterion, and Bayesian $D$-optimality. To construct designs minimizing $Φ_{\mathrm{BCD}}$, we develop a coordinate-exchange (CE) algorithm with efficient incremental updates, together with a simulation-based procedure for calibrating the criterion weights to the intended downstream task. Numerical experiments confirm that the proposed method compares favorably with existing alternatives across a range of problem sizes and constraint strengths.

stat.ME

Adaptive Bi-Level Variable Selection of Conditional Main Effects for Generalized Linear Models

Understanding interaction effects among variables is important for regression modeling in various applications. The conventional approach of quantifying interactions as the product of variables often lacks clear interpretability, especially in complex systems. The concept of conditional main effects (CME) provides a more intuitive and interpretable framework for capturing interaction effects by quantifying the effect of one variable conditional on the level of another. A recent method called cmenet further considered the bi-level selection of CMEs by leveraging their natural grouping structure (e.g., sibling and cousin groups) through penalization. However, there are several limitations in the cmenet method, including the coupling ability of penalties for within-group CMEs, lack of adaptiveness for between-group penalties, and restriction to linear models with continuous responses. To overcome these limitations, we propose an adaptive cmenet method for CME selection under the generalized linear model (GLM) framework. The proposed method considers a penalized likelihood approach with adaptive weights to enable effective bi-level variable selection, improving both between-group and within-group selection. An efficient algorithm for parameter estimation is also developed by employing an iteratively reweighted least squares procedure. The performance of the proposed method is evaluated by both simulation studies and real-data studies in gene association analysis.

stat.ME

StatLLM: A Dataset for Evaluating the Performance of Large Language Models in Statistical Analysis

The coding capabilities of large language models (LLMs) have opened up new opportunities for automatic statistical analysis in machine learning and data science. However, before their widespread adoption, it is crucial to assess the accuracy of code generated by LLMs. A major challenge in this evaluation lies in the absence of a benchmark dataset for statistical code (e.g., SAS and R). To fill in this gap, this paper introduces StatLLM, an open-source dataset for evaluating the performance of LLMs in statistical analysis. The StatLLM dataset comprises three key components: statistical analysis tasks, LLM-generated SAS code, and human evaluation scores. The first component includes statistical analysis tasks spanning a variety of analyses and datasets, providing problem descriptions, dataset details, and human-verified SAS code. The second component features SAS code generated by ChatGPT 3.5, ChatGPT 4.0, and Llama 3.1 for those tasks. The third component contains evaluation scores from human experts in assessing the correctness, effectiveness, readability, executability, and output accuracy of the LLM-generated code. We also illustrate the unique potential of the established benchmark dataset for (1) evaluating and enhancing natural language processing metrics, (2) assessing and improving LLM performance in statistical coding, and (3) developing and testing of next-generation statistical software - advancements that are crucial for data science and machine learning research.

stat.AP

Performance Evaluation of Large Language Models in Statistical Programming

The programming capabilities of large language models (LLMs) have revolutionized automatic code generation and opened new avenues for automatic statistical analysis. However, the validity and quality of these generated codes need to be systematically evaluated before they can be widely adopted. Despite their growing prominence, a comprehensive evaluation of statistical code generated by LLMs remains scarce in the literature. In this paper, we assess the performance of LLMs, including two versions of ChatGPT and one version of Llama, in the domain of SAS programming for statistical analysis. Our study utilizes a set of statistical analysis tasks encompassing diverse statistical topics and datasets. Each task includes a problem description, dataset information, and human-verified SAS code. We conduct a comprehensive assessment of the quality of SAS code generated by LLMs through human expert evaluation based on correctness, effectiveness, readability, executability, and the accuracy of output results. The analysis of rating scores reveals that while LLMs demonstrate usefulness in generating syntactically correct code, they struggle with tasks requiring deep domain understanding and may produce redundant or incorrect results. This study offers valuable insights into the capabilities and limitations of LLMs in statistical programming, providing guidance for future advancements in AI-assisted coding systems for statistical analysis.

stat.AP

Single Inclusive Jet Production in $pA$ Collisions at NLO in the small-$x$ regime

We present the first complete NLO prediction with full jet algorithm implementation for the single inclusive jet production in $pA$ collisions within the CGC effective theory. Our prediction is fully differential over the final state physical kinematics, which allows the implementation of any IR safe observable including the jet clustering procedure. The NLO calculation is organized with the aid of the power counting proposed in [1] which gives rise to the novel soft contributions in the CGC factorization. We achieve the fully-differential calculation by constructing suitable subtraction terms to handle the singularities in the real corrections. The subtraction contributions can be exactly integrated analytically. We present the NLO cross section with the jets constructed using the anti-$k_T$ algorithm. The NLO calculation demonstrates explicitly the validity of the CGC factorization in jet production. Furthermore, as a byproduct of the subtraction method, we also derive the fully analytic cross section for the forward jet production in the small-$R$ limit. We show that in the small-$R$ approximation, the forward jet cross section can be factorized into a semi-hard cross section that produces a parton and the semi-inclusive jet functions. We argue that this feature holds for generic jet production and jet substructure observables in the CGC framework. Last, we show numerical analyses of the derived formula to validate our calculations. We justify when the small-$R$ approximation is appropriate. Like forward hadron production, the obtained NLO result also exhibits the negativity of the cross section in the large jet transverse regime, which signals the need for the threshold resummation. A sketch of the threshold resummation in the CGC framework is presented based on the multiple emission picture.

hep-ph

Forward and Reverse Parking in a Parking Lot

The choice of forward and reverse parking in a parking lot is studied as a stochastic process. An $M/M/c/c$ queueing system is used as an initial framework. We use Monte Carlo simulation to get the relationship between vehicle orientation and vehicle entry and exit rates, as well as the most likely parking states at each specific rate. We view the change in parking status over time.

physics.soc-ph