SearcharxivSearch

arXiv subjects

Ying-Ju Chen

Publications and source records attributed to Ying-Ju Chen.

9 recordsLinked to original sources

Reliable Decision Support with LLMs: A Framework for Evaluating Consistency in Binary Text Classification Applications

This study introduces a framework for evaluating consistency in large language model (LLM) binary text classification, addressing the lack of established reliability assessment methods. Adapting psychometric principles, we determine sample size requirements, develop metrics for invalid responses, and evaluate intra- and inter-rater reliability. Our case study examines financial news sentiment classification across 14 LLMs (including claude-3-7-sonnet, gpt-4o, deepseek-r1, gemma3, llama3.2, phi4, and command-r-plus), with five replicates per model on 1,350 articles. Models demonstrated high intra-rater consistency, achieving perfect agreement on 90-98% of examples, with minimal differences between expensive and economical models from the same families. When validated against StockNewsAPI labels, models achieved strong performance (accuracy 0.76-0.88), with smaller models like gemma3:1B, llama3.2:3B, and claude-3-5-haiku outperforming larger counterparts. All models performed at chance when predicting actual market movements, indicating task constraints rather than model limitations. Our framework provides systematic guidance for LLM selection, sample size planning, and reliability assessment, enabling organizations to optimize resources for classification tasks.

cs.CL

Adapting OpenAI's CLIP Model for Few-Shot Image Inspection in Manufacturing Quality Control: An Expository Case Study with Multiple Application Examples

This expository paper introduces a simplified approach to image-based quality inspection in manufacturing using OpenAI's CLIP (Contrastive Language-Image Pretraining) model adapted for few-shot learning. While CLIP has demonstrated impressive capabilities in general computer vision tasks, its direct application to manufacturing inspection presents challenges due to the domain gap between its training data and industrial applications. We evaluate CLIP's effectiveness through five case studies: metallic pan surface inspection, 3D printing extrusion profile analysis, stochastic textured surface evaluation, automotive assembly inspection, and microstructure image classification. Our results show that CLIP can achieve high classification accuracy with relatively small learning sets (50-100 examples per class) for single-component and texture-based applications. However, the performance degrades with complex multi-component scenes. We provide a practical implementation framework that enables quality engineers to quickly assess CLIP's suitability for their specific applications before pursuing more complex solutions. This work establishes CLIP-based few-shot learning as an effective baseline approach that balances implementation simplicity with robust performance, demonstrated in several manufacturing quality control applications.

cs.CV

Flexible Modeling of Multivariate Skewed and Heavy-Tailed Data via a Non-Central Skew t Distribution: Application to Tumor Shape Data

We propose a flexible formulation of the multivariate non-central skew t (NCST) distribution, defined by scaling skew-normal random vectors with independent chi-squared variables. This construction extends the classical multivariate t family by allowing both asymmetry and non-centrality, which provides an alternative to existing skew t models that often rely on restrictive assumptions for tractability. We derive key theoretical properties of the NCST distribution, which includes its moment structure, affine transformation behavior, and the distribution of quadratic forms. Due to the lack of a closed-form density, we implement a Monte Carlo likelihood approximation to enable maximum likelihood estimation and evaluate its performance through simulation studies. To demonstrate practical utility, we apply the NCST model to breast cancer diagnostic data, modeling multiple features of tumor shape. The NCST model achieves a superior fit based on information criteria and visual diagnostics, particularly in the presence of skewness and heavy tails compared to standard alternatives, including the multivariate normal, skew normal, and Azzalini's skew $t$ distribution. Our findings suggest that the NCST distribution offers a useful and interpretable choice for modeling complex multivariate data, which highlights promising directions for future development in likelihood inference, Bayesian computation, and applications involving asymmetry and non-Gaussian dependence.

stat.ME

ChatISA: A Prompt-Engineered, In-House Multi-Modal Generative AI Chatbot for Information Systems Education

As generative AI ('GenAI') continues to evolve, educators face the challenge of preparing students for a future where AI-assisted work is integral to professional success. This paper introduces ChatISA, an in-house, multi-model AI chatbot designed to support students and faculty in an Information Systems and Analytics (ISA) department. ChatISA comprises four primary modules: Coding Companion, Project Coach, Exam Ally, and Interview Mentor, each tailored to enhance different aspects of the educational experience. Through iterative development, student feedback, and leveraging open-source frameworks, we created a robust tool that addresses coding inquiries, project management, exam preparation, and interview readiness. The implementation of ChatISA provided valuable insights and highlighted key challenges. Our findings demonstrate the benefits of ChatISA for ISA education while underscoring the need for adaptive pedagogy and proactive engagement with AI tools to fully harness their educational potential. To support broader adoption and innovation, all code for ChatISA is made publicly available on GitHub, enabling other institutions to customize and integrate similar AI-driven educational tools within their curricula.

cs.CY

Introducing ChatSQC: Enhancing Statistical Quality Control with Augmented AI

We introduce ChatSQC, an innovative chatbot system that combines the power of OpenAI's Large Language Models (LLM) with a specific knowledge base in Statistical Quality Control (SQC). Our research focuses on enhancing LLMs using specific SQC references, shedding light on how data preprocessing parameters and LLM selection impact the quality of generated responses. By illustrating this process, we hope to motivate wider community engagement to refine LLM design and output appraisal techniques. We also highlight potential research opportunities within the SQC domain that can be facilitated by leveraging ChatSQC, thereby broadening the application spectrum of SQC. A primary goal of our work is to provide a template and proof-of-concept on how LLMs can be utilized by our community. To continuously improve ChatSQC, we ask the SQC community to provide feedback, highlight potential issues, request additional features, and/or contribute via pull requests through our public GitHub repository. Additionally, the team will continue to explore adding supplementary reference material that would further improve the contextual understanding of the chatbot. Overall, ChatSQC serves as a testament to the transformative potential of AI within SQC, and we hope it will spur further advancements in the integration of AI in this field.

cs.HC

A Markov-Modulated (s, S) Inventory System with Repeated Calls and Blocked Demands

In this article, we consider a continuous review (s, S) inventory system with failures of demand fulfillment (service) modeled as a Markov-modulated retrial queueing system. The inventory system features a single product that experiences Markovian inter-demand and service intervals with random service interruptions and instantaneous replenishments. A recently developed criterion for the ergodicity of a class of discrete-time level-dependent-quasi-birth-and-death (LDQBD) processes with convergent transition matrix rows is applied to the jump chain of the process in order to elicit a closed-form traffic-intensity formula. An analytic solution for the steady-state average minimum cost is provided.

math.PR

How Generative AI models such as ChatGPT can be (Mis)Used in SPC Practice, Education, and Research? An Exploratory Study

Generative Artificial Intelligence (AI) models such as OpenAI's ChatGPT have the potential to revolutionize Statistical Process Control (SPC) practice, learning, and research. However, these tools are in the early stages of development and can be easily misused or misunderstood. In this paper, we give an overview of the development of Generative AI. Specifically, we explore ChatGPT's ability to provide code, explain basic concepts, and create knowledge related to SPC practice, learning, and research. By investigating responses to structured prompts, we highlight the benefits and limitations of the results. Our study indicates that the current version of ChatGPT performs well for structured tasks, such as translating code from one language to another and explaining well-known concepts but struggles with more nuanced tasks, such as explaining less widely known terms and creating code from scratch. We find that using new AI tools may help practitioners, educators, and researchers to be more efficient and productive. However, in their current stages of development, some results are misleading and wrong. Overall, the use of generative AI models in SPC must be properly validated and used in conjunction with other methods to ensure accurate results.

cs.LG

Combining Spot and Futures Markets: A Hybrid Market Approach to Dynamic Spectrum Access

Dynamic spectrum access is a new paradigm of secondary spectrum utilization and sharing. It allows unlicensed secondary users (SUs) to exploit opportunistically the under-utilized licensed spectrum. Market mechanism is a widely-used promising means to regulate the consuming behaviours of users and, hence, achieves the efficient allocation and consumption of limited resources. In this paper, we propose and study a hybrid secondary spectrum market consisting of both the futures market and the spot market, in which SUs (buyers) purchase under-utilized licensed spectrum from a spectrum regulator, either through predefined contracts via the futures market, or through spot transactions via the spot market. We focus on the optimal spectrum allocation among SUs in an exogenous hybrid market that maximizes the secondary spectrum utilization efficiency. The problem is challenging due to the stochasticity and asymmetry of network information. To solve this problem, we first derive an off-line optimal allocation policy that maximizes the ex-ante expected spectrum utilization efficiency based on the stochastic distribution of network information. We then propose an on-line VickreyCClarkeCGroves (VCG) auction that determines the real-time allocation and pricing of every spectrum based on the realized network information and the pre-derived off-line policy. We further show that with the spatial frequency reuse, the proposed VCG auction is NP-hard; hence, it is not suitable for on-line implementation, especially in a large-scale market. To this end, we propose a heuristics approach based on an on-line VCG-like mechanism with polynomial-time complexity, and further characterize the corresponding performance loss bound analytically. We finally provide extensive numerical results to evaluate the performance of the proposed solutions.

cs.NI

Adjusted Jackknife Empirical Likelihood

Jackknife empirical likelihood (JEL) is an effective modified version of empirical likelihood method (EL). Through the construction of the jackknife pseudo-values, JEL overcomes the computational difficulty of EL method when its constraints are nonlinear while maintaining the same asymptotic results for one sample and two-sample U statistics. In this paper, we propose an adjusted version of JEL to guarantee that the adjusted jackknife empirical likelihood (AJEL) statistic is well-defined for all the values of the parameter, instead of restricting on the convex hull of the estimation equation. The properties of JEL have been preserved for AJEL.

stat.ME