SearcharxivSearch

arXiv subjects

Yunjin Choi

Publications and source records attributed to Yunjin Choi.

9 recordsLinked to original sources

Interpretable Water Level Forecaster with Spatiotemporal Causal Attention Mechanisms

Accurate forecasting of river water levels is vital for effectively managing traffic flow and mitigating the risks associated with natural disasters. This task presents challenges due to the intricate factors influencing the flow of a river. Recent advances in machine learning have introduced numerous effective forecasting methods. However, these methods lack interpretability due to their complex structure, resulting in limited reliability. Addressing this issue, this study proposes a deep learning model that quantifies interpretability, with an emphasis on water level forecasting. This model focuses on generating quantitative interpretability measurements, which align with the common knowledge embedded in the input data. This is facilitated by the utilization of a transformer architecture that is purposefully designed with masking, incorporating a multi-layer network that captures spatiotemporal causation. We perform a comparative analysis on the Han River dataset obtained from Seoul, South Korea, from 2016 to 2021. The results illustrate that our approach offers enhanced interpretability consistent with common knowledge, outperforming competing methods and also enhances robustness against distribution shift.

cs.LG

Does a Large Language Model Really Speak in Human-Like Language?

Large Language Models (LLMs) have recently emerged, attracting considerable attention due to their ability to generate highly natural, human-like text. This study compares the latent community structures of LLM-generated text and human-written text within a hypothesis testing procedure. Specifically, we analyze three text sets: original human-written texts ($\mathcal{O}$), their LLM-paraphrased versions ($\mathcal{G}$), and a twice-paraphrased set ($\mathcal{S}$) derived from $\mathcal{G}$. Our analysis addresses two key questions: (1) Is the difference in latent community structures between $\mathcal{O}$ and $\mathcal{G}$ the same as that between $\mathcal{G}$ and $\mathcal{S}$? (2) Does $\mathcal{G}$ become more similar to $\mathcal{O}$ as the LLM parameter controlling text variability is adjusted? The first question is based on the assumption that if LLM-generated text truly resembles human language, then the gap between the pair ($\mathcal{O}$, $\mathcal{G}$) should be similar to that between the pair ($\mathcal{G}$, $\mathcal{S}$), as both pairs consist of an original text and its paraphrase. The second question examines whether the degree of similarity between LLM-generated and human text varies with changes in the breadth of text generation. To address these questions, we propose a statistical hypothesis testing framework that leverages the fact that each text has corresponding parts across all datasets due to their paraphrasing relationship. This relationship enables the mapping of one dataset's relative position to another, allowing two datasets to be mapped to a third dataset. As a result, both mapped datasets can be quantified with respect to the space characterized by the third dataset, facilitating a direct comparison between them. Our results indicate that GPT-generated text remains distinct from human-authored text.

cs.CL

Capturing usage patterns in bike sharing system via multilayer network fused Lasso

Data collected from a bike-sharing system exhibit complex temporal and spatial features. We analyze shared-bike usage data collected in three large cities at the level of individual stations, accounting for station-specific behavior and covariate effects. For this, we adopt a penalized regression approach with a multilayer network fused Lasso penalty. These fusion penalties are imposed on networks which embed spatio-temporal linkages, and capture the homogeneity in bike usage that is attributed to intricate spatio-temporal features without arbitrarily partitioning the data. On the real-life datasets, we demonstrate that the proposed approach yields competitive predictive performance and provides a new interpretation of the data.

stat.AP

Enhancing Social Media Post Popularity Prediction with Visual Content

Our study presents a framework for predicting image-based social media content popularity that focuses on addressing complex image information and a hierarchical data structure. We utilize the Google Cloud Vision API to effectively extract key image and color information from users' postings, achieving 6.8% higher accuracy compared to using non-image covariates alone. For prediction, we explore a wide range of prediction models, including Linear Mixed Model, Support Vector Regression, Multi-layer Perceptron, Random Forest, and XGBoost, with linear regression as the benchmark. Our comparative study demonstrates that models that are capable of capturing the underlying nonlinear interactions between covariates outperform other methods.

cs.LG

On the early-time behavior of quantum subharmonic generation

A few years ago Avetissian {\it et al.} \cite{Avetissian2014,Avetissian2015} discovered that the exponential growth rate of the stimulated annihilation photons from a singlet positronium Bose-Einstein condensate should be proportional to the square root of the positronium number density, not to the number density itself. In order to elucidate this surprising result obtained via a field-theoretical analysis, we point out that the basic physics involved is the same as that of resonant subharmonic transitions between two quantum oscillators. Using this model, we show that nonlinearities of the type discovered by Avetissian {\it et al.} are not unique to positronium and in fact will be encountered in a wide range of systems that can be modeled as nonlinearly coupled quantum oscillators.

quant-ph

Three-terminal heat engine and refrigerator based on superlattices

We propose a three terminal heat engine based on semiconductor superlattices for energy harvesting. The periodicity of the superlattice structure creates an energy miniband, giving an energy window for allowed electron transport. We find that this device delivers a large power, nearly twice than the heat engine based on quantum wells, with a small reduction of efficiency. This engine also works as a refrigerator in a different regime of the system's parameters. The thermoelectric performance of the refrigerator is analyzed, including the cooling power and coefficient of performance in the optimized condition. We also calculate phonon heat current through the system, and explore the reduction of phonon heat current compared to the bulk material. The direct phonon heat current is negligible at low temperatures, but dominates over the electronic at room temperature and we discuss ways to reduce it.

cond-mat.mes-hall

Selecting the number of principal components: estimation of the true rank of a noisy matrix

Principal component analysis (PCA) is a well-known tool in multivariate statistics. One significant challenge in using PCA is the choice of the number of components. In order to address this challenge, we propose an exact distribution-based method for hypothesis testing and construction of confidence intervals for signals in a noisy matrix. Assuming Gaussian noise, we use the conditional distribution of the singular values of a Wishart matrix and derive exact hypothesis tests and confidence intervals for the true signals. Our paper is based on the approach of Taylor, Loftus and Tibshirani (2013) for testing the global null: we generalize it to test for any number of principal components, and derive an integrated version with greater power. In simulation studies we find that our proposed methods compare well to existing approaches.

stat.ME

An operational approach to indirectly measuring tunneling time

The tunneling time through an arbitrary bounded one-dimensional barrier is investigated using the dwell time operator. We relate the tunneling time to the conditioned average of the dwell time operator because of the natural post-selection in the case of successful tunneling. We discuss an indirect measurement by timing the particle, and show we are able to reconstruct the conditioned average value of the dwell time operator by applying the contextual values formalism for generalized measurements based on the physics of Larmor precession. The experimentally measurable tunneling time in the weak interaction limit is given by the weak value of the dwell time operator plus a measurement-context dependent disturbance term. We show how the expectation value and higher moments of the dwell time operator can be extracted from measurement data of the particle's spin.

quant-ph

An Investigation of Methods for Handling Missing Data with Penalized Regression

We investigate methods for penalized regression in the presence of missing observations. This paper introduces a method for estimating the parameters which compensates for the missing observations. We first, derive an unbiased estimator of the objective function with respect to the missing data and then, modify the criterion to ensure convexity. Finally, we extend our approach to a family of models that embraces the mean imputation method. These approaches are compared to the mean imputation method, one of the simplest methods for dealing with missing observations problem, via simulations. We also investigate the problem of making predictions when there are missing values in the test set.

stat.AP