SearcharxivSearch

arXiv subjects

Byol Kim

Publications and source records attributed to Byol Kim.

5 recordsLinked to original sources

Spectral Differential Network Analysis for High-Dimensional Time Series

Spectral networks derived from multivariate time series data arise in many domains, from brain science to Earth science. Often, it is of interest to study how these networks change under different conditions. For instance, to better understand epilepsy, it would be interesting to capture the changes in the brain connectivity network as a patient experiences a seizure, using electroencephalography data. A common approach relies on estimating the networks in each condition and calculating their difference. Such estimates may behave poorly in high dimensions as the networks themselves may not be sparse in structure while their difference may be. We build upon this observation to develop an estimator of the difference in inverse spectral densities across two conditions. Using an L1 penalty on the difference, consistency is established by only requiring the difference to be sparse. We illustrate the method on synthetic data experiments and on experiments with electroencephalography data.

stat.ME

Black-box tests for algorithmic stability

Algorithmic stability is a concept from learning theory that expresses the degree to which changes to the input data (e.g., removal of a single data point) may affect the outputs of a regression algorithm. Knowing an algorithm's stability properties is often useful for many downstream applications -- for example, stability is known to lead to desirable generalization properties and predictive inference guarantees. However, many modern algorithms currently used in practice are too complex for a theoretical analysis of their stability properties, and thus we can only attempt to establish these properties through an empirical exploration of the algorithm's behavior on various data sets. In this work, we lay out a formal statistical framework for this kind of "black-box testing" without any assumptions on the algorithm or the data distribution and establish fundamental bounds on the ability of any black-box test to identify algorithmic stability.

cs.LG

Two-sample inference for high-dimensional Markov networks

Markov networks are frequently used in sciences to represent conditional independence relationships underlying observed variables arising from a complex system. It is often of interest to understand how an underlying network differs between two conditions. In this paper, we develop methods for comparing a pair of high-dimensional Markov networks where we allow the number of observed variables to increase with the sample sizes. By taking the density ratio approach, we are able to learn the network difference directly and avoid estimating the individual graphs. Our methods are thus applicable even when the individual networks are dense as long as their difference is sparse. We prove finite-sample Gaussian approximation error bounds for the estimator we construct under significantly weaker assumptions than are typically required for model selection consistency. Furthermore, we propose bootstrap procedures for estimating quantiles of a max-type statistics based on our estimator, and show how they can be used to test the equality of two Markov networks or construct simultaneous confidence intervals. The performance of our methods is demonstrated through extensive simulations. The scientific usefulness is illustrated with an analysis of a new fMRI dataset.

stat.ME

Predictive Inference Is Free with the Jackknife+-after-Bootstrap

Ensemble learning is widely used in applications to make predictions in complex decision problems---for example, averaging models fitted to a sequence of samples bootstrapped from the available training data. While such methods offer more accurate, stable, and robust predictions and model estimates, much less is known about how to perform valid, assumption-lean inference on the output of these types of procedures. In this paper, we propose the jackknife+-after-bootstrap (J+aB), a procedure for constructing a predictive interval, which uses only the available bootstrapped samples and their corresponding fitted models, and is therefore "free" in terms of the cost of model fitting. The J+aB offers a predictive coverage guarantee that holds with no assumptions on the distribution of the data, the nature of the fitted model, or the way in which the ensemble of models are aggregated---at worst, the failure rate of the predictive interval is inflated by a factor of 2. Our numerical experiments verify the coverage and accuracy of the resulting predictive intervals on real data.

stat.ME

Interface Engineering in Hybrid Iodide CH3NH3PbI3 Perovskite Using Lewis Base and Graphene towards High Performance Solar Cells

Perovskite solar cells have achieved a substantial breakthrough via advanced interface engineerings. Reports have emphasized that combining the hybrid perovskites with Lewis base and graphene improve the performance; the underlying mechanisms are not yet fully understood. Here, using density functional theory, we show that upon the formation of CH3NH3PbI3 interfaces with three different Lewis base molecules and graphene, the binding strength with S-donors thiocarbamide and thioacetamide is higher than with O-donor dimethyl sulfoxide, while the interface dipole and work function reduction tend to increase from S-donors to O-donor. Furthermore, we provide evidences of deep trap states elimination in the S-donor perovskite interfaces through the analysis of defect formation on CH3NH3PbI3(110) surface, and of stability enhancement by estimating activation barriers for iodine atom migrations. These theoretical predictions are in line with the experimental observation of performance enhancement in the perovskites prepared using thiocarbamide.

physics.app-ph