Searcharxiv⌕ Search

arXiv subjects

Nathakhun Wiroonsri

Publications and source records attributed to Nathakhun Wiroonsri.

14 recordsLinked to original sources

An approximate zero bias transformation for random sums: Applications to sampling with outliers, auto insurance, and generative AI

We develop $L^1$ bounds for the difference between a test function of a random sum and a standard normal random variable, where the summands are assumed to be independent but not necessarily identically distributed. The bounds are obtained through a new version of the approximate zero bias transformation specifically developed for random sums. Although the identical distribution assumption is relaxed, the bounds are of order $1/\sqrt{n}$, matching the order of existing bounds in the literature under the same distributional assumption on the number of summands. The main results are then applied to three real-world settings: random sums obtained from simple random sampling with outliers, total insurance claims, and generative AI response times.

math.PR↗

4TaStiC: Time and trend traveling time series clustering for classifying long-term type 2 diabetes patients

Diabetes is one of the most prevalent diseases worldwide, characterized by persistently high blood sugar levels, capable of damaging various internal organs and systems. Diabetes patients require routine check-ups, resulting in a time series of laboratory records, such as hemoglobin A1c, which reflects each patient's health behavior over time and informs their doctor's recommendations. Clustering patients into groups based on their entire time series data assists doctors in making recommendations and choosing treatments without the need to review all records. However, time series clustering of this type of dataset introduces some challenges; patients visit their doctors at different time points, making it difficult to capture and match trends, peaks, and patterns. Additionally, two aspects must be considered: differences in the levels of laboratory results and differences in trends and patterns. To address these challenges, we introduce a new clustering algorithm called Time and Trend Traveling Time Series Clustering (4TaStiC), using a base dissimilarity measure combined with Euclidean and Pearson correlation metrics. We evaluated this algorithm on artificial datasets, comparing its performance with that of seven existing methods. The results show that 4TaStiC outperformed the other methods on the targeted datasets. Finally, we applied 4TaStiC to cluster a cohort of 1,989 type 2 diabetes patients at Siriraj Hospital. Each group of patients exhibits clear characteristics that will benefit doctors in making efficient clinical decisions. Furthermore, the proposed algorithm can be applied to contexts outside the medical field.

cs.LG↗

Ranked differences Pearson correlation dissimilarity with an application to electricity users time series clustering

Time series clustering is an unsupervised learning method for classifying time series data into groups with similar behavior. It is used in applications such as healthcare, finance, economics, energy, and climate science. Several time series clustering methods have been introduced and used for over four decades. Most of them focus on measuring either Euclidean distances or association dissimilarities between time series. In this work, we propose a new dissimilarity measure called ranked Pearson correlation dissimilarity (RDPC), which combines a weighted average of a specified fraction of the largest element-wise differences with the well-known Pearson correlation dissimilarity. It is incorporated into hierarchical clustering. The performance is evaluated and compared with existing clustering algorithms. The results show that the RDPC algorithm outperforms others in complicated cases involving different seasonal patterns, trends, and peaks. Finally, we demonstrate our method by clustering a random sample of customers from a Thai electricity consumption time series dataset into seven groups with unique characteristics.

stat.ML↗

A correlation-based fuzzy cluster validity index with secondary options detector

The optimal number of clusters is one of the main concerns when applying cluster analysis. Several cluster validity indexes have been introduced to address this problem. However, in some situations, there is more than one option that can be chosen as the final number of clusters. This aspect has been overlooked by most of the existing works in this area. In this study, we introduce a correlation-based fuzzy cluster validity index known as the Wiroonsri-Preedasawakul (WP) index. This index is defined based on the correlation between the actual distance between a pair of data points and the distance between adjusted centroids with respect to that pair. We evaluate and compare the performance of our index with several existing indexes, including Xie-Beni, Pakhira-Bandyopadhyay-Maulik, Tang, Wu-Li, generalized C, and Kwon2. We conduct this evaluation on four types of datasets: artificial datasets, real-world datasets, simulated datasets with ranks, and image datasets, using the fuzzy c-means algorithm. Overall, the WP index outperforms most, if not all, of these indexes in terms of accurately detecting the optimal number of clusters and providing accurate secondary options. Moreover, our index remains effective even when the fuzziness parameter $m$ is set to a large value. Our R package called UniversalCVI used in this work is available at https://CRAN.R-project.org/package=UniversalCVI.

stat.ML↗

A Bayesian cluster validity index

Selecting the appropriate number of clusters is a critical step in applying clustering algorithms. To assist in this process, various cluster validity indices (CVIs) have been developed. These indices are designed to identify the optimal number of clusters within a dataset. However, users may not always seek the absolute optimal number of clusters but rather a secondary option that better aligns with their specific applications. This realization has led us to introduce a Bayesian cluster validity index (BCVI), which builds upon existing indices. The BCVI utilizes either Dirichlet or generalized Dirichlet priors, resulting in the same posterior distribution. We evaluate our BCVI using the Wiroonsri index for hard clustering and the Wiroonsri-Preedasawakul index for soft clustering as underlying indices. We compare the performance of our proposed BCVI with that of the original underlying indices and several other existing CVIs, including Davies-Bouldin, Starczewski, Xie-Beni, and KWON2 indices. Our BCVI offers clear advantages in situations where user expertise is valuable, allowing users to specify their desired range for the final number of clusters. To illustrate this, we conduct experiments classified into three different scenarios. Additionally, we showcase the practical applicability of our approach through real-world datasets, such as MRI brain tumor images. These tools will be published as a new R package 'BayesCVI'.

stat.ML↗

The Expected Values of Hosoya Index and Merrifield-Simmons Index of Random Hexagonal Cacti

Hosoya index and Merrifield-Simmons index are two well-known topological descriptors that reflex some physical properties, boiling point or heat of formation for instance, of bezenoid hydrocarbon compounds. In this paper, we establish the generating functions of the expected values of these two indices of random hexagonal cacti. This generalizes the results of Doslic and Maloy, published in Discrete Mathemaics, in 2010. By applying the ideas on meromorphic functions and the growth of power series coefficients, the asymptotic behaviors of these indices on the random cacti have been established.

math.CO↗

Clustering performance analysis using a new correlation-based cluster validity index

There are various cluster validity indices used for evaluating clustering results. One of the main objectives of using these indices is to seek the optimal unknown number of clusters. Some indices work well for clusters with different densities, sizes, and shapes. Yet, one shared weakness of those validity indices is that they often provide only one optimal number of clusters. That number is unknown in real-world problems, and there might be more than one possible option. We develop a new cluster validity index based on a correlation between an actual distance between a pair of data points and a centroid distance of clusters that the two points occupy. Our proposed index constantly yields several local peaks and overcomes the previously stated weakness. Several experiments in different scenarios, including UCI real-world data sets, have been conducted to compare the proposed validity index with several well-known ones. An R package related to this new index called NCvalid is available at https://github.com/nwiroonsri/NCvalid.

stat.ML↗

Normal approximation for fire incident simulation using permanental Cox processes

Estimating the number of natural disasters benefits the insurance industry in terms of risk management. However, the estimation process is complicated due to the fact that there are many factors affecting the number of such incidents. In this work, we propose a Normal approximation technique for associated point processes for estimating the number of natural disasters under the following two assumptions: 1) the incident counts in any two distinct areas are positively associated and 2) the association between these counts in two distinct areas decays exponentially with respect to distance outside some small local neighborhood. Under the stated assumptions, we extend previous results for the Normal approximation technique for associated point processes, i.e., the establishment of non-asymptotic $L^1$ bounds for the functionals of these processes [Wiroonsri (2019)]. Then we apply this new result to permanental Cox processes that are known to be positively associated. Finally, we apply our Normal approximation results for permanental Cox processes to Thailand's fire data from 2007 to 2020, which was collected by the Geo-Informatics and Space Technology Development Agency of Thailand.

math.PR↗

Do elderly want to work? Modeling elderly's decision to fight aging Thailand

Thailand has entered into an aging society since the year 2000. Using the 2017 Survey of the Older Persons in Thailand collected by Thailand National Statistical Office, this study uses cross tabulation, random forest with variable importance measure and lasso logistic regression to examine factors that have effects on the elderly's decision to remain in the labor market after retirement. This study reveals that these following variables: age, education level, healthcare eligibility, marital status, health condition, total assets, gender, residential type, percent of elderly in the household, and number of children have strong influences on an elderly's desire to continue work. By knowing which factors contribute to the elderly wish to continue work in the market, this research allows for future prediction of the labor market that can accommodate elderly in Thailand. Our final models of random forest and lasso logistic regression provide prediction accuracy of 68.19 and 69.58 percent on the elderly's desire to work, respectively. This study has a significant impact as policymakers can utilize our models in predicting elderly's desire to work after retirement age and design a labor market that can accommodate elderly in Thailand in the future.

stat.AP↗

Concentration inequalities using approximate zero bias couplings with applications to Hoeffding's statistic under the Ewens distribution

We prove concentration inequalities of the form $P(Y \ge t) \le \exp(-B(t))$ for a random variable $Y$ with mean zero and variance $σ^2$ using a coupling technique from Stein's method that is so-called approximate zero bias couplings. Applications to the Hoeffding's statistic where the random permutation has the Ewens distribution with parameter $θ>0$ are also presented. A few simulation experiments are then provided to visualize the tail probability of the Hoeffding's statistic and our bounds. Based on the simulation results, our bounds work well especially when $θ\le 1$.

math.PR↗

Normal approximation for associated point processes via Stein's method with applications to determinantal point processes

We use Stein's method to provide non asymptotic $L^1$ bounds to the normal for functionals of associated point processes. As for supporting tools, we use the connection between association and $α$-mixing properties that was recently uncovered by [PDL17]. We apply our main results to determinantal point processes which are known to be negatively associated. A potential application to point processes in the Laguerre-Gaussian family is also presented.

math.PR↗

Stein's method for negatively associated random variables with applications to second order stationary random fields

Let $\boldsymbolξ=(ξ_1,\ldots,ξ_m)$ be a negatively associated mean zero random vector with components that obey the bound $|ξ_i| \le B, i=1,\ldots,m$, and whose sum $W = \sum_{i=1}^m ξ_i$ has variance 1, the bound \[ d_1\big({\cal L}(W),{\cal L}(Z)\big) \le 5B - 5.2\sum_{i \not = j} σ_{ij}. \] is obtained where $Z$ has the standard normal distribution and $d_1(\cdot,\cdot)$ is the $L^1$ metric. The result is extended to the multidimensional case with the $L^1$ metric replaced by a smooth functions metric. Applications to second order stationary random fields with exponential decreasing covariance are also presented.

math.PR↗

Stein's method using approximate zero bias couplings with applications to combinatorial central limit theorems under the Ewens distribution

We generalize the well-known zero bias distribution and the $λ$-Stein pair to an approximate zero bias distribution and an approximate $λ,R$-Stein pair, respectively. Berry Esseen type bounds to the normal, based on approximate zero bias couplings and approximate $λ,R$-Stein pairs, are obtained using Stein's method. The bounds are then applied to combinatorial central limit theorems where the random permutation has the Ewens $\mathcal{E}_θ$ distribution with $θ>0$ which can be specialized to the uniform distribution by letting $θ=1$. The family of the Ewens distributions appears in the context of population genetics in biology.

math.PR↗

Stein's method for positively associated random variables with applications to the Ising and voter models, bond percolation, and contact process

We provide non-asymptotic $L^1$ bounds to the normal for four well-known models in statistical physics and particle systems in $\mathbb{Z}^d$; the ferromagnetic nearest-neighbor Ising model, the supercritical bond percolation model, the voter model and the contact process. In the Ising model, we obtain an $L^1$ distance bound between the total magnetization and the normal distribution at any temperature when the magnetic moment parameter is nonzero, and when the inverse temperature is below critical and the magnetic moment parameter is zero. In the percolation model we obtain such a bound for the total number of points in a finite region belonging to an infinite cluster in dimensions $d \ge 2$, in the voter model for the occupation time of the origin in dimensions $d \ge 7$, and for finite time integrals of non-constant increasing cylindrical functions evaluated on the one dimensional supercritical contact process started in its unique invariant distribution. The tool developed for these purposes is a version of Stein's method adapted to positively associated random variables. In one dimension, letting $\boldsymbolξ=(ξ_1,\ldots,ξ_m)$ be a positively associated mean zero random vector with components that obey the bound $|ξ_i| \le B, i=1,\ldots,m$, and whose sum $W = \sum_{i=1}^m ξ_i$ has variance 1, it holds that $$ d_1 \left(\mathcal{L}(W),\mathcal{L}(Z) \right) \leq 5B + \sqrt{\frac{8}π}\sum_{i \neq j} \mathbb{E}[ξ_i ξ_j] $$ where $Z$ has the standard normal distribution and $d_1(\cdot,\cdot)$ is the $L^1$ metric. Our methods apply in the multidimensional case with the $L^1$ metric replaced by a smooth function metric.

math.PR↗