SearcharxivSearch

arXiv subjects

Mikko Kuronen

Publications and source records attributed to Mikko Kuronen.

7 recordsLinked to original sources

Mispronunciation Detection Without L2 Pronunciation Dataset in Low-Resource Setting: A Case Study in Finland Swedish

Mispronunciation detection (MD) models are the cornerstones of many language learning applications. Unfortunately, most systems are built for English and other major languages, while low-resourced language varieties, such as Finland Swedish (FS), lack such tools. In this paper, we introduce our MD model for FS, trained on 89 hours of first language (L1) speakers' spontaneous speech and tested on 33 minutes of L2 transcribed read-aloud speech. We trained a multilingual wav2vec 2.0 model with entropy regularization, followed by temperature scaling and top-k normalization after the inference to better adapt it for MD. The main novelty of our method lies in its simplicity, requiring minimal L2 data. The process is also language-independent, making it suitable for other low-resource languages. Our proposed algorithm allows us to balance Recall (43.2%) and Precision (29.8%), compared with the baseline model's Recall (77.5%) and Precision (17.6%).

cs.CL

What you see is not what is there: Mechanisms, models, and methods for point pattern deviations

Many natural systems are observed as point patterns in time, space, or space and time. Examples include plant and cellular systems, animal colonies, earthquakes, and wildfires. In practice the locations of the points are not always observed correctly. However, in the point process literature, there has been relatively scant attention paid to the issue of errors in the location of points. In this paper, we discuss how the observed point pattern may deviate from the actual point pattern and review methods and models that exist to handle such deviations. The discussion is supplemented with several scientific illustrations.

stat.ME

Global quantile regression

Quantile regression is used to study effects of covariates on a particular quantile of the data distribution. Here we are interested in the question whether a covariate has any effect on the entire data distribution, i.e., on any of the quantiles. To this end, we treat all the quantiles simultaneously and consider global tests for the existence of the covariate effect in the presence of nuisance covariates. This global quantile regression can be used as the extension of linear regression or as the extension of distribution comparison in the sense of Kolmogorov-Smirnov test. The proposed method is based on pointwise coefficients, permutations and global envelope tests. The global envelope test serves as the multiple test adjustment procedure under the control of the family-wise error rate and provides the graphical interpretation which automatically shows the quantiles or the levels of categorical covariate responsible for the rejection. The Freedman-Lane permutation strategy showed liberality of the test for extreme quantiles, therefore we propose four alternatives that work well even for extreme quantiles and are suitable in different conditions. We present a simulation study to inspect the performance of these strategies, and we apply the chosen strategies to two data examples.

stat.ME

Hierarchical log Gaussian Cox process for regeneration in uneven-aged forests

We propose a hierarchical log Gaussian Cox process (LGCP) for point patterns, where a set of points x affects another set of points y but not vice versa. We use the model to investigate the effect of large trees to the locations of seedlings. In the model, every point in x has a parametric influence kernel or signal, which together form an influence field. Conditionally on the parameters, the influence field acts as a spatial covariate in the intensity of the model, and the intensity itself is a non-linear function of the parameters. Points outside the observation window may affect the influence field inside the window. We propose an edge correction to account for this missing data. The parameters of the model are estimated in a Bayesian framework using Markov chain Monte Carlo (MCMC) where a Laplace approximation is used for the Gaussian field of the LGCP model. The proposed model is used to analyze the effect of large trees on the success of regeneration in uneven-aged forest stands in Finland.

stat.ME

New methods for multiple testing in permutation inference for the general linear model

Permutation methods are commonly used to test significance of regressors of interest in general linear models (GLMs) for functional (image) data sets, in particular for neuroimaging applications as they rely on mild assumptions. Permutation inference for GLMs typically consists of three parts: choosing a relevant test statistic, computing pointwise permutation tests and applying a multiple testing correction. We propose new multiple testing methods as an alternative to the commonly used maximum value of test statistics across the image. The new methods improve power and robustness against inhomogeneity of the test statistic across its domain. The methods rely on sorting the permuted functional test statistics based on pointwise rank measures; still they can be implemented even for large brain data. The performance of the methods is demonstrated through a designed simulation experiment, and an example of brain imaging data. We developed the R package GET which can be used for computation of the proposed procedures.

stat.ME

Point process models for sweat gland activation observed with noise

The aim of the paper is to construct spatial models for the activation of sweat glands for healthy subjects and subjects suffering from peripheral neuropathy by using videos of sweating recorded from the subjects. The sweat patterns are regarded as realizations of spatial point processes and two point process models for the sweat gland activation and two methods for inference are proposed. Several image analysis steps are needed to extract the point patterns from the videos and some incorrectly identified sweat gland locations may be present in the data. To take into account the errors we either include an error term in the point process model or use an estimation procedure that is robust with respect to the errors.

stat.ME

Hard-core thinnings of germ-grain models with power-law grain sizes

Random sets with long-range dependence can be generated using a Boolean model with power-law grain sizes. We study thinnings of such Boolean models which have the hard-core property that no grains overlap in the resulting germ-grain model. A fundamental question is whether long-range dependence is preserved under such thinnings. To answer this question we study four natural thinnings of a Poisson germ-grain model where the grains are spheres with a regularly varying size distribution. We show that a thinning which favors large grains preserves the slow correlation decay of the original model, whereas a thinning which favors small grains does not. Our most interesting finding concerns the case where only disjoint grains are retained, which corresponds to the well-known Matérn type I thinning. In the resulting germ-grain model, typical grains have exponentially small sizes, but rather surprisingly, the long-range dependence property is still present. As a byproduct, we obtain new mechanisms for generating homogeneous and isotropic random point configurations having a power-law correlation decay.

math.PR