Searcharxiv⌕ Search

arXiv subjects

Mari Myllymäki

Publications and source records attributed to Mari Myllymäki.

At least 19 recordsLinked to original sources

A Formal Graphical Inference Framework for Combining Effect Size with Statistical Significance: Application to Multivariate and Functional Linear Models

To address the critical need for statistical methods that evaluate practical magnitude alongside statistical significance, we introduce a formal graphical inference framework that intrinsically combines effect size with uncertainty. Built upon resampling and global envelopes, this approach provides universality and direct visual interpretability. While existing envelope tests provide only weak family-wise error rate (FWER) or false discovery rate control, we advance the methodology by introducing novel step-down global envelopes. We theoretically prove that several of these envelopes achieve exact strong FWER control under exchangeability, guaranteeing rigorous inference for each individual local hypothesis. Beyond yielding adjusted $p$-values, the framework provides adjusted subset $p$-values for blocks of hypotheses, such as the functional effect of a covariate within a specific category. By explicitly visualizing the size and direction of the effect relative to its variability under the null hypothesis, our approach overcomes the dichotomous nature of traditional testing and provides profound informational value. The proposed framework is applied to multivariate and functional linear models.

stat.ME↗

Mark distance correlation functions: from moment-based to distributional mark summary characteristics in spatial point processes

With the rapid advancement in data collection devices and storage capacities, we have access to increasing amount of spatial point pattern data where each event location is augmented by multiple, potentially non-scalar marks. Therefore, there is need for efficient analysis techniques to investigate the structural relationships between the marks. In this paper, we recall the distance covariance and distance correlation and adjust them to the marked point process setting. As a result, we introduce a novel class of mark characteristics for single real-valued marks as well as multivariate combinations of marks, including mixtures of integer- and real-valued quantities, and non-scalar marks.

stat.ME↗

Multispectral airborne laser scanning for tree species classification: a benchmark of machine learning and deep learning algorithms

Climate-smart and biodiversity-preserving forestry demands precise information on forest resources, extending to the individual tree level. Multispectral airborne laser scanning (ALS) has shown promise in automated point cloud processing, but challenges remain in leveraging deep learning techniques and identifying rare tree species in class-imbalanced datasets. This study addresses these gaps by conducting a comprehensive benchmark of deep learning and traditional shallow machine learning methods for tree species classification. For the study, we collected high-density multispectral ALS data ($>1000$ $\mathrm{pts}/\mathrm{m}^2$) at three wavelengths using the FGI-developed HeliALS system, complemented by existing Optech Titan data (35 $\mathrm{pts}/\mathrm{m}^2$), to evaluate the species classification accuracy of various algorithms in a peri-urban study area located in southern Finland. We established a field reference dataset of 6326 segments across nine species using a newly developed browser-based crowdsourcing tool, which facilitated efficient data annotation. The ALS data, including a training dataset of 1065 segments, was shared with the scientific community to foster collaborative research and diverse algorithmic contributions. Based on 5261 test segments, our findings demonstrate that point-based deep learning methods, particularly a point transformer model, outperformed traditional machine learning and image-based deep learning approaches on high-density multispectral point clouds. For the high-density ALS dataset, a point transformer model provided the best performance reaching an overall (macro-average) accuracy of 87.9% (74.5%) with a training set of 1065 segments and 92.0% (85.1%) with a larger training set of 5000 segments.

cs.CV↗

On spatial point processes with composition-valued marks

Methods for marked spatial point processes with scalar marks have seen extensive development in recent years. While the impressive progress in data collection and storage capacities has yielded an immense increase in spatial point process data with highly challenging non-scalar marks, methods for their analysis are not equally well developed. In particular, there are no methods for composition-valued marks, i.e. vector-valued marks with a sum-to-constant constrain (typically 1 or 100). Prompted by the need for a suitable methodological framework, we extend existing methods to spatial point processes with composition-valued marks and adapt common mark characteristics to this context. The proposed methods are applied to analyse spatial correlations in data on tree crown-to-base and business sector compositions.

stat.ME↗

The power of visualizing distributional differences: Formal graphical $n$-sample tests

Classical tests are available for the two-sample test of correspondence of distribution functions. From these, the Kolmogorov-Smirnov test provides also the graphical interpretation of the test results, in different forms. Here, we propose modifications of the Kolmogorov-Smirnov test with higher power. The proposed tests are based on the so-called global envelope test which allows for graphical interpretation, similarly as the Kolmogorov-Smirnov test. The tests are based on rank statistics and are suitable also for the comparison of $n$ samples, with $n \geq 2$. We compare the alternatives for the two-sample case through an extensive simulation study and discuss their interpretation. Finally, we apply the tests to real data. Specifically, we compare the height distributions between boys and girls at different ages, the sepal length distributions of different flower species, and distributions of standardized residuals from a time series model for different exchange courses using the proposed methodologies.

stat.ME↗

GET: Global envelopes in R

This work describes the R package GET that implements global envelopes for a general set of $d$-dimensional vectors $T$ in various applications. A $100(1-α)$% global envelope is a band bounded by two vectors such that the probability that $T$ falls outside this envelope in any of the $d$ points is equal to $α$. The term 'global' means that this probability is controlled simultaneously for all the $d$ elements of the vectors. The global envelopes can be employed for central regions of functional or multivariate data, for graphical Monte Carlo and permutation tests where the test statistic is multivariate or functional, and for global confidence and prediction bands. Intrinsic graphical interpretation property is introduced for global envelopes. The global envelopes included in the GET package have this property, which particularly helps to interpret test results, by providing a graphical interpretation that shows the reasons of rejection of the tested hypothesis. Examples of different uses of global envelopes and their implementation in the GET package are presented, including global envelopes for single and several one- or two-dimensional functions, Monte Carlo goodness-of-fit tests for simple and composite hypotheses, comparison of distributions, functional analysis of variance, functional linear model, and confidence bands in polynomial regression.

stat.ME↗

What you see is not what is there: Mechanisms, models, and methods for point pattern deviations

Many natural systems are observed as point patterns in time, space, or space and time. Examples include plant and cellular systems, animal colonies, earthquakes, and wildfires. In practice the locations of the points are not always observed correctly. However, in the point process literature, there has been relatively scant attention paid to the issue of errors in the location of points. In this paper, we discuss how the observed point pattern may deviate from the actual point pattern and review methods and models that exist to handle such deviations. The discussion is supplemented with several scientific illustrations.

stat.ME↗

Global quantile regression

Quantile regression is used to study effects of covariates on a particular quantile of the data distribution. Here we are interested in the question whether a covariate has any effect on the entire data distribution, i.e., on any of the quantiles. To this end, we treat all the quantiles simultaneously and consider global tests for the existence of the covariate effect in the presence of nuisance covariates. This global quantile regression can be used as the extension of linear regression or as the extension of distribution comparison in the sense of Kolmogorov-Smirnov test. The proposed method is based on pointwise coefficients, permutations and global envelope tests. The global envelope test serves as the multiple test adjustment procedure under the control of the family-wise error rate and provides the graphical interpretation which automatically shows the quantiles or the levels of categorical covariate responsible for the rejection. The Freedman-Lane permutation strategy showed liberality of the test for extreme quantiles, therefore we propose four alternatives that work well even for extreme quantiles and are suitable in different conditions. We present a simulation study to inspect the performance of these strategies, and we apply the chosen strategies to two data examples.

stat.ME↗

Spatio-temporal determinantal point processes

Determinantal point processes are models for regular spatial point patterns, with appealing probabilistic properties. We present their spatio-temporal counterparts and give examples of these models, based on spatio-temporal covariance functions which are separable and non-separable in space and time.

math.ST↗

False discovery rate envelopes

False discovery rate (FDR) is a common way to control the number of false discoveries in multiple testing. There are a number of approaches available for controlling FDR. However, for functional test statistics, which are discretized into $m$ highly correlated hypotheses, the methods must account for changes in distribution across the functional domain and correlation structure. Further, it is of great practical importance to visualize the test statistic together with its rejection or acceptance region. Therefore, the aim of this paper is to find, based on resampling principles, a graphical envelope that controls FDR and detects the outcomes of all individual hypotheses by a simple rule: the hypothesis is rejected if and only if the empirical test statistic is outside of the envelope. Such an envelope offers a straightforward interpretation of the test results, similarly as the recently developed global envelope testing which controls the family-wise error rate. Two different adaptive single threshold procedures are developed to fulfill this aim. Their performance is studied in an extensive simulation study. The new methods are illustrated by three real data examples.

stat.ME↗

Hierarchical log Gaussian Cox process for regeneration in uneven-aged forests

We propose a hierarchical log Gaussian Cox process (LGCP) for point patterns, where a set of points x affects another set of points y but not vice versa. We use the model to investigate the effect of large trees to the locations of seedlings. In the model, every point in x has a parametric influence kernel or signal, which together form an influence field. Conditionally on the parameters, the influence field acts as a spatial covariate in the intensity of the model, and the intensity itself is a non-linear function of the parameters. Points outside the observation window may affect the influence field inside the window. We propose an edge correction to account for this missing data. The parameters of the model are estimated in a Bayesian framework using Markov chain Monte Carlo (MCMC) where a Laplace approximation is used for the Gaussian field of the LGCP model. The proposed model is used to analyze the effect of large trees on the success of regeneration in uneven-aged forest stands in Finland.

stat.ME↗

Tree species, crown cover, and age as determinants of the vertical distribution of airborne LiDAR returns

Light detection and ranging (LiDAR) provides information on the vertical structure of forest stands enabling detailed and extensive ecosystem study. The vertical structure is often summarized by scalar features and data-reduction techniques that limit the interpretation of results. Instead, we quantified the influence of three variables, species, crown cover, and age, on the vertical distribution of airborne LiDAR returns from forest stands. We studied 5,428 regular, even-aged stands in Quebec (Canada) with five dominant species: balsam fir (Abies balsamea (L.) Mill.), paper birch (Betula papyrifera Marsh), black spruce (Picea mariana (Mill.) BSP), white spruce (Picea glauca Moench) and aspen (Populus tremuloides Michx.). We modeled the vertical distribution against the three variables using a functional general linear model and a novel nonparametric graphical test of significance. Results indicate that LiDAR returns from aspen stands had the most uniform vertical distribution. Balsam fir and white birch distributions were similar and centered at around 50% of the stand height, and black spruce and white spruce distributions were skewed to below 30% of stand height (p<0.001). Increased crown cover concentrated the distributions around 50% of stand height. Increasing age gradually shifted the distributions higher in the stand for stands younger than 70-years, before plateauing and slowly declining at 90-120 years. Results suggest that the vertical distributions of LiDAR returns depend on the three variables studied.

q-bio.QM↗

Testing the first-order separability hypothesis for spatio-temporal point patterns

First-order separability of a spatio-temporal point process plays a fundamental role in the analysis of spatio-temporal point pattern data. While it is often a convenient assumption that simplifies the analysis greatly, existing non-separable structures should be accounted for in the model construction. We propose three different tests to investigate this hypothesis as a step of preliminary data analysis. The first two tests are exact or asymptotically exact for Poisson processes. The first test based on permutations and global envelopes allows us to detect at which spatial and temporal locations or lags the data deviate from the null hypothesis. The second test is a simple and computationally cheap $χ^2$-test. The third test is based on statistical reconstruction method and can be generally applied for non-Poisson processes. The performance of the first two tests is studied in a simulation study for Poisson and non-Poisson models. The third test is applied to the real data of the UK 2001 epidemic foot and mouth disease.

stat.ME↗

Comparison of non-parametric global envelopes

This study presents a simulation study to compare different non-parametric global envelopes that are refinements of the rank envelope proposed by Myllymäki et al. (2017, Global envelope tests for spatial processes, J. R. Statist. Soc. B 79, 381-404, doi: 10.1111/rssb.12172). The global envelopes are constructed for a set of functions or vectors. For a large number of vectors, all the refinements lead to the same outcome as the global rank envelope. For smaller numbers of vectors the refinement playes a role, where different refinements are sensitive to different types of extremeness of a vector among the set of vectors. The performance of the different alternatives are compared in a simulation study with respect to the numbers of available vectors, the dimensionality of the vectors, the amount of dependence between the vector elements and the expected type of extremeness.

stat.ME↗

Point process models for sweat gland activation observed with noise

The aim of the paper is to construct spatial models for the activation of sweat glands for healthy subjects and subjects suffering from peripheral neuropathy by using videos of sweating recorded from the subjects. The sweat patterns are regarded as realizations of spatial point processes and two point process models for the sweat gland activation and two methods for inference are proposed. Several image analysis steps are needed to extract the point patterns from the videos and some incorrectly identified sweat gland locations may be present in the data. To take into account the errors we either include an error term in the point process model or use an estimation procedure that is robust with respect to the errors.

stat.ME↗

Spatial analysis of airborne laser scanning point clouds for predicting forest variables

With recent developments in remote sensing technologies, plot-level forest resources can be predicted utilizing airborne laser scanning (ALS). The prediction is often assisted by mostly vertical summaries of the ALS point clouds. We present a spatial analysis of the point cloud by studying the horizontal distribution of the pulse returns through canopy height models thresholded at different height levels. The resulting patterns of patches of vegetation and gabs on each layer are summarized to spatial ALS features. We propose new features based on the Euler number, which is the number of patches minus the number of gaps, and the empty-space function, which is a spatial summary function of the gab space. The empty-space function is also used to describe differences in the gab structure between two different layers. We illustrate usefulness of the proposed spatial features for predicting different forest variables that summarize the spatial structure of forests or their breast height diameter distribution. We employ the proposed spatial features, in addition to commonly used features from literature, in the well-known k-nn estimation method to predict the forest variables. We present the methodology on the example of a study site in Central Finland.

stat.AP↗

Global envelope tests for spatial processes

Envelope tests are a popular tool in spatial statistics, where they are used in goodness-of-fit testing. These tests graphically compare an empirical function $T(r)$ with its simulated counterparts from the null model. However, the type I error probability $α$ is conventionally controlled for a fixed distance $r$ only, whereas the functions are inspected on an interval of distances $I$. In this study, we propose two approaches related to Barnard's Monte Carlo test for building global envelope tests on $I$:(1) ordering the empirical and simulated functions based on their $r$-wise ranks among each other, and (2) the construction of envelopes for a deviation test. These new tests allow the a priori selection of the global $α$ and they yield $p$-values. We illustrate these tests using simulated and real point pattern data.

stat.ME↗

Multiple Monte Carlo Testing with Applications in Spatial Point Processes

The rank envelope test (Myllymäki et al., Global envelope tests for spatial processes, arXiv:1307.0239 [stat.ME]) is proposed as a solution to multiple testing problem for Monte Carlo tests. Three different situations are recognized: 1) a few univariate Monte Carlo tests, 2) a Monte Carlo test with a function as the test statistic, 3) several Monte Carlo tests with functions as test statistics. The rank test has correct (global) type I error in each case and it is accompanied with a $p$-value and with a graphical interpretation which shows which subtest or which distances of the used test function(s) lead to the rejection at the prescribed significance level of the test. Examples of null hypothesis from point process and random set statistics are used to demonstrate the strength of the rank envelope test. The examples include goodness-of-fit test with several test functions, goodness-of-fit test for one group of point patterns, comparison of several groups of point patterns, test of dependence of components in a multi-type point pattern, and test of Boolean assumption for random closed sets.

stat.ME↗