Searcharxiv⌕ Search

arXiv subjects

Luigi Salmaso

Publications and source records attributed to Luigi Salmaso.

5 recordsLinked to original sources

Design choice and machine learning model performances

An increasing number of publications present the joint application of Design of Experiments (DOE) and machine learning (ML) as a methodology to collect and analyze data on a specific industrial phenomenon. However, the literature shows that the choice of the design for data collection and model for data analysis is often not driven by statistical or algorithmic advantages, thus there is a lack of studies which provide guidelines on what designs and ML models to jointly use for data collection and analysis. This article discusses the choice of design in relation to the ML model performances. A study is conducted that considers 12 experimental designs, 7 families of predictive models, 7 test functions that emulate physical processes, and 8 noise settings, both homoscedastic and heteroscedastic. The results of the research can have an immediate impact on the work of practitioners, providing guidelines for practical applications of DOE and ML.

stat.ML↗

A Combined Approach To Detect Key Variables In Thick Data Analytics

In machine learning one of the strategic tasks is the selection of only significant variables as predictors for the response(s). In this paper an approach is proposed which consists in the application of permutation tests on the candidate predictor variables in the aim of identifying only the most informative ones. Several industrial problems may benefit from such an approach, and an application in the field of chemical analysis is presented. A comparison is carried out between the approach proposed and Lasso, that is one of the most common alternatives for feature selection available in the literature.

stat.ML↗

An Empirical Comparison of Parametric and Permutation Tests for Regression Analysis of Randomized Experiments

Hypothesis tests based on linear models are widely accepted by organizations that regulate clinical trials. These tests are derived using strong assumptions about the data-generating process so that the resulting inference can be based on parametric distributions. Because these methods are well understood and robust, they are sometimes applied to data that depart from assumptions, such as ordinal integer scores. Permutation tests are a nonparametric alternative that require minimal assumptions which are often guaranteed by the randomization that was conducted. We compare analysis of covariance (ANCOVA), a special case of linear regression that incorporates stratification, to several permutation tests based on linear models that control for pretreatment covariates. In simulations of randomized experiments using models which violate some of the parametric regression assumptions, the permutation tests maintain power comparable to ANCOVA. We illustrate the use of these permutation tests alongside ANCOVA using data from a clinical trial comparing the effectiveness of two treatments for gastroesophageal reflux disease. Given the considerable costs and scientific importance of clinical trials, an additional nonparametric method, such as a linear model permutation test, may serve as a robustness check on the statistical inference for the main study endpoints.

stat.AP↗

A permutation approach for ranking of multivariate populations

The subject of this paper is to introduce a novel permutation-based nonparametric approach for the problem of ranking several multivariate populations with respect to both experimental and observation studies to be referred to the most useful design such as MANOVA (multivariate independent samples) and MRCB (multivariate randomized complete block design, i.e. multivariate dependent samples also known as repeated measures). This topic is not only of theoretical interest but also have a practical relevance, especially to business and industrial research where a reliable global ranking in terms of performance of all investigated products/prototypes is a very natural goal. In fact, the need to define an appropriate ranking of items (products, services, teaching courses, degree programs, and so on) is very common in both experimental and observational studies within the areas of business and industrial research.

stat.ME↗