SearcharxivSearch

arXiv subjects

Angela Montanari

Publications and source records attributed to Angela Montanari.

8 recordsLinked to original sources

A Poisson Factor Mixture Model for the Analysis of Linguistic Competence in Italian University Students' Writing

Public debate on the alleged decline of language skills among younger generations often focuses on university students, the most highly educated segment of the population. Rather than addressing the ill posed question of linguistic decline, this paper examines how formal written Italian is currently used by university students and whether systematic patterns of competence and heterogeneity can be identified. The analysis is based on data from the UniversITA project, which collected formal texts written by a large and nationally representative sample of Italian university students. Texts were annotated for linguistically motivated features covering orthography, lexicon, syntax, morphosyntax, coherence, register, and sentence structure, yielding low frequency multivariate count data. To analyse these data, we propose a novel model-based clustering approach based on a Poisson factor mixture model that accounts for dependence among linguistic features and unobserved population heterogeneity. The results identify two correlated dimensions of writing competence, interpretable as communicative competence and linguistic grammatical competence. When educational and socio demographic information is incorporated, distinct student profiles emerge that are associated with field of study and educational background. These findings provide quantitative evidence on contemporary writing and offer insights relevant for language education and higher education policy.

stat.ME

Large factor model estimation by nuclear norm plus $l_1$ norm penalization

This paper provides a comprehensive estimation framework via nuclear norm plus $l_1$ norm penalization for high-dimensional approximate factor models with a sparse residual covariance. The underlying assumptions allow for non-pervasive latent eigenvalues and a prominent residual covariance pattern. In that context, existing approaches based on principal components may lead to misestimate the latent rank, due to the numerical instability of sample eigenvalues. On the contrary, the proposed optimization problem retrieves the latent covariance structure and exactly recovers the latent rank and the residual sparsity pattern. Conditioning on them, the asymptotic rates of the subsequent ordinary least squares estimates of loadings and factor scores are provided, the recovered latent eigenvalues are shown to be maximally concentrated and the estimates of factor scores via Bartlett's and Thompson's methods are proved to be the most precise given the data. The validity of outlined results is highlighted in an exhaustive simulation study and in a real financial data example.

math.ST

High-dimensional clustering via Random Projections

In this work, we address the unsupervised classification issue by exploiting the general idea of Random Projection Ensemble. Specifically, we propose to generate a set of low dimensional independent random projections and to perform model-based clustering on each of them. The top $B^*$ projections, i.e. the projections which show the best grouping structure are then retained. The final partition is obtained by aggregating the clusters found in the projections via consensus. The performances of the method are assessed on both real and simulated datasets. The obtained results suggest that the proposal represents a promising tool for high-dimensional clustering.

stat.ME

Matrix sketching for supervised classification with imbalanced classes

Matrix sketching is a recently developed data compression technique. An input matrix A is efficiently approximated with a smaller matrix B, so that B preserves most of the properties of A up to some guaranteed approximation ratio. In so doing numerical operations on big data sets become faster. Sketching algorithms generally use random projections to compress the original dataset and this stochastic generation process makes them amenable to statistical analysis. The statistical properties of sketching algorithms have been widely studied in the context of multiple linear regression. In this paper we propose matrix sketching as a tool for rebalancing class sizes in supervised classification with imbalanced classes. It is well-known in fact that class imbalance may lead to poor classification performances especially as far as the minority class is concerned.

stat.ML

One-class classification with application to forensic analysis

The analysis of broken glass is forensically important to reconstruct the events of a criminal act. In particular, the comparison between the glass fragments found on a suspect (recovered cases) and those collected on the crime scene (control cases) may help the police to correctly identify the offender(s). The forensic issue can be framed as a one-class classification problem. One-class classification is a recently emerging and special classification task, where only one class is fully known (the so-called target class), while information on the others is completely missing. We propose to consider classic Gini's transvariation probability as a measure of typicality, i.e. a measure of resemblance between an observation and a set of well-known objects (the control cases). The aim of the proposed Transvariation-based One-Class Classifier (TOCC) is to identify the best boundary around the target class, that is, to recognise as many target objects as possible while rejecting all those deviating from this class.

stat.AP

A bootstrap test to detect prominent Granger-causalities across frequencies

Granger-causality in the frequency domain is an emerging tool to analyze the causal relationship between two time series. We propose a bootstrap test on unconditional and conditional Granger-causality spectra, as well as on their difference, to catch particularly prominent causality cycles in relative terms. In particular, we consider a stochastic process derived applying independently the stationary bootstrap to the original series. Our null hypothesis is that each causality or causality difference is equal to the median across frequencies computed on that process. In this way, we are able to disambiguate causalities which depart significantly from the median one obtained ignoring the causality structure. Our test shows power one as the process tends to non-stationarity, thus being more conservative than parametric alternatives. As an example, we infer about the relationship between money stock and GDP in the Euro Area via our approach, considering inflation, unemployment and interest rates as conditioning variables. We point out that during the period 1999-2017 the money stock aggregate M1 had a significant impact on economic output at all frequencies, while the opposite relationship is significant only at high frequencies.

q-fin.ST

A large covariance matrix estimator under intermediate spikiness regimes

The present paper concerns large covariance matrix estimation via composite minimization under the assumption of low rank plus sparse structure. In this approach, the low rank plus sparse decomposition of the covariance matrix is recovered by least squares minimization under nuclear norm plus $l_1$ norm penalization. This paper proposes a new estimator of that family based on an additional least-squares re-optimization step aimed at un-shrinking the eigenvalues of the low rank component estimated at the first step. We prove that such un-shrinkage causes the final estimate to approach the target as closely as possible in Frobenius norm while recovering exactly the underlying low rank and sparsity pattern. Consistency is guaranteed when $n$ is at least $O(p^{\frac{3}{2}δ})$, provided that the maximum number of non-zeros per row in the sparse component is $O(p^δ)$ with $δ\leq \frac{1}{2}$. Consistent recovery is ensured if the latent eigenvalues scale to $p^α$, $α\in[0,1]$, while rank consistency is ensured if $δ\leq α$. The resulting estimator is called UNALCE (UNshrunk ALgebraic Covariance Estimator) and is shown to outperform state of the art estimators, especially for what concerns fitting properties and sparsity pattern detection. The effectiveness of UNALCE is highlighted on a real example regarding ECB banking supervisory data.

stat.ME

The Importance of Being Clustered: Uncluttering the Trends of Statistics from 1970 to 2015

In this paper we retrace the recent history of statistics by analyzing all the papers published in five prestigious statistical journals since 1970, namely: Annals of Statistics, Biometrika, Journal of the American Statistical Association, Journal of the Royal Statistical Society, series B and Statistical Science. The aim is to construct a kind of "taxonomy" of the statistical papers by organizing and by clustering them in main themes. In this sense being identified in a cluster means being important enough to be uncluttered in the vast and interconnected world of the statistical research. Since the main statistical research topics naturally born, evolve or die during time, we will also develop a dynamic clustering strategy, where a group in a time period is allowed to migrate or to merge into different groups in the following one. Results show that statistics is a very dynamic and evolving science, stimulated by the rise of new research questions and types of data.

stat.AP