SearcharxivSearch

arXiv subjects

Antonio Forcina

Publications and source records attributed to Antonio Forcina.

14 recordsLinked to original sources

The Marginal Likelihood of two-way tables and Ecological Inference

The paper derives new results on the marginal likelihood of a two-way table which clarify the conditions under which Ecological inference is possible and lead to an efficient algorithm for maximizing the exact multinomial likelihood. The first part generalizes the work of Placket(1977} on the marginal likelihood of a 2 x 2 table to a general R x C table. In doing so, new conceptual tools are introduced and new insights on the geometry of the collection of tables having fixed row and column margins and the extended hypergeometric distribution are derived. In the second part, when observations on the row and the column marginal distributions are available for a collection of two-way tables sharing the same association structure, an efficient Fisher scoring algorithm for maximizing the exact likelihood under multinomial sampling is introduced and a small simulation study is used to compare the performance of the proposed method with two well established ones.

stat.ME

A small-area ecological approach for estimating vote changes and their determinants

Empirical analyses on the factors driving vote switching are rare, usually conducted at the national level without considering the parties of origin and destination, and often unreliable due to the severe inaccuracy of recall survey data. To overcome the problem of lack of adequate data and to incorporate the increasingly relevant role of local factors, we propose an ecological inference methodology to estimate the number of vote transitions within small homogeneous areas and to assess the relationships between these counts and local characteristics through multinomial logistic models. This approach allows for a disaggregate analysis of contextual factors behind vote switching, distinguishing between their different origins and destinations. We apply this methodology to the Italian region of Umbria, divided into 19 small areas. To explain the number of transitions toward the right-wing nationalist party that won the elections and towards increasing abstentionism, we focused on measures of geographical, economic, and cultural disadvantages of local communities. Among the main findings, the economic disadvantages mainly pushed previous abstainers and far-right Lega voters to change their choices in favor of the rising right-wing party, while transitions from the opposite political camp were mostly influenced by cultural factors such as a lack of social capital, negative attitude towards the EU, and political tradition.

stat.ME

Marginal log-linear models and mediation analysis

We review some not well known results about marginal log-linear models, derive some new ones and show how they might be relevant in mediation analysis within logistic regression. In particular, we elaborate on the relation between interaction parameters defined within different marginal distributions and describe an algorithm for estimating the sane interaction parameters within different marginals.

stat.ME

Estimating the size of a closed population by modeling latent and observed heterogeneity

The paper describes a new class of capture-recapture models for closed populations when individual covariates are available. The novelty consists in combining a latent class model for the distribution of the capture history, where the class weights and the conditional distributions given the latent may depend on covariates, with a model for the marginal distribution of the available covariates as in \cite{Liu2017}. In addition, any general form of serial dependence is allowed when modeling capture histories conditionally on the latent and covariates. A Fisher-scoring algorithm for maximum likelihood estimation is proposed, and the Implicit Function Theorem is used to show that the mapping between the marginal distribution of the observed covariates and the probabilities of being never captured is one-to-one. Asymptotic results are outlined, and a procedure for constructing likelihood based confidence intervals for the population size is presented. Two examples based on real data are used to illustrate the proposed approach

stat.ME

An extended class of RC association models: estimation and main properties

The extended class of multiplicative row-column (RC) association models, introduced in this paper for two-way contingency tables, allows users to select both the type of logit (local, global, continuation, reverse continuation) suitable for the row and column classification variables and the scale on which interactions are measured. As in \cite{Kateri95} for the case of local logits, our extended class of bivariate interactions is linked to divergence measures and, by means of a representation theorem, we provide reconstruction formulas for the joint probabilities depending on pairs of logit types. These results are the key to show that, given marginal logits, our extended interactions determine uniquely the bivariate distribution. We also determine the kind of positive association which is implied by our extended interactions being non negative. Quick model selection within this wide class can be performed by an efficient algorithm for computing maximum likelihood estimates which exploits the properties of a reduced rank constraint imposed on the matrix of extended interactions and allows for additional linear constraint on marginal logits. An application to social mobility data is presented and discussed.

stat.CO

Ecological fallacy and covariates in the estimation of voters transitions

We provide a simple formulation of the conditions under which ecological bias should be expected and argue that the bias will affect any method of ecological inference; our claim is supported by formal derivations and several examples where individual data are available. The conditions which we highlight imply that, when they are violated, ecological bias cannot be avoided unless the a suitable model for the effect of specific covariates is incorporated into ecological inference. We also detect situations where the ecological bias cannot be corrected even if the effect of covariates is incorporated into the model. In particular, when the association in the individual data is rather weak and certain transition probabilities are similar functions of a given covariate, ecological inference methods may be unable to disentangle the individual components from their aggregate. In any case, the value of the covariates measured at the level of local units, when these are rather extensive, do not provide enough information on the within unit heterogeneity. Our findings are applied and tested on several data sets, both real and simulated, where individual observations are available in addition to the aggregated ones.,.

stat.ME

Modelling dark current and hot pixels in imaging sensors

A Gaussian mixture model with a complex covariance structure was used to analyse experimental data from images recorded by a digital sensor under darkness, to model the effects of temperature and duration of exposure on artificial signals (dark current), on ordinary and possibly defective (hot) pixels. The model accounts for two components of variance within each latent type: random noise in each image and lack of uniformity within the sensor; both components are allowed to depend on experimental conditions. The results seem to indicate that the way dark current grows with the duration of exposure and temperature cannot be represented by a simple parametric model. The latent class model detects the presence of at least two types of hot pixels, where the less frequent ones have also a more extreme behaviour. Though the lack of uniformity of the sensor is amplified by duration of exposure and temperature, pixels characteristics seem to deviate in the same direction and with the same relative size.

stat.AP

Multiplicative models for frequency data, estimation and testing

This paper is about models for a vector of probabilities whose elements must have a multiplicative structure and sum to 1 at the same time; in certain applications, as basket analysis, these models may be seen as a constrained version of quasi-independence. After reviewing the basic properties of these models, their geometric features as a curved exponential family are investigated. A new algorithm for computing maximum likelihood estimates is presented and new insights are provided on the underlying geometry. The asymptotic distribution of three statistics for hypothesis testing are derived and a small simulation study is presented to investigate the accuracy of asymptotic approximations.

math.ST

An efficient Fisher-scoring algorithm for fitting latent class models with individual covariates

For latent class models where the class weights depend on individual covariates, we derive a simple expression for computing the score vector and a convenient hybrid between the observed and the expected information matrices which is always positive defnite. These ingredients, combined with a maximization algorithm based on line search, provides an efficient tool for maximum likelihood estimation. In particular, the proposed algorithm is such that the log-likelihood never decreases from one step to the next and the choice of starting values is not crucial for reaching a local maximum. We show how the same algorithm may be used for numerical investigation of the effect of model mispecifications. An application to education transmission is used as an illustration.

stat.CO

Fitting directed acyclic graphs with latent nodes as finite mixtures models, with application to education transmission

This paper describes an efficient EM algorithm for maximum likelihood estimation of a system of nonlinear structural equations corresponding to a directed acyclic graph model that can contain an arbitrary number of latent variables. The endogenous variables in the model must be categorical, while the exogenous variables may be arbitrary. The models discussed in this paper are an extended version of finite mixture models suitable for causal inference. An application to the problem of education transmission is presented as an illustration.

stat.CO

Ecological fallacy and covariates: new insights based on multilevel modelling of individual data

This paper deals with the issue of ecological bias in ecological inference. We provide an explicit formulation of the conditions required for the ordinary ecological regression to produce unbiased estimates and argue that, when these conditions are violated, any method of ecological inference is going to produce biased estimates. These findings are clarified and supported by empirical evidence provided by comparing the results of three main ecological inference methods with those of multilevel logistic regression applied to a unique set of individual data on voting behaviour. The main findings of our study have two important implications that apply to all situations where the conditions for no ecological bias are violated: (i) only ecological inference methods that allow to model the effect of covariates have a chance to produce unbiased estimates; (ii) the set of covariates to be included in the model to remove bias is limited to the marginal proportions. Finally, our results suggest that, when the association between two ecological variables is very weak, it is not possible to obtain unbiased estimates even by an appropriate model that accounts for the effect of relevant covariates.

stat.AP

Testing order restrictions in contingency tables

Several interesting models for contingency tables are defined by a system of equality and inequality constraints on a suitable set of marginal log-linear parameters. After reviewing the most common difficulties which are intrinsic to order restricted testing problems, we propose two new families of testing procedures, based on similar attempts appeared in the econometric literature, in order to increase the probability of detecting several relevant violations of the supposed order relations. One set of procedures is based on the decomposition of the log-likelihood ratio when testing the given set of inequalities and the nested model derived by forcing inequalities into strict equalities. The other set uses the asymptotic joint normal distribution of the estimates of the marginal log-linear parameters to be constrained.

math.ST

Two algorithms for fitting constrained marginal models

We study in detail the two main algorithms which have been considered for fitting constrained marginal models to discrete data, one based on Lagrange multipliers and the other on a regression model. We show that the updates produced by the two methods are identical, but that the Lagrangian method is more efficient in the case of identically distributed observations. We provide a generalization of the regression algorithm for modelling the effect of exogenous individual-level covariates, a context in which the use of the Lagrangian algorithm would be infeasible for even moderate sample sizes. An extension of the method to likelihood-based estimation under $L_1$-penalties is also considered.

stat.CO

Modelling sources of ecological fallacy within a revised Brown and Payne model of voting transitions

We present a model of voting behaviour based on a version of aggregated overdispersed multinomial distributions; relative to a similar model by \citet{BP86}, our model is based on more realistic assumptions and free from certain shortcomings of the previous model. We show that, within this model, it is possible to test for certain confounding effects due to observable covariates measured at the aggregate level; such effects, if ignored, might cause substantial bias in the estimated relation between voting decisions in two close in time elections, a phenomenon known as {\em Ecological Fallacy}. An application to a referendum following an election for the major in the town of Milan, which was interpreted as a defeat for the Berlusconi gouvernment, is used as an illustration.

stat.ME