SearcharxivSearch

arXiv subjects

Sylvie Huet

Publications and source records attributed to Sylvie Huet.

At least 19 recordsLinked to original sources

RKHSMetaMod: An R package to estimate the Hoeffding decomposition of a complex model by solving RKHS ridge group sparse optimization problem

In this paper, we propose an R package, called RKHSMetaMod, that implements a procedure for estimating a meta-model of a complex model. The meta-model approximates the Hoeffding decomposition of the complex model and allows us to perform sensitivity analysis on it. It belongs to a reproducing kernel Hilbert space that is constructed as a direct sum of Hilbert spaces. The estimator of the meta-model is the solution of a penalized empirical least-squares minimization with the sum of the Hilbert norm and the empirical L^2-norm. This procedure, called RKHS ridge group sparse, allows both to select and estimate the terms in the Hoeffding decomposition, and therefore, to select and estimate the Sobol indices that are non-zero. The RKHSMetaMod package provides an interface from R statistical computing environment to the C++ libraries Eigen and GSL. In order to speed up the execution time and optimize the storage memory, except for a function that is written in R, all of the functions of this package are written using the efficient C++ libraries through RcppEigen and RcppGSL packages. These functions are then interfaced in the R environment in order to propose a user-friendly package.

stat.ML

Risk upper bounds for RKHS ridge group sparse estimator in the regression model with non-Gaussian and non-bounded error

We consider the problem of estimating a meta-model of an unknown regression model with non-Gaussian and non-bounded error. The meta-model belongs to a reproducing kernel Hilbert space constructed as a direct sum of Hilbert spaces leading to an additive decomposition including the variables and interactions between them. The estimator of this meta-model is calculated by minimizing an empirical least-squares criterion penalized by the sum of the Hilbert norm and the empirical $L^2$-norm. In this context, the upper bounds of the empirical $L^2$ risk and the $L^2$ risk of the estimator are established.

math.ST

Metamodel construction for sensitivity analysis

We propose to estimate a metamodel and the sensitivity indices of a complex model m in the Gaussian regression framework. Our approach combines methods for sensitivity analysis of complex models and statistical tools for sparse non-parametric estimation in multivariate Gaussian regression model. It rests on the construction of a metamodel for aproximating the Hoeffding-Sobol decomposition of m. This metamodel belongs to a reproducing kernel Hilbert space constructed as a direct sum of Hilbert spaces leading to a functional ANOVA decomposition. The estimation of the metamodel is carried out via a penalized least-squares minimization allowing to select the subsets of variables that contribute to predict the output. It allows to estimate the sensitivity indices of m. We establish an oracle-type inequality for the risk of the estimator, describe the procedure for estimating the metamodel and the sensitivity indices, and assess the performances of the procedure via a simulation study.

math.ST

Few self-involved agents among BC agents can lead to polarized local or global consensus

Social issues are generally discussed by highly-involved and less-involved people to build social norms defining what has to be thought and done about them. As self-involved agents share different attitude dynamics to other agents Wood, Pool et al, 1996, we study the emergence and evolution of norms through an individual-based model involving these two types of agents. The dynamics of self-involved agents is drawn from Huet and Deffuant, 2010, and the dynamics of others, from Deffuant et al, 2001. The attitude of an agent is represented as a segment on a continuous attitudinal space. Two agents are close if their attitude segments share sufficient overlap. Our agents discuss two different issues, one of which, called main issue, is more important for the self-involved agents than the other, called secondary issue. Self-involved agents are attracted on both issues if they are close on main issue, but shift away from their peer's opinion if they are only close on secondary issue. Differently, non-self-involved agents are attracted by other agents when they are close on both the main and secondary issues. We observe the emergence of various types of extreme minor clusters. In one or different groups of attitudes, they can lead to an already-built moderate norm or a norm polarized on secondary and/or main issues. They can also push disagreeing agents gathered in different groups to a global moderate consensus.

cs.MA

Resisting hostility generated by terror: An agent-based study

We aim to study through an agent-based model the cultural conditions leading to a decrease or an increase of discrimination between groups after a major cultural threat such as a terrorist attack. We propose an agent-based model of cultural dynamics inspired from the social psychological theories. An agent has a cultural identity comprised of the most acceptable positions about each of the different cultural worldviews corresponding to the main cultural groups of the considered society and a margin of acceptance around each of these most acceptable positions. An agent forms an attitude about another agent depending on the similarity between their cultural identities. When a terrorist attack is perpetrated in the name of an extreme cultural identity, the negatively perceived agents from this extreme cultural identity modify their margins of acceptance in order to differentiate themselves more from the threatening cultural identity. We generated a set of populations with cultural identities compatible with data given by a survey on groups' attitudes among a large sample representative of the population of France; we then simulated the reaction of these agents facing a threat. For most populations, the average attitude toward agents with the same preferred worldview as the terrorists becomes more negative; however, when the population shows some cultural properties, we noticed the opposite effect as the average attitude of the population becomes less negative. This particular context requires that the agents sharing the same preferred worldview with the terrorists strongly differentiate themselves from the terrorists' extreme cultural identity and that the other agents be aware of these changes.

cs.MA

The anatomy of a Web of Trust: the Bitcoin-OTC market

Bitcoin-otc is a peer to peer (over-the-counter) marketplace for trading with bit- coin crypto-currency. To mitigate the risks of the p2p unsupervised exchanges, the establishment of a reliable reputation systems is needed: for this reason, a web of trust is implemented on the website. The availability of all the historic of the users interaction data makes this dataset a unique playground for studying reputation dynamics through others evaluations. We analyze the structure and the dynamics of this web of trust with a multilayer network approach distin- guishing the rewarding and the punitive behaviors. We show that the rewarding and the punitive behavior have similar emergent topological properties (apart from the clustering coefficient being higher for the rewarding layer) and that the resultant reputation originates from the complex interaction of the more regular behaviors on the layers. We show which are the behaviors that correlate (i.e. the rewarding activity) or not (i.e. the punitive activity) with reputation. We show that the network activity presents bursty behaviors on both the layers and that the inequality reaches a steady value (higher for the rewarding layer) with the network evolution. Finally, we characterize the reputation trajectories and we identify prototypical behaviors associated to three classes of users: trustworthy, untrusted and controversial.

cs.CY

Testing k-monotonicity of a discrete distribution. Application to the estimation of the number of classes in a population

We develop here several goodness-of-fit tests for testing the k-monotonicity of a discrete density, based on the empirical distribution of the observations. Our tests are non-parametric, easy to implement and are proved to be asymptotically of the desired level and consistent. We propose an estimator of the degree of k-monotonicity of the distribution based on the non-parametric goodness-of-fit tests. We apply our work to the estimation of the total number of classes in a population. A large simulation study allows to assess the performances of our procedures.

stat.ME

A Universal Model of Commuting Networks

We test a recently proposed model of commuting networks on 80 case studies from different regions of the world (Europe and United-States) and with geographic units of different sizes (municipality, county, region). The model takes as input the number of commuters coming in and out of each geographic unit and generates the matrix of commuting flows betwen the geographic units. We show that the single parameter of the model, which rules the compromise between the influence of the distance and job opportunities, follows a universal law that depends only on the average surface of the geographic units. We verified that the law derived from a part of the case studies yields accurate results on other case studies. We also show that our model significantly outperforms the two other approaches proposing a universal commuting model (Balcan et al. (2009); Simini et al. (2012)), particularly when the geographic units are small (e.g. municipalities).

math.ST

Generating French virtual commuting network at municipality level

We aim to generate virtual commuting networks in the French rural regions in order to study the dynamics of their municipalities. Since we have to model small commuting flows between municipalities with a few hundreds or thousands inhabitants, we opt for a stochastic model presented by Gargiulo et al. 2012. It reproduces the various possible complete networks using an iterative process, stochastically choosing a workplace in the region for each commuter living in the municipality of a region. The choice is made considering the job offers in each municipality of the region and the distance to all the possible destinations. This paper presents how to adapt and implement this model to generate French regions commuting networks between municipalities. We address three different questions: How to generate a reliable virtual commuting network for a region highly dependant of other regions for the satisfaction of its resident's demand for employment? What about a convenient deterrence function? How to calibrate the model when detailed data is not available? We answer proposing an extended job search geographical base for commuters living in the municipalities, we compare two different deterrence functions and we show that the parameter is a constant for network linking French municipalities.

math.ST

Deriving the number of jobs in proximity services from the number of inhabitants in French rural municipalities

We use a minimum requirement approach to derive the number of jobs in proximity services per inhabitant in French rural municipalities. We first classify the municipalities according to their time distance to the municipality where the inhabitants go the most frequently to get services (called MFM). For each set corresponding to a range of time distance to MFM, we perform a quantile regression estimating the minimum number of service jobs per inhabitant, that we interpret as an estimation of the number of proximity jobs per inhabitant. We observe that the minimum number of service jobs per inhabitant is smaller in small municipalities. Moreover, for municipalities of similar sizes, when the distance to the MFM increases, we find that the number of jobs of proximity services per inhabitant increases.

stat.AP

Rejection Mechanism in 2D Bounded Confidence Provides more Conformity

We add a rejection mechanism (negative influence) into a two-dimensions bounded confidence model. The principle is that one shifts aways from a close attitude of one's interlocutor, when there is a strong disagreement on the other attitude. The model shows metastable clusters, which maintain themselves through opposite influences of competitor clusters. Our analysis and first experiments support the hypothesis that the number of clusters grows linearly with the inverse of the uncertainty, whereas this growth is quadratic in the bounded confidence model.

physics.soc-ph

Openness leads to opinion stability and narrowness to volatility

We propose a new opinion dynamic model based on the experiments and results of Wood et al (1996). We consider pairs of individuals discussing on two attitudinal dimensions, and we suppose that one dimension is important, the other secondary. The dynamics are mainly ruled by the level of agreement on the main dimension. If two individuals are close on the main dimension, then they attract each other on the main and on the secondary dimensions, whatever their disagreement on the secondary dimension. If they are far from each other on the main dimension, then too much proximity on the secondary dimension is uncomfortable, and generates rejection on this dimension. The proximity is defined by comparing the opinion distance with a threshold called attraction threshold on the main dimension and rejection threshold on the secondary dimension. With such dynamics, a population with opinions initially uniformly drawn evolves to a set of clusters, inside which secondary opinions fluctuate more or less depending on threshold values. We observe that a low attraction threshold favours fluctuations on the secondary dimension, especially when the rejection threshold is high. The opinion evolutions of the model can be related to some stylised facts.

physics.soc-ph

Nonparametric species richness estimation under convexity constraint

We consider the estimation of the total number $N$ of species based on the abundances of species that have been observed. We adopt a non parametric approach where the true abundance distribution $p$ is only supposed to be convex. From this assumption, we propose a definition for convex abundance distributions. We use a least-squares estimate of the truncated version of $p$ under the convexity constraint. We deduce two estimators of the total number of species, the asymptotic distribution of which are derived. We propose three different procedures, including a bootstrap one, to obtain a confidence interval for $N$. The performances of the estimators are assessed in a simulation study and compared with competitors. The proposed method is illustrated on several examples.

stat.ME

The Leviathan model: Absolute dominance, generalised distrust, small worlds and other patterns emerging from combining vanity with opinion propagation

We propose an opinion dynamics model that combines processes of vanity and opinion propagation. The interactions take place between randomly chosen pairs. During an interaction, the agents propagate their opinions about themselves and about other people they know. Moreover, each individual is subject to vanity: if her interlocutor seems to value her highly, then she increases her opinion about this interlocutor. On the contrary she tends to decrease her opinion about those who seem to undervalue her. The combination of these dynamics with the hypothesis that the opinion propagation is more efficient when coming from highly valued individuals, leads to different patterns when varying the parameters. For instance, for some parameters the positive opinion links between individuals generate a small world network. In one of the patterns, absolute dominance of one agent alternates with a state of generalised distrust, where all agents have a very low opinion of all the others (including themselves). We provide some explanations of the mechanisms behind these emergent behaviors and finally propose a discussion about their interest

physics.soc-ph

Estimation of a convex discrete distribution

Non-parametric estimation of a convex discrete distribution may be of interest in several applications, such as the estimation of species abundance distribution in ecology. In this paper we study the least squares estimator of a discrete distribution under the constraint of convexity. We show that this estimator exists and is unique, and that it always outperforms the classical empirical estimator in terms of the $\ell_{2}$-distance. We provide an algorithm for its computation, based on the support reduction algorithm. We compare its performance to those of the empirical estimator, on the basis of a simulation study.

stat.ME

High-dimensional regression with unknown variance

We review recent results for high-dimensional sparse linear regression in the practical case of unknown variance. Different sparsity settings are covered, including coordinate-sparsity, group-sparsity and variation-sparsity. The emphasis is put on non-asymptotic analyses and feasible procedures. In addition, a small numerical study compares the practical performance of three schemes for tuning the Lasso estimator and some references are collected for some more general models, including multivariate regression and nonparametric regression.

math.ST

Graph selection with GGMselect

Applications on inference of biological networks have raised a strong interest in the problem of graph estimation in high-dimensional Gaussian graphical models. To handle this problem, we propose a two-stage procedure which first builds a family of candidate graphs from the data, and then selects one graph among this family according to a dedicated criterion. This estimation procedure is shown to be consistent in a high-dimensional setting, and its risk is controlled by a non-asymptotic oracle-like inequality. The procedure is tested on a real data set concerning gene expression data, and its performances are assessed on the basis of a large numerical study. The procedure is implemented in the R-package GGMselect available on the CRAN.

math.ST

Estimator selection in the Gaussian setting

We consider the problem of estimating the mean $f$ of a Gaussian vector $Y$ with independent components of common unknown variance $σ^{2}$. Our estimation procedure is based on estimator selection. More precisely, we start with an arbitrary and possibly infinite collection $\FF$ of estimators of $f$ based on $Y$ and, with the same data $Y$, aim at selecting an estimator among $\FF$ with the smallest Euclidean risk. No assumptions on the estimators are made and their dependencies with respect to $Y$ may be unknown. We establish a non-asymptotic risk bound for the selected estimator. As particular cases, our approach allows to handle the problems of aggregation and model selection as well as those of choosing a window and a kernel for estimating a regression function, or tuning the parameter involved in a penalized criterion. We also derive oracle-type inequalities when $\FF$ consists of linear estimators. For illustration, we carry out two simulation studies. One aims at comparing our procedure to cross-validation for choosing a tuning parameter. The other shows how to implement our approach to solve the problem of variable selection in practice.

math.ST