SearcharxivSearch

arXiv subjects

Amir Sepehri

Publications and source records attributed to Amir Sepehri.

9 recordsLinked to original sources

Wasserstein mixing of a systematic-scan random rotation sampler

We study the mixing time of a systematic-scan analogue of Kac's walk that was proposed as a fast surrogate for Haar-distributed orthogonal matrices in randomized high-dimensional algorithms and was conjectured to approach Haar measure after only logarithmically many sweeps. We show that this conjectured speed-up does not occur for convergence of the full matrix law to Haar measure in Frobenius Wasserstein distance. At fixed normalized accuracy, the mixing time lies between order $n/\log n$ and order $n$ sweeps; at fixed absolute Frobenius accuracy, the corresponding bounds are between order $n$ and order $n\log n$. More strongly, below the scale $n/\log n$, the normalized Wasserstein distance remains asymptotically at its extremal value. We also show that the output law is singular with respect to Haar measure for fewer than $n/2$ sweeps. Thus the sampler may provide effective application-specific randomization without exhibiting the much faster full-Haar mixing.

stat.CO

Interpretable Assessment of Fairness During Model Evaluation

For companies developing products or algorithms, it is important to understand the potential effects not only globally, but also on sub-populations of users. In particular, it is important to detect if there are certain groups of users that are impacted differently compared to others with regard to business metrics or for whom a model treats unequally along fairness concerns. In this paper, we introduce a novel hierarchical clustering algorithm to detect heterogeneity among users in given sets of sub-populations with respect to any specified notion of group similarity. We prove statistical guarantees about the output and provide interpretable results. We demonstrate the performance of the algorithm on real data from LinkedIn.

cs.LG

Fairness through Experimentation: Inequality in A/B testing as an approach to responsible design

As technology continues to advance, there is increasing concern about individuals being left behind. Many businesses are striving to adopt responsible design practices and avoid any unintended consequences of their products and services, ranging from privacy vulnerabilities to algorithmic bias. We propose a novel approach to fairness and inclusiveness based on experimentation. We use experimentation because we want to assess not only the intrinsic properties of products and algorithms but also their impact on people. We do this by introducing an inequality approach to A/B testing, leveraging the Atkinson index from the economics literature. We show how to perform causal inference over this inequality measure. We also introduce the concept of site-wide inequality impact, which captures the inclusiveness impact of targeting specific subpopulations for experiments, and show how to conduct statistical inference on this impact. We provide real examples from LinkedIn, as well as an open-source, highly scalable implementation of the computation of the Atkinson index and its variance in Spark/Scala. We also provide over a year's worth of learnings -- gathered by deploying our method at scale and analyzing thousands of experiments -- on which areas and which kinds of product innovations seem to inherently foster fairness through inclusiveness.

cs.SI

Non-reversible, tuning- and rejection-free Markov chain Monte Carlo via iterated random functions

In this work we present a non-reversible, tuning- and rejection-free Markov chain Monte Carlo which naturally fits in the framework of hit-and-run. The sampler only requires access to the gradient of the log-density function, hence the normalizing constant is not needed. We prove the proposed Markov chain is invariant for the target distribution and illustrate its applicability through a wide range of examples. We show that the sampler introduced in the present paper is intimately related to the continuous sampler of Peters and de With (2012), Bouchard-Cote et al. (2017). In particular, the computation is quite similar in the sense that both are centered around simulating an inhomogenuous Poisson process. The computation can be simplified when the gradient of the log-density admits a computationally efficient directional decomposition into a sum of two monotone functions. We apply our sampler in selective inference, gaining significant improvement over the formerly used sampler (Tian et al. 2016).

stat.CO

New Tests of Uniformity on the Compact Classical Groups as Diagnostics for Weak-star Mixing of Markov Chains

This paper introduces two new families of non-parametric tests of goodness-of-fit on the compact classical groups. One of them is a family of tests for the eigenvalue distribution induced by the uniform distribution, which is consistent against all fixed alternatives. The other is a family of tests for the uniform distribution on the entire group, which is again consistent against all fixed alternatives. We find the asymptotic distribution under the null and general alternatives. The tests are proved to be asymptotically admissible. Local power is derived and the global properties of the power function against local alternatives are explored. The new tests are validated on two random walks for which the mixing-time is studied in the literature. The new tests, and several others, are applied to the Markov chain sampler proposed by \cite{jones2011randomized}, providing strong evidence supporting the claim that the sampler mixes quickly.

math.ST

Bouncy Hybrid Sampler as a Unifying Device

This work introduces a class of rejection-free Markov chain Monte Carlo (MCMC) samplers, named the Bouncy Hybrid Sampler, which unifies several existing methods from the literature. Examples include the Bouncy Particle Sampler of Peters and de With (2012), Bouchard-Cote et al. (2015) and the Hamiltonian MCMC. Following the introduced general framework, we derive a new sampler called the Quadratic Bouncy Hybrid Sampler. We apply this novel sampler to the problem of sampling from a truncated Gaussian distribution.

stat.CO

Cauchy Identities for the Characters of the Compact Classical Groups

Motivated by statistical applications, this paper introduces Cauchy identities for characters of the compact classical groups. These identities generalize the well-known Cauchy identity for characters of the unitary group, which are Schur functions of symmetric function theory. Application to statistical hypothesis testing is briefly sketched.

math.RT

The Bayesian SLOPE

The SLOPE estimates regression coefficients by minimizing a regularized residual sum of squares using a sorted-$\ell_1$-norm penalty. The SLOPE combines testing and estimation in regression problems. It exhibits suitable variable selection and prediction properties, as well as minimax optimality. This paper introduces the Bayesian SLOPE procedure for linear regression. The classical SLOPE estimate is the posterior mode in the normal regression problem with an appropriate prior on the coefficients. The Bayesian SLOPE considers the full Bayesian model and has the advantage of offering credible sets and standard error estimates for the parameters. Moreover, the hierarchical Bayesian framework allows for full Bayesian and empirical Bayes treatment of the penalty coefficients; whereas it is not clear how to choose these coefficients when using the SLOPE on a general design matrix. A direct characterization of the posterior is provided which suggests a Gibbs sampler that does not involve latent variables. An efficient hybrid Gibbs sampler for the Bayesian SLOPE is introduced. Point estimation using the posterior mean is highlighted, which automatically facilitates the Bayesian prediction of future observations. These are demonstrated on real and synthetic data.

stat.ME

The Accessible Lasso Models

A new line of research on the lasso exploits the beautiful geometric fact that the lasso fit is the residual from projecting the response vector $y$ onto a certain convex polytope. This geometric picture also allows an exact geometric description of the set of accessible lasso models for a given design matrix, that is, which configurations of the signs of the coefficients it is possible to realize with some choice of $y$. In particular, the accessible lasso models are those that correspond to a face of the convex hull of all the feature vectors together with their negations. This convex hull representation then permits the enumeration and bounding of the number of accessible lasso models, which in turn provides a direct proof of model selection inconsistency when the size of the true model is greater than half the number of observations.

math.ST