SearcharxivSearch

arXiv subjects

Japneet Singh

Publications and source records attributed to Japneet Singh.

9 recordsLinked to original sources

Doeblin Curves

Recent research on Doeblin coefficients has shed light on their usefulness as a multi-way generalization of the Dobrushin contraction coefficient for TV distance, in a separate vein from their classic role in the theory of Markov chain ergodicity. However, strong conditions, such as being bounded away from 0, are typically necessary for Doeblin coefficients to establish the existence of information contraction. Building on recently formulated concepts of nonlinear information contraction, we aim to propose a finer-grained Doeblin-based characterization of multi-way contraction behavior which yields non-vacuous contraction guarantees even for channels whose Doeblin coefficient is 0. To this end, we introduce the notion of a Doeblin curve -- a nonlinear function which quantifies the contraction behavior of a Markov kernel on collections of input distributions at specific levels of divergence and power. Through the course of our analysis, we develop a new variational characterization of Doeblin coefficients, present several properties of Doeblin curves, define several versions of power-constrained Doeblin curves, and derive upper and lower bounds using our aforementioned variational characterization. We then utilize these results in diverse areas, including generalization bounds for noisy iterative optimization, error bounds for reliable computation with noisy circuits, and differential privacy guarantees for online iterative algorithms. In particular, we extend results in these areas to broader domains or group settings, leveraging Doeblin curves to reveal finer-grained contraction phenomena than Doeblin coefficients.

cs.IT

Entrywise Error Bounds for Spectral Ranking with Semi-Random Adversaries

Bradley-Terry-Luce (BTL) model estimation is a well-established strategy to rank a collection of items given a dataset of pairwise comparisons. Although the theoretical performance of BTL estimation methods, such as spectral and maximum likelihood estimation, is well studied in the regime of uniformly sampled graphs, generalizing such results to a wider class of random graphs has proved challenging. In this work, we investigate the entry-wise error of spectral algorithms against a semi-random adversary that can arbitrarily boost the sampling probabilities of certain edges. We find that the performance of the unweighted spectral method is heavily dependent on the spectral properties of the generated graph. Furthermore, we show that asymptotic performance approaching that of uniformly sampled graphs can be recovered by appropriately reweighting the observed edges to counteract the adversary and restore the spectral gap. Finally, we provide numerical simulations that support our theoretical findings.

cs.LG

Fair Division Under Inaccurate Preferences

The fair allocation of scarce resources is a central problem in mathematics, computer science, operations research, and economics. While much of the fair-division literature assumes that individuals have underlying cardinal preferences, eliciting exact numerical values is often cognitively burdensome and prone to inaccuracies. A growing body of work in fair division addresses this challenge by assuming access only to ordinal preferences. However, the restricted expressiveness of ordinal preferences makes it challenging to quantify and optimize cardinal fairness objectives such as envy. In this paper, we explore the broad landscape of fair division of indivisible items given inaccurate cardinal preferences, with a focus on minimizing envy. We consider various settings based on whether the true preferences of the agents are stochastic or worst-case, and whether the inaccuracies, modeled as additive noise, are stochastic or worst-case. When the true preferences are stochastic, we show that envy-free allocations can be computed with high probability; this is achieved both in the setting with stochastic and worst-case noise. This generalizes a notable result in stochastic fair division, which establishes a similar guarantee, albeit in the absence of any noise. When the true preferences are worst-case, and the noise is bounded, we analyze the maximum envy achieved by the Round-Robin algorithm. This bound is shown to be tight for deterministic algorithms, and applications of this bound are provided. Lastly, we consider a setting with worst-case preferences and noise, where the true preferences for each item are revealed upon its allocation. Here, we give an efficient online algorithm that guarantees logarithmic maximum envy with high probability. This result generalizes a known result from algorithmic discrepancy to a setting with noisy input data.

cs.GT

Bounds on Maximal Leakage over Bayesian Networks

Maximal leakage quantifies the leakage of information from data $X \in \mathcal{X}$ due to an observation $Y$. While fundamental properties of maximal leakage, such as data processing, sub-additivity, and its connection to mutual information, are well-established, its behavior over Bayesian networks is not well-understood and existing bounds are primarily limited to binary $\mathcal{X}$. In this paper, we investigate the behavior of maximal leakage over Bayesian networks with finite alphabets. Our bounds on maximal leakage are established by utilizing coupling-based characterizations which exist for channels satisfying certain conditions. Furthermore, we provide more general conditions under which such coupling characterizations hold for $|\mathcal{X}| = 4$. In the course of our analysis, we also present a new simultaneous coupling result on maximal leakage exponents. Finally, we illustrate the effectiveness of the proposed bounds with some examples.

cs.IT

Hypothesis Testing for Generalized Thurstone Models

In this work, we develop a hypothesis testing framework to determine whether pairwise comparison data is generated by an underlying \emph{generalized Thurstone model} $\mathcal{T}_F$ for a given choice function $F$. While prior work has predominantly focused on parameter estimation and uncertainty quantification for such models, we address the fundamental problem of minimax hypothesis testing for $\mathcal{T}_F$ models. We formulate this testing problem by introducing a notion of separation distance between general pairwise comparison models and the class of $\mathcal{T}_F$ models. We then derive upper and lower bounds on the critical threshold for testing that depend on the topology of the observation graph. For the special case of complete observation graphs, this threshold scales as $Θ((nk)^{-1/2})$, where $n$ is the number of agents and $k$ is the number of comparisons per pair. Furthermore, we propose a hypothesis test based on our separation distance, construct confidence intervals, establish time-uniform bounds on the probabilities of type I and II errors using reverse martingale techniques, and derive minimax lower bounds using information-theoretic methods. Finally, we validate our results through experiments on synthetic and real-world datasets.

cs.LG

Minimax Hypothesis Testing for the Bradley-Terry-Luce Model

The Bradley-Terry-Luce (BTL) model is one of the most widely used models for ranking a collection of items or agents based on pairwise comparisons among them. Given $n$ agents, the BTL model endows each agent $i$ with a latent skill score $α_i > 0$ and posits that the probability that agent $i$ is preferred over agent $j$ is $α_i/(α_i + α_j)$. In this work, our objective is to formulate a hypothesis test that determines whether a given pairwise comparison dataset, with $k$ comparisons per pair of agents, originates from an underlying BTL model. We formalize this testing problem in the minimax sense and define the critical threshold of the problem. We then establish upper bounds on the critical threshold for general induced observation graphs (satisfying mild assumptions) and develop lower bounds for complete induced graphs. Our bounds demonstrate that for complete induced graphs, the critical threshold scales as $Θ((nk)^{-1/2})$ in a minimax sense. In particular, our test statistic for the upper bounds is based on a new approximation we derive for the separation distance between general pairwise comparison models and the class of BTL models. To further assess the performance of our statistical test, we prove upper bounds on the type I and type II probabilities of error. Much of our analysis is conducted within the context of a fixed observation graph structure, where the graph possesses certain ``nice'' properties, such as expansion and bounded principal ratio. Additionally, we derive several auxiliary results, such as bounds on principal ratios of graphs, $\ell^2$-bounds on BTL parameter estimation under model mismatch, stability of rankings under the BTL model, etc. We validate our theoretical results through experiments on synthetic and real-world datasets and propose a data-driven permutation testing approach to determine test thresholds.

cs.LG

Attenuation from the optical to the extreme ultraviolet by dust associated with broad absorption line quasars: the driving force for outflows

We derive a mean attenuation curve out to the rest-frame extreme ultraviolet (EUV) for 'BAL dust' - the dust causing the additional extinction of active galactic nuclei (AGNs) with broad absorption lines (BALQSOs). In contrast to the normal, relatively flat, mean AGN attenuation curve, BAL dust is well fit by a steeply rising, SMC-like curve. We confirm the shape of the theoretical Weingartner & Draine SMC curve out to 700 Angstroms but the drop in attenuation at still shorter wavelengths is less than predicted. The similar SMC-like attenuation curve for low-ionization BALQSOs (LoBALs) does not support the idea that they are an early phase in the life of an AGN when it is breaking out of a cocoon of star-forming dust. Although the attenuation is only E(B - V) ~ 0.03 - 0.05 in the optical, it rises to one magnitude in the EUV, which is an optimum value for radiative acceleration of dusty gas. Because the spectral energy distribution of AGNs peaks in the EUV, the force on the dust dominates the acceleration of BAL gas. Although the shape of the attenuation curve for LoBALs is similar to the shape for HiBALs, the LoBALs on average show negative attenuation in the optical. This is naturally explained if there is more light scattered into our line of sight in LoBALs compared with non-BALQSOs. We suggest that this and partial covering are causes when attenuation curves appear to be steeper in the UV than an SMC curve.

astro-ph.GA

Doeblin Coefficients and Related Measures

Doeblin coefficients are a classical tool for analyzing the ergodicity and exponential convergence rates of Markov chains. Propelled by recent works on contraction coefficients of strong data processing inequalities, we investigate whether Doeblin coefficients also exhibit some of the notable properties of canonical contraction coefficients. In this paper, we present several new structural and geometric properties of Doeblin coefficients. Specifically, we show that Doeblin coefficients form a multi-way divergence, exhibit tensorization, and possess an extremal trace characterization. We then show that they also have extremal coupling and simultaneously maximal coupling characterizations. By leveraging these characterizations, we demonstrate that Doeblin coefficients act as a nice generalization of the well-known total variation (TV) distance to a multi-way divergence, enabling us to measure the "distance" between multiple distributions rather than just two. We then prove that Doeblin coefficients exhibit contraction properties over Bayesian networks similar to other canonical contraction coefficients. We additionally derive some other results and discuss an application of Doeblin coefficients to distribution fusion. Finally, in a complementary vein, we introduce and discuss three new quantities: max-Doeblin coefficient, max-DeGroot distance, and min-DeGroot distance. The max-Doeblin coefficient shares a connection with the concept of maximal leakage in information security; we explore its properties and provide a coupling characterization. On the other hand, the max-DeGroot and min-DeGroot measures extend the concept of DeGroot distance to multiple distributions.

cs.IT

Generative models for sampling and phase transition indication in spin systems

Recently, generative machine-learning models have gained popularity in physics, driven by the goal of improving the efficiency of Markov chain Monte Carlo techniques and of exploring their potential in capturing experimental data distributions. Motivated by their ability to generate images that look realistic to the human eye, we here study generative adversarial networks (GANs) as tools to learn the distribution of spin configurations and to generate samples, conditioned on external tuning parameters, such as temperature. We propose ways to efficiently represent the physical states, e.g., by exploiting symmetries, and to minimize the correlations between generated samples. We present a detailed evaluation of the various modifications, using the two-dimensional XY model as an example, and find considerable improvements in our proposed implicit generative model. It is also shown that the model can reliably generate samples in the vicinity of the phase transition, even when it has not been trained in the critical region. On top of using the samples generated by the model to capture the phase transition via evaluation of observables, we show how the model itself can be employed as an unsupervised indicator of transitions, by constructing measures of the model's susceptibility to changes in tuning parameters.

cond-mat.stat-mech