SearcharxivSearch

arXiv subjects

Chris Hays

Publications and source records attributed to Chris Hays.

At least 19 recordsLinked to original sources

Strategic Candidacy in Generative AI Arenas

AI arenas, which rank generative models from pairwise preferences of users, are a popular method for measuring the relative performance of models in the course of their organic use. Because rankings are computed from noisy preferences, there is a concern that model producers can exploit this randomness by submitting many models (e.g., multiple variants of essentially the same model) and thereby artificially improve the rank of their top models. This can lead to degradations in the quality, and therefore the usefulness, of the ranking. In this paper, we begin by establishing, both theoretically and in simulations calibrated to data from the platform Arena (formerly LMArena, Chatbot Arena), conditions under which producers can benefit from submitting clones when their goal is to be ranked highly. We then propose a new mechanism for ranking models from pairwise comparisons, called You-Rank-We-Rank (YRWR). It requires that producers submit rankings over their own models and uses these rankings to correct statistical estimates of model quality. We prove that this mechanism is approximately clone-robust, in the sense that a producer cannot improve their rank much by doing anything other than submitting each of their unique models exactly once. Moreover, to the extent that model producers are able to correctly rank their own models, YRWR improves overall ranking accuracy. In further simulations, we show that indeed the mechanism is approximately clone-robust and quantify improvements to ranking accuracy, even under producer misranking.

cs.LG

Statistical Guarantees in the Search for Less Discriminatory Algorithms

U.S. discrimination law can impose liability on firms that fail to adopt a less discriminatory alternative (LDA): a decision policy that achieves the same business objectives while reducing disparate impact on legally protected groups. Recent scholarship argues that this doctrine has direct implications for algorithmic decision-making in high-stakes domains such as employment, lending, and housing, potentially obligating firms to search for "less discriminatory algorithms" (Black et al., 2024). Regulators have at times encouraged proactive LDA searches, reinforcing the expectation of a good-faith effort to identify equally performant models with lower disparate impact. Model multiplicity makes such searches plausible: retraining with different random seeds can yield models with comparable predictive performance but materially different disparate impacts. Yet firms cannot retrain indefinitely, raising a central question: when is the search sufficient to demonstrate good faith? We formalize LDA search under multiplicity as an optimal stopping problem in which a developer seeks to produce evidence that further search is unlikely to yield meaningful improvements. Our main contribution is an adaptive stopping algorithm that provides a high-probability upper bound on the best disparate-impact gains attainable through continued retraining, enabling developers to certify (e.g., to a court) that additional search is unlikely to help. We also show how stronger distributional assumptions over the model space can yield tighter bounds, and we validate the approach on real-world credit and housing datasets.

cs.CY

The Impossibility of Inverse Permutation Learning in Transformer Models

In this technical note, we study the problem of inverse permutation learning in decoder-only transformers. Given a permutation and a string to which that permutation has been applied, the model is tasked with producing the original (``canonical'') string. We argue that this task models a natural robustness property across a variety of reasoning tasks, including long-context retrieval, multiple choice QA and in-context learning. Our primary contribution is an impossibility result: we show that an arbitrary depth, decoder-only transformer cannot learn this task. This result concerns the expressive capacity of decoder-only transformer models and is agnostic to training dynamics or sample complexity. We give a pair of alternative constructions under which inverse permutation learning is feasible. The first of these highlights the fundamental role of the causal attention mask, and reveals a gap between the expressivity of encoder-decoder transformers and the more popular decoder-only architecture. The latter result is more surprising: we show that simply padding the input with ``scratch tokens" yields a construction under which inverse permutation learning is possible. We conjecture that this may suggest an alternative mechanism by which chain-of-thought prompting or, more generally, intermediate ``thinking'' tokens can enable reasoning in large language models, even when these tokens encode no meaningful semantic information (e.g., the results of intermediate computations).

cs.LG

Sensitivity of $W$-boson measurements to low-mass right-handed neutrinos

A low-mass right-handed neutrino could interact with electroweak bosons via mixing, a mediator particle, or loop corrections. Using an effective field theory, we determine constraints on these interactions from $W$-boson measurements at hadron colliders. Due to the difference in the initial states at the Tevatron and the LHC, $W$-boson decays to a right-handed neutrino would artificially increase the mass measured at the Tevatron while only affecting the difference between $W^+$ and $W^-$ mass measurements at the LHC. Measurements from CDF and the LHC are used to infer the corresponding parameter values, which are found to be inconsistent between the two. The LHC experiments can improve sensitivity to these interactions by measuring the cosine of the helicity angle using $W$ bosons produced with transverse momentum above $\approx 50$ GeV.

hep-ph

Double Machine Learning for Causal Inference under Shared-State Interference

Researchers and practitioners often wish to measure treatment effects in settings where units interact via markets and recommendation systems. In these settings, units are affected by certain shared states, like prices, algorithmic recommendations or social signals. We formalize this structure, calling it shared-state interference, and argue that our formulation captures many relevant applied settings. Our key modeling assumption is that individuals' potential outcomes are independent conditional on the shared state. We then prove an extension of a double machine learning (DML) theorem providing conditions for achieving efficient inference under shared-state interference. We also instantiate our general theorem in several models of interest where it is possible to efficiently estimate the average direct effect (ADE) or global average treatment effect (GATE).

stat.ML

Inducing Efficient and Equitable Professional Networks through Link Recommendations

Professional networks are a key determinant of individuals' labor market outcomes. They may also play a role in either exacerbating or ameliorating inequality of opportunity across demographic groups. In a theoretical model of professional network formation, we show that inequality can increase even without exogenous in-group preferences, confirming and complementing existing theoretical literature. Increased inequality emerges from the differential leverage privileged and unprivileged individuals have in forming connections due to their asymmetric ex ante prospects. This is a formalization of a source of inequality in the labor market which has not been previously explored. We next show how inequality-aware platforms may reduce inequality by subsidizing connections, through link recommendations that reduce costs, between privileged and unprivileged individuals. Indeed, mixed-privilege connections turn out to be welfare improving, over all possible equilibria, compared to not recommending links or recommending some smaller fraction of cross-group links. Taken together, these two findings reveal a stark reality: professional networking platforms that fail to foster integration in the link formation process risk reducing the platform's utility to its users and exacerbating existing labor market inequality.

cs.GT

From Fairness to Infinity: Outcome-Indistinguishable (Omni)Prediction in Evolving Graphs

Professional networks provide invaluable entree to opportunity through referrals and introductions. A rich literature shows they also serve to entrench and even exacerbate a status quo of privilege and disadvantage. Hiring platforms, equipped with the ability to nudge link formation, provide a tantalizing opening for beneficial structural change. We anticipate that key to this prospect will be the ability to estimate the likelihood of edge formation in an evolving graph. Outcome-indistinguishable prediction algorithms ensure that the modeled world is indistinguishable from the real world by a family of statistical tests. Omnipredictors ensure that predictions can be post-processed to yield loss minimization competitive with respect to a benchmark class of predictors for many losses simultaneously, with appropriate post-processing. We begin by observing that, by combining a slightly modified form of the online K29 star algorithm of Vovk (2007) with basic facts from the theory of reproducing kernel Hilbert spaces, one can derive simple and efficient online algorithms satisfying outcome indistinguishability and omniprediction, with guarantees that improve upon, or are complementary to, those currently known. This is of independent interest. We apply these techniques to evolving graphs, obtaining online outcome-indistinguishable omnipredictors for rich -- possibly infinite -- sets of distinguishers that capture properties of pairs of nodes, and their neighborhoods. This yields, inter alia, multicalibrated predictions of edge formation with respect to pairs of demographic groups, and the ability to simultaneously optimize loss as measured by a variety of social welfare functions.

cs.LG

Equilibria, Efficiency, and Inequality in Network Formation for Hiring and Opportunity

Professional networks -- the social networks among people in a given line of work -- can serve as a conduit for job prospects and other opportunities. Here we propose a model for the formation of such networks and the transfer of opportunities within them. In our theoretical model, individuals strategically connect with others to maximize the probability that they receive opportunities from them. We explore how professional networks balance connectivity, where connections facilitate opportunity transfers to those who did not get them from outside sources, and congestion, where some individuals receive too many opportunities from their connections and waste some of them. We show that strategic individuals are over-connected at equilibrium relative to a social optimum, leading to a price of anarchy for which we derive nearly tight asymptotic bounds. We also show that, at equilibrium, individuals form connections to those who provide similar benefit to them as they provide to others. Thus, our model provides a microfoundation in professional networking contexts for the fundamental sociological principle of homophily, that "similarity breeds connection," which in our setting is realized as a form of status homophily based on alignment in individual benefit. We further explore how, even if individuals are a priori equally likely to receive opportunities from outside sources, equilibria can be unequal, and we provide nearly tight bounds on how unequal they can be. Finally, we explore the ability for online platforms to intervene to improve social welfare and show that natural heuristics may result in adverse effects at equilibrium. Our simple model allows for a surprisingly rich analysis of coordination problems in professional networks and suggests many directions for further exploration.

cs.GT

Focus topics for the ECFA study on Higgs / Top / EW factories

In order to stimulate new engagement and trigger some concrete studies in areas where further work would be beneficial towards fully understanding the physics potential of an $e^+e^-$ Higgs / Top / Electroweak factory, we propose to define a set of focus topics. The general reasoning and the proposed topics are described in this document.

hep-ph

Content Moderation and the Formation of Online Communities: A Theoretical Framework

We study the impact of content moderation policies in online communities. In our theoretical model, a platform chooses a content moderation policy and individuals choose whether or not to participate in the community according to the fraction of user content that aligns with their preferences. The effects of content moderation, at first blush, might seem obvious: it restricts speech on a platform. However, when user participation decisions are taken into account, its effects can be more subtle $\unicode{x2013}$ and counter-intuitive. For example, our model can straightforwardly demonstrate how moderation policies may increase participation and diversify content available on the platform. In our analysis, we explore a rich set of interconnected phenomena related to content moderation in online communities. We first characterize the effectiveness of a natural class of moderation policies for creating and sustaining stable communities. Building on this, we explore how resource-limited or ideological platforms might set policies, how communities are affected by differing levels of personalization, and competition between platforms. Our model provides a vocabulary and mathematically tractable framework for analyzing platform decisions about content moderation.

cs.DS

Compatibility and combination of world W-boson mass measurements

The compatibility of W-boson mass measurements performed by the ATLAS, LHCb, CDF, and D0 experiments is studied using a coherent framework with theory uncertainty correlations. The measurements are combined using a number of recent sets of parton distribution functions (PDF), and are further combined with the average value of measurements from the Large Electron-Positron collider. The considered PDF sets generally have a low compatibility with a suite of global rapidity-sensitive Drell-Yan measurements. The most compatible set is CT18 due to its larger uncertainties. A combination of all mW measurements yields a value of mW = 80394.6 +- 11.5 MeV with the CT18 set, but has a probability of compatibility of 0.5% and is therefore disfavoured. Combinations are performed removing each measurement individually, and a 91% probability of compatibility is obtained when the CDF measurement is removed. The corresponding value of the W boson mass is 80369.2 +- 13.3 MeV, which differs by 3.6 sigma from the CDF value determined using the same PDF set.

hep-ex

Simplistic Collection and Labeling Practices Limit the Utility of Benchmark Datasets for Twitter Bot Detection

Accurate bot detection is necessary for the safety and integrity of online platforms. It is also crucial for research on the influence of bots in elections, the spread of misinformation, and financial market manipulation. Platforms deploy infrastructure to flag or remove automated accounts, but their tools and data are not publicly available. Thus, the public must rely on third-party bot detection. These tools employ machine learning and often achieve near perfect performance for classification on existing datasets, suggesting bot detection is accurate, reliable and fit for use in downstream applications. We provide evidence that this is not the case and show that high performance is attributable to limitations in dataset collection and labeling rather than sophistication of the tools. Specifically, we show that simple decision rules -- shallow decision trees trained on a small number of features -- achieve near-state-of-the-art performance on most available datasets and that bot detection datasets, even when combined together, do not generalize well to out-of-sample datasets. Our findings reveal that predictions are highly dependent on each dataset's collection and labeling procedures rather than fundamental differences between bots and humans. These results have important implications for both transparency in sampling and labeling procedures and potential biases in research using existing bot detection tools for pre-processing.

cs.LG

Electroweak input parameters

Different sets of electroweak input parameters are discussed for SMEFT predictions at the LHC. The {Gmu,mZ,mW} one is presently recommended.

hep-ph

Prospects for direct CP tests of $hqq$ interactions

We study the prospects for probing the CP structure of $hqq$ interactions using the decays of the lightest baryon $Λ_q$ formed in the quark's hadronization. The low yields of reconstructible events make it unlikely for tests to be performed with the next generation of colliders. In $h\to b\bar b\to Λ_b \barΛ_b$ decays a CP-sensitive distribution could be measured with a high-luminosity $e^+ e^-$ collider, while in both $h\to b\bar b\to Λ_b \barΛ_b$ and $h\to c\bar c\to Λ_c \barΛ_c$ decays such a distribution could be measured with a very high luminosity $μ^+ μ^-$ collider. However, we find that only the $μ^+ μ^-$ collider can produce enough $h\to b\bar b\to Λ_b \barΛ_b$ decays to probe a physical CP asymmetry in the $hbb$ vertex.

hep-ph

The Effect of the Rooney Rule on Implicit Bias in the Long Term

A robust body of evidence demonstrates the adverse effects of implicit bias in various contexts--from hiring to health care. The Rooney Rule is an intervention developed to counter implicit bias and has been implemented in the private and public sectors. The Rooney Rule requires that a selection panel include at least one candidate from an underrepresented group in their shortlist of candidates. Recently, Kleinberg and Raghavan proposed a model of implicit bias and studied the effectiveness of the Rooney Rule when applied to a single selection decision. However, selection decisions often occur repeatedly over time. Further, it has been observed that, given consistent counterstereotypical feedback, implicit biases against underrepresented candidates can change. We consider a model of how a selection panel's implicit bias changes over time given their hiring decisions either with or without the Rooney Rule in place. Our main result is that, when the panel is constrained by the Rooney Rule, their implicit bias roughly reduces at a rate that is the inverse of the size of the shortlist--independent of the number of candidates, whereas without the Rooney Rule, the rate is inversely proportional to the number of candidates. Thus, when the number of candidates is much larger than the size of the shortlist, the Rooney Rule enables a faster reduction in implicit bias, providing an additional reason in favor of using it as a strategy to mitigate implicit bias. Towards empirically evaluating the long-term effect of the Rooney Rule in repeated selection decisions, we conduct an iterative candidate selection experiment on Amazon MTurk. We observe that, indeed, decision-makers subject to the Rooney Rule select more minority candidates in addition to those required by the rule itself than they would if no rule is in effect, and do so without considerably decreasing the utility of candidates selected.

cs.CY

Exact SMEFT formulation and expansion to $\mathcal{O}(v^4/Λ^4)$

The Standard Model Effective Field Theory (SMEFT) theoretical framework is increasingly used to interpret particle physics measurements and constrain physics beyond the Standard Model. We investigate the truncation of the effective-operator expansion using the geometric formulation of the SMEFT, which allows exact solutions, up to mass-dimension eight. Using this construction, we compare the exact solution to the expansion at ${\mathcal{O}}(v^2/Λ^2)$, partial ${\mathcal{O}}(v^4/Λ^4)$ using a subset of terms with dimension-6 operators, and full ${\mathcal{O}}(v^4/Λ^4)$, where $v$ is the vacuum expectation value and $Λ$ is the scale of new physics. This comparison is performed for general values of the coefficients, and for the specific model of a heavy U(1) gauge field kinetically mixed with the Standard Model. We additionally determine the input-parameter scheme dependence at all orders in $v/Λ$, and show that this dependence increases at higher orders in $v/Λ$.

hep-ph

Angles on CP-violation in Higgs boson interactions

CP-violation in the Higgs sector remains a possible source of the baryon asymmetry of the universe. Recent differential measurements of signed angular distributions in Higgs boson production provide a general experimental probe of the CP structure of Higgs boson interactions. We interpret these measurements using the Standard Model Effective Field Theory and show that they do not distinguish the various CP-violating operators that couple the Higgs and gauge fields. However, the constraints can be sharpened by measuring additional CP-sensitive observables and exploiting phase-space-dependent effects. Using these observables, we demonstrate that perturbatively meaningful constraints on CP-violating operators can be obtained at the LHC with luminosities of ${\cal{O}}$(100/fb). Our results provide a roadmap to a global Higgs boson coupling analysis that includes CP-violating effects.

hep-ph

On the impact of dimension-eight SMEFT operators on Higgs measurements

Using the production of a Higgs boson in association with a $W$ boson as a test case, we assess the impact of dimension-8 operators within the context of the Standard Model Effective Field Theory. Dimension-8--SM-interference and dimension-6-squared terms appear at the same order in an expansion in $1/Λ$, hence dimension-8 effects can be treated as a systematic uncertainty on the new physics inferred from analyses using dimension-6 operators alone. To study the phenomenological consequences of dimension-8 operators, one must first determine the complete set of operators that can contribute to a given process. We accomplish this through a combination of Hilbert series methods, which yield the number of invariants and their field content, and a step-by-step recipe to convert the Hilbert series output into a phenomenologically useful format. The recipe we provide is general and applies to any other process within the dimension $\le 8$ Standard Model Effective Theory. We quantify the effects of dimension-8 by turning on one dimension-6 operator at a time and setting all dimension-8 operator coefficients to the same magnitude. Under this procedure and given the current accuracy on $σ(pp \to h\,W^+)$, we find the effect of dimension-8 operators on the inferred new physics scale to be small, $\mathcal O(\text{few}\,\%)$, with some variation depending on the relative signs of the dimension-8 coefficients and on which dimension-6 operator is considered. The impact of the dimension-8 terms grows as $σ(pp \to h\,W^+)$ is measured more accurately or (more significantly) in high-mass kinematic regions. We provide a FeynRules implementation of our operator set to be used for further more detailed analyses.

hep-ph