SearcharxivSearch

arXiv subjects

Daria Boratyn

Publications and source records attributed to Daria Boratyn.

8 recordsLinked to original sources

Is Textual Similarity Invariant under Machine Translation? Evidence Based on the Political Manifesto Corpus

We investigate the extent to which cosine similarity between paragraph embeddings is invariant under machine translation, using the Manifesto Corpus of over 2,800 political party platforms in 28 languages translated to English via the EU eTranslation service. Rather than measuring translation-induced semantic shift directly we measure the stability of pairwise similarity relationships across embedding models, and use inter-model disagreement on original-language text as a calibrated invariance threshold. This yields a per-language non-inferiority test for four hypotheses about how translation interacts with embedding choice, with verdicts that distinguish languages where translation demonstrably preserves semantic structure from those where it demonstrably degrades it and from those where the available evidence does not resolve the question. The framework is corpus- and pipeline-agnostic and extends naturally to downstream tasks. Applied to our data, it identifies ten languages with translation invariance and four with detectable distortion.

cs.CL

Allocation Proportionality of OWA--Based Committee Scoring Rules

While proportionality is frequently named as a desirable property of voting rules, its interpretation in multiwinner voting differs significantly from that in apportionment. We aim to bridge these two distinct notions of proportionality by introducing the concept of allocation proportionality, founded upon the framework of party elections, where each candidate in a multiwinner election is assigned to a party. A voting rule is allocation proportional if each party's share of elected candidates equals that party's aggregate score. Recognizing that no committee scoring rule can universally satisfy allocation proportionality in practice, we introduce a new measure of allocation proportionality degree and discuss how it relates to other quantitative measures of proportionality. This measure allows us to compare OWA-based committee scoring rules according to how much they diverge from the ideal of allocation proportionality. We present experimental results for several common rules: SNTV, $k$-Borda, Chamberlin-Courant, Harmonic Borda, Proportional $k$-Approval Voting, and Bloc Voting.

cs.GT

Effect of Electoral Seat Bias on Political Polarization: A Computational Perspective

Research on the causes of political polarization points towards multiple drivers of the problem, from social and psychological to economic and technological. However, political institutions stand out, because -- while capable of exacerbating or alleviating polarization -- they can be re-engineered more readily than others. Accordingly, we analyze one class of such institutions -- electoral systems -- investigating whether the large-party seat bias found in many common systems (particularly plurality and Jefferson-D'Hondt) exacerbates polarization. Cross-national empirical data being relatively sparse and heavily confounded, we use computational methods: an agent-based Monte Carlo simulation. We model voter behavior over multiple electoral cycles, building upon the classic spatial model, but incorporating other known voter behavior patterns, such as the bandwagon effect, strategic voting, preference updating, retrospective voting, and the thermostatic effect. We confirm our hypothesis that electoral systems with a stronger large-party bias exhibit significantly higher polarization, as measured by the Mehlhaff index.

physics.soc-ph

Spoiler Susceptibility in Party Elections

An electoral spoiler is usually defined as a losing candidate whose removal would affect the outcome by changing the winner. So far, spoiler effects have been analyzed primarily for single-winner electoral systems. We consider this subject in the context of party elections, where there is no longer a sharp distinction between winners and losers. Hence, we propose a more general definition, under which a party is a spoiler if their elimination causes any other party's share in the outcome to decrease. We characterize spoiler-proof electoral allocation rules for zero-sum voting methods. In particular, we prove that for seats-votes functions only identity is spoiler-proof. However, even if spoilers are unavoidable under common electoral rules, their expected impact can vary depending on the rule. Hence, we introduce a measure of spoilership, which allows us to experimentally compare a number of multiwinner social choice rules according to their spoiler susceptibility. Since the probabilistic models used in COMSOC have been developed for nonparty elections, we extend them to generate multidistrict party elections.

cs.GT

Seat Allocation and Seat Bias under the Jefferson--D'Hondt Method

We prove that under the Jefferson--D'Hondt method of apportionment, given certain distributional assumptions regarding mean rounding residuals, as well as absence of correlations between party vote shares, district sizes (in votes), and multipliers, the seat share of each relevant party is an affine function of the aggregate vote share, the number of relevant parties, and the mean district magnitude. We further show that the first of those assumptions follows approximately from more general ones regarding smoothness, vanishing at the extremes, and total variation of the density of the distribution of vote shares. We also discuss how our main result differs from the simple generalization of the single-district asymptotic seat bias formulae, and how it can be used to derive an estimate of the natural threshold and certain properties thereof.

physics.soc-ph

Machine Learning and Statistical Approaches to Measuring Similarity of Political Parties

Mapping political party systems to metric policy spaces is one of the major methodological problems in political science. At present, in most political science project this task is performed by domain experts relying on purely qualitative assessments, with all the attendant problems of subjectivity and labor intensiveness. We consider how advances in natural language processing, including large transformer-based language models, can be applied to solve that issue. We apply a number of texts similarity measures to party political programs, analyze how they correlate with each other, and -- in the absence of a satisfactory benchmark -- evaluate them against other measures, including those based on expert surveys, voting records, electoral patterns, and candidate networks. Finally, we consider the prospects of relying on those methods to correct, supplement, and eventually replace expert judgments.

cs.CL

Average Weights and Power in Weighted Voting Games

We investigate a class of weighted voting games for which weights are randomly distributed over the standard probability simplex. We provide close-formed formulae for the expectation and density of the distribution of weight of the $k$-th largest player under the uniform distribution. We analyze the average voting power of the $k$-th largest player and its dependence on the quota, obtaining analytical and numerical results for small values of $n$ and a general theorem about the functional form of the relation between the average Penrose--Banzhaf power index and the quota for the uniform measure on the simplex. We also analyze the power of a collectivity to act (Coleman efficiency index) of random weighted voting games, obtaining analytical upper bounds therefor.

cs.GT

A Formal Model of the Relationship between the Number of Parties and the District Magnitude

On the basis of a formula for calculating seat shares and natural thresholds in multidistrict elections under the Jefferson-D'Hondt system and a probabilistic model of electoral behavior based on Pólya's urn model, we propose a new model of the relationship between the district magnitude and the number / effective number of relevant parties. We test that model on both electoral results from multiple countries employing the D'Hondt method (relatively small number of elections, but wide diversity of political configurations) and data based on hundreds of Polish local elections (large number of elections, but much higher degree of parameter uniformity). We also explore some applications of the proposed model, demonstrating how it can be used to estimate the potential effects of electoral engineering.

physics.soc-ph