SearcharxivSearch

arXiv subjects

Priyanka Dey

Publications and source records attributed to Priyanka Dey.

15 recordsLinked to original sources

PALMs: Using Multi Construct-Grounded Rationales for Modeling Population Preferences in LLMs

Large language models are being extensively used to simulate individual user behavior, yet faithfully representing a population requires capturing the systematic variation in values, beliefs, and cultural norms that distinguish one group from another. We introduce Population Aligned Language Models (PALMs), a suite of models each aligned to specific populations, covering five countries: USA, India, Brazil, France and Italy. PALMs are created by synthesizing rationales grounded in psychological and cultural constructs and using these as latent supervision during preference tuning for population-specific alignment. Evaluated across four dimensions: personality, values and beliefs, cultural norms, and morality, PALMs consistently outperform baselines, including culture-specialized models, achieving an average of 8.59% relative improvement over the best baseline across all five populations. Notably, construct-grounded rationales outperform both demographic prompting and survey-based fine-tuning, suggesting that grounding preference learning in psychology and culture provides a richer inductive signal than surface-level response distributions. We further demonstrate strong generalization to downstream applications with- out task-specific supervision: outperforming best baselines by 5.19% in personalized reward modeling, 6.34% in population simulation, and showing strong transfer to social reasoning tasks. Datasets and code are available at: https://github.com/limenlp/PALMs.

cs.CL

Rademacher-type formula and higher order Tur\'{a}n inequalities for $\ell$-regular overpartitions

For $\ell\geq 2$, let $\overline{A}_\ell(n)$ count the number of overpartitions of $n$ with no parts divisible by $\ell$. In this article, we employ the circle method to derive a Rademacher-type formula for $\overline{A}_\ell(n)$, when $\ell$ is a squarefree odd integer. As an application, we derive higher order Tu\'{r}an inequalities for the $\ell$-regular overpartition function using a result of Griffin, Ono, Rolen, and Zagier.

math.NT

Explicit Ensemble Mean Synchronization for Time Scale Generation with Mixed Atomic Clock Ensembles

In this paper, we consider a mixed ensemble containing a mixture of cesium-type and hydrogen maser-type atomic clocks. For the mixed ensemble, the conventional Kalman filtering algorithm has certain limitations due to divergence of the error covariance matrix. To overcome these limitations, we obtain a Kalman filtering algorithm based on observable canonical decomposition that does not have any diverging terms. We use the estimates from the transformed Kalman filter to propose a time scale generation algorithm called explicit ensemble mean synchronization algorithm for the mixed ensemble. In this algorithm, we synchronize the time deviation of each clock from the ideal clock behavior to the unobservable ensemble mean of the phases where the weighting can be decided by the user. By regulating the free-running dynamics associated with the unobservable state, through choosing an appropriate weight vector, the frequency stability of the generated time scale or the synchronized time shared by the clocks is optimized over shorter (resp. longer) intervals, as measured by Hadamard variance. An illustrative example is given to demonstrate the efficiency of our algorithm.

eess.SY

Israel-Hamas War on X: A Case Study of Coordinated Campaigns and Information Integrity

Coordinated campaigns on social media play a critical role in shaping crisis information environments, particularly during the onset of conflicts when uncertainty is high and verified information is scarce. We study the interplay between coordinated campaigns and information integrity through a case study of the 2023 Israel-Hamas War on Twitter (X). We analyze 4.5~million tweets and employ established coordination detection methods to identify 11 coordinated groups involving 541 accounts. We characterize these groups through a multimodal analysis that includes topics, account amplification, toxicity, emotional tone, visual themes, and misleading claims. Our analysis reveal that coordinated campaigns rely predominantly on low-complexity tactics, such as retweet amplification and copy-paste diffusion, and promote distinct narratives consistent with a fragmented manipulation landscape, without centralized control. Widely amplified misleading claims concentrate within just three of the identified coordinated groups; the remaining groups primarily engage in advocacy, religious solidarity, or humanitarian mobilization. Claim-level integrity, toxicity, and emotional signals are mutually uncorrelated: no single behavioral signal is a reliable proxy for the others. Targeting the most prolific spreaders of misleading content for moderation would be effective in reducing such content. However, targeting prolific amplifiers in general would not achieve the same mitigation effect. These findings suggest that evaluating coordination structures jointly with their specific content footprints is needed to effectively prioritize moderation interventions.

cs.SI

RLHF May Not Reflect Genuine Preferences

Reinforcement Learning from Human Feedback (RLHF) assumes that annotation responses reflect genuine human preferences. They often do not. Behavioral scientists have documented for sixty years that people produce responses without holding genuine opinions, construct preferences on the spot from contextual cues, and interpret identical questions differently. Importantly, these failures are common for the judgments on values that matter most for AI alignment. We argue that measurement validity is logically prior to preference aggregation. Before asking how to combine annotations, the field must ask whether the responses being combined are preferences at all. We organize annotation responses along a spectrum, from non-attitudes (no signal) to genuine preferences (full signal), and develop diagnostics that locate responses on this spectrum. In two RLHF datasets, we show that inconsistency is systematic and directionally biased. Filtering high-inconsistency annotators flips majority harm classifications for 18.6% of prompts and shifts mean ratings by over 13 points on a 100-point scale. As such, much of the current RLHF practice models noise as signal and elicitation artifacts as human values.

cs.HC

GRAVITY: A Framework for Personalized Text Generation via Profile-Grounded Synthetic Preferences

Personalization in LLMs often relies on costly human feedback or interaction logs, limiting scalability and neglecting deeper user attributes. To reduce the reliance on human annotations, we introduce GRAVITY (Generative Response with Aligned Values, Interests, and Traits of You), a framework for generating synthetic, profile-grounded preference data that captures users' interests, values, beliefs, and personality traits. By integrating demographic, cultural, and psychological frameworks -- including Hofstede's cultural dimensions, Schwartz's basic values, the World Values Survey, and Big Five OCEAN traits -- GRAVITY synthesizes preference pairs to guide personalized content generation. We evaluate GRAVITY on book descriptions for 400 Amazon users, comparing it to prompt-based conditioning, standard fine-tuning, and naive synthetic pair generation. Profile-grounded synthetic data consistently improves generation, especially across multiple cultures (USA, Brazil, Japan, India), achieving over 4% higher preference gains across baselines, with user studies showing that GRAVITY outputs are preferred over 86% of the time. Our results show that scenario-grounded synthetic data can capture richer user variation, reduce reliance on costly annotation, and produce more engaging, user-centered content, offering a scalable path for LLM personalization.

cs.CL

Can LLMs Express Personality Across Cultures? Introducing CulturalPersonas for Evaluating Trait Alignment

As LLMs become central to interactive applications, ranging from tutoring to mental health, the ability to express personality in culturally appropriate ways is increasingly important. While recent works have explored personality evaluation of LLMs, they largely overlook the interplay between culture and personality. To address this, we introduce CulturalPersonas, the first large-scale benchmark with human validation for evaluating LLMs' personality expression in culturally grounded, behaviorally rich contexts. Our dataset spans 3,000 scenario-based questions across six diverse countries, designed to elicit personality through everyday scenarios rooted in local values. We evaluate three LLMs, using both multiple-choice and open-ended response formats. Our results show that CulturalPersonas improves alignment with country-specific human personality distributions (over a 20% reduction in Wasserstein distance across models and countries) and elicits more expressive, culturally coherent outputs compared to existing benchmarks. CulturalPersonas surfaces meaningful modulated trait outputs in response to culturally grounded prompts, offering new directions for aligning LLMs to global norms of behavior. By bridging personality expression and cultural nuance, we envision that CulturalPersonas will pave the way for more socially intelligent and globally adaptive LLMs.

cs.CL

Can LLMs Grasp Implicit Cultural Values? Benchmarking LLMs' Cultural Intelligence with CQ-Bench

Cultural Intelligence (CQ) refers to the ability to understand unfamiliar cultural contexts, a crucial skill for large language models (LLMs) to effectively engage with globally diverse users. Existing studies often focus on explicitly stated cultural norms, but fail to capture the subtle, implicit values that are common in daily conversation. To address this gap, we introduce CQBench, a benchmark specifically designed to assess LLMs' capability to infer implicit cultural values from natural conversational contexts. CQBench consists of multi character conversation based stories using values from the World Value Survey and the GlobalOpinions, with topics including ethical, religious, social, etc. Our automatic dataset construction pipeline integrates rigorous validation procedures (incorporation, consistency, and implicitness checks), achieving a 94.5% human model agreement in the final validation. To leverage CQBench data, we design three tasks of increasing complexity: attitude detection, value selection, and value extraction. These tasks evaluate whether models can detect attitude and recognize values embedded within natural dialogues rather than relying on explicit cultural knowledge. We find that while frontier models like o1 reach human level performance in value selection (0.809 F1), they still fall short in nuanced attitude detection (0.622 F1). Notably, finetuning a smaller LLaMA-3.2-3B on only 500 culturally rich examples improves performance by over 10%, even outperforming o3-mini in some cases. Using CQ-Bench, we provide insights into the current challenges in LLMs' CQ research and suggest practical pathways for enhancing LLMs' cross-cultural reasoning abilities.

cs.CL

On the detection of the presence of malicious components in cyber-physical systems in the almost sure sense

This article studies a fundamental problem of security of cyber-physical systems (CPSs): that of detecting, almost surely, the presence of malicious components in the CPS. We assume that some of the actuators may be malicious while all sensors are honest. We introduce a novel idea of separability of state trajectories generated by CPSs in two situations: those under the nominal no-attack situation and those under the influence of an attacker. We establish its connection to security of CPSs in the context of detecting the presence of malicious actuators (if any) in them. As primary contributions we establish necessary and sufficient conditions for the aforementioned detection in CPSs modeled as Markov decision processes (MDPs). Moreover, we focus on the mechanism of perturbing the pre-determined control policies of the honest agents in CPSs modeled as stochastic linear systems, by injecting a certain class of random process called private excitation; sufficient conditions for detectability and non-detectability of the presence of malicious actuators assuming that the policies are randomized history dependent and randomized Markovian, are established. Several technical aspects of our results are discussed extensively.

math.OC

Coordinated Activity Modulates the Behavior and Emotions of Organic Users: A Case Study on Tweets about the Gaza Conflict

Social media has become a crucial conduit for the swift dissemination of information during global crises. However, this also paves the way for the manipulation of narratives by malicious actors. This research delves into the interaction dynamics between coordinated (malicious) entities and organic (regular) users on Twitter amidst the Gaza conflict. Through the analysis of approximately 3.5 million tweets from over 1.3 million users, our study uncovers that coordinated users significantly impact the information landscape, successfully disseminating their content across the network: a substantial fraction of their messages is adopted and shared by organic users. Furthermore, the study documents a progressive increase in organic users' engagement with coordinated content, which is paralleled by a discernible shift towards more emotionally polarized expressions in their subsequent communications. These results highlight the critical need for vigilance and a nuanced understanding of information manipulation on social media platforms.

cs.SI

Investigating Stylistic Profiles for the Task of Empathy Classification in Medical Narrative Essays

One important aspect of language is how speakers generate utterances and texts to convey their intended meanings. In this paper, we bring various aspects of the Construction Grammar (CxG) and the Systemic Functional Grammar (SFG) theories in a deep learning computational framework to model empathic language. Our corpus consists of 440 essays written by premed students as narrated simulated patient-doctor interactions. We start with baseline classifiers (state-of-the-art recurrent neural networks and transformer models). Then, we enrich these models with a set of linguistic constructions proving the importance of this novel approach to the task of empathy classification for this dataset. Our results indicate the potential of such constructions to contribute to the overall empathy profile of first-person narrative essays.

cs.CL

Strong Quantum Confinement Effects and Chiral Excitons in Bio-Inspired ZnO-Amino Acid Co-Crystals

Elucidating the underlying principles behind band gap engineering is paramount for the successful implementation of semiconductors in photonic and optoelectronic devices. Recently it has been shown that the band gap of a wide and direct band gap semiconductor, such as ZnO, can be modified upon co-crystallization with amino acids, with the role of the biomolecules remaining unclear. Here, by probing and modeling the light emitting properties of ZnO-amino acid co-crystals, we identify the amino acids role on this band gap modulation and demonstrate their effective chirality transfer to the inter-band excitations in ZnO. Our 3D quantum model suggests that the strong band edge emission blue shift in the co-crystals can be explained by a quasi-periodic distribution of amino acid potential barriers within the ZnO crystal lattice. Overall, our findings indicate that biomolecule co-crystallization can be used as a truly bio-inspired means to induce chiral quantum confinement effects in quasi-bulk semiconductors.

cond-mat.mtrl-sci

On Minimum Cost Sparsest Input-Connectivity for Controllability of Linear Systems

We deal with algorithmic techniques for minimal cost input-connectivity while maintaining controllability of linear systems. The input matrix is assumed to be constrained in the sense that the set of states that each input (if present) can influence is known a priori, and that each interconnection between an input and a state is associated with a certain cost. In this setting we determine a set of input-connections that lead to the minimum cost and ensures that the resulting system is structurally controllable. We also identify a sparsest set of input-connections with minimum cost while maintaining structural controllability of the system. A large class of systems are identified for which these problems are solvable in polynomial time using efficient algorithms. A 2-approximation solution is presented for the general case. Graph-theoretic tools are employed to tackle the above class of constrained design problems. Illustrative examples are included to demonstrate the efficacy of the techniques developed here.

math.OC

Efficient constrained sensor placement for observability of linear systems

This article studies two problems related to observability and efficient constrained sensor placement in linear time-invariant discrete-time systems with partial state observations. (i) We impose the condition that both the set of outputs and the state that each output can measure are pre-specified. We establish that for any fixed \(k > 2\), the problem of placing the minimum number of sensors/outputs required to ensure that the structural observability index is at most \(k\), is NP-complete. Conversely, we identify a subclass of systems whose structures are directed trees with self-loops at every state vertex, for which the problem can be solved in linear time. (ii) Assuming that the set of states that each given output can measure is given, we prove that the problem of selecting a pre-assigned number of sensors in order to maximize the number of states of the system that are structurally observable is also NP-hard. As an application, we identify suitable conditions on the system structure under which there exists an efficient greedy strategy, which we provide, to obtain a \((1-\frac{1}{e})\)-approximate solution. An illustration of the techniques developed for this problem is given on the benchmark IEEE 118-bus power network containing roughly \(400\) states in its linearized model.

math.OC

Resilience of Complex Networks

This article determines and characterizes the minimal number of actuators needed to ensure structural controllability of a linear system under structural alterations that can severe the connection between any two states. We assume that initially the system is structurally controllable with respect to a given set of controls, and propose an efficient system-synthesis mechanism to find the minimal number of additional actuators required for resilience of the system w.r.t such structural changes. The effectiveness of this approach is demonstrated by using standard IEEE power networks.

eess.SY