SearcharxivSearch

arXiv subjects

Harang Ju

Publications and source records attributed to Harang Ju.

16 recordsLinked to original sources

Pairit: A Platform for Live Experiments on Human-AI Collaboration

Organizational design in the era of artificial intelligence requires experimental methods that can test how human-AI groups coordinate, delegate, and make decisions. Programmable platforms coordinate live human-to-human sessions or real-time human-AI chat, but researchers cannot easily declare experiment protocols in which AI participants both communicate and act on shared work within one auditable configuration. Here we introduce Pairit, an online platform that facilitates the design, testing, and deployment of experiments that test human-AI organizational designs and interventions. Through a single YAML configuration file, researchers declare an executable experiment graph (pages, routing, randomization, matchmaking, chat, shared workspaces, server-hosted agents, surveys, timers, and custom HTML components) and combine any number of humans and AI agents in live sessions. We have validated the feasibility of the platform through multiple live deployments, including peer-reviewed published studies, capturing high-resolution process traces of communication, negotiation, and collaborative work in live human-AI dyads. By representing complex interactive protocols as standardized, auditable configuration files, Pairit provides reusable infrastructure for specifying, deploying, and sharing live human-AI organizational experiments.

cs.HC

SMolLM: Small Language Models Learn Small Molecular Grammar

Language models for molecular design have scaled to hundreds of millions of parameters, yet how they learn chemical grammar is poorly understood. We train SMolLM, a 53K-parameter weight-shared transformer, to generate novel SMILES with 95% validity on the ZINC-250K drug-like-molecule benchmark, outperforming a standard GPT with 10 times more parameters. Mechanistically, the same block resolves SMILES constraints across passes in a fixed hierarchy: brackets first, rings second, and valence last, as shown by error classification and linear probing, with ablation isolating the bracket-matching head. Together, these results yield a compact, mechanistically interpretable molecular generator and a testbed for studying iterative computation in formal-language domains.

cs.LG

Act or Escalate? Evaluating Escalation Behavior in Automation with Language Models

Effective automation hinges on deciding when to act and when to escalate. We model this as a decision under uncertainty: an LLM forms a prediction, estimates its probability of being correct, and compares the expected costs of acting and escalating. Using this framework across five domains of recorded human decisions-demand forecasting, content recommendation, content moderation, loan approval, and autonomous driving-and across multiple model families, we find marked differences in the implicit thresholds models use to trade off these costs. These thresholds vary substantially and are not predicted by architecture or scale, while self-estimates are miscalibrated in model-specific ways. We then test interventions that target this decision process by varying cost ratios, providing accuracy signals, and training models to follow the desired escalation rule. Prompting helps mainly for reasoning models. SFT on chain-of-thought targets yields the most robust policies, which generalize across datasets, cost ratios, prompt framings, and held-out domains. These results suggest that escalation behavior is a model-specific property that should be characterized before deployment, and that robust alignment benefits from training models to reason explicitly about uncertainty and decision costs.

cs.LG

When Coordination Is Avoidable: A Monotonicity Analysis of Organizational Tasks

Organizations devote substantial resources to coordination, yet which tasks actually require it for correctness remains unclear. The problem is acute in multi-agent AI systems, where coordination cost is directly measurable and can exceed the cost of the work itself. Distributed systems theory provides a precise criterion: coordination is required when a task specification is non-monotonic, meaning that as histories grow, new information can invalidate prior conclusions. Here we show that Thompson's classic taxonomy of interdependence maps to that criterion, yielding a decision rule for when coordination is required for correctness. We formalize the correspondence in a bridge theorem, apply the rule to 65 workflows from the American Productivity & Quality Center (APQC) and, with a calibrated large language model (LLM), 13,417 Occupational Information Network (O*NET) tasks, and illustrate it in multi-agent AI simulations. Under our decompositions, 74% of workflows and 42% of O*NET tasks are monotonic, implying that up to 24-57% of coordination spending is unnecessary for correctness.

cs.MA

Personality pairing improves human-AI collaboration

Here we examine how AI agent "personalities" interact with human personalities to shape human-AI collaboration and performance. In a large-scale, preregistered randomized experiment, we paired 1,258 participants with AI agents prompted to exhibit varying levels of the Big Five personality traits. These human-AI teams produced 7,266 display ads for a real think tank, which we evaluated using 1,168 independent human raters, and a field experiment on X that generated nearly 5 million impressions. We found that human and AI personalities individually shaped ad quality and teamwork and that human-AI personality pairings directly influenced ad quality. For example, extraverted humans paired with conscientious AI produced the lowest quality ads, followed by conscientious humans paired with agreeable AI and neurotic humans paired with conscientious AI. In the field experiment, ad quality significantly influenced ad performance, measured by click-through rates and cost-per-click. Together, these results demonstrate that personality pairing can improve human-AI collaboration and performance. They also motivate future research on the complex implications of AI personalization for human-AI collaboration, teamwork, and performance.

cs.HC

Are Crypto Ecosystems (De)centralizing? A Framework for Longitudinal Analysis

Blockchain technology relies on decentralization to resist faults and attacks while operating without trusted intermediaries. Although industry experts have touted decentralization as central to their promise and disruptive potential, it is still unclear whether the crypto ecosystems built around blockchains are becoming more or less decentralized over time. As crypto plays an increasing role in facilitating economic transactions and peer-to-peer interactions, measuring their decentralization becomes even more essential. We thus propose a systematic framework for measuring the decentralization of crypto ecosystems over time and compare commonly used decentralization metrics. We applied this framework to seven prominent crypto ecosystems, across five distinct subsystems and across their lifetime for over 15 years. Our analysis revealed that while crypto has largely become more decentralized over time, recent trends show a shift toward centralization in the consensus layer, NFT marketplaces, and developers. Our framework and results inform researchers, policymakers, and practitioners about the design, regulation, and implementation of crypto ecosystems and provide a systematic, replicable foundation for future studies.

cs.CR

Explaining Sustained Blockchain Decentralization with Quasi-Experiments: The Resource Flexibility of Consensus Mechanisms

Decentralization is a fundamental design element of the Web3 economy. Blockchains and distributed consensus mechanisms are touted as fault-tolerant, attack-resistant, and collusion-proof because they are decentralized. Recent analyses, however, find some blockchains are decentralized, others are centralized, and that there are trends towards both centralization and decentralization in the blockchain economy. Despite the importance and variability of decentralization across blockchains, we still know little about what enables or constrains blockchain decentralization. We hypothesize that the resource flexibility of consensus mechanisms is a key enabler of the sustained decentralization of blockchain networks. We test this hypothesis using three quasi-experimental shocks -- policy-related, infrastructure-related, and technical -- to resources used in consensus. We find strong suggestive evidence that the resource flexibility of consensus mechanisms enables sustained blockchain decentralization and discuss the implications for the design, regulation, and implementation of blockchains.

cs.NI

Advertising Spillovers in Mobile Apps: Evidence from Ad Shutoffs and Store Rankings

Using advertising campaign data from a large US-based mobile game developer, the authors study a global advertising shutoff in the context of mobile app install ads. Contrary to prior studies in search advertising, which reveal a major over-attribution problem where paid advertising takes credit for organic traffic that would have occurred otherwise, this study shows the opposite: paid ads generate positive spillovers to organic installs. Event study analysis shows that the shutoff decreased organic installs by 20-30%. Fixed-effects panel models estimated on longer-term data find that every $100 spent is associated with 32 paid installs and 2.2 organic installs, highly consistent with the event study estimates. Further analysis strongly suggests that this positive paid-to-organic spillover operates through a ranking mechanism: paid installs boost app store category rankings, thereby increasing organic visibility. Combining campaign and ranking data, the authors find that (1) ad spend has a statistically and economically significant relationship with store rankings; and (2) the relationship between organic installs and ad spend disappears once these rankings are factored in, indicating that they absorb the relationship. These findings demonstrate that mobile app install ads are more effective than paid install metrics alone indicate, implying that developers may systematically underinvest in marketing.

econ.GN

Collaborating with AI Agents: Field Experiments on Teamwork, Productivity, and Performance

We examined the mechanisms underlying productivity and performance gains from AI agents using a large-scale experiment on Pairit, a platform we developed to study human-AI collaboration. We randomly assigned 2,234 participants to human-human and human-AI teams that produced 11,024 ads for a think tank. We evaluated the ads using independent human ratings and a field experiment on X which garnered ~5M impressions. We found human-AI teams produced 50% more ads per worker and higher text quality, while human-human teams produced higher image quality, suggesting a jagged frontier of AI agent capability. Human-AI teams also produced more homogeneous, or self-similar, outputs. The field experiment revealed higher text quality improved click-through rates and view-through duration, while higher image quality improved cost-per-click rates. We found three mechanisms explained these effects. First, human-AI collaboration was more task-oriented, with 25% more task-oriented messages and 18% fewer interpersonal messages. Second, human-AI collaboration displayed more delegation, as participants delegated 17% more work to AI agents than to human partners and performed 62% fewer direct text edits when working with AI. Third, recognition that the collaborator was an AI moderated these effects as participants who correctly identified they were working with AI were more task-oriented and more likely to delegate work. These mechanisms then explained performance as task-oriented communication improved ad quality, specifically when working with AI, while interpersonal communication reduced ad quality; delegation improved text quality but had no effect on image quality and was positively associated with diversity collapse, creating homogeneous outputs of higher average quality. The results suggest AI agents drive changes in productivity, performance, and output diversity by reshaping teamwork.

cs.CY

Advancing AI Negotiations: A Large-Scale Autonomous Negotiation Competition

We conducted an International AI Negotiation Competition in which participants designed and refined prompts for AI negotiation agents. We then facilitated over 180,000 negotiations between these agents across multiple scenarios with diverse characteristics and objectives. Our findings revealed that principles from human negotiation theory remain crucial even in AI-AI contexts. Surprisingly, warmth -- a traditionally human relationship-building trait -- was consistently associated with superior outcomes across all key performance metrics. Dominant agents, meanwhile, were especially effective at claiming value. Our analysis also revealed unique dynamics in AI-AI negotiations not fully explained by existing theory, including AI-specific technical strategies like chain-of-thought reasoning and prompt injection. When we applied natural language processing (NLP) methods to the full transcripts of all negotiations, we found positivity, gratitude, and question-asking (associated with warmth) were strongly associated with reaching deals as well as objective and subjective value, whereas conversation lengths (associated with dominance) were strongly associated with impasses. The results suggest the need to establish a new theory of AI negotiation, which integrates classic negotiation theory with AI-specific negotiation theories to better understand autonomous negotiations and optimize agent performance.

cs.AI

Teaching AI to Handle Exceptions: Supervised Fine-Tuning with Human-Aligned Judgment

Large language models (LLMs), initially developed for generative AI, are now evolving into agentic AI systems, which make decisions in complex, real-world contexts. Unfortunately, while their generative capabilities are well-documented, their decision-making processes remain poorly understood. This is particularly evident when testing targeted decision-making: for instance, how models handle exceptions, a critical and challenging aspect of decision-making made relevant by the inherent incompleteness of contracts. Here we demonstrate that LLMs, even ones that excel at reasoning, deviate significantly from human judgments because they adhere strictly to policies, even when such adherence is impractical, suboptimal, or even counterproductive. We then evaluate three approaches to tuning AI agents to handle exceptions: ethical framework prompting, chain-of-thought reasoning, and supervised fine-tuning. We find that while ethical framework prompting fails and chain-of-thought prompting provides only slight improvements, supervised fine-tuning - specifically with human explanations - yields markedly better results. Surprisingly, in our experiments, supervised fine-tuning even enabled models to generalize human-like decision-making to novel scenarios, demonstrating transfer learning of human-aligned decision-making across contexts. Furthermore, fine-tuning with explanations, not just labels, was critical for alignment, suggesting that aligning LLMs with human judgment requires explicit training on how decisions are made, not just which decisions are made. These findings highlight the need to address LLMs' shortcomings in handling exceptions in order to guide the development of agentic AI toward models that can effectively align with human judgment and simultaneously adapt to novel contexts.

cs.AI

Curiosity as filling, compressing, and reconfiguring knowledge networks

Due to the significant role that curiosity plays in our lives, several theoretical constructs, such as the information gap theory and compression progress theory, have sought to explain how we engage in its practice. According to the former, curiosity is the drive to acquire information that is missing from our understanding of the world. According to the latter, curiosity is the drive to construct an increasingly parsimonious mental model of the world. To complement the densification processes inherent to these theories, we propose the conformational change theory, wherein we posit that curiosity results in mental models with marked conceptual flexibility. We formalize curiosity as the process of building a growing knowledge network to quantitatively investigate information gap theory, compression progress theory, and the conformational change theory of curiosity. In knowledge networks, gaps can be identified as topological cavities, compression progress can be quantified using network compressibility, and flexibility can be measured as the number of conformational degrees of freedom. We leverage data acquired from the online encyclopedia Wikipedia to determine the degree to which each theory explains the growth of knowledge networks built by individuals and by collectives. Our findings lend support to a pluralistic view of curiosity, wherein intrinsically motivated information acquisition fills knowledge gaps and simultaneously leads to increasingly compressible and flexible knowledge networks. Across individuals and collectives, we determine the contexts in which each theoretical account may be explanatory, thereby clarifying their complementary and distinct explanations of curiosity. Our findings offer a novel network theoretical perspective on intrinsically motivated information acquisition that may harmonize with or compel an expansion of the traditional taxonomy of curiosity.

q-bio.NC

The network structure of scientific revolutions

Philosophers of science have long postulated how collective scientific knowledge grows. Empirical validation has been challenging due to limitations in collecting and systematizing large historical records. Here, we capitalize on the largest online encyclopedia to formulate knowledge as growing networks of articles and their hyperlinked inter-relations. We demonstrate that concept networks grow not by expanding from their core but rather by creating and filling knowledge gaps, a process which produces discoveries that are more frequently awarded Nobel prizes than others. Moreover, we operationalize paradigms as network modules to reveal a temporal signature in structural stability across scientific subjects. In a network formulation of scientific discovery, data-driven conditions underlying breakthroughs depend just as much on identifying uncharted gaps as on advancing solutions within scientific communities.

cs.DL

The control of brain network dynamics across diverse scales of space and time

The human brain is composed of distinct regions that are each associated with particular functions and distinct propensities for the control of neural dynamics. However, the relation between these functions and control profiles is poorly understood, as is the variation in this relation across diverse scales of space and time. Here we probe the relation between control and dynamics in brain networks constructed from diffusion tensor imaging data in a large community based sample of young adults. Specifically, we probe the control properties of each brain region and investigate their relationship with dynamics across various spatial scales using the Laplacian eigenspectrum. In addition, through analysis of regional modal controllability and partitioning of modes, we determine whether the associated dynamics are fast or slow, as well as whether they are alternating or monotone. We find that brain regions that facilitate the control of energetically easy transitions are associated with activity on short length scales and slow time scales. Conversely, brain regions that facilitate control of difficult transitions are associated with activity on long length scales and fast time scales. Built on linear dynamical models, our results offer parsimonious explanations for the activity propagation and network control profiles supported by regions of differing neuroanatomical structure.

q-bio.NC

Models of communication and control for brain networks: distinctions, convergence, and future outlook

Recent advances in computational models of signal propagation and routing in the human brain have underscored the critical role of white matter structure. A complementary approach has utilized the framework of network control theory to better understand how white matter constrains the manner in which a region or set of regions can direct or control the activity of other regions. Despite the potential for both of these approaches to enhance our understanding of the role of network structure in brain function, little work has sought to understand the relations between them. Here, we seek to explicitly bridge computational models of communication and principles of network control in a conceptual review of the current literature. By drawing comparisons between communication and control models in terms of the level of abstraction, the dynamical complexity, the dependence on network attributes, and the interplay of multiple spatiotemporal scales, we highlight the convergence of and distinctions between the two frameworks. Based on the understanding of the intertwined nature of communication and control in human brain networks, this work provides an integrative perspective for the field and outlines exciting directions for future work.

q-bio.NC

Network structure of cascading neural systems predicts stimulus propagation and recovery

Many neural systems display cascading behavior characterized by uninterrupted sequences of neuronal firing. This gap precludes an understanding of how variations in network structure manifest in neural dynamics and either support or impinge upon information processing. Here, we develop a theoretical understanding of how network structure supports information processing through network dynamics, and we validate our theory with empirical data. Using a generalized spiking model and mathematical tools from linear systems theory, network control theory, and information theory, we show how network structure can be designed to temporally extend the propagation and recovery of certain stimulus patterns. Moreover, we observe cycles as structural and dynamic motifs that are prevalent in such networks. Broadly, our results demonstrate how cascading neural networks could contribute to cognitive faculties that require lasting activation of neuronal patterns, such as working memory or attention.

q-bio.NC