Searcharxiv⌕ Search

subject

cs.SI

cs.SI: explore 109 source-linked works published from 2013 to 2026, with original documents and citations.

This collection is a preview while coverage and quality are evaluated.

Search within this collection

Coverage and selection

Includes records with this source-supplied label or an explicit phrase match in their metadata. Matches indicate a mention, not proof that a paper uses a method or tests a material. Source versions are consolidated by DOI.

Sources: arxiv. Collection updated 2026-09-16. Counts describe this index, not the complete source archives.

Past-aware game-theoretic centrality: a framework for cardinality-constrained set-function maximization on networks

We consider cardinality-constrained optimization of set functions over the nodes of a graph. The standard greedy algorithm selects each node according to its immediate marginal contribution, a local criterion that may fail to anticipate the synergies within the final set. We introduce past-aware game-theoretic centrality (PAGTC), which evaluates a candidate node through its expected marginal contribution over the possible completions of the current partial solution to the prescribed target size. This yields a sequential selection strategy that explicitly accounts for the final budget. For nonnegative monotone submodular objectives and a budget $r$, we prove an approximation guarantee of $r/(2r-1)$ and derive computable a posteriori bounds. Since direct PAGTC evaluation involves averaging over a large number of coalitions, we extend an exact computation framework for game-theoretic centrality and derive efficiently computable expressions for two classes of graph-optimization problems, namely facility location and influence in complex contagion, covering both submodular and non-submodular cases. The numerical results show that the benefits depend on the objective and are most pronounced for complex contagion, where submodularity does not hold.

cs.SI↗

Jointly Optimizing Debiased CTR and Uplift for Coupons Marketing: A Unified Causal Framework

In online advertising, marketing interventions such as coupons introduce significant confounding bias into Click-Through Rate (CTR) prediction. Observed clicks reflect a mixture of users' intrinsic preferences and the uplift induced by these interventions. This causes conventional models to miscalibrate base CTRs, which distorts downstream ranking and billing decisions. Furthermore, marketing interventions often operate as multi-valued treatments with varying magnitudes, introducing additional complexity to CTR prediction. To address these issues, we propose the Unified Multi-Valued Treatment Network (UniMVT). Specifically, UniMVT disentangles confounding factors from treatment-sensitive representations, enabling a full-space counterfactual inference module to jointly reconstruct the debiased base CTR and intensity-response curves. To handle the complexity of multi-valued treatments, UniMVT employs an auxiliary intensity estimation task to capture treatment propensities and devise a unit uplift objective that normalizes the intervention effect. This ensures comparable estimation across the continuous coupon-value spectrum. UniMVT simultaneously achieves debiased CTR prediction for accurate system calibration and precise uplift estimation for incentive allocation. Extensive experiments on synthetic and industrial datasets demonstrate UniMVT's superiority in both predictive accuracy and calibration. Furthermore, real-world A/B tests confirm that UniMVT significantly improves business metrics through more effective coupon distribution.

cs.SI↗

Mapping the Reddit Bot Ecosystem: Taxonomy and Evolution

Automated agents increasingly participate in online communities, yet their population structure and roles remain poorly understood. Using a dataset of 3,389 identified bots and their full activity histories, we construct a taxonomy of bot "species" on the news aggregation and social media platform Reddit based on temporal, community, linguistic, and semantic features. Clustering analysis reveals 18 distinct bot types spanning content-specialized, behavior-driven, and infrastructural roles such as moderation and utility support. In addition, temporal analysis shows that bot numbers and activity expanded rapidly before peaking around the COVID-19 period, then started declining even before Reddit's 2023 API policy changes. However, the overall diversity of bot species has remained remarkably stable. These findings suggest that online bot populations form evolving digital ecosystems.

cs.SI↗

Topology-induced Operators Reveal Complementary Graph Representations without Training

Graph representation learning has largely focused on designing increasingly sophisticated models to transform graph topology into vector representations, or embeddings. However, the extent to which embedding quality depends on model learning, rather than on the underlying topological transformations, remains unclear. Here, we show that informative embeddings can be derived without complicated model design and gradient-based training. Propagating random features through implicit hierarchical structures induced by random walks and anonymous walks yields embeddings that capture node proximity and structural role, respectively. These two training-free embeddings preserve complementary aspects of graph organization and perform competitively with classic and recent methods across various node-, edge-, and graph-level tasks. They often require substantially less computation, resulting in a favorable quality-efficiency trade-off. Combining the two types of embeddings further improves inference quality of some tasks compared with using either embedding type alone. Our results suggest that informative graph embeddings can arise from carefully chosen topological transformations before any learning operation is applied.

cs.LG↗

The conservative turn in science: The changing character of knowledge recombination

This study examines how the dominant mode of cross-disciplinary knowledge combination has changed over time. It addresses a limitation of existing studies that document declining disruptiveness at the level of individual works but do not reveal whether the mechanisms of cross-disciplinary knowledge flow have themselves shifted. Drawing on 56 million publications and 816 million citations from OpenAlex (1960-2025), the study classifies cross-disciplinary citations into five types based on the set relationship between the subfield portfolios of citing and cited publications and introduces a semantic-distance measure derived from SPECTER-v2 embeddings of publication abstracts. It subsequently demonstrates that cross-disciplinary citations have shifted markedly away from configurations with no shared disciplinary ground toward types built on partially overlapping knowledge bases, a reorientation further reflected in a steady decline in average semantic distance between citing and cited fields. This shift co-occurs with a diffusion of brokerage positions across the network and a greater reliance on older literature, consistent with a system that has become broader and more decentralized in its interdisciplinary reach. Finally, this study discusses the implications of these findings for science policy and for the design of cross-disciplinary funding mechanisms.

cs.SI↗

Endorsement Without New Evidence: How Sequential Voting Inflates Mandates in Online Community Governance

Online communities often treat large support margins in public elections as strong mandates. We argue that such margins can overstate the independent scrutiny behind a decision. Using 198,275 free-text rationales from Wikipedia admin elections, we introduce vote-text divergence, a measure that flags a decisive vote paired with a thin, deferential rationale. Divergence rises as voters arrive later, even after controlling for voter and election fixed effects. The pattern is consistent with information saturation: once prior text is accounted for, arrival order no longer predicts divergence, while accumulated prior evidence does. The effect is strongest among peripheral voters in the co-voting network. Yet divergence does not predict worse post-promotion outcomes, such as administrative activity or survival. Public tallies can therefore weaken the scrutiny signal even while selecting capable administrators: a margin may appear to reflect more consensus and support than it actually contains.

cs.SI↗

Ephemeral Feeds and Enduring Rituals: RushTok and the Formation of Event-Based Algorithmic Communities

Each August, TikTok's For You page turns the University of Alabama's sorority recruitment into RushTok. We examine RushTok as an event-based algorithmic community: a collective assembled around a bounded offline ritual and sustained by recommendation. Using a mixed-methods survey (n=71) and a reflexive account of creator outreach, we ask who participates, how, and with what stakes. Findings show an ambiguous and entertainment based throughline; many called it a community (51/71) but few claimed membership (11/71). Affiliation centered on creators rather than shared practices, with parasocial attention clustering around a small set of potential new members (PNMs) and returning figures. Higher content exposure tracked with self-identification as a community member; those members commented, followed creators, and engaged across videos. Attempts to interview creators were met with silence or refusals, reflecting community boundary-work despite viral visibility. We outline implications for platform governance, including time-bounded context, graduated visibility, and aftercare.

cs.HC↗

Democracy Needs Reach: Political Equality, Online Speech, and Algorithmic Recommendation

Within democracies, the capacity to influence political outcomes through speech depends not only on the right to express oneself, but also on the opportunity to reach relevant audiences. In this paper, I argue that the unequal distribution of algorithmic reach on social media platforms undermines equality of opportunity for political influence (EOPI), which is a central democratic ideal. Drawing on Niko Kolodny's work, I contend that current recommendation algorithms create and perpetuate informal inequalities by concentrating attention among a small minority of already-amplified speakers while systematically marginalizing others. To address this problem, I propose recommendation floors as a mechanism for equalizing political speech. To help users achieve meaningful participation, each verified account would receive guaranteed minimum recommendation for up to a limited number of political posts per week. Although this measure represents one component of the structural reforms needed to move the digital public sphere closer to democratic ideals, it offers a feasible pathway to reducing informal inequalities in political influence online.

cs.CY↗

Emergence of polarization in networks of large language model agents

Rapid advances in large language models (LLMs) have not only empowered autonomous agents to generate social networks, communicate, and form shared and diverging opinions on political issues, but have also begun to play a growing role in shaping human political deliberation. Our understanding of their collective behaviours and underlying mechanisms remains incomplete, however, posing unexpected risks to human society. In this paper, we simulate networked systems involving thousands of LLM agents across different backbone models (GPT-3.5, GPT-4o, ChatGLM, Llama-3, and DeepSeek-V3), in which agents interact through LLM-guided conversations and update their opinions over time, resulting in the emergence of opinion polarization. We discover that these agents spontaneously develop their own social network with properties characteristic of human social networks, including homophilic clustering. The collective opinions of these LLM agents evolve in ways that exhibit behavioural patterns consistent with social phenomena and mechanisms widely discussed in studies of human behaviour. Overall, these behaviours, patterns, and emergent phenomena produced by LLM agents are consistent with real-world observations and established opinion-dynamics models. This consistency suggests that LLM agents can serve as a valuable synthetic testbed for exploring hypothetical intervention strategies in networked LLM-agent systems.

cs.SI↗

Rewarding Engagement and Personalization in Popularity-Based Rankings Amplifies Extremism and Polarization

Despite extensive research, the mechanisms through which online platforms shape extremism and polarization remain poorly understood. We identify and test a mechanism, grounded in empirical evidence, that explains how ranking algorithms can amplify both phenomena. This mechanism is based on well-documented assumptions: (i) users exhibit position bias and tend to prefer items displayed higher in the ranking, (ii) users prefer like-minded content, (iii) users with more extreme views are more likely to engage actively, and (iv) ranking algorithms are popularity-based, assigning higher positions to items that attract more clicks. Under these conditions, when platforms additionally reward active engagement and implement personalized rankings, users are inevitably driven toward more extremist and polarized news consumption. We formalize this mechanism in a dynamical model, which we evaluate by means of simulations and interactive experiments with hundreds of human participants, where the rankings are updated dynamically in response to user activity.

cs.SI↗

Co-evolution of the global research collaboration network and the performance of nations in science and technology

Researchers have long suspected that international research collaboration (IRC) and scientific and technological (S&T) performance are subject to reciprocal causality, yet the endogenous co-evolution of these twin phenomena has yet to be tested by large-scale empirical analysis. This study tests these effects simultaneously using a longitudinal co-evolution model on three decades of global network and national performance data. Stochastic actor-oriented models (SAOM) are used to analyze data on 172 countries from 1993 to 2022. Yearly IRC networks are constructed from Web of Science's XML database, and performance data are gathered from Elsevier's fractional field-weighted citation impact (FWCI). The models also account for geographic, economic, demographic, and political factors, as well as endogenous network processes. The results provide support for co-evolution. Distance and shared language moderate this relationship in contrasting ways. The selection effect of performance on tie formation is amplified across distance and dampened by a shared language, whereas the influence effect of performance on centrality is attenuated by remoteness and strengthened by linguistic reach. This pattern suggests that information asymmetry shapes partner selection, while communication and coordination shape the returns to collaboration, pointing to a signaling role for citation-based performance metrics in collaborator selection.

cs.SI↗

Less Is Moral: A CHARMing Framework for Moral Foundations Detection in Endorsement Behaviour

Moral language plays a central role in shaping online endorsement and the diffusion of information, yet existing moral foundation detection systems often suffer from poor cross-domain generalization, weak rationale grounding, and reliance on costly prompting-based large language models (LLMs). We introduce CHARM, a MAC- and Hate-speech-Aware Rationalealigned Moral foundation detection framework built on a lightweight fine-tuned LLM, which integrates complementary moral grounding, rationale alignment, and polarity-aware hate speech signals to support more robust and faithful moral prediction. Unlike prior dictionary-, fine-tune-, or prompt-based detectors, which decouple computation from psychological theory, CHARM is built so that each component -- MAC cross-attention, rationale alignment, and hate-speech modulation -- operationalizes a distinct psychological construct. Using a 30\% subsample of the MFTC, MFRC, and News training pools together with the richer supervision in MFTCXplain, CHARM improves AUC by up to 15.3\% in-domain, surpasses the supervised baselines on every out-of-domain dataset in both AUC and F1, and offers a scalable, low-cost alternative to prompting-based LLM detectors. We further apply CHARM to large-scale COVID-19 discourse on Twitter and show that moral value alignment is strongly associated with online endorsement behavior. By making moral framing measurable at scale, CHARM offers a practical tool for studying the spread of morally charged misinformation. Code and additional materials: https://github.com/HuixiangF/CHARM/.

cs.CL↗

How Well Do LLMs Simulate Survey Responses Following a Breast Cancer Screening Intervention?

Collecting survey data is laborious and limited by privacy constraints. Large language models (LLMs) have shown promise as predictive social simulations. It is unclear whether they can replicate population-level response distributions before and after a healthcare intervention. Using information derived from 4125 women aged 35-59 years, we evaluate whether agents informed solely by pre-intervention profile information can reproduce post-intervention response distributions. Groups of LLM agents (n=50) were created with Gemma 4 E4B and Qwen3.5 9B; conditions ranged from zero-shot prompting to agent profiles enriched with aggregate or individual-level demographic characteristics and pre-intervention questionnaire responses. We compared predicted and observed response distributions with Total Variation Distance (TVD) and Normalized Wasserstein Distance (NWD). Across both LLMs, profile-based agents improved distributional accuracy relative to zero-shot and random baselines. Nevertheless, direct sampling of 50 real participants remained more accurate. Prediction errors were also higher among participants aged 55-59 years and those living in private property. Errors also varied by question theme and LLM model, with the highest errors observed for cancer fatalism and post intervention attitudes toward genetics. Sensitivity analyses showed that performance was influenced by prompt template changes and temperature hyperparameter. Our results show the potential of LLM-based agents to model behavioral responses to interventions in silico. However, profiles containing additional information beyond demographics did not consistently outperform simpler ones. Certain cultural constructs and population groups also remain inadequately represented by the LLM models evaluated. Future work may include building behaviorally grounded and locally validated virtual populations.

cs.SI↗

Ex Ante Estimation of Payable Relief and Compensation Timing for Supplier Selection

Supplier selection affects not only operating performance but also the payable network entered by a new obligation. This paper develops the Compensability Capacity Assessment (CCA), an ex ante buyer-supplier measure of expected gross payable relief and its likely timing. CPM provides the bounded structural kernel; concave CCA variants add bilateral invoice capacity. The measure is tested on 749,952 analytical invoices issued during 2012-2023, using frozen nine-month histories and future weekly, monthly, quarterly, semester, and annual windows. Cycle-restricted and path-enabled clearing are independent outcome-generating environments used to validate the measure, not technologies compared by this study. Across seven fully observed quarters in 2022-2023, log-CCA has a median Spearman correlation of 0.638 with future integrated relief; persistent relations carry 93.5% of relief, and the highest-scoring relation captures 92.5% of buyer-specific best relief. Predictability remains positive from week to year, with quarterly recalibration providing the best operating balance between signal, coverage, and timeliness. A timing analysis shows that, across the two validation environments, 89-91% of attributed relief occurs within seven days of invoice issue and 94-96% within thirty days, on average about sixteen days before contractual maturity. Log-CCA correlates 0.564 with thirty-day relief and 0.566 with relief-days. Its highest quartile has a 96.5% median probability of thirty-day compensation, compared with 57.6% in the lowest quartile. Conditional waiting-time prediction is weaker, so CCA should rank timely compensation opportunity rather than forecast an exact payment date. The findings link supplier placement, financial circularity, and working-capital exposure while motivating deployment, causal testing, and quarterly drift monitoring.

cs.SI↗

How Bitcoin Forms Its Network: Peer-Table Sampling and Structural Properties

The global structure of P2P networks underlying modern cryptocurrencies is hidden by design: each node only knows its neighbors and maintains a local \textit{peer table} of IP addresses. The peer table of a node can be seen as the node's ``local view'' of the set of nodes currently in the network: it is constantly updated with information collected from the neighbors and it is used by the node to establish new connections when needed. The maintenance rules of the nodes' peer tables determine the global structure of the network and its evolution. Even though the global structure is unknown to any node of the network as well as to any external observer, the rules for the exchange of information between neighbors are compatible with the design and use of network \textit{crawlers} that can query the nodes and extract some information about their peer tables. In this paper we first analyze the data we crawled from nodes of the Bitcoin and Dogecoin P2P networks, to collect information about the distribution of IP addresses in the peer tables and to estimate the evolution of network size and churn rate; we then use the estimates to set up the parameters for a simulation of the Bitcoin network and we analyze the evolution of the network structure that we get from the simulation. Overall, our results show that Bitcoin's peer-table maintenance rules induce non-uniform, heavy-tailed visibility patterns while nevertheless generating a well-connected and structurally robust network.

cs.NI↗

Latent geometry organizes higher-order interactions

Higher-order structures offer a natural representation of complex systems that involve interactions between groups of different sizes. A widespread feature of their higher-order structure is nestedness, whereby interactions involving smaller groups are contained within larger ones. Yet, why interactions of different orders organise into nested structures remains largely unexplained. Here, we introduce an analytically tractable geometric model of higher-order networks in which a single latent geometric space couples interactions across orders, leading to the spontaneous emergence of nestedness. We show analytically and numerically that nestedness undergoes a transition between a nested geometric regime, where it remains finite in the thermodynamic limit, and a regime where it vanishes with system size. In this regime, we uncover a weakly geometric range characterized by an anomalously slow finite-size decay, allowing substantial nestedness to persist in finite systems even when its asymptotic value vanishes. Finally, with a single geometric coupling parameter, the model reproduces the nestedness profiles observed in real-world hypergraphs across different domains. Our results reveal latent geometry as a simple organising principle underlying the nested organisation of higher-order interactions.

physics.soc-ph↗

Designing for Healthy, Affordable, and Sustainable Human-HVAC Interactions for Heating in Smart Homes

As geopolitical tensions, energy crises, and energy-intensive AI infrastructure intensify concerns about demand, affordability, and resilience, communities increasingly encounter these challenges through everyday energy practices, particularly winter heating. Against this background, the doctoral exposé, "Designing Human-HVAC Interaction for Healthy, Affordable, and Sustainable Heating in Smart Homes", is structured around four chapters. First, a multidisciplinary literature review defines and positions Human-HVAC Interaction, focusing on heating in smart homes. Second, longitudinal living lab studies with design probes examine everyday heating practices, thermal comfort, and indoor environmental quality, with attention to thermally vulnerable groups such as older adults, pregnant or menopausal women, parents with infants, and people affected by allergies or airborne pollutants. Third, a VR-based smart home demonstrator explores how heating and IEQ scenarios can be prototyped and evaluated as a virtual living lab, while critically examining the limits of representing bodily indoor climate conditions through VR. Fourth, follow-up design studies examine how VR-based insights can be translated into physical-digital prototypes that combine digital fabrication, distributed environmental sensing, and diverse interface forms for critical heating and IEQ contexts. The thesis aims to contribute a design-oriented understanding of Human-HVAC Interaction by building from a multidisciplinary literature review to empirical living lab and co-design studies, VR-based prototyping, and physical system development, examining how smart home users make sense of, negotiate, and respond to smart HVAC system.

cs.HC↗

"Shut Up and Let Me Enjoy My Otome": Understanding and Measuring the Toxicity in Otome Game Communities

Otome games, a romance simulation genre primarily targeting female, have emerged as a major force in the global gaming market, attracting hundreds of millions of players and billions in revenue. Despite their popularity, otome game communities face pervasive online toxicity, which has been largely unexplored. In this work, we present the first large-scale measurement of toxicity in otome game communities across social platforms. We introduce OtomeSCAN, a framework for collecting, evaluating, and analyzing 620,045 posts from Weibo and Reddit spanning 18 months. To support robust analysis, we manually annotated a ground-truth dataset of 4,308 posts, identifying eight target groups such as players and game developers. We evaluate seven toxicity detectors on the dataset, including general-purpose models and our proposed LLM-based detectors, with our best model achieving F1-scores of 0.82 (Weibo) and 0.78 (Reddit). Our analysis reveals significant platform-based differences in toxicity: 22.20% of otome-related posts on Weibo are toxic, compared to 3.71% on Reddit. Besides, real-world events like in-community conflicts can rapidly escalate toxicity, with toxicity ratios increasing to 37.09% in just 72 hours during an external attack on Weibo. We also flag 191 potential-coordination clusters in otome game communities, 64.40% of which target game developers, with several accounts participating repeatedly across multiple clusters. We hope our work inspires further research on community-specific toxicity and contributes to building healthier online spaces for marginalized gaming communities.

cs.CR↗
Compare source metadata on this page
WorkPublishedSource identifierSource
Past-aware game-theoretic centrality: a framework for cardinality-constrained set-function maximization on networks2026-09-082511.07157arxiv
Jointly Optimizing Debiased CTR and Uplift for Coupons Marketing: A Unified Causal Framework2026-09-082602.12972arxiv
Mapping the Reddit Bot Ecosystem: Taxonomy and Evolution2026-09-082607.23941arxiv
Topology-induced Operators Reveal Complementary Graph Representations without Training2026-09-082609.08152arxiv
The conservative turn in science: The changing character of knowledge recombination2026-09-082609.08468arxiv
Endorsement Without New Evidence: How Sequential Voting Inflates Mandates in Online Community Governance2026-09-082609.09321arxiv
Ephemeral Feeds and Enduring Rituals: RushTok and the Formation of Event-Based Algorithmic Communities2026-09-082609.09331arxiv
Democracy Needs Reach: Political Equality, Online Speech, and Algorithmic Recommendation2026-09-082609.09465arxiv
Emergence of polarization in networks of large language model agents2026-09-072501.05171arxiv
Rewarding Engagement and Personalization in Popularity-Based Rankings Amplifies Extremism and Polarization2026-09-072510.24354arxiv
Co-evolution of the global research collaboration network and the performance of nations in science and technology2026-09-072606.18549arxiv
Less Is Moral: A CHARMing Framework for Moral Foundations Detection in Endorsement Behaviour2026-09-072609.03330arxiv
How Well Do LLMs Simulate Survey Responses Following a Breast Cancer Screening Intervention?2026-09-072609.07141arxiv
Ex Ante Estimation of Payable Relief and Compensation Timing for Supplier Selection2026-09-072609.07293arxiv
How Bitcoin Forms Its Network: Peer-Table Sampling and Structural Properties2026-09-072609.07609arxiv
Latent geometry organizes higher-order interactions2026-09-072609.07906arxiv
Designing for Healthy, Affordable, and Sustainable Human-HVAC Interactions for Heating in Smart Homes2026-09-072609.07936arxiv
"Shut Up and Let Me Enjoy My Otome": Understanding and Measuring the Toxicity in Otome Game Communities2026-09-072609.08009arxiv

These are bibliographic comparisons, not experimental rankings. Follow the original document for methods and conditions.