SearcharxivSearch

arXiv subjects

Salvatore Romano

Publications and source records attributed to Salvatore Romano.

11 recordsLinked to original sources

ProtoGuide: Prototype-Driven Guidance for Class-Conditional Graph Generation

Discrete diffusion models are a prominent family for graph generation, but standard class-conditional mechanisms embed the class signal in the denoiser during training, tying the conditioning mechanism to the trained model. Classifier guidance avoids this coupling in continuous domains by steering a frozen model with a classifier's gradient, but discrete graph diffusion samples discrete edge states, so gradients cannot propagate through the sampled graph. We introduce ProtoGuide, a post-hoc, backbone-agnostic framework that recovers an analogous mechanism. At each reverse step the denoiser's per-edge output is relaxed into a differentiable soft adjacency, embedded by a frozen Siamese graph neural network, and scored against a target-class prototype and its nearest competitor; the resulting per-edge gradient, damped by a cosine schedule, is injected back into the denoiser output. All components stay frozen, so guidance is retargeted by supplying a different prototype. On five classes of real-world networks and two architecturally different backbones, EDGE and DiGress, ProtoGuide raises macro classification accuracy from 50.7% to 73.5% and from 73.6% to 83.8%, and outperforms DiGress's built-in conditional training under our configuration. Gains are largest where the unguided models are weakest, and are not uniform across classes. Per-graph coverage remains high in most settings, while distributional effects are class-dependent. A Best-of-N selection baseline matches this accuracy given enough oversampling, but at a substantial cost in graph diversity. An independently initialized classifier, a directionality test, and a few-shot analysis support target-directed steering and robustness to very small support sets.

cs.LG

The Reweighting Principle in Statistical Mechanics

Reweighting of probability measures provides a unifying perspective on ensemble transformations in statistical mechanics. We distinguish two complementary classes of reweighting: soft constraints, which redistribute probability while preserving the support of the reference measure, and hard constraints, which impose support restrictions through conditioning. We show that exponential tilting and conditioning on an exact observable value arise as the minimum relative entropy updates associated with soft expectation constraints and hard exact-value constraints, respectively. Their relative entropies naturally inherit complementary thermodynamic structures: exponential tilting gives rise to the Legendre structure of the canonical ensemble and reduces, for a uniform reference measure, to Gibbs entropy, whereas conditioning reduces to Boltzmann entropy through the surprisal of the constrained macrostate. By introducing an enlarged probability space in which observables are treated as explicit random variables, we further show that canonical and microcanonical ensembles arise as marginal and conditional distributions of a common joint reference measure. In the thermodynamic limit, large-deviation concentration makes soft and hard constraints macroscopically equivalent, providing a probabilistic interpretation of canonical--microcanonical ensemble equivalence. Finally, we outline how the same information-theoretic framework naturally extends to path space, suggesting a unified probabilistic description of equilibrium statistical mechanics and conditioned stochastic dynamics.

cond-mat.stat-mech

Uncertainty-Aware Estimation of Mis/Disinformation Prevalence on Social Media

Estimation of mis/disinformation prevalence in social media is crucial for designing mitigation strategies to limit its impact. Yet, such estimations are subject to several uncertainties that are rarely quantified jointly. In this study, we present a methodological contribution in which confidence intervals were used to quantify uncertainties related to mis/disinformation prevalence. The analysis draws on a multi-platform, multilingual dataset annotated by professional fact-checkers. Data were collected between March and April 2025 from Facebook, Instagram, LinkedIn, TikTok, X/Twitter, and YouTube across four EU Member States (France, Poland, Slovakia, and Spain). We account for different causes of uncertainty: (i) sample uncertainty, (ii) annotation uncertainty arising from human disagreement and misclassification, and (iii) data retrieval uncertainty induced by keyword-based data collection. First, we estimate the uncertainty arising from the different causes separately using confidence intervals, simulation-based methods, and bootstrapping. Finally, we combined multinomial simulations of annotator behaviour with keyword and post-resampling to capture the joint impact of measurement uncertainty on mis/disinformation prevalence estimates. The proposed methodological approach highlights the importance of uncertainty-aware estimation of mis/disinformation prevalence for robust analysis. The empirical results of this study show that keyword-based data retrieval can exceed baseline variability, leading to wider confidence intervals around prevalence estimates.

cs.SI

Beyond MMD: Evaluating Graph Generative Models with Geometric Deep Learning

Graph generation is a crucial task in many fields, including network science and bioinformatics, as it enables the creation of synthetic graphs that mimic the properties of real-world networks for various applications. Graph Generative Models (GGMs) have emerged as a promising solution to this problem, leveraging deep learning techniques to learn the underlying distribution of real-world graphs and generate new samples that closely resemble them. Examples include approaches based on Variational Auto-Encoders, Recurrent Neural Networks, and more recently, diffusion-based models. However, the main limitation often lies in the evaluation process, which typically relies on Maximum Mean Discrepancy (MMD) as a metric to assess the distribution of graph properties in the generated ensemble. This paper introduces a novel methodology for evaluating GGMs that overcomes the limitations of MMD, which we call RGM (Representation-aware Graph-generation Model evaluation). As a practical demonstration of our methodology, we present a comprehensive evaluation of two state-of-the-art Graph Generative Models: Graph Recurrent Attention Networks (GRAN) and Efficient and Degree-guided graph GEnerative model (EDGE). We investigate their performance in generating realistic graphs and compare them using a Geometric Deep Learning model trained on a custom dataset of synthetic and real-world graphs, specifically designed for graph classification tasks. Our findings reveal that while both models can generate graphs with certain topological properties, they exhibit significant limitations in preserving the structural characteristics that distinguish different graph domains. We also highlight the inadequacy of Maximum Mean Discrepancy as an evaluation metric for GGMs and suggest alternative approaches for future research.

cs.LG

DSA, AIA, and LLMs: Approaches to conceptualizing and auditing moderation in LLM-based chatbots across languages and interfaces in the electoral contexts

The integration of Large Language Models (LLMs) into chatbot-like search engines poses new challenges for governing, assessing, and scrutinizing the content output by these online entities, especially in light of the Digital Service Act (DSA). In what follows, we first survey the regulation landscape in which we can situate LLM-based chatbots and the notion of moderation. Second, we outline the methodological approaches to our study: a mixed-methods audit across chatbots, languages, and elections. We investigated Copilot, ChatGPT, and Gemini across ten languages in the context of the 2024 European Parliamentary Election and the 2024 US Presidential Election. Despite the uncertainty in regulatory frameworks, we propose a set of solutions on how to situate, study, and evaluate chatbot moderation.

cs.CY

AI-Generated Algorithmic Virality

There is a growing discussion about social media feeds being increasingly filled with AI-generated content. Due to its visual plausibility, low cost, and fast production speed, AI-generated content is said to be highly effective in "gaming the algorithm" and going viral. Popularly referred to as "AI slop," this phenomenon arguably leads to the presence of sloppy and potentially deceptive content at a scale unseen before. This investigation offers a systematic analysis of AI-generated content and its labelling in TikTok's and Instagram's search results across 13 hashtags (see Appendix) in three European countries (Spain, Germany, and Poland) over the course of June 2025. We manually annotated and analyzed the 30 top search results on political (#trump, #zelensky, #pope) and broader topics (e.g.,#health, #history) to understand the relation between synthetic (content that is partially or entirely made using generative AI) and non-synthetic content across languages and countries. We then explored the emerging phenomenon of accounts producing generative AI content at scale by analyzing 153 accounts and proposing a new categorization schema of what we termed Agentic AI Accounts. Our main findings are:

cs.CY

TikTok's Research API: Problems Without Explanations

Following the Digital Services Act of 2023, which requires Very Large Online Platforms (VLOPs) and Very Large Online Search Engines (VLOSEs) to facilitate data accessibility for independent research, TikTok augmented its Research API access within Europe in July 2023. This action was intended to ensure compliance with the DSA, bolster transparency, and address systemic risks. Nonetheless, research findings reveal that despite this expansion, notable limitations and inconsistencies persist within the data provided. Our experiment reveals that the API fails to provide metadata for one in eight videos provided through data donations, including official TikTok videos, advertisements, and content from specific accounts, without an apparent reason. The API data is incomplete, making it unreliable when working with data donations, a prominent methodology for algorithm audits and research on platform accountability. To monitor the functionality of the API and eventual fixes implemented by TikTok, we publish a dashboard with a daily check of the availability of 10 videos that were not retrievable in the last month. The video list includes very well-known accounts, notably that of Taylor Swift. The current API lacks the necessary capabilities for thorough independent research and scrutiny. It is crucial to support and safeguard researchers who utilize data scraping to independently validate the platform's data quality.

cs.CY

Structure of the water/magnetite interface from sum frequency generation experiments and neural network based molecular dynamics simulations

Magnetite, a naturally abundant mineral, frequently interacts with water in both natural settings and various technical applications, making the study of its surface chemistry highly relevant. In this work, we investigate the hydrogen bonding dynamics and the presence of hydroxyl species at the magnetite-water interface using a combination of neural network potential-based molecular dynamics simulations and sum frequency generation vibrational spectroscopy. Our simulations, which involved large water systems, allowed us to identify distinct interfacial species, such as dissociated hydrogen and hydroxide ions formed by water dissociation. Notably, water molecules near the interface exhibited a preference for dipole orientation towards the surface, with bulk-like water behavior only re-emerging beyond 60 Å from the surface. The vibrational spectroscopy results aligned well with the simulations, confirming the presence of a hydrogen bond network in the surface ad-layers. The analysis revealed that surface-adsorbed hydroxyl groups orient their hydrogen atoms towards the water bulk. In contrast, hydrogen-bonded water molecules align with their hydrogen atoms pointing towards the magnetite surface.

cond-mat.mtrl-sci

Structure and dynamics of the magnetite(001)/water interface from molecular dynamics simulations based on a neural network potential

The magnetite/water interface is commonly found in nature and plays a crucial role in various technological applications. However, our understanding of its structural and dynamical properties at the molecular scale remains still limited. In this study, we develop an efficient Behler-Parrinello neural network potential (NNP) for the magnetite/water system, paying particular attention to the accurate generation of reference data with density functional theory. Using this NNP, we performed extensive molecular dynamics simulations of the magnetite (001) surface across a wide range of water coverages, from the single molecule to bulk water. Our simulations revealed several new ground states of low coverage water on the Subsurface Cation Vacancy (SCV) model and yielded a density profile of water at the surface that exhibits marked layering. By calculating mean square displacements, we obtained quantitative information on the diffusion of water molecules on the SCV for different coverages, revealing significant anisotropy. Additionally, our simulations provided qualitative insights into the dissociation mechanisms of water molecules at the surface.

physics.comp-ph

Conditioning Normalizing Flows for Rare Event Sampling

Understanding the dynamics of complex molecular processes is often linked to the study of infrequent transitions between long-lived stable states. The standard approach to the sampling of such rare events is to generate an ensemble of transition paths using a random walk in trajectory space. This, however, comes with the drawback of strong correlations between subsequently sampled paths and with an intrinsic difficulty in parallelizing the sampling process. We propose a transition path sampling scheme based on neural-network generated configurations. These are obtained employing normalizing flows, a neural network class able to generate statistically independent samples from a given distribution. With this approach, not only are correlations between visited paths removed, but the sampling process becomes easily parallelizable. Moreover, by conditioning the normalizing flow, the sampling of configurations can be steered towards regions of interest. We show that this approach enables the resolution of both the thermodynamics and kinetics of the transition region.

physics.comp-ph

The kinetics of the ice-water interface from ab initio machine learning simulations

Molecular simulations employing empiric force fields have provided valuable knowledge about the ice growth process in the last decade. The development of novel computational techniques allows us to study this process, which requires long simulations of relatively large systems, with ab initio accuracy. In this work, we use a neural-network potential for water trained on the Revised Perdew-Burke-Ernzerhof functional to describe the kinetics of the ice-water interface. We study both ice melting and growth processes. Our results for the ice growth rate are in reasonable agreement with previous experiments and simulations. We find that the kinetics of ice melting presents a different behavior (monotonic) than that of ice growth (non-monotonic). In particular, a maximum in the ice growth rate of 6.5 Å/ns is found at 14 K of supercooling. The effect of the surface structure is explored by investigating the basal and primary and secondary prismatic facets. We use the Wilson-Frenkel relation to explain these results in terms of the mobility of molecules and the thermodynamic driving force. Moreover, we study the effect of pressure by complementing the standard isobar with simulations at negative pressure (-1000 bar) and at high pressure (2000 bar). We find that prismatic facets grow faster than the basal one, and that pressure does not play an important role when the speed of the interface is considered as a function of the difference between the melting temperature and the actual one, i.e. to the degree of either supercooling or overheating.

cond-mat.soft