SearcharxivSearch

arXiv subjects

Ian Roberts

Publications and source records attributed to Ian Roberts.

17 recordsLinked to original sources

The star formation history of NGC 2276: Comparison between SED modeling and hydrodynamical simulations

Analyzing environmental effects in galaxy groups is pivotal for understanding how galaxies evolve in these moderate-density settings. The degree to which ram pressure versus tidal interactions drive the unusual morphologies and kinematics seen in group galaxies remains a subject of active debate. This study focuses on the nearby galaxy NGC 2276, which has the longest radio continuum tail in galactic groups. It is a member of the NGC 2300 group that possibly experiences both processes, and we aim to determine which has the greatest impact on its overall structure. We combined broadband images with synthetic narrow band filters around emission line maps from integral field spectrograph to construct spatially resolved spectral energy distributions (SED), which we modeled using the BAGPIPES software package to reconstruct kpc-scale star formation histories. These are compared with those derived from adaptive mesh refinement wind tunnel simulations of an NGC 2276-like system. We show that the spatial distribution of the oldest stellar populations ($>1.1$ Gyr) is highly symmetric compared to the youngest ones. This is the most compelling evidence so far pointing towards ram pressure being the only morphological disturber of this system, as it primarily affects the gaseous component. Simulated data show similar results, with older stellar populations being more symmetric. Our RPS-only model showed no significant differences in the morphology of the stellar populations compared to the RPS+tidal model with initial separation of 50 kpc.

astro-ph.GA

A Browser-based Open Source Assistant for Multimodal Content Verification

Disinformation and false content produced by generative AI pose a significant challenge for journalists and fact-checkers who must rapidly verify digital media information. While there is an abundance of NLP models for detecting credibility signals such as persuasion techniques, subjectivity, or machine-generated text, such methods often remain inaccessible to non-expert users and are not integrated into their daily workflows as a unified framework. This paper demonstrates the VERIFICATION ASSISTANT, a browser-based tool designed to bridge this gap. The VERIFICATION ASSISTANT, a core component of the widely adopted VERIFICATION PLUGIN (140,000+ users), allows users to submit URLs or media files to a unified interface. It automatically extracts content and routes it to a suite of backend NLP classifiers, delivering actionable credibility signals, estimating AI-generated content, and providing other verification guidance in a clear, easy-to-digest format. This paper showcases the tool architecture, its integration of multiple NLP services, and its real-world application to detecting disinformation.

cs.CL

Multiphase Astrophysics to Unveil the Virgo Environment (MAUVE)

The Multiphase Astrophysics to Unveil the Virgo Environment (MAUVE) project is a multi-facility programme exploring how dense environments transform galaxies. Combining a VLT/MUSE P110 Large Programme and ALMA observations of 40 late-type Virgo Cluster galaxies, MAUVE resolves star formation, kinematics, and chemical enrichment within their molecular gas discs. A key goal is to track the evolution of cold gas that survives in the inner regions of satellites after entering the cluster, and how it evolves across different infall stages. With its high spatial resolution -- probing down to the physical scales of giant molecular cloud complexes -- and multiphase synergy, MAUVE aims to offer a time-resolved view of environmental quenching and set a new benchmark for cluster galaxy studies.

astro-ph.GA

Tailored untruths: How personalisation challenges LLM safeguards

Large Language Models (LLMs) can generate highly persuasive disinformation, yet little is known about how effectively they personalise it across languages and demographic groups. We present the first large-scale multilingual study of persona-targeted disinformation generation by LLMs. Using a red-teaming methodology, we prompted eight leading models with 324 false narratives and 150 demographic personas in four languages (English, Russian, Portuguese, and Hindi), creating AI-TRAITS, a dataset of 1.6 million personalised disinformation texts. We treat safeguards as compromised whenever a model generates the requested falsehood, whether directly or accompanied by a safety disclaimer. Across models, safeguards failed for 80% of non-personalised prompts and 77.7% of personalised ones, with Grok producing disinformation in over 94% of cases. All models effectively tailored outputs to target personas, employing substantially more persuasive techniques than in non-personalised content. Additional analyses reveal persona-specific linguistic and psychological patterns and show that safeguard effectiveness varies markedly across languages. Together, these findings expose significant weaknesses in current LLM safety mechanisms and highlight the need for more robust, multilingual safeguards against personalised AI-generated disinformation.

cs.CL

Efficient Annotator Reliability Assessment with EffiARA

Data annotation is an essential component of the machine learning pipeline; it is also a costly and time-consuming process. With the introduction of transformer-based models, annotation at the document level is increasingly popular; however, there is no standard framework for structuring such tasks. The EffiARA annotation framework is, to our knowledge, the first project to support the whole annotation pipeline, from understanding the resources required for an annotation task to compiling the annotated dataset and gaining insights into the reliability of individual annotators as well as the dataset as a whole. The framework's efficacy is supported by two previous studies: one improving classification performance through annotator-reliability-based soft-label aggregation and sample weighting, and the other increasing the overall agreement among annotators through removing identifying and replacing an unreliable annotator. This work introduces the EffiARA Python package and its accompanying webtool, which provides an accessible graphical user interface for the system. We open-source the EffiARA Python package at https://github.com/MiniEggz/EffiARA and the webtool is publicly accessible at https://effiara.gate.ac.uk.

cs.CL

Ancestral process for infectious disease outbreaks with superspreading

When an infectious disease outbreak is of a relatively small size, describing the ancestry of a sample of infected individuals is difficult because most ancestral models assume large population sizes. Given a set of infected individuals, we show that it is possible to express exactly the probability that they have the same infector, either inclusively (so that other individuals may have the same infector too) or exclusively (so that they may not). To compute these probabilities requires knowledge of the offspring distribution, which determines how many infections each infected individual causes. We consider transmission both without and with superspreading, in the form of a Poisson and a Negative-Binomial offspring distribution, respectively. We show how our results can be incorporated into a new lambda-coalescent model which allows multiple lineages to coalesce together. We call this new model the omega-coalescent, we compare it with previously proposed alternatives, and advocate its use in future studies of infectious disease outbreaks.

q-bio.PE

Do machine learning methods lead to similar individualized treatment rules? A comparison study on real data

Identifying patients who benefit from a treatment is a key aspect of personalized medicine, which allows the development of individualized treatment rules (ITRs). Many machine learning methods have been proposed to create such rules. However, to what extent the methods lead to similar ITRs, i.e., recommending the same treatment for the same individuals is unclear. In this work, we compared 22 of the most common approaches in two randomized control trials. Two classes of methods can be distinguished. The first class of methods relies on predicting individualized treatment effects from which an ITR is derived by recommending the treatment evaluated to the individuals with a predicted benefit. In the second class, methods directly estimate the ITR without estimating individualized treatment effects. For each trial, the performance of ITRs was assessed by various metrics, and the pairwise agreement between all ITRs was also calculated. Results showed that the ITRs obtained via the different methods generally had considerable disagreements regarding the patients to be treated. A better concordance was found among akin methods. Overall, when evaluating the performance of ITRs in a validation sample, all methods produced ITRs with limited performance, suggesting a high potential for optimism. For non-parametric methods, this optimism was likely due to overfitting. The different methods do not lead to similar ITRs and are therefore not interchangeable. The choice of the method strongly influences for which patients a certain treatment is recommended, drawing some concerns about their practical use.

stat.AP

A New Method to Constrain the Appearance and Disappearance of Observed Jellyfish Galaxy Tails

We present a new approach to observationally constrain where the tails of Jellyfish (JF) galaxies in groups and clusters first appear and how long they remain visible with respect to the moment of their orbital pericenter. This is accomplished by measuring the distribution of their tail directions with respect to their host's center, and their distribution in a projected velocity-radius phase-diagram. We then model these observed distributions using a fast and flexible approach where JF tails are painted onto dark matter halos according to a simple parameterised prescription, and perform a Bayesian analysis to estimate the parameters. We demonstrate the effectiveness of our approach using observational mocks, and then apply it to a known observational sample of 106 JF galaxies with radio continuum tails located inside 68 hosts such as groups and clusters. We find that, typically, the radio continuum tails become visible on first infall when the galaxy reaches roughly three quarters of r$_{200}$, and the tails remain visible for a few hundred Myr after pericenter passage. Lower mass galaxies in more massive hosts tend to form visible tails further out and their tails disappear more quickly after pericenter. We argue that this indicates they are more sensitive to ram pressure stripping. With upcoming large area surveys of JF galaxies in progress, this is a promising new method to constrain the environmental conditions in which visible JF tails exist.

astro-ph.GA

Towards an Interoperable Ecosystem of AI and LT Platforms: A Roadmap for the Implementation of Different Levels of Interoperability

With regard to the wider area of AI/LT platform interoperability, we concentrate on two core aspects: (1) cross-platform search and discovery of resources and services; (2) composition of cross-platform service workflows. We devise five different levels (of increasing complexity) of platform interoperability that we suggest to implement in a wider federation of AI/LT platforms. We illustrate the approach using the five emerging AI/LT platforms AI4EU, ELG, Lynx, QURATOR and SPEAKER.

cs.CL

European Language Grid: An Overview

With 24 official EU and many additional languages, multilingualism in Europe and an inclusive Digital Single Market can only be enabled through Language Technologies (LTs). European LT business is dominated by hundreds of SMEs and a few large players. Many are world-class, with technologies that outperform the global players. However, European LT business is also fragmented, by nation states, languages, verticals and sectors, significantly holding back its impact. The European Language Grid (ELG) project addresses this fragmentation by establishing the ELG as the primary platform for LT in Europe. The ELG is a scalable cloud platform, providing, in an easy-to-integrate way, access to hundreds of commercial and non-commercial LTs for all European languages, including running tools and services as well as data sets and resources. Once fully operational, it will enable the commercial and non-commercial European LT community to deposit and upload their technologies and data sets into the ELG, to deploy them through the grid, and to connect with other resources. The ELG will boost the Multilingual Digital Single Market towards a thriving European LT community, creating new jobs and opportunities. Furthermore, the ELG project organises two open calls for up to 20 pilot projects. It also sets up 32 National Competence Centres (NCCs) and the European LT Council (LTC) for outreach and coordination purposes.

cs.CL

Online Abuse toward Candidates during the UK General Election 2019: Working Paper

The 2019 UK general election took place against a background of rising online hostility levels toward politicians and concerns about its impact on democracy. We collected 4.2 million tweets sent to or from election candidates in the six week period spanning from the start of November until shortly after the December 12th election. We found abuse in 4.46\% of replies received by candidates, up from 3.27\% in the matching period for the 2017 UK general election. Abuse levels have also been climbing month on month throughout 2019. Abuse also escalated throughout the campaign period. Abuse focused mainly on a small number of high profile politicians. Abuse is "spiky", triggered by external events such as debates, or certain tweets. Abuse increases when politicians discuss inflammatory topics such as borders and immigration. There may also be a backlash on topics such as social justice. Some tweets may become viral targets for personal abuse. On average, men received more general and political abuse; women received more sexist abuse. MPs choosing not to stand again had received more abuse during 2019.

cs.CY

Deep Bidirectional Transformers for Relation Extraction without Supervision

We present a novel framework to deal with relation extraction tasks in cases where there is complete lack of supervision, either in the form of gold annotations, or relations from a knowledge base. Our approach leverages syntactic parsing and pre-trained word embeddings to extract few but precise relations,which are then used to annotate a larger cor-pus, in a manner identical to distant supervision. The resulting data set is employed to fine tune a pre-trained BERT model in order to perform relation extraction. Empirical evaluation on four data sets from the biomedical domain shows that our method significantly outperforms two simple baselines for unsupervised relation extraction and, even if not using any supervision at all, achieves slightly worse results than the state-of-the-art in three out of four data sets. Importantly, we show that it is possible to successfully fine tune a large pre-trained language model with noisy data, as op-posed to previous works that rely on gold data for fine tuning.

cs.LG

Reasoning Over Paths via Knowledge Base Completion

Reasoning over paths in large scale knowledge graphs is an important problem for many applications. In this paper we discuss a simple approach to automatically build and rank paths between a source and target entity pair with learned embeddings using a knowledge base completion model (KBC). We assembled a knowledge graph by mining the available biomedical scientific literature and extracted a set of high frequency paths to use for validation. We demonstrate that our method is able to effectively rank a list of known paths between a pair of entities and also come up with plausible paths that are not present in the knowledge graph. For a given entity pair we are able to reconstruct the highest ranking path 60% of the time within the the top 10 ranked paths and achieve 49% mean average precision. Our approach is compositional since any KBC model that can produce vector representations of entities can be used.

cs.AI

Race and Religion in Online Abuse towards UK Politicians: Working Paper

Against a backdrop of tensions related to EU membership, we find levels of online abuse toward UK MPs reach a new high. Race and religion have become pressing topics globally, and in the UK this interacts with "Brexit" and the rise of social media to create a complex social climate in which much can be learned about evolving attitudes. In 8 million tweets by and to UK MPs in the first half of 2019, religious intolerance scandals in the UK's two main political parties attracted significant attention. Furthermore, high profile ethnic minority MPs started conversations on Twitter about race and religion, the responses to which provide a valuable source of insight. We found a significant presence for disturbing racial and religious abuse. We also explore metrics relating to abuse patterns, which may affect its impact. We find "burstiness" of abuse doesn't depend on race or gender, but individual factors may lead to politicians having very different experiences online.

cs.CY

Online Abuse of UK MPs from 2015 to 2019: Working Paper

We extend previous work about general election-related abuse of UK MPs with two new time periods, one in late 2018 and the other in early 2019, allowing previous observations to be extended to new data and the impact of key stages in the UK withdrawal from the European Union on patterns of abuse to be explored. The topics that draw abuse evolve over the four time periods are reviewed, with topics relevant to the Brexit debate and campaign tone showing a varying pattern as events unfold, and a suggestion of a "bubble" of topics emphasized in the run-up to the highly Brexit-focused 2017 general election. Brexit stance shows a variable relationship with abuse received. We find, as previously, that in quantitative terms, Conservatives and male politicians receive more abuse. Gender difference remains significant even when accounting for prominence, as gauged from Google Trends data, but prominence, or other factors related to being in power, as well as gender, likely account for the difference associated with party membership. No clear relationship between ethnicity and abuse is found in what remains a very small sample (BAME and mixed heritage MPs). Differences are found in the choice of abuse terms levelled at female vs. male MPs.

cs.CY

Partisanship, Propaganda and Post-Truth Politics: Quantifying Impact in Online Debate

The recent past has highlighted the influential role of social networks and online media in shaping public debate on current affairs and political issues. This paper is focused on studying the role of politically-motivated actors and their strategies for influencing and manipulating public opinion online: partisan media, state-backed propaganda, and post-truth politics. In particular, we present quantitative research on the presence and impact of these three `Ps' in online Twitter debates in two contexts: (i) the run up to the UK EU membership referendum (`Brexit'); and (ii) the information operations of Russia-backed online troll accounts. We first compare the impact of highly partisan versus mainstream media during the Brexit referendum, specifically comparing tweets by half a million `leave' and `remain' supporters. Next, online propaganda strategies are examined, specifically left- and right-wing troll accounts. Lastly, we study the impact of misleading claims made by the political leaders of the leave and remain campaigns. This is then compared to the impact of the Russia-backed partisan media and propaganda accounts during the referendum. In particular, just two of the many misleading claims made by politicians during the referendum were found to be cited in 4.6 times more tweets than the 7,103 tweets related to Russia Today and Sputnik and in 10.2 times more tweets than the 3,200 Brexit-related tweets by the Russian troll accounts.

cs.CY

Online Abuse of UK MPs in 2015 and 2017: Perpetrators, Targets, and Topics

Concerns have reached the mainstream about how social media are affecting political outcomes. One trajectory for this is the exposure of politicians to online abuse. In this paper we use 1.4 million tweets from the months before the 2015 and 2017 UK general elections to explore the abuse directed at politicians. This collection allows us to look at abuse broken down by both party and gender and aimed at specific Members of Parliament. It also allows us to investigate the characteristics of those who send abuse and their topics of interest. Results show that in both absolute and proportional terms, abuse increased substantially in 2017 compared with 2015. Abusive replies are somewhat less directed at women and those not in the currently governing party. Those who send the abuse may be issue-focused, or they may repeatedly target an individual. In the latter category, accounts are more likely to be throwaway. Those sending abuse have a wide range of topical triggers, including borders and terrorism.

cs.CY