Searcharxiv⌕ Search

arXiv subjects

Hywel T. P. Williams

Publications and source records attributed to Hywel T. P. Williams.

12 recordsLinked to original sources

Automated classification of natural habitats using ground-level imagery

Accurate classification of terrestrial habitats is critical for biodiversity conservation, ecological monitoring, and land-use planning. Several habitat classification schemes are in use, typically based on analysis of satellite imagery with validation by field ecologists. Here we present a methodology for classification of habitats based solely on ground-level imagery (photographs), offering improved validation and the ability to classify habitats at scale (for example using citizen-science imagery). In collaboration with Natural England, a public sector organisation responsible for nature conservation in England, this study develops a classification system that applies deep learning to ground-level habitat photographs, categorising each image into one of 18 classes defined by the 'Living England' framework. Images were pre-processed using resizing, normalisation, and augmentation; re-sampling was used to balance classes in the training data and enhance model robustness. We developed and fine-tuned a DeepLabV3-ResNet101 classifier to assign a habitat class label to each photograph. Using five-fold cross-validation, the model demonstrated strong overall performance across 18 habitat classes, with accuracy and F1-scores varying between classes. Across all folds, the model achieved a mean F1-score of 0.61, with visually distinct habitats such as Bare Soil, Silt and Peat (BSSP) and Bare Sand (BS) reaching values above 0.90, and mixed or ambiguous classes scoring lower. These findings demonstrate the potential of this approach for ecological monitoring. Ground-level imagery is readily obtained, and accurate computational methods for habitat classification based on such data have many potential applications. To support use by practitioners, we also provide a simple web application that classifies uploaded images using our model.

cs.CV↗

Timescale-agnostic characterisation for collective attention events

Online communications, and in particular social media, are a key component of how society interacts with and promotes content online. Collective attention on such content can vary wildly. The majority of breaking topics quickly fade into obscurity after only a handful of interactions, while the possibility exists for content to ``go viral'', seeing sustained interaction by large audiences over long periods. In this paper we investigate the mechanisms behind such events and introduce a new representation that enables direct comparison of events over diverse time and volume scales. We find four characteristic behaviours in the usage of hashtags on Twitter that are indicative of different patterns of attention to topics. We go on to develop an agent-based model for generating collective attention events to test the factors affecting emergence of these phenomena. This model can reproduce the characteristic behaviours seen in the Twitter dataset using a small set of parameters, and reveal that three of these behaviours instead represent a continuum determined by model parameters rather than discrete categories. These insights suggest that collective attention in social systems develops in line with a set of universal principles independent of effects inherent to system scale, and the techniques we introduce here present a valuable opportunity to infer the possible mechanisms of attention flow in online communications.

cs.SI↗

Artificial Intelligence for Collective Intelligence: A National-Scale Research Strategy

Advances in artificial intelligence (AI) have great potential to help address societal challenges that are both collective in nature and present at national or trans-national scale. Pressing challenges in healthcare, finance, infrastructure and sustainability, for instance, might all be productively addressed by leveraging and amplifying AI for national-scale collective intelligence. The development and deployment of this kind of AI faces distinctive challenges, both technical and socio-technical. Here, a research strategy for mobilising inter-disciplinary research to address these challenges is detailed and some of the key issues that must be faced are outlined.

cs.AI↗

The Language of Weather: Social Media Reactions to Weather Accounting for Climatic and Linguistic Baselines

This study explores how different weather conditions influence public sentiment on social media, focusing on Twitter data from the UK. By considering climate and linguistic baselines, we improve the accuracy of weather-related sentiment analysis. Our findings show that emotional responses to weather are complex, influenced by combinations of weather variables and regional language differences. The results highlight the importance of context-sensitive methods for better understanding public mood in response to weather, which can enhance impact-based forecasting and risk communication in the context of climate change.

cs.HC↗

CIDER: Context sensitive sentiment analysis for short-form text

Researchers commonly perform sentiment analysis on large collections of short texts like tweets, Reddit posts or newspaper headlines that are all focused on a specific topic, theme or event. Usually, general-purpose sentiment analysis methods are used. These perform well on average but miss the variation in meaning that happens across different contexts, for example, the word "active" has a very different intention and valence in the phrase "active lifestyle" versus "active volcano". This work presents a new approach, CIDER (Context Informed Dictionary and sEmantic Reasoner), which performs context-sensitive linguistic analysis, where the valence of sentiment-laden terms is inferred from the whole corpus before being used to score the individual texts. In this paper, we detail the CIDER algorithm and demonstrate that it outperforms state-of-the-art generalist unsupervised sentiment analysis techniques on a large collection of tweets about the weather. CIDER is also applicable to alternative (non-sentiment) linguistic scales. A case study on gender in the UK is presented, with the identification of highly gendered and sentiment-laden days. We have made our implementation of CIDER available as a Python package: https://pypi.org/project/ciderpolarity/.

cs.CL↗

The structure of segregation in co-authorship networks and its impact on scientific production

Co-authorship networks, where nodes represent authors and edges represent co-authorship relations, are key to understanding the production and diffusion of knowledge in academia. Social constructs, biases (implicit and explicit), and constraints (e.g. spatial, temporal) affect who works with whom and cause co-authorship networks to organise into tight communities with different levels of segregation. We aim to look at aspects of the co-authorship network structure that lead to segregation and its impact on scientific production. We measure segregation using the Spectral Segregation Index (SSI) and find 4 ordered segregation categories: completely segregated, highly segregated, moderately segregated and non-segregated communities. We direct our attention to the non-segregated and highly segregated communities, quantifying and comparing their structural topologies and k-core positions. When considering communities of both categories (controlling for size), our results show no differences in density and clustering but substantial variability in core position. Larger non-segregated communities are more likely to occupy cores near the network nucleus, while the highly segregated ones tend to be closer to the network periphery. Finally, we analyse differences in citations gained by researchers within communities showing different segregation categories. Researchers in highly segregated communities get more citations from their community members in middle cores and gain more citations per publication in middle/periphery cores. Those in non-segregated communities get more citations per publication in the nucleus. To our knowledge, this work is the first to characterise community segregation in co-authorship networks and investigate the relationship between community segregation and author citations.

cs.SI↗

Using Semantic Similarity and Text Embedding to Measure the Social Media Echo of Strategic Communications

Online discourse covers a wide range of topics and many actors tailor their content to impact online discussions through carefully crafted messages and targeted campaigns. Yet the scale and diversity of online media content make it difficult to evaluate the impact of a particular message. In this paper, we present a new technique that leverages semantic similarity to quantify the change in the discussion after a particular message has been published. We use a set of press releases from environmental organisations and tweets from the climate change debate to show that our novel approach reveals a heavy-tailed distribution of response in online discourse to strategic communications.

cs.SI↗

Sponsored messaging about climate change on Facebook: Actors, content, frames

Online communication about climate change is central to public discourse around this contested issue. Facebook is a dominant social media platform known to be a major source of information and online influence, yet discussion of climate change on the platform has remained largely unstudied due to difficulties in accessing data. This paper utilises Facebook's repository of social/political ads to study how climate change is framed as an issue in adverts placed by different actors. Sponsored content is a strategic investment and presumably intended to be persuasive, so patterns of who pays for adverts and how those adverts frame the issue can reveal large-scale trends in public discourse. We show that most money spent on climate-related messaging is targeted at users in the US, GB and CA. While the number of advert impressions correlates with total spend by an actor, there is a secondary effect of unpaid social sharing which can substantially affect the number of impressions per dollar spent. Most spend in the US is by political actors, while environmental non-governmental organisations dominate spend in GB. Analysis shows that climate change solutions are well represented in GB, while climate change impacts such as extreme weather events are strongly represented in the US and CA. Different actor types frame the issue of climate change in different ways; political actors position the issue as party political and a point of difference between candidates, whereas environmental NGOs frame climate change as the focus of collective action and social mobilisation. Overall, our study provides a first empirical exploration of climate-related advertising on Facebook. It shows the diversity of actors seeking to use Facebook as a platform for their campaigns and how they utilise different topic frames to persuade users to act.

stat.AP↗

Complex networks for event detection in heterogeneous high volume news streams

Detecting important events in high volume news streams is an important task for a variety of purposes.The volume and rate of online news increases the need for automated event detection methods thatcan operate in real time. In this paper we develop a network-based approach that makes the workingassumption that important news events always involve named entities (such as persons, locationsand organizations) that are linked in news articles. Our approach uses natural language processingtechniques to detect these entities in a stream of news articles and then creates a time-stamped seriesof networks in which the detected entities are linked by co-occurrence in articles and sentences. Inthis prototype, weighted node degree is tracked over time and change-point detection used to locateimportant events. Potential events are characterized and distinguished using community detectionon KeyGraphs that relate named entities and informative noun-phrases from related articles. Thismethodology already produces promising results and will be extended in future to include a widervariety of complex network analysis techniques.

cs.SI↗

The Human Geography of Twitter

Given the centrality of regions in social movements, politics and public administration we aim to quantitatively study inter- and intra-regional communication for the first time. This work uses social media posts to first identify contiguous geographical regions with a shared social identity and then investigate patterns of communication within and between them. Our case study uses over 150 days of located Twitter data from England and Wales. In contrast to other approaches, (e.g. phone call data records or online friendship networks) we have the message contents as well as the social connection. This allows us to investigate not only the volume of communication but also the sentiment and vocabulary. We find that the South-East and North-West regions are the most talked about; regions tend to be more positive about themselves than about others; people talk politics much more between regions than within. This methodology gives researchers a powerful tool to study identity and interaction within and between social-geographic regions.

cs.SI↗

Gaian bottlenecks and planetary habitability maintained by evolving model biospheres: The ExoGaia model

The search for habitable exoplanets inspires the question - how do habitable planets form? Planet habitability models traditionally focus on abiotic processes and neglect a biotic response to changing conditions on an inhabited planet. The Gaia hypothesis postulates that life influences the Earth's feedback mechanisms to form a self-regulating system, and hence that life can maintain habitable conditions on its host planet. If life has a strong influence, it will have a role in determining a planet's habitability over time. We present the ExoGaia model - a model of simple 'planets' host to evolving microbial biospheres. Microbes interact with their host planet via consumption and excretion of atmospheric chemicals. Model planets orbit a 'star' which provides incoming radiation, and atmospheric chemicals have either an albedo, or a heat-trapping property. Planetary temperatures can therefore be altered by microbes via their metabolisms. We seed multiple model planets with life while their atmospheres are still forming and find that the microbial biospheres are, under suitable conditions, generally able to prevent the host planets from reaching inhospitable temperatures, as would happen on a lifeless planet. We find that the underlying geochemistry plays a strong role in determining long-term habitability prospects of a planet. We find five distinct classes of model planets, including clear examples of 'Gaian bottlenecks' - a phenomenon whereby life either rapidly goes extinct leaving an inhospitable planet, or survives indefinitely maintaining planetary habitability. These results suggest that life might play a crucial role in determining the long-term habitability of planets.

astro-ph.EP↗

Social Sensing of Floods in the UK

"Social sensing" is a form of crowd-sourcing that involves systematic analysis of digital communications to detect real-world events. Here we consider the use of social sensing for observing natural hazards. In particular, we present a case study that uses data from a popular social media platform (Twitter) to detect and locate flood events in the UK. In order to improve data quality we apply a number of filters (timezone, simple text filters and a naive Bayes `relevance' filter) to the data. We then use place names in the user profile and message text to infer the location of the tweets. These two steps remove most of the irrelevant tweets and yield orders of magnitude more located tweets than we have by relying on geo-tagged data. We demonstrate that high resolution social sensing of floods is feasible and we can produce high-quality historical and real-time maps of floods using Twitter.

cs.HC↗