SearcharxivSearch

arXiv subjects

Andreas Züfle

Publications and source records attributed to Andreas Züfle.

At least 19 recordsLinked to original sources

Synthetic Wastewater Epidemiology Data Generation using Patterns-of-Life Simulation

Wastewater contains rich biological signals that can be used to monitor population health, track infectious diseases, and detect emerging outbreaks. Pathogens and other biomarkers in sewage provide a unique, noninvasive view of disease prevalence at the community level. However, extracting these signals requires extensive field sampling, laboratory analysis, and expert interpretation. Consequently, wastewater-based epidemiology (WBE) datasets are scarce, geographically fragmented, and rarely released as open data. Even when available, existing datasets typically cover short time periods and limited geographic regions, restricting their usefulness for method development, benchmarking, and large-scale modeling. To address this gap, we present an application of an existing patterns-of-life simulation framework for generating synthetic infectious disease and wastewater pathogen datasets for wastewater-based epidemiology. Our approach extends a patterns-of-life simulation framework by using it as a model of human mobility and behavior while incorporating disease transmission and pathogen shedding dynamics. The resulting framework generates high-resolution spatial and temporal datasets capturing infection dynamics, mobility behavior, and wastewater-associated pathogen signals. We release a fully simulated dataset containing check-in records, social network links, infection states, pathogen loads, and ground-truth disease transmission information. These data support controlled experimentation for outbreak detection, source localization, resource allocation, surveillance strategy design, mobility-aware wastewater analysis, and targeted public health interventions. We provide datasets for Fulton County and demonstrate that the framework generalizes to other regions by generating data for any city with available OpenStreetMap information.

q-bio.PE

Staypoint Detection from Noisy Trajectory Data [Experiment Paper]

Detecting staypoints from raw trajectory data is fundamental to numerous spatial computing applications. This process transforms raw numeric sequences of geolocations into semantically meaningful locations, such as homes, workplaces, or restaurants. Despite its importance for semantic trajectory analysis, staypoint detection lacks standard benchmarks, and existing algorithms have never been systematically evaluated. This gap persists because no publicly available datasets provide both raw individual trajectories and ground-truth staypoint annotations. This benchmark paper addresses this limitation with two key contributions: (1) we introduce 16 large-scale simulated datasets capturing thousands of agents with annotated staypoints across varying trajectory noise levels, and (2) we evaluate nine staypoint detection algorithms-including both state-of-the-art and novel methods-to analyze their robustness to noise. Our evaluation reveals that existing state-of-the-art algorithms perform poorly under realistic noise conditions. Conversely, our proposed unsupervised methods yield substantial improvements, while supervised approaches drastically outperform existing baselines. While these results are very promising, these datasets and methods are only meant as starting points for future research in staypoint detection.

cs.LG

HD-GEN: A High-Performance Software System for Human Mobility Data Generation Based on Patterns of Life

Understanding individual-level human mobility is critical for a wide range of applications. Real-world trajectory datasets provide valuable insights into movement behaviors and patterns of life but are often constrained by data sparsity and participation bias. Synthetic data, by contrast, offers scalability and flexibility but frequently lacks realism. % To address this gap, we introduce a comprehensive software pipeline for generating, calibrating, processing, and visualizing large-scale individual-level human mobility datasets that combine the realism of empirical data with the control and extensibility simulations. % Our system consists of four integrated components: (1) a data generation engine that constructs geographically grounded simulations using OpenStreetMap data to produce diverse mobility logs; (2) a genetic algorithm--based calibration module that fine-tunes simulation parameters to align with real-world mobility characteristics; (3) a data processing suite that transforms raw simulation logs into structured formats suitable for downstream applications; and (4) a visualization module that extracts and presents key mobility patterns and insights from the processed datasets for improved interpretability. Evaluation of generated trajectory datasets for the Atlanta, Georgia, USA region show realistic behavior that, despite emerging from a simulation without any reference to real human individuals, exhibits realistic human behavior that closely matches aggregate metrics of real-world datasets. We also provide a sensitivity analysis to study what simulation parameters affect simulation runtime. Code and simulated datasets are shared to provide the broad research community with large-scale dataset that, albeit not real, exhibit realistic human behavior while being orders of magnitudes larger than any open real-world mobility dataset.

cs.SE

Where Do We Poop? City-Wide Simulation of Defecation Behavior for Wastewater-Based Epidemiology

Wastewater surveillance, which regularly measures pathogen biomarkers in wastewater samples, is a valuable tool for monitoring infectious diseases circulating in communities. Yet, most wastewater-based epidemiology methods that use wastewater surveillance results to infer disease trends implicitly assume that individuals excrete only at their residential locations and that the populations contributing to wastewater samples are static. These simplifying assumptions ignore daily mobility, social interactions, and heterogeneous toilet-use patterns, which can bias the interpretation of wastewater results, especially at upstream sampling locations such as neighborhoods, institutions, or buildings. Here, we introduce an agent-based geospatial simulation framework. Building on an established Patterns of Life model, we simulate daily human activities within a realistic urban environment and extend the framework with a physiologically motivated defecation cycle and toilet-use patterns. We couple this behavioral model with an infectious disease model to simulate transmission through spatial and social interactions. When an infected agent defecates, a pathogen-shedding model determines the amount of pathogen released in the feces. By integrating population mobility, disease transmission, toilet-use behavior, and pathogen shedding, the framework can simulate the spatiotemporal dynamics of wastewater pathogen loads. Using a case study of 10,000 simulated agents in Fulton County, Georgia, we examine how varying infection rates alter epidemic trajectories, wastewater pathogen loads, and the spatial distribution of pathogen shedding over time. Our results show that mobility and toilet use can substantially decouple residential disease prevalence from wastewater pathogen loads and demonstrate how behaviorally grounded simulations can support interpretation, scenario analysis, and wastewater surveillance

physics.soc-ph

Simulating Public Transit Fare Policies in NYC: An Efficient, Socioeconomic-Aware Framework

Designing equitable and effective public transit fare policies is challenging due to complex interactions among traveler behavior, multimodal networks, and socioeconomic heterogeneity. This paper presents a scalable, data-driven simulation framework for evaluating transit fare policies in New York City (NYC), integrating a synthetic population, agent-based simulation, multimodal travel-time estimation, and fare-sensitive mode choice modeling. We evaluate multiple fare scenarios, including distance-based pricing, fare increases, and fare-free bus policies. Results show that pricing changes modestly affect total ridership but significantly alter modal composition and produce heterogeneous impacts across income groups. In particular, fare-free bus policies generate substantial benefits for lower-income riders by increasing bus usage and reducing fare burden, while introducing trade-offs in revenue. To support city-scale analysis, we introduce a sampling-based approach that reduces computational cost while preserving aggregate accuracy. The proposed framework provides a practical tool for assessing trade-offs between ridership, revenue, and equity, enabling more informed and equitable transit policy design.

cs.CE

Mobility Anomaly Generation using LLM-Driven Behavior with Kinematic Constraints

Although the study of human trajectory anomalies is critical for advancing spatial data mining, empirical research remains severely hindered by a pervasive lack of ground-truth datasets. Despite the availability of several real-world and simulated human trajectory collections, these datasets exclusively capture normal mobility patterns and lack annotated anomalies. This specific scarcity is fundamentally driven by the inherent statistical rarity of anomalous events, precluding the feasibility of conventional observational methods. Compounding this challenge, the systematic acquisition of large-scale mobility data is strictly bottlenecked by prohibitive costs and stringent privacy regulations. To overcome these fundamental limitations and establish a reliable human trajectory anomalies dataset with annotated ground truth, we introduce a novel, end-to-end generative framework designed to synthesize realistic trajectory anomalies at scale. Our architecture bridges the gap between purely synthetic mobility data and complex real-world physical constraints by operating directly on baseline simulated trajectories. We employ Large Language Model (LLM) agents to systematically inject semantically meaningful behavioral anomalies such as irregular out-of-distribution check-ins and skipped routine visits. To ensure rigorous spatial validity, the system leverages map-constrained routing reconstruction to recalculate the physical transitions between these LLM agent-modified staypoints. Moreover, to narrow the simulation-to-reality gap, we augment the resulting trajectories with a context-aware spatial noise model, parameterized by environmental and location-specific variables, to accurately emulate heterogeneous GPS sensor degradation.

cs.AI

SF-LIFE: A Large-Scale Simulated Movement Dataset for the San Francisco Bay Area

We introduce SF-LIFE, a large-scale simulated movement dataset designed to accelerate research in transportation, mobility, and machine learning. The dataset contains 3,024,000,000,000 location records capturing complete, noise-free, multi-modality trajectories of 500,000 simulated agents observed at a 1Hz frequency navigating the San Francisco Bay Area network over a 70-day period. The data captures (1) needs-driven daily agendas of individual agents generated by an agent-based simulation of human patterns of life and (2) detailed kinematic trajectories moving agents across the OpenStreetMap representation of San Francisco using data from 40+ transit agencies across 9 counties. SF-LIFE provides unprecedented scale and detail as trajectories are based on real transit infrastructure using San Francisco General Transit Feed Specification (GTFS) data, having agent movements across multiple modalities, including bus, rail, bike, automobile, and walking. For this high-fidelity simulated representation of San Francisco, we provide (1) the full trajectory data annotated with transportation mode labels, (2) reduced-size versions of the trajectory data with reduced temporal frequency, (3) agent activity information describing the causal activity why an agent visits a place, (4) agent demographic data, and (5) the underlying OSM road network and building data. As the first dataset of its scale and level of detail, SF-LIFE overcomes the privacy, noise, and completeness limitations inherent in real-world tracking data, providing a robust and ethically sourced resource for research in transit optimization, human mobility analysis, and urban computing.

physics.soc-ph

Genomic-Informed Heterogeneous Graph Learning for Spatiotemporal Avian Influenza Outbreak Forecasting

Accurate forecasting of Avian Influenza Virus (AIV) outbreaks within wild bird populations necessitates models that account for complex, multi-scale transmission patterns driven by diverse factors. While conventional spatiotemporal epidemic models are robust for human-centric diseases, they rely on spatial homophily and diffusive transmission between geographic regions. This simplification is incomplete for AIV as it neglects valuable genomic information critical for capturing dynamics like high-frequency reassortment and lineage turnover at the case level (e.g., genetic descent across regions), which are essential for understanding AIV spread. To address these limitations, we systematically formulate the AIV forecasting problem and propose a Bi-Layer genomic-aware heterogeneous graph fusion pipeline. This pipeline integrates genetic, spatial, and ecological data to achieve highly accurate outbreak forecasting. It 1) defines a multi-layered graph structure incorporating information from diverse sources and multiple layers (case and location), 2) applies cross-relation smoothing to smooth information flow across edge types, 3) performs graph fusion that preserves critical structural patterns backed by theoretical spectral guarantees, and 4) forecasts future outbreaks using an autoregressive graph sequence model to capture transmission dynamics. To support research, we release the Avian-US dataset, which provides comprehensive genetic, spatial, and ecological data on US avian influenza outbreaks. BLUE demonstrates superior performance over existing baselines, highlighting the efficacy of integrating multi-layer information for infectious disease forecasting. The code is available at: https://github.com/cruiseresearchgroup/BLUE.

cs.SI

From Ecological Connectivity to Outbreak Risk: A Heterogeneous Graph Network for Epidemiological Reasoning under Sparse Spatiotemporal Data

Estimating population-level prevalence and transmission dynamics of wildlife pathogens can be challenging, partly because surveillance data is sparse, detection-driven, and unevenly sequenced. Using highly pathogenic avian influenza A/H5 clade 2.3.4.4b as a case study, we develop zooNet, a graph-based epidemiological framework that integrates mechanistic transmission simulation, metadata-driven genetic distance imputation, and spatiotemporal graph learning to reconstruct outbreak dynamics from incomplete observations. Applied to wild bird surveillance data from the United States during 2022, zooNet recovered coherent spatiotemporal structure despite intermittent detections, revealing sustained regional circulation across multiple migratory flyways. The framework consistently identified counties with ongoing transmission weeks to months before confirmed detections, including persistent activity in northeastern regions prior to documented re-emergence. These signals were detectable even in areas with sparse sequencing and irregular reporting. These results show that explicitly representing ecological processes and inferred genomic connectivity within a unified graph structure allows persistence and spatial risk structure to be inferred from detection-driven wildlife surveillance data.

q-bio.PE

World-POI: Global Point-of-Interest Data Enriched from Foursquare and OpenStreetMap as Tabular and Graph Data

Recently, Foursquare released a global dataset with more than 100 million points of interest (POIs), each representing a real-world business on its platform. However, many entries lack complete metadata such as addresses or categories, and some correspond to non-existent or fictional locations. In contrast, OpenStreetMap (OSM) offers a rich, user-contributed POI dataset with detailed and frequently updated metadata, though it does not formally verify whether a POI represents an actual business. In this data paper, we present a methodology that integrates the strengths of both datasets: Foursquare as a comprehensive baseline of commercial POIs and OSM as a source of enriched metadata. The combined dataset totals approximately 1 TB. While this full version is not publicly released, we provide filtered releases with adjustable thresholds that reduce storage needs and make the data practical to download and use across domains. We also provide step-by-step instructions to reproduce the full 631 GB build. Record linkage is achieved by computing name similarity scores and spatial distances between Foursquare and OSM POIs. These measures identify and retain high-confidence matches that correspond to real businesses in Foursquare, have representations in OSM, and show strong name similarity. Finally, we use this filtered dataset to construct a graph-based representation of POIs enriched with attributes from both sources, enabling advanced spatial analyses and a range of downstream applications.

cs.DB

A Probabilistic Framework for Imputing Genetic Distances in Spatiotemporal Pathogen Models

Pathogen genome data offers valuable structure for spatial models, but its utility is limited by incomplete sequencing coverage. We propose a probabilistic framework for inferring genetic distances between unsequenced cases and known sequences within defined transmission chains, using time-aware evolutionary distance modeling. The method estimates pairwise divergence from collection dates and observed genetic distances, enabling biologically plausible imputation grounded in observed divergence patterns, without requiring sequence alignment or known transmission chains. Applied to highly pathogenic avian influenza A/H5 cases in wild birds in the United States, this approach supports scalable, uncertainty-aware augmentation of genomic datasets and enhances the integration of evolutionary information into spatiotemporal modeling workflows.

q-bio.GN

Exploring Economic Sectoral Dynamics Through High-resolution Mobility Data

We present a comprehensive dataset capturing patterns of human mobility across the United States from January 2019 to January 2023, based on anonymized mobile device data. Aggregated weekly, the dataset reports visits, travel distances, and time spent at public locations organized by economic sector for approximately 12 million Points of Interest (POIs). This resource enables the study of how mobility and economic activity changed over time, particularly during major events such as the COVID-19 pandemic. By disaggregating patterns across different types of businesses, it provides valuable insights for researchers in economics, urban studies, and public health. To protect privacy, all data have been aggregated and anonymized. This dataset offers an opportunity to explore the dynamics of human behavior across sectors over an extended time period, supporting studies of mobility, resilience, and recovery.

cs.CY

Training Machine Learning Models on Human Spatio-temporal Mobility Data: An Experimental Study [Experiment Paper]

Individual-level human mobility prediction has emerged as a significant topic of research with applications in infectious disease monitoring, child, and elderly care. Existing studies predominantly focus on the microscopic aspects of human trajectories: such as predicting short-term trajectories or the next location visited, while offering limited attention to macro-level mobility patterns and the corresponding life routines. In this paper, we focus on an underexplored problem in human mobility prediction: determining the best practices to train a machine learning model using historical data to forecast an individuals complete trajectory over the next days and weeks. In this experiment paper, we undertake a comprehensive experimental analysis of diverse models, parameter configurations, and training strategies, accompanied by an in-depth examination of the statistical distribution inherent in human mobility patterns. Our empirical evaluations encompass both Long Short-Term Memory and Transformer-based architectures, and further investigate how incorporating individual life patterns can enhance the effectiveness of the prediction. We show that explicitly including semantic information such as day-of-the-week and user-specific historical information can help the model better understand individual patterns of life and improve predictions. Moreover, since the absence of explicit user information is often missing due to user privacy, we show that the sampling of users may exacerbate data skewness and result in a substantial loss in predictive accuracy. To mitigate data imbalance and preserve diversity, we apply user semantic clustering with stratified sampling to ensure that the sampled dataset remains representative. Our results further show that small-batch stochastic gradient optimization improves model performance, especially when human mobility training data is limited.

cs.LG

Evaluating the Bias in LLMs for Surveying Opinion and Decision Making in Healthcare

Generative agents have been increasingly used to simulate human behaviour in silico, driven by large language models (LLMs). These simulacra serve as sandboxes for studying human behaviour without compromising privacy or safety. However, it remains unclear whether such agents can truly represent real individuals. This work compares survey data from the Understanding America Study (UAS) on healthcare decision-making with simulated responses from generative agents. Using demographic-based prompt engineering, we create digital twins of survey respondents and analyse how well different LLMs reproduce real-world behaviours. Our findings show that some LLMs fail to reflect realistic decision-making, such as predicting universal vaccine acceptance. However, Llama 3 captures variations across race and Income more accurately but also introduces biases not present in the UAS data. This study highlights the potential of generative agents for behavioural research while underscoring the risks of bias from both LLMs and prompting strategies.

cs.CL

Chatting with Logs: An exploratory study on Finetuning LLMs for LogQL

Logging is a critical function in modern distributed applications, but the lack of standardization in log query languages and formats creates significant challenges. Developers currently must write ad hoc queries in platform-specific languages, requiring expertise in both the query language and application-specific log details -- an impractical expectation given the variety of platforms and volume of logs and applications. While generating these queries with large language models (LLMs) seems intuitive, we show that current LLMs struggle with log-specific query generation due to the lack of exposure to domain-specific knowledge. We propose a novel natural language (NL) interface to address these inconsistencies and aide log query generation, enabling developers to create queries in a target log query language by providing NL inputs. We further introduce ~\textbf{NL2QL}, a manually annotated, real-world dataset of natural language questions paired with corresponding LogQL queries spread across three log formats, to promote the training and evaluation of NL-to-loq query systems. Using NL2QL, we subsequently fine-tune and evaluate several state of the art LLMs, and demonstrate their improved capability to generate accurate LogQL queries. We perform further ablation studies to demonstrate the effect of additional training data, and the transferability across different log formats. In our experiments, we find up to 75\% improvement of finetuned models to generate LogQL queries compared to non finetuned models.

cs.DB

Neural Collaborative Filtering to Detect Anomalies in Human Semantic Trajectories

Human trajectory anomaly detection has become increasingly important across a wide range of applications, including security surveillance and public health. However, existing trajectory anomaly detection methods are primarily focused on vehicle-level traffic, while human-level trajectory anomaly detection remains under-explored. Since human trajectory data is often very sparse, machine learning methods have become the preferred approach for identifying complex patterns. However, concerns regarding potential biases and the robustness of these models have intensified the demand for more transparent and explainable alternatives. In response to these challenges, our research focuses on developing a lightweight anomaly detection model specifically designed to detect anomalies in human trajectories. We propose a Neural Collaborative Filtering approach to model and predict normal mobility. Our method is designed to model users' daily patterns of life without requiring prior knowledge, thereby enhancing performance in scenarios where data is sparse or incomplete, such as in cold start situations. Our algorithm consists of two main modules. The first is the collaborative filtering module, which applies collaborative filtering to model normal mobility of individual humans to places of interest. The second is the neural module, responsible for interpreting the complex spatio-temporal relationships inherent in human trajectory data. To validate our approach, we conducted extensive experiments using simulated and real-world datasets comparing to numerous state-of-the-art trajectory anomaly detection approaches.

cs.LG

Kinematic Detection of Anomalies in Human Trajectory Data

Historically, much of the research in understanding, modeling, and mining human trajectory data has focused on where an individual stays. Thus, the focus of existing research has been on where a user goes. On the other hand, the study of how a user moves between locations has great potential for new research opportunities. Kinematic features describe how an individual moves between locations and can be used for tasks such as identification of individuals or anomaly detection. Unfortunately, data availability and quality challenges make kinematic trajectory mining difficult. In this paper, we leverage the Geolife dataset of human trajectories to investigate the viability of using kinematic features to identify individuals and detect anomalies. We show that humans have an individual "kinematic profile" which can be used as a strong signal to identify individual humans. We experimentally show that, for the two use-cases of individual identification and anomaly detection, simple kinematic features fed to standard classification and anomaly detection algorithms significantly improve results.

cs.LG

Spatial Transfer Learning for Estimating PM2.5 in Data-poor Regions

Air pollution, especially particulate matter 2.5 (PM2.5), is a pressing concern for public health and is difficult to estimate in developing countries (data-poor regions) due to a lack of ground sensors. Transfer learning models can be leveraged to solve this problem, as they use alternate data sources to gain knowledge (i.e., data from data-rich regions). However, current transfer learning methodologies do not account for dependencies between the source and the target domains. We recognize this transfer problem as spatial transfer learning and propose a new feature named Latent Dependency Factor (LDF) that captures spatial and semantic dependencies of both domains and is subsequently added to the feature spaces of the domains. We generate LDF using a novel two-stage autoencoder model that learns from clusters of similar source and target domain data. Our experiments show that transfer learning models using LDF have a 19.34% improvement over the baselines. We additionally support our experiments with qualitative findings.

cs.LG