SearcharxivSearch

arXiv subjects

David Stern

Publications and source records attributed to David Stern.

10 recordsLinked to original sources

BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance

As pathogen genomic surveillance scales, the bottleneck is shifting from data generation to analysis. We present BioSecBench-Surveillance, a verifiable benchmark of 100 evaluations testing whether AI agents can infer the right analysis pipeline from raw sequencing data and surveillance context. Each evaluation gives an agent only the data and context a human analyst would have, then grades its structured answer deterministically. The tasks span seven categories, from taxonomic classification to genetic-engineering detection, across diverse sample types and sequencing technologies. Across 3,962 gradable attempts from sixteen model-harness pairs, the strongest configuration cleared only about half. Opus 4.8 with PI led at 50.2 percent, with a 95 percent confidence interval of 40.1 to 60.3 percent across 83 evaluations, tied with GPT-5.5 with Codex at 50.2 percent, with a 95 percent confidence interval of 40.8 to 59.6 percent, followed by Opus 4.7 with PI at 49.6 percent, with a 95 percent confidence interval of 40.0 to 59.2 percent, and Sonnet 4.6 with PI at 48.6 percent, with a 95 percent confidence interval of 38.9 to 58.3 percent. Even when agents invoked the correct workflows, their mistakes came from the choices around them, such as which references, thresholds, filters, and normalization to apply. BioSecBench-Surveillance provides a standard for measuring whether agents can be trusted to perform genomic surveillance when the next outbreak arrives.

cs.AI

Improving Bias Correction Methods for Daily Rainfall Using a Markov Chain Approach

Accurate, localised rainfall information is essential for agricultural planning, climate risk assessment, and water resources management. Gridded climate products provide rainfall information over large areas but can lack the accuracy needed at local scales, often requiring bias correction before use in local impact studies. Local intensity scaling (LOCI) and quantile mapping (QM) are two widely used bias correction methods which adjust both rainfall frequency and intensity, but do not account for the temporal structure of daily rainfall. This can lead to biases in the representation of wet and dry spells. This study proposes integrating a two-state first-order Markov chain into existing bias correction methods through state-dependent rain day thresholds and rainfall adjustments, aimed at improving temporal structure. Two implementations of this framework are presented: Markov chain local intensity scaling (MC LOCI) and Markov chain quantile mapping (MC QM). The proposed methods were applied to AgERA5 reanalysis data with rainfall data from five stations in Zimbabwe. Results showed that the Markov chain methods improved the representation of rainfall persistence, onset, and wet and dry spell characteristics compared to LOCI and QM, while maintaining improvements in rain day frequency, mean and total rainfall. Improvements in event timing and daily rainfall amounts were limited. Results from five locations in Zimbabwe demonstrate that the proposed methods could be beneficial for crop simulation, hydrological modelling and other applications requiring accurate rainfall sequencing. Evaluation across additional regions and gridded products would establish the broader applicability of the proposed methods under a range of conditions.

stat.AP

Spatio-temporal evolution of surface temperature trends in Ghana (1983-2021): a multi-station approach

Surface temperature is a fundamental Essential Climate Variable, serving as a primary indicator of climate change and exerting a profound influence on ecosystems, agriculture, and human livelihoods. Although existing research provides a foundation for understanding the climate of Ghana, there remains an opportunity to enhance this landscape with granular station-level analysis. Such high-resolution analysis complements existing studies by capturing localised climatic nuances. This study conducts a detailed spatio-temporal analysis of temperature trends across 22 meteorological stations from 1983 to 2021. Using daily maximum (Tmax) and minimum (Tmin) observations, data were subjected to quality control, homogeneity testing, and homogenisation according to World Meteorological Organisation (WMO) standards, using AgERA5 reanalysis as a reference.The significance and magnitude of trends were determined using the Modified Mann-Kendall test, which is robust in handling potential effects of autocorrelation, and Sen's slope estimator. Results revealed that temperature trends in Ghana are highly localised and seasonal, highlighting the necessity for more studies of this nature. A critical finding is the asymmetric warming across the country, with minimum temperatures rising at an accelerated rate compared to maximum temperatures. This narrowing of the diurnal temperature range poses significant threats to agricultural stability and public health because nocturnal cooling is diminished. These findings underscore the urgent need for site-specific, seasonal climate monitoring to inform customised adaptation strategies. To mitigate these impacts, the study recommends a robust policy framework focusing on afforestation and the transition to green energy.

stat.AP

Bias correction of satellite and reanalysis products for daily rainfall occurrence and intensity

In data-sparse regions, satellite and reanalysis rainfall estimates (SREs) are vital but limited by inherent biases. This study evaluates bias correction (BC) methods, including traditional statistical (LOCI, QM) and machine learning (SVR, GPR), applied to seven SREs across 38 stations in Ghana and Zambia. We introduce a constrained LOCI method to prevent the unrealistically high rainfall values produced by the original approach. Results indicate that statistical methods generally outperformed machine learning, though QM tended to inflate rainfall. Corrected SREs showed high capability in detecting dry days (POD $\ge$ 0.80). The ENACTS product, which integrates numerous station records, was the most amenable to correction in Zambia; most BC methods reduced mean error at >70% of stations. However, ENACTS performed less reliably at an independent station (Moorings), highlighting the need for broader validation at locations not incorporated into the product. Crucially, even after correction, most SREs (except ENACTS) failed to improve the detection of heavy and violent rainfall (POD $\le$ 0.2). This limits their utility for flood risk assessment and highlights a vital research gap regarding extreme event estimation.

stat.AP

Evaluating satellite and reanalysis rainfall estimates for climate services in agriculture: a comprehensive methodology

High-resolution rainfall estimates from satellite and reanalysis sources (SRE) could play a major role in improving climate services for agriculture. This is particularly relevant in regions that rely on rain-fed farming but lack a dense network of ground-based measurements to provide localised historical climate information, as in most of the Global South. However, there is a need for a framework which practitioners can use to determine the suitability of these estimated data for specific agricultural applications. This paper presents a comprehensive methodology for evaluating the ability of SRE to provide historical rainfall information for agricultural applications, primarily through comparison with ground-based measurements. The methodology comprises five main steps: data selection and pre-processing, spatial and temporal consistency checks, quantitative SRE-gauge comparisons, bias correction, and application specific summaries. The methodology makes use of graphical summaries, standard comparison metrics, and Markov chain models. We describe how users can apply this methodology to evaluate rainfall estimates for specific applications, complementing existing validation studies. Evaluation cases are presented to demonstrate the methodology using five widely used satellite and reanalysis rainfall products and ground-based measurements from 12 stations in Africa and the Caribbean. The case studies demonstrate how the methodology can be applied to examine multiple aspects of the rainfall estimates. While previous validation studies ask "Does the SRE estimate the true rainfall well?", this methodology provides means of establishing "To what extent can an SRE be used for this specific purpose?" and a comprehensive framework for this. This meets a major need for location specific rainfall information to improve climate information services for millions of small-holder farming households.

physics.ao-ph

Validation of satellite and reanalysis rainfall products against rain gauge observations in Ghana and Zambia

Accurate rainfall data are crucial for effective climate services, especially in Sub-Saharan Africa, where agriculture depends heavily on rain-fed systems. The sparse distribution of rain-gauge networks necessitates reliance on satellite and reanalysis rainfall products (REs). This study evaluated eight REs -- CHIRPS, TAMSAT, CHIRP, ENACTS, ERA5, AgERA5, PERSIANN-CDR, and PERSIANN-CCS-CDR -- in Zambia and Ghana using a point-to-pixel validation approach. The analysis covered spatial consistency, annual rainfall summaries, seasonal patterns, and rainfall intensity detection across 38 ground stations. Results showed no single product performed optimally across all contexts, highlighting the need for application-specific recommendations. All products exhibited a high probability of detection (POD) for dry days in Zambia and northern Ghana (70% < POD < 100%, and 60% < POD < 85%, respectively), suggesting their utility for drought-related studies. However, all products showed limited skill in detecting heavy and violent rains (POD close to 0%), making them unsuitable for analyzing such events (e.g., floods) in their current form. Products integrated with station data (ENACTS, CHIRPS, and TAMSAT) outperformed others in many contexts, emphasizing the importance of local observation calibration. Bias correction is strongly recommended due to varying bias levels across rainfall summaries. A critical area for improvement is the detection of heavy and violent rains, with which REs currently struggle. Future research should focus on this aspect.

stat.AP

Partitions in real quadratic fields

We study partitions of totally positive integers in real quadratic fields. We develop an algorithm for computing the number of partitions, prove a result about the parity of the partition function, and characterize the quadratic fields such that there exists an element with exactly 1-5, 7, and 11 partitions.

math.NT

A mobile web for enhancing statistics and mathematics education

A freely available educational application (a mobile website) is presented. This provides access to educational material and drilling on selected topics within mathematics and statistics with an emphasis on tablets and mobile phones. The application adapts to the student's performance, selecting from easy to difficult questions, or older material etc. These adaptations are based on statistical models and analyses of data from testing precursors of the system within several courses, from calculus and introductory statistics through multiple linear regression. The application can be used in both on-line and off-line modes. The behavior of the application is determined by parameters, the effects of which can be estimated statistically. Results presented include analyses of how the internal algorithms relate to passing a course and general incremental improvement in knowledge during a semester.

stat.OT

Kernel Topic Models

Latent Dirichlet Allocation models discrete data as a mixture of discrete distributions, using Dirichlet beliefs over the mixture weights. We study a variation of this concept, in which the documents' mixture weight beliefs are replaced with squashed Gaussian distributions. This allows documents to be associated with elements of a Hilbert space, admitting kernel topic models (KTM), modelling temporal, spatial, hierarchical, social and other structure between documents. The main challenge is efficient approximate inference on the latent Gaussian. We present an approximate algorithm cast around a Laplace approximation in a transformed basis. The KTM can also be interpreted as a type of Gaussian process latent variable model, or as a topic model conditional on document features, uncovering links between earlier work in these areas.

cs.LG

Helices on del Pezzo surfaces and tilting Calabi-Yau algebras

We study tilting for a class of Calabi-Yau algebras associated to helices on Fano varieties. We do this by relating the tilting operation to mutations of exceptional collections. For helices on del Pezzo surfaces the algebras are of dimension three, and using an argument of Herzog, together with results of Kuleshov and Orlov, we obtain a complete description of the tilting process in terms of quiver mutations.

math.RA