SearcharxivSearch

arXiv · 1403.3659

Construction d'une plate-forme intégrée pour la cartographie de l'exposition des populations aux substances chimiques de l'environnement

Abstract

L'analyse du lien entre l'environnement et la santé est devenue une préoccupation majeure de santé publique comme en témoigne l'émergence des deux Plans nationaux santé environnement. Pour ce faire, les décideurs sont confrontés au besoin de développement d'outils nécessaires à l'identification des zones géographiques dans lesquelles une surexposition potentielle à des substances toxiques est observée. L'objectif du projet Système d'information géographique (SIG), facteurs de risques environnementaux et décès par cancer (SIGFRIED 1) est de construire une plate-forme de modélisation permettant d'évaluer, par une approche spatiale, l'exposition de la population française aux substances chimiques et d'en identifier ses déterminants. L'évaluation des expositions est réalisée par le biais d'une modélisation multimédia probabiliste. Les problèmes épistémologiques liés à l'absence de données sont palliés par la mise en œuvre d'outils utilisant les techniques d'analyse spatiale. Un exemple est fourni sur la région Nord-Pas-de-Calais et Picardie, pour le cadmium, le nickel et le plomb. Le calcul de l'exposition est réalisé sur une durée de 70 ans sur la base des données disponibles autour de l'année 2004 sur une maille de 1 km de côté. Par exemple pour le Nord-Pas-de-Calais, les indicateurs permettent de définir deux zones pour le cadmium et trois zones pour le plomb. Celles-ci sont liées à l'historique industriel de la région : le bassin minier, les activités métallurgiques et l'agglomération lilloise. La contribution des différentes voies d'exposition varie sensiblement d'un polluant à l'autre. Les cartes d'exposition ainsi obtenues permettent d'identifier les zones géographiques dans lesquelles conduire en priorité des études environnementales de terrains. Le SIG construit constitue la base d'une plate-forme où les données d'émission à la source, de mesures environnementales, d'exposition, puis sanitaires et socio-économiques pourront être associées. -- Analysis of the association between the environment and health has become a major public health concern, as shown by the development of two national environmental health plans. For such an analysis, policy-makers need tools to identify the geographic areas where overexposure to toxic agents may be observed. The objective of the SIGFRIED 1 project is to build a work station for spatial modeling of the exposure of the French population to chemical substances and for identifying the determinants of this exposure. Probabilistic multimedia modeling is used to assess exposure. The epistemological problems associated with the absence of data are overcome by the implementation of tools that apply spatial analysis techniques. An example is furnished for the region of Nord-Pas-de-Calais and Picardie, for cadmium, nickel and lead exposure. The calculation of exposure is performed for duration of 70 years on the basis of data collected around 2004 fora grid of squares 1 km in length. For example, for Nord-Pas-de-Calais, the indicators allow us to define two areas for cadmium and three for lead. They are linked to the region's industrial history: mining basin, metallurgy activities, and the Lille metropolitan area. The contribution of various exposure pathways varied substantially from one pollutant to another. The exposure maps thus obtained allow us to identify the geographic area where environmental studies must be conducted in priority. The GIS thus constructed is the foundation of a workstation where source emission data, environmental exposure measurements, and finally health and socioeconomic measurements can be combined.

Explore related subjects

Keep this discovery

BibTeXRIS

Julien Caudeville, Céline Boudet, Gérard Govaert, Roseline Bonnard, André Cicollela. 2014-02-12. Construction d'une plate-forme intégrée pour la cartographie de l'exposition des populations aux substances chimiques de l'environnement. https://arxiv.org/abs/1403.3659

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

MarkerScout: A Disease-Agnostic Machine Learning Framework for Biomarker Prediction from Multi-Scale Mechanistic Models

We demonstrate the framework on three infectious diseases derived from a companion mechanistic immune-simulation platform: SARS-CoV-2, Influenza A Virus, and Plasmodium falciparum. Each disease was evaluated across hospitalization and intensive care unit cohorts, yielding six cohorts in total. Best-pipeline cross-validated macro F1 ranged from 0.82 for IAV-HOSP to 0.99 for COV-ICU, and the framework produced tiered, direction-aware biomarker lists for each disease and phase. Interleukin-18 (IL-18) reached the strongest tier in both SARS-CoV-2 phases with consistent direction. When benchmarked against three separate, independently collected clinical ICU datasets, MarkerScout's top-ranked features outperformed 94.4% of randomly selected feature sets of equivalent size for SARS-CoV-2, with a weaker but directionally consistent advantage for Influenza A Virus (66.7%) and Plasmodium falciparum (60.7%).

q-bio.OT

Enhancing Clinical Decision Support and Differential Diagnosis with Knowledge Graphs, and Retrieval Augmented Generation in Generative AI

Diagnostic error carries a burden, while unconstrained large language models (LLMs) remain vulnerable to hallucination and weak integration of quantitative laboratory dynamics. We developed a decision-support pipeline combining disease-specific biomarker correlation graphs, ordinary differential equations (ODEs), deep sequence classification, and retrieval-augmented generation (RAG). For 103 disease classes from a full blood count (FBC) repository, biomarker networks were used as coupling matrices to generate 30 trajectories per disease (3,090 total). A one-dimensional convolutional neural network (CNN) and long short-term memory (LSTM) network classified disease trajectories and six dynamical clusters. A constrained GPT-4o-mini RAG layer used a 19-pattern BMJ Best Practice/NICE corpus to generate differential diagnoses evaluated for diagnostic suitability, evidential grounding, and clinical plausibility. Across five random-seed runs, disease-level accuracy was $0.940 \pm 0.006$ for the CNN (95\% CI 0.933--0.948) and $0.852 \pm 0.019$ for the LSTM (95\% CI 0.828--0.875); the CNN advantage was 8.87 percentage points (95\% CI 6.47--11.27; $t(4)=10.26$, $p=5.1\times10^{-4}$; Hedges' $g=3.67$). Among 100 sampled RAG cases, 96 parsed successfully; evidence was cited in 97.9\%, the true diagnosis was mentioned in 71.9\%, and the composite score was 3.82/5 with a 47.9\% strict pass rate. The central finding was a decoupling between grounding and diagnostic correctness: classifier-correct versus classifier-wrong outputs differed in diagnostic suitability but not evidential grounding. Post-hoc analysis confirmed a 1.02-point diagnostic-score difference (Mann--Whitney $p=0.0024$; Hedges' $g=0.72$), whereas grounding differed by only $-0.02$ points ($p=0.839$; $g=-0.04$).

q-bio.OT

Expanding the Human Ancestry Ontology to include under-represented populations and ethnicities for broader utility in annotations

Successful discovery, integration and reuse of data relies on the availability of rich, well-structured and machine-readable metadata to describe every aspect of the data, from sample sources to collection processes to experimental protocols. The use of standardised terminologies to express concepts in a harmonised fashion lies at the core of high-quality data annotation, increasing the FAIRness of the data, facilitating data integration and promoting reproducibility. Here, we describe the Human Ancestry Ontology (HANCESTRO), originally developed to improve standardised reporting of genetic ancestry genomic resources such as the NHGRI-EBI GWAS Catalog and the Human Cell Atlas through high-level population descriptors, and more recently expanded to include diverse and previously under-represented populations in genomics and genetics research. HANCESTRO provides a framework for population descriptors that includes both ancestry based on the analysis of genetic information and self-reported ethnicity, which is based on social and cultural factors that don't necessarily align with genetic populations. By enabling the accurate and interoperable representation of population-related data, it promotes inclusive, representative and reproducible science.

q-bio.OT