SearcharxivSearch

arXiv subjects

Hankun He

Publications and source records attributed to Hankun He.

7 recordsLinked to original sources

Superstatistical Analysis of PDFs and autocorrelation functions for air pollution concentrations in the UK

Conventional statistical models often struggle to fully capture the complex spatio-temporal dynamics, intermittent fluctuations, and heavy-tailed distributions characteristic of real-world air pollution data. Furthermore, existing literature frequently focuses on extreme events, overlooking the persistence of low-pollution states and temporal memory effects. To address these gaps, we apply superstatistical frameworks from non-equilibrium statistical physics to analyse a comprehensive five-year dataset (2020-2025) of hourly air pollutant concentrations across the United Kingdom. Excellent fits of experimentally measured distributions are obtained from our theoretical models. We observe large heterogeneities of the best fitting parameters depending on the locations where the measurements are performed. These parameters form characteristic patterns in the 3-dimensional parameter space and depend on the type of pollutant considered, as well as on the environmental conditions (high traffic, industrial, or rural surroundings). We also investigate autocorrelation functions and provide evidence for differences in day-time and night-time decays of the autocorrelation function. Our investigation mainly focuses onto the dynamics of NO, NO2, PM2.5, PM10, but we also report on some anomalous distributions observed for O3.

physics.ao-ph

Optimising Temporary Accommodation Placement Across London with AI-Powered SaaS in E-Governance Systems

Temporary accommodation has become a major fiscal and administrative pressure for English local authorities, particularly in London, where demand and costs have risen sharply. This paper documents the creation and use of DOMUS, a cloud-based, AI-enabled decision-support system built from scratch at the University of East London and customised for the needs of London Borough of Newham to support statutory Temporary accommodation placement. DOMUS integrates household case records, policy-constrained affordability and suitability rules, and live private-rental listings within a single governance-aligned workflow. The system combines transparent, rule-based filtering with large language model-assisted search to standardise the application of bedroom need, affordability thresholds, geographic preferences, and accessibility requirements, while preserving officer discretion and audibility. Household and property attributes are encoded into policy-consistent representations prior to AI-assisted ranking and explanation. A pilot deployment in Newham's secure environment evaluated operational performance relative to manual workflows. Results indicate substantial reductions in search time, improved adherence to key placement constraints, and high staff satisfaction, while maintaining statutory compliance and role-based accountability. Beyond TA, the paper frames DOMUS as replicable digital public infrastructure: a modular, cloud-native Software-as-a-Service architecture that can be deployed across other UK boroughs and adapted to other public administration tasks characterised by scarcity, rule-bound eligibility, and high stakes. The findings demonstrate the feasibility of scalable, ethically governed AI deployment in local government and contribute to debates on AI-enabled public value creation in e-governance.

cs.CY

Measuring Fairness in Financial Transaction Machine Learning Models

Mastercard, a global leader in financial services, develops and deploys machine learning models aimed at optimizing card usage and preventing attrition through advanced predictive models. These models use aggregated and anonymized card usage patterns, including cross-border transactions and industry-specific spending, to tailor bank offerings and maximize revenue opportunities. Mastercard has established an AI Governance program, based on its Data and Tech Responsibility Principles, to evaluate any built and bought AI for efficacy, fairness, and transparency. As part of this effort, Mastercard has sought expertise from the Turing Institute through a Data Study Group to better assess fairness in more complex AI/ML models. The Data Study Group challenge lies in defining, measuring, and mitigating fairness in these predictions, which can be complex due to the various interpretations of fairness, gaps in the research literature, and ML-operations challenges.

cs.LG

Analyzing Spatio-Temporal Dynamics of Dissolved Oxygen for the River Thames using Superstatistical Methods and Machine Learning

By employing superstatistical methods and machine learning, we analyze time series data of water quality indicators for the River Thames, with a specific focus on the dynamics of dissolved oxygen. After detrending, the probability density functions of dissolved oxygen fluctuations exhibit heavy tails that are effectively modeled using $q$-Gaussian distributions. Our findings indicate that the multiplicative Empirical Mode Decomposition method stands out as the most effective detrending technique, yielding the highest log-likelihood in nearly all fittings. We also observe that the optimally fitted width parameter of the $q$-Gaussian shows a negative correlation with the distance to the sea, highlighting the influence of geographical factors on water quality dynamics. In the context of same-time prediction of dissolved oxygen, regression analysis incorporating various water quality indicators and temporal features identify the Light Gradient Boosting Machine as the best model. SHapley Additive exPlanations reveal that temperature, pH, and time of year play crucial roles in the predictions. Furthermore, we use the Transformer to forecast dissolved oxygen concentrations. For long-term forecasting, the Informer model consistently delivers superior performance, achieving the lowest MAE and SMAPE with the 192 historical time steps that we used. This performance is attributed to the Informer's ProbSparse self-attention mechanism, which allows it to capture long-range dependencies in time-series data more effectively than other machine learning models. It effectively recognizes the half-life cycle of dissolved oxygen, with particular attention to key intervals. Our findings provide valuable insights for policymakers involved in ecological health assessments, aiding in accurate predictions of river water quality and the maintenance of healthy aquatic ecosystems.

cs.LG

Spatial analysis of tails of air pollution PDFs in Europe

Outdoor air pollution is estimated to cause a huge number of premature deaths worldwide, it catalyses many diseases on a variety of time scales, and it has a detrimental effect on the environment. In light of these impacts it is necessary to obtain a better understanding of the dynamics and statistics of measured air pollution concentrations, including temporal fluctuations of observed concentrations and spatial heterogeneities. Here we present an extensive analysis for measured data from Europe. The observed probability density functions (PDFs) of air pollution concentrations depend very much on the spatial location and on the pollutant substance. We analyse a large number of time series data from 3544 different European monitoring sites and show that the PDFs of nitric oxide ($NO$), nitrogen dioxide ($NO_{2}$) and particulate matter ($PM_{10}$ and $PM_{2.5}$) concentrations generically exhibit heavy tails. These are asymptotically well approximated by $q$-exponential distributions with a given entropic index $q$ and width parameter $\lambda$. We observe that the power-law parameter $q$ and the width parameter $\lambda$ vary widely for the different spatial locations. We present the results of our data analysis in the form of a map that shows which parameters $q$ and $\lambda$ are most relevant in a given region. A variety of interesting spatial patterns is observed that correlate to properties of the geographical region. We also present results on typical time scales associated with the dynamical behaviour.

physics.ao-ph

Spatial heterogeneity of air pollution statistics

Air pollution is one of the leading causes of death globally, and continues to have a detrimental effect on our health. In light of these impacts, an extensive range of statistical modelling approaches has been devised in order to better understand air pollution statistics. However, the time-varying statistics of different types of air pollutants are far from being fully understood. The observed probability density functions (PDFs) of concentrations depend very much on the spatial location and on the pollutant substance. In this paper, we analyse a large variety of data from 3544 different European monitoring sites and show that the PDFs of nitric oxide ($NO$), nitrogen dioxide ($NO2$) and particulate matter ($PM10$ and $PM2.5$) concentrations generically exhibit heavy tails and are asymptotically well approximated by $q$-exponential distributions with a given width parameter $\lambda$. We observe that the power-law parameter $q$ and the width parameter $\lambda$ vary widely for the different spatial locations. For each substance, we find different patterns of parameter clouds in the $(q, \lambda)$ plane. These depend on the type of pollutants and on the environmental characteristics (urban/suburban/rural/traffic/industrial/background). This means the effective statistical physics description of air pollution exhibits a strong degree of spatial heterogeneity.

physics.ao-ph

Lockdown effects on air quality in megacities during the first and second waves of COVID-19 pandemic

Air pollution is among the highest contributors to mortality worldwide, especially in urban areas. During spring 2020, many countries enacted social distancing measures in order to slow down the ongoing COVID-19 pandemic. A particularly drastic measure, the 'lockdown', urged people to stay at home and thereby prevent new COVID-19 infections during the first (2020) and second wave (2021) of the pandemic. In turn, it also reduced traffic and industrial activities. But how much did these lockdown measures improve air quality in large cities, and are there differences in how air quality was affected? Here, we analyse data from two megacities: London as an example for Europe and Delhi as an example for Asia. We consider data during first and second wave lockdowns and compare them to 2019 values. Overall, we find a reduction in almost all air pollutants with intriguing differences between the two cities except Delhi in 2021. In London, despite smaller average concentrations, we still observe high-pollutant states and an increased tendency towards extreme events (a higher kurtosis of the probability density during lockdown) during 2020 and low pollution levels during 2021. For Delhi, we observe a much stronger decrease of pollution concentrations, including high pollution states during 2020 and higher pollution levels in 2021. These results could help to design policies to improve long-term air quality in megacities.

physics.soc-ph