SearcharxivSearch

arXiv subjects

Hailiang Du

Publications and source records attributed to Hailiang Du.

8 recordsLinked to original sources

Amortized Bandwidth Learning for Kernel Density Estimation under Logarithmic Score

Kernel density estimation converts finite samples into probability densities, but its performance depends critically on bandwidth selection. Classical selectors prescribe the sample-to-bandwidth rule analytically or asymptotically, or solve a new optimization for each sample. An amortized framework is proposed that instead learns this mapping across a distribution of density-estimation tasks by optimizing the logarithmic score. A truncated-and-renormalized bounded-support formulation enables stable learning across heterogeneous tasks, while affine standardization allows a selector trained on a single reference interval to transfer across bounded intervals. Experiments under Gaussian sampling, a multi-family benchmark, and randomized Gaussian-mixture training show that the amortized selector consistently and substantially outperforms Silverman's rule, the Sheather--Jones selector, and least-squares cross-validation, with especially large gains in small and heterogeneous samples. Finite Gaussian mixtures provide a generic training mechanism supported by their $L^1$ approximation property. Selectors trained in this way generalize strongly across different density structures, allowing the same trained selector to be applied directly to finite samples from unknown densities without specifying or fitting a distributional family. This combination of broad applicability and strong empirical performance makes the framework attractive for a wide range of applications in which finite samples or ensembles must be converted into continuous probability densities.

cs.LG

A Distribution-to-Distribution Neural Probabilistic Forecasting Framework for Dynamical Systems

Probabilistic forecasting provides a principled framework for uncertainty quantification in dynamical systems by representing predictions as probability distributions rather than deterministic trajectories. However, existing forecasting approaches, whether physics-based or neural-network-based, remain fundamentally trajectory-oriented: predictive distributions are usually accessed through ensembles or sampling, rather than evolved directly as dynamical objects. A distribution-to-distribution (D2D) neural probabilistic forecasting framework is developed to operate directly on predictive distributions. The framework introduces a distributional encoding and decoding structure around a replaceable neural forecasting module, using kernel mean embeddings to represent input distributions and mixture density networks to parameterise output predictive distributions. This design enables recursive propagation of predictive uncertainty within a unified end-to-end neural architecture, with model training and evaluation carried out directly in terms of probabilistic forecast skill. The framework is demonstrated on the Lorenz63 chaotic dynamical system. Results show that the D2D model captures nontrivial distributional evolution under nonlinear dynamics, produces skillful probabilistic forecasts without explicit ensemble simulation, and remains competitive with, and in some cases outperforms, a simplified perfect model benchmark. These findings point to a new paradigm for probabilistic forecasting, in which predictive distributions are learned and evolved directly rather than reconstructed indirectly through ensemble-based uncertainty propagation.

stat.ML

Centroid Decision Forest

This paper introduces the centroid decision forest (CDF), a novel ensemble learning framework that redefines the splitting strategy and tree building in the ordinary decision trees for high-dimensional classification. The splitting approach in CDF differs from the traditional decision trees in theat the class separability score (CSS) determines the selection of the most discriminative features at each node to construct centroids of the partitions (daughter nodes). The splitting criterion uses the Euclidean distance measurements from each class centroid to achieve a splitting mechanism that is more flexible and robust. Centroids are constructed by computing the mean feature values of the selected features for each class, ensuring a class-representative division of the feature space. This centroid-driven approach enables CDF to capture complex class structures while maintaining interpretability and scalability. To evaluate CDF, 23 high-dimensional datasets are used to assess its performance against different state-of-the-art classifiers through classification accuracy and Cohen's kappa statistic. The experimental results show that CDF outperforms the conventional methods establishing its effectiveness and flexibility for high-dimensional classification problems.

stat.ML

Beyond Strictly Proper Scoring Rules: The Importance of Being Local

The evaluation of probabilistic forecasts plays a central role both in the interpretation and in the use of forecast systems and their development. Probabilistic scores (scoring rules) provide statistical measures to assess the quality of probabilistic forecasts. Often, many probabilistic forecast systems are available while evaluations of their performance are not standardized, with different scoring rules being used to measure different aspects of forecast performance. Even when the discussion is restricted to strictly proper scoring rules, there remains considerable variability between them; indeed strictly proper scoring rules need not rank competing forecast systems in the same order when none of these systems are perfect. The locality property is explored to further distinguish scoring rules. The nonlocal strictly proper scoring rules considered are shown to have a property that can produce "unfortunate" evaluations. Particularly the fact that Continuous Rank Probability Score prefers the outcome close to the median of the forecast distribution regardless the probability mass assigned to the value at/near the median raises concern to its use. The only local strictly proper scoring rules, the logarithmic score, has direct interpretations in terms of probabilities and bits of information. The nonlocal strictly proper scoring rules, on the other hand, lack meaningful direct interpretation for decision support. The logarithmic score is also shown to be invariant under smooth transformation of the forecast variable, while the nonlocal strictly proper scoring rules considered may, however, change their preferences due to the transformation. It is therefore suggested that the logarithmic score always be included in the evaluation of probabilistic forecasts.

stat.ME

Rising Above Chaotic Likelihoods

Berliner (Likelihood and Bayesian prediction for chaotic systems, J. Am. Stat. Assoc. 1991) identified a number of difficulties in using the likelihood function within the Bayesian paradigm which arise both for state estimation and for parameter estimation of chaotic systems. Even when the equations of the system are given, he demonstrated "chaotic likelihood functions" both of initial conditions and of parameter values in the Logistic Map. Chaotic likelihood functions, while ultimately smooth, have such complicated small scale structure as to cast doubt on the possibility of identifying high likelihood states in practice. In this paper, the challenge of chaotic likelihoods is overcome by embedding the observations in a higher dimensional sequence-space; this allows good state estimation with finite computational power. An importance sampling approach is introduced, where Pseudo-orbit Data Assimilation is employed in the sequence-space, first to identify relevant pseudo-orbits and then relevant trajectories. Estimates are identified with likelihoods orders of magnitude higher than those previously identified in the examples given by Berliner. Pseudo-orbit Data Assimilation importance sampler exploits the information both from the model dynamics and from the observations. While sampling from the relevant prior (here, the natural measure) will, of course, eventually yield an accountable sample, given the realistic computational resource this traditional approach would provide no high likelihood points at all. While one of the challenges Berliner posed is overcome, his central conclusion is supported. "Chaotic likelihood functions" for parameter estimation still pose a challenge; this fact helps clarify why physical scientists maintain a strong distinction between the initial condition uncertainty and parameter uncertainty.

physics.data-an

On the Design and use of Ensembles of Multi-model Simulations for Forecasting

Probability forecasting is common in the geosciences, the finance sector, and elsewhere. It is sometimes the case that one has multiple probability-forecasts for the same target. How is the information in these multiple forecast systems best "combined"? Assuming stationary, then in the limit of a very large forecast-outcome archive, each model-based probability density function can be weighted to form a "multi-model forecast" which will, in expectation, provide the most information. In the case that one of the forecast systems yields a probability distribution which reflects the distribution from which the outcome will be drawn, then Bayesian Model Averaging will identify this model as the number of forecast-outcome pairs goes to infinity. In many applications, like those of seasonal forecasting, data are precious: the archive is often limited to fewer than $2^6$ entries. And no perfect model is in hand. In this case, it is shown that forming a single "multi-model probability forecast" can be expected to prove misleading. These issues are investigated using probability forecasts of a simple mathematical system, which allows most limiting behaviours to be quantified.

stat.ME

Multi-model Cross Pollination in Time

Predictive skill of complex models is often not uniform in model-state space; in weather forecasting models, for example, the skill of the model can be greater in populated regions of interest than in "remote" regions of the globe. Given a collection of models, a multi-model forecast system using the cross pollination in time approach can be generalised to take advantage of instances where some models produce systematically more accurate forecast of some components of the model-state. This generalisation is stated and then successfully demonstrated in a moderate ~40 dimensional nonlinear dynamical system suggested by Lorenz. In this demonstration four imperfect models, each with similar global forecast skill, are used. Future applications in weather forecasting and in economic forecasting are discussed. The demonstration establishes that cross pollinating forecast trajectories to enrich the collection of simulations upon which the forecast is built can yield a new forecast system with significantly more skills than the original multi-model forecast system.

physics.data-an

Parameter Estimation Through Ignorance

Dynamical modelling lies at the heart of our understanding of physical systems. Its role in science is deeper than mere operational forecasting, in that it allows us to evaluate the adequacy of the mathematical structure of our models. Despite the importance of model parameters, there is no general method of parameter estimation outside linear systems. A new relatively simple method of parameter estimation for nonlinear systems is presented, based on variations in the accuracy of probability forecasts. It is illustrated on the Logistic Map, the Henon Map and the 12-D Lorenz96 flow, and its ability to outperform linear least squares in these systems is explored at various noise levels and sampling rates. As expected, it is more effective when the forecast error distributions are non-Gaussian. The new method selects parameter values by minimizing a proper, local skill score for continuous probability forecasts as a function of the parameter values. This new approach is easier to implement in practice than alternative nonlinear methods based on the geometry of attractors or the ability of the model to shadow the observations. New direct measures of inadequacy in the model, the "Implied Ignorance" and the information deficit are introduced.

physics.data-an