SearcharxivSearch

arXiv · 0909.1281

Gibrat's law for cities: uniformly most powerful unbiased test of the Pareto against the lognormal

Abstract

We address the general problem of testing a power law distribution versus a log-normal distribution in statistical data. This general problem is illustrated on the distribution of the 2000 US census of city sizes. We provide definitive results to close the debate between Eeckhout (2004, 2009) and Levy (2009) on the validity of Zipf's law, which is the special Pareto law with tail exponent 1, to describe the tail of the distribution of U.S. city sizes. Because the origin of the disagreement between Eeckhout and Levy stems from the limited power of their tests, we perform the {\em uniformly most powerful unbiased test} for the null hypothesis of the Pareto distribution against the lognormal. The $p$-value and Hill's estimator as a function of city size lower threshold confirm indubitably that the size distribution of the 1000 largest cities or so, which include more than half of the total U.S. population, is Pareto, but we rule out that the tail exponent, estimated to be $1.4 \pm 0.1$, is equal to 1. For larger ranks, the $p$-value becomes very small and Hill's estimator decays systematically with decreasing ranks, qualifying the lognormal distribution as the better model for the set of smaller cities. These two results reconcile the opposite views of Eeckhout (2004, 2009) and Levy (2009). We explain how Gibrat's law of proportional growth underpins both the Pareto and lognormal distributions and stress the key ingredient at the origin of their difference in standard stochastic growth models of cities \cite{Gabaix99,Eeckhout2004}.

Explore related subjects

Keep this discovery

BibTeXRIS

Y. Malevergne, V. Pisarenko, D. Sornette. 2009-09-07. Gibrat's law for cities: uniformly most powerful unbiased test of the Pareto against the lognormal. https://doi.org/10.1103/physreve.83.036111

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Combinations of measurements that are simultaneous fits of parameters of interest and systematic uncertainties using the BLUE method

Combining estimates of the same physics parameter obtained from different measurements improves the precision and robustness of the parameter determination. Modern particle physics measurements are often performed using likelihood fits that include the physics parameter(s) of interest together with nuisance parameters representing systematic uncertainties. In high-statistics analyses, typical of many analyses at the Large Hadron Collider, these nuisance parameters can be constrained in the likelihood fits. We describe how the Best Linear Unbiased Estimator method for combinations can be applied to both the estimates of the parameters of interest and the estimates of the nuisance parameters. We show with concrete example combinations that including the nuisance parameters can improve the precision on the parameters of interest. We show with pseudo-experiments that the uncertainty reported by the combination is reliable and that the approximate likelihood combination proposed in a previous publication and implemented in the Convino software reports an underestimated uncertainty. The method is implemented in an open-source software tool, Combiner.

physics.data-an

Quantity, quality, and timing: Guiding glacier data assimilation strategies in the high Arctic

Accurate simulation of glacier surface mass balance is essential for predicting sea level rise and freshwater resources, but it is constrained by uncertainties in meteorological forcing and model parameters. Here, we deploy glacier data assimilation strategies to assess the value of observations for improving surface mass balance simulation, focusing on observation quantity, quality, and timing. We perform synthetic twin experiments on Kongsvegen glacier, Svalbard, using a Particle Batch Smoother with 1000 ensemble members. Synthetic observations of albedo, snow depth, and surface temperature are assimilated at two quality levels, under two climatic scenarios, and over 12 years. Assimilation benefit is measured as the percentage improvement in the continuous ranked probability score of the posterior glacier surface mass balance relative to the prior. A single optimally timed high quality observation yields mean improvements of up to 80\%. Larger numbers of low quality observations partially compensate for lower improvement. In the accumulation zone, however, additional snow depth observations degrade performance through particle degeneracy. Optimal timing is governed by the seasonal transitions of the truth trajectory rather than by prior ensemble spread alone. The optimal windows shift by up to six weeks between early and late melting years. Joint assimilation adds value through temporal diversity rather than observational diversity, while independently timed observations outperform same day combinations. The asynchronously optimally timed combined assimilation of three variables sustains improvements of 85 to 97\% across all years in the ablation zone. These findings provide guidelines for adaptive observation scheduling in glacier monitoring and reanalysis.

physics.data-an

The Greedy Bump Bias: Local Profiling Geometry and the Look-Elsewhere Effect

When fitting a localized signal whose position or shape is not known in advance, one typically allows these parameters to vary together with the signal amplitude and chooses the values that maximize the likelihood. This freedom has two related statistical consequences. If a genuine signal is present, its fitted amplitude will be affected by a positive bias; otherwise, the same freedom increases the chance of finding an unusually signal-like background fluctuation, giving rise to the look-elsewhere effect. We show that these two effects can be understood as consequences of the same local geometry of the family of signal templates. We study this connection in a Gaussian matched-filter model, where a smooth D-dimensional family of normalized templates describes the unknown signal location or shape. In the normalized matched-filter problem, the curvature of a genuine signal peak and the fluctuations that determine the curvature of a high background peak are governed by the same template metric. This allows us to derive an explicit asymptotic relation. We then follow the problem away from the strong-signal and high-threshold limits. Separating the signal-associated maximum from the best competing maximum gives an exact decomposition of the global bias into a local profiling contribution and a contribution from remote-peak competition. In a one-dimensional Gaussian example, the second factorial cumulant accounts for most of this correction, while the third brings the prediction into close agreement with simulation. A two-point Kac--Rice calculation reproduces the second cumulant and reveals a quartic short-distance suppression of nearby maxima. The resulting picture separates the roles of local dimension, model-dependent curvature, and global extremal competition within a common framework.

physics.data-an