Searcharxiv⌕ Search

arXiv subjects

V. P. Roychowdhury

Publications and source records attributed to V. P. Roychowdhury.

At least 19 recordsLinked to original sources

Estimating the number of serial killers that were never caught

Many serial killers commit tens of murders. At the same time inter-murder intervals can be decades long. This suggests that some serial killers can die of an accident or a disease, having been never caught. We use the distribution of the killers by the number of murders, the distribution of the length of inter-murder intervals and USA life tables to estimate the number of the uncaught killers. The result is that in 20th century there were about seven of such killers. The most prolific of them likely committed over sixty murders.

physics.soc-ph↗

Statistical study of time intervals between murders for serial killers

We study the distribution of 2,837 inter-murder intervals (cooling off periods) for 1,012 American serial killers. The distribution is smooth, following a power law in the region of 10-10,000 days. The power law cuts off where inter-murder intervals become comparable with the length of human life. Otherwise there is no other characteristic scale in the distribution. In particular, we do not see any characteristic spree-killer interval or serial-killer interval, but only a monotonous smooth distribution lacking any features. This suggests that there is only a quantitative difference between serial killers and spree-killers, representing different samples generated by the same underlying phenomenon. The over decade long inter-murder intervals are not anomalies, but rare events described by the same power-law distribution and therefore should not necessarily be looked upon with suspicion, as has been done in a recent case involving a serial killer dubbed as the "Grim Sleeper." This large-scale study supports the conclusions of a previous study, involving three prolific serial killers, and the associated neural net model, which can explain the observed power law distribution.

physics.soc-ph↗

Chess players' fame versus their merit

We investigate a pool of international chess title holders born between 1901 and 1943. Using Elo ratings we compute for every player his expected score in a game with a randomly selected player from the pool. We use this figure as player's merit. We measure players' fame as the number of Google hits. The correlation between fame and merit is 0.38. At the same time the correlation between the logarithm of fame and merit is 0.61. This suggests that fame grows exponentially with merit.

physics.soc-ph↗

Why does attention to web articles fall with time?

We analyze access statistics of a hundred and fifty blog entries and news articles, for periods of up to three years. Access rate falls as an inverse power of time passed since publication. The power law holds for periods of up to thousand days. The exponents are different for different blogs and are distributed between 0.6 and 3.2. We argue that the decay of attention to a web article is caused by the link to it first dropping down the list of links on the website's front page, and then disappearing from the front page and its subsequent movement further into background. The other proposed explanations that use a decaying with time novelty factor, or some intricate theory of human dynamics cannot explain all of the experimental observations.

cs.IR↗

Stochastic modeling of a serial killer

We analyze the time pattern of the activity of a serial killer, who during twelve years had murdered 53 people. The plot of the cumulative number of murders as a function of time is of "Devil's staircase" type. The distribution of the intervals between murders (step length) follows a power law with the exponent of 1.4. We propose a model according to which the serial killer commits murders when neuronal excitation in his brain exceeds certain threshold. We model this neural activity as a branching process, which in turn is approximated by a random walk. As the distribution of the random walk return times is a power law with the exponent 1.5, the distribution of the inter-murder intervals is thus explained. We illustrate analytical results by numerical simulation. Time pattern activity data from two other serial killers further substantiate our analysis.

physics.soc-ph↗

A mathematical theory of fame

We study empirically how the fame of WWI fighter-pilot aces, measured in numbers of web pages mentioning them, is related to their achievement, measured in numbers of opponent aircraft destroyed. We find that on the average fame grows exponentially with achievement; the correlation coefficient between achievement and the logarithm of fame is 0.72. The number of people with a particular level of achievement decreases exponentially with the level, leading to a power-law distribution of fame. We propose a stochastic model that can explain the exponential growth of fame with achievement. Next, we hypothesize that the same functional relation between achievement and fame that we found for the aces holds for other professions. This allows us to estimate achievement for professions where an unquestionable and universally accepted measure of achievement does not exist. We apply the method to Nobel Prize winners in Physics. For example, we obtain that Paul Dirac, who is a hundred times less famous than Einstein contributed to physics only two times less. We compare our results with Landau's ranking.

physics.soc-ph↗

Stochastic modeling of Congress

We analyze the dynamics of growth of the number of congressmen supporting the resolution HR1207 to audit the Federal Reserve. The plot of the total number of co-sponsors as a function of time is of "Devil's staircase" type. The distribution of the numbers of new co-sponsors joining during a particular day (step height) follows a power law. The distribution of the length of intervals between additions of new co-sponsors (step length) also follows a power law. We use a modification of Bak-Tang-Wiesenfeld sandpile model to simulate the dynamics of Congress and obtain a good agreement with the data.

physics.soc-ph↗

Theory of citing

We present empirical data on misprints in citations to twelve high-profile papers. The great majority of misprints are identical to misprints in articles that earlier cited the same paper. The distribution of the numbers of misprint repetitions follows a power law. We develop a stochastic model of the citation process, which explains these findings and shows that about 70-90% of scientific citations are copied from the lists of references used in other papers. Citation copying can explain not only why some misprints become popular, but also why some papers become highly cited. We show that a model where a scientist picks few random papers, cites them, and copies a fraction of their references accounts quantitatively for empirically observed distribution of citations.

physics.soc-ph↗

Estimating achievement from fame

We report a method for estimating people's achievement based on their fame. Earlier we discovered (cond-mat/0310049) that fame of fighter pilot aces (measured as number of Google hits) grows exponentially with their achievement (number of victories). We hypothesize that the same functional relation between achievement and fame holds for other professions. This allows us to estimate achievement for professions where an unquestionable and universally accepted measure of achievement does not exist. We apply the method to Nobel Prize winners in Physics. For example, we obtain that Paul Dirac, who is hundred times less famous than Einstein contributed to physics only two times less. We compare our results with Landau's ranking.

physics.soc-ph↗

Re-inventing Willis

Scientists often re-invent things that were long known. Here we review these activities as related to the mechanism of producing power law distributions, originally proposed in 1922 by Yule to explain experimental data on the sizes of biological genera, collected by Willis. We also review the history of re-invention of closely related branching processes, random graphs and coagulation models.

physics.soc-ph↗

An explanation of the distribution of inter-seizure intervals

Recently Osorio et al (Eur. J. Neurosci., 30 (2009) 1554) reported that probability distribution of intervals between successive epileptic seizures follows a power law with exponent 1.5. We theoretically explain this finding by modeling epileptic activity as a branching process, which we in turn approximate by a random walk. We confirm the theoretical conclusion by numerical simulation.

q-bio.NC↗

Theory of aces: high score by skill or luck?

We studied the distribution of World War I fighter pilots by the number of victories they were credited with, along with casualty reports. Using the maximum entropy method we obtained the underlying distribution of pilots by their skill. We find that the variance of this skill distribution is not very large, and that the top aces achieved their victory scores mostly by luck. For example, the ace of aces, Manfred von Richthofen, most likely had a skill in the top quarter of the active WWI German fighter pilots and was no more special than that. When combined with our recent study (cond-mat/0310049), showing that fame grows exponentially with victory scores, these results (derived from real data) show that both outstanding achievement records and resulting fame are mostly due to chance.

physics.soc-ph↗

A theory of web traffic

We analyze access statistics of several popular webpages for a period of several years. The graphs of daily downloads are highly non-homogeneous with long periods of low activity interrupted by bursts of heavy traffic. These bursts are due to avalanches of blog entries, referring to the page. We quantitatively explain this behavior using the theory of branching processes. We extrapolate these findings to construct a model of the entire web. According to the model, the competition between webpages for viewers pushes the web into a self-organized critical state. In this regime, the most interesting webpages are in a near-critical state, with a power law distribution of traffic intensity.

physics.soc-ph↗

A mathematical theory of citing

Recently we proposed a model in which when a scientist writes a manuscript, he picks up several random papers, cites them and also copies a fraction of their references (cond-mat/0305150). The model was stimulated by our discovery that a majority of scientific citations are copied from the lists of references used in other papers (cond-mat/0212043). It accounted quantitatively for several properties of empirically observed distribution of citations. However, important features, such as power-law distribution of citations to papers published during the same year and the fact that the average rate of citing decreases with aging of a paper, were not accounted for by that model. Here we propose a modified model: when a scientist writes a manuscript, he picks up several random recent papers, cites them and also copies some of their references. The difference with the original model is the word recent. We solve the model using methods of the theory of branching processes, and find that it can explain the aforementioned features of citation distribution, which our original model couldn't account for. The model can also explain "sleeping beauties in science", i.e., papers that are little cited for a decade or so, and later "awake" and get a lot of citations. Although much can be understood from purely random models, we find that to obtain a good quantitative agreement with empirical citation data one must introduce Darwinian fitness parameter for the papers.

physics.soc-ph↗

An introduction to the theory of citing

Statistical analysis of repeat misprints in scientific citations leads to the conclusion that about 80% of scientific citations are copied from the lists of references used in othe papers. Based on this finding a mathematical theory of citing is constructed. It leads to the conclusion that a large number of citations does not have to be a result of paper's extraordinary qualities, but can be explained by the ordinary law of chances.

math.ST↗

Theory of Aces: Fame by chance or merit?

We study empirically how fame of WWI fighter-pilot aces, measured in numbers of web pages mentioning them, is related to their achievement or merit, measured in numbers of opponent aircraft destroyed. We find that on the average fame grows exponentially with achievement; to be precise, there is a strong correlation (~0.7) between achievement and the logarithm of fame. At the same time, the number of individuals achieving a particular level of merit decreases exponentially with the magnitude of the level, leading to a power-law distribution of fame. A stochastic model that can explain the exponential growth of fame with merit is also proposed.

cond-mat.dis-nn↗

Stochastic modeling of citation slips

We present empirical data on frequency and pattern of misprints in citations to twelve high-profile papers. We find that the distribution of misprints, ranked by frequency of their repetition, follows Zipf's law. We propose a stochastic model of citation process, which explains these findings, and leads to the conclusion that 70-90% of scientific citations are copied from the lists of references used in other papers.

cond-mat.dis-nn↗

Copied citations create renowned papers?

Recently we discovered (cond-mat/0212043) that the majority of scientific citations are copied from the lists of references used in other papers. Here we show that a model, in which a scientist picks three random papers, cites them,and also copies a quarter of their references accounts quantitatively for empirically observed citation distribution. Simple mathematical probability, not genius, can explain why some papers are cited a lot more than the other.

cond-mat.dis-nn↗