SearcharxivSearch

arXiv subjects

Maurizio Serva

Publications and source records attributed to Maurizio Serva.

At least 19 recordsLinked to original sources

Evolution of the lexicon: a probabilistic point of view

The Swadesh approach for determining the temporal separation between two languages relies on the stochastic process of words replacement (when a complete new word emerges to represent a given concept). It is well known that the basic assumptions of the Swadesh approach are often unrealistic due to various contamination phenomena and misjudgments (horizontal transfers, variations over time and space of the replacement rate, incorrect assessments of cognacy relationships, presence of synonyms, and so on). All of this means that the results cannot be completely correct. More importantly, even in the unrealistic case that all basic assumptions are satisfied, simple mathematics places limits on the accuracy of estimating the temporal separation between two languages. These limits, which are purely probabilistic in nature and which are often neglected in lexicostatistical studies, are analyzed in detail in this article. Furthermore, in this work we highlight that the evolution of a language's lexicon is also driven by another stochastic process: gradual lexical modification of words. We show that this process equally also represents a major contribution to the reshaping of the vocabulary of languages over the centuries and we also show, from a purely probabilistic perspective, that taking into account this second random process significantly increases the precision in determining the temporal separation between two languages.

cs.CL

Brownian motion at the speed of light: a Lorentz invariant family of processes

We recently introduced a new family of processes which describe particles which only can move at the speed of light c in the ordinary 3D physical space. The velocity, which randomly changes direction, can be represented as a point on the surface of a sphere of radius c and its trajectories only may connect the points of this variety. A process can be constructed both by considering jumps from one point to another (velocity changes discontinuously) and by continuous velocity trajectories on the surface. We followed this second new strategy assuming that the velocity is described by a Wiener process (which is isotropic only in the 'rest frame') on the surface of the sphere. Using both Ito calculus and Lorentz boost rules, we succeed here in characterizing the entire Lorentz-invariant family of processes. Moreover, we highlight and describe the short-term ballistic behavior versus the long-term diffusive behavior of the particles in the 3D physical space.

physics.class-ph

Stability of meanings versus rate of replacement of words: an experimental test

The words of a language are randomly replaced in time by new ones, but it has long been known that words corresponding to some items (meanings) are less frequently replaced than others. Usually, the rate of replacement for a given item is not directly observable, but it is inferred by the estimated stability which, on the contrary, is observable. This idea goes back a long way in the lexicostatistical literature, nevertheless nothing ensures that it gives the correct answer. The family of Romance languages allows for a direct test of the estimated stabilities against the replacement rates since the proto-language (Latin) is known and the replacement rates can be explicitly computed. The output of the test is threefold:first, we prove that the standard approach which tries to infer the replacement rates trough the estimated stabilities is sound; second, we are able to rewrite the fundamental formula of Glottochronology for a non universal replacement rate (a rate which depends on the item); third, we give indisputable evidence that the stability ranking is far from being the same for different families of languages. This last result is also supported by comparison with the Malagasy family of dialects. As a side result we also provide some evidence that Vulgar Latin and not Late Classical Latin is at the root of modern Romance languages.

cs.CL

The origins of the Malagasy people, some certainties and a few mysteries

The Malagasy language belongs to the Greater Barito East group of the Austronesian family, the language most closely connected to Malagasy dialects is Maanyan (Kalimantan), but Malay as well other Indonesian and Philippine languages are also related. The African contribution is very high in the Malagasy genetic make-up (about 50%) but negligible in the language. Because of the linguistic link, it is widely accepted that the island was settled by Indonesian sailors after a maritime trek but date and place of landing are still debated. The 50% Indonesian genetic contribution to present Malagasy points in a different direction then Maanyan for the Asian ancestry, therefore, the ethnic composition of the Austronesian settlers is also still debated. In this talk I mainly review the joint research of Filippo Petroni, Dima Volchenkov, Sören Wichmann and myself which tries to shed new light on these problems. The key point is the application of a new quantitative methodology which is able to find out the kinship relations among languages (or dialects). New techniques are also introduced in order to extract the maximum information from these relations concerning time and space patterns.

q-bio.PE

Linear and anomalous front propagation in system with non Gaussian diffusion: the importance of tails

We investigate front propagation in systems with diffusive and sub-diffusive behavior. The scaling behavior of moments of the diffusive problem, both in the standard and in the anomalous cases, is not enough to determine the features of the reactive front. In fact, the shape of the bulk of the probability distribution of the transport process, which determines the diffusive properties, is important just for pre-asymptotic behavior of front propagation, while the precise shape of the tails of the probability distribution determines asymptotic behavior of front propagation.

cond-mat.stat-mech

Observability of Market Daily Volatility

We study the price dynamics of 65 stocks from the Dow Jones Composite Average from 1973 until 2014. We show that it is possible to define a Daily Market Volatility $σ(t)$ which is directly observable from data. This quantity is usually indirectly defined by $r(t)=σ(t) ω(t)$ where the $r(t)$ are the daily returns of the market index and the $ω(t)$ are i.i.d. random variables with vanishing average and unitary variance. The relation $r(t)=σ(t) ω(t)$ alone is unable to give an operative definition of the index volatility, which remains unobservable. On the contrary, we show that using the whole information available in the market, the index volatility can be operatively defined and detected.

q-fin.ST

Asymptotic properties of a bold random walk

In a recent paper we proposed a non-Markovian random walk model with memory of the maximum distance ever reached from the starting point (home). The behavior of the walker is at variance with respect to the simple symmetric random walk (SSRW) only when she is at this maximum distance, where, having the choice to move either farther or closer, she decides with different probabilities. If the probability of a forward step is higher then the probability of a backward step, the walker is bold and her behavior turns out to be super-diffusive, otherwise she is timorous and her behavior turns out to be sub-diffusive. The scaling behavior vary continuously from sub-diffusive (timorous) to super-diffusive (bold) according to a single parameter $γ\in R$. We investigate here the asymptotic properties of the bold case in the non ballistic region $γ\in [0,1/2]$, a problem which was left partially unsolved in \cite{S}. The exact results proved in this paper require new probabilistic tools which rely on the construction of appropriate martingales of the random walk and its hitting times.

math-ph

Exactly solvable tight-binding model on the RAN: fractal energy spectrum and Bose-Einstein condensation

We initially consider a single-particle tight-binding model on the Regularized Apollonian Network (RAN). The RAN is defined starting from a tetrahedral structure with four nodes all connected (generation 0). At any successive generations, new nodes are added and connected with the surrounding three nodes. As a result, a power-law cumulative distribution of connectivity $P(k)\propto {1}/{k^η}$ with $η=\ln(3)/\ln(2) \approx 1.585$ is obtained. The eigenvalues of the Hamiltonian are exactly computed by a recursive approach for any size of the network. In the infinite size limit, the density of states and the cumulative distribution of states (integrated density of states) are also exactly determined. The relevant scaling behavior of the cumulative distribution close to the band bottom is shown to be power law with an exponent depending on the spectral dimension and not on the embedding dimension. We then consider a gas made by an infinite number of non-interacting bosons each of them described by the tight-binding Hamiltonian on the RAN and we prove that, for sufficiently large bosonic density and sufficiently small temperature, a macroscopic fraction of the particles occupy the lowest single-particle energy state forming the Bose-Einstein condensate. We determine no only the transition temperature as a function of the bosonic density, but also the fraction of condensed particle, the fugacity, the energy and the specific heat for any temperature and bosonic density.

cond-mat.quant-gas

Effect of migration in a diffusion model for template coexistence in protocells

The compartmentalization of distinct templates in protocells and the exchange of templates between them (migration) are key elements of a modern scenario for prebiotic evolution. Here we use the diffusion approximation of population genetics to study analytically the steady-state properties of such prebiotic scenario. The coexistence of distinct template types inside a protocell is achieved by a selective pressure at the protocell level (group selection) favoring protocells with a mixed template composition. In the degenerate case, where the templates have the same replication rate, we find that a vanishingly small migration rate suffices to eliminate the segregation effect of random drift and so to promote coexistence. In the non-degenerate case, a small migration rate greatly boosts coexistence as compared with the situation where there is no migration. However, increase of the migration rate beyond a critical value leads to the complete dominance of the more efficient template type (homogeneous regime). In this case, we find a continuous phase transition separating the homogeneous and the coexistence regimes, with the order parameter vanishing linearly with the distance to the transition point.

q-bio.PE

Nonlinear group survival in Kimura's model for the evolution of altruism

Establishing the conditions that guarantee the spreading or the sustenance of altruistic traits in a population is the main goal of intergroup selection models. Of particular interest is the balance of the parameters associated to group size, migration and group survival against the selective advantage of the non-altruistic individuals. Here we use Kimura's diffusion model of intergroup selection to determine those conditions in the case the group survival probability is a nonlinear non-decreasing function of the proportion of altruists in a group. In the case this function is linear, there are two possible steady states which correspond to the non-altruistic and the altruistic phases. At the discontinuous transition line separating these phases there is a non-ergodic coexistence phase. For a continuous concave survival function, we find an ergodic coexistence phase that occupies a finite region of the parameter space in between the altruistic and the non-altruistic phases, and is separated from these phases by continuous transition lines. For a convex survival function, the coexistence phase disappears altogether but a bistable phase appears for which the choice of the initial condition determines whether the evolutionary dynamics leads to the altruistic or the non-altruistic steady state.

q-bio.PE

The bold/timorous walker on the trek from home

We study a one-dimensional random walk with memory. The behavior of the walker is modified with respect to the simple symmetric random walk (SSRW) only when he is at the maximum distance ever reached from his starting point (home). In this case, having the choice to move farther or to move closer, he decides with different probabilities. If the probability of a forward step is higher then the probability of a backward step, the walker is bold, otherwise he is timorous. We investigate the asymptotic properties of this bold/timorous random walk (BTRW) showing that the scaling behavior vary continuously from subdiffusive (timorous) to superdiffusive (bold). The scaling exponents are fully determined with a new mathematical approach.

physics.data-an

Automated languages phylogeny from Levenshtein distance

Languages evolve over time in a process in which reproduction, mutation and extinction are all possible, similar to what happens to living organisms. Using this similarity it is possible, in principle, to build family trees which show the degree of relatedness between languages. The method used by modern glottochronology, developed by Swadesh in the 1950s, measures distances from the percentage of words with a common historical origin. The weak point of this method is that subjective judgment plays a relevant role. Recently we proposed an automated method that avoids the subjectivity, whose results can be replicated by studies that use the same database and that doesn't require a specific linguistic knowledge. Moreover, the method allows a quick comparison of a large number of languages. We applied our method to the Indo-European and Austronesian families, considering in both cases, fifty different languages. The resulting trees are similar to those of previous studies, but with some important differences in the position of few languages and subgroups. We believe that these differences carry new information on the structure of the tree and on the phylogenetic relationships within families.

cs.CL

The settlement of Madagascar: what dialects and languages can tell

The dialects of Madagascar belong to the Greater Barito East group of the Austronesian family and it is widely accepted that the Island was colonized by Indonesian sailors after a maritime trek which probably took place around 650 CE. The language most closely related to Malagasy dialects is Maanyan but also Malay is strongly related especially for what concerns navigation terms. Since the Maanyan Dayaks live along the Barito river in Kalimantan (Borneo) and they do not possess the necessary skill for long maritime navigation, probably they were brought as subordinates by Malay sailors. In a recent paper we compared 23 different Malagasy dialects in order to determine the time and the landing area of the first colonization. In this research we use new data and new methods to confirm that the landing took place on the south-east coast of the Island. Furthermore, we are able to state here that it is unlikely that there were multiple settlements and, therefore, colonization consisted in a single founding event. To reach our goal we find out the internal kinship relations among all the 23 Malagasy dialects and we also find out the different kinship degrees of the 23 dialects versus Malay and Maanyan. The method used is an automated version of the lexicostatistic approach. The data concerning Madagascar were collected by the author at the beginning of 2010 and consist of Swadesh lists of 200 items for 23 dialects covering all areas of the Island. The lists for Maanyan and Malay were obtained from published datasets integrated by author's interviews.

cs.CL

Extremely rare interbreeding events can explain Neanderthal DNA in modern humans

Considering the recent experimental discovery of Green et al that present day non-Africans have 1 to 4% of their nuclear DNA of Neanderthal origin, we propose here a model which is able to quantify the interbreeding events between Africans and Neanderthals at the time they coexisted in the Middle East. The model consists of a solvable system of deterministic ordinary differential equations containing as a stochastic ingredient a realization of the neutral Wright-Fisher drift process. By simulating the stochastic part of the model we are able to apply it to the interbreeding of African and Neanderthal subpopulations and estimate the only parameter of the model, which is the number of individuals per generation exchanged between subpopulations. Our results indicate that the amount of Neanderthal DNA in non-Africans can be explained with maximum probability by the exchange of a single pair of individuals between the subpopulations at each 77 generations, but larger exchange frequencies are also allowed with sizeable probability. The results are compatible with a total interbreeding population of order 10,000 individuals and with all living humans being descendents of Africans both for mitochondrial DNA and Y chromosome.

q-bio.PE

Phylogeny and geometry of languages from normalized Levenshtein distance

The idea that the distance among pairs of languages can be evaluated from lexical differences seems to have its roots in the work of the French explorer Dumont D'Urville. He collected comparative words lists of various languages during his voyages aboard the Astrolabe from 1826 to 1829 and, in his work about the geographical division of the Pacific, he proposed a method to measure the degree of relation between languages. The method used by the modern lexicostatistics, developed by Morris Swadesh in the 1950s, measures distances from the percentage of shared cognates, which are words with a common historical origin. The weak point of this method is that subjective judgment plays a relevant role. Recently, we have proposed a new automated method which is motivated by the analogy with genetics. The new approach avoids any subjectivity and results can be easily replicated by other scholars. The distance between two languages is defined by considering a renormalized Levenshtein distance between pair of words with the same meaning and averaging on the words contained in a list. The renormalization, which takes into account the length of the words, plays a crucial role, and no sensible results can be found without it. In this paper we give a short review of our automated method and we illustrate it by considering the cluster of Malagasy dialects. We show that it sheds new light on their kinship relation and also that it furnishes a lot of new information concerning the modalities of the settlement of Madagascar.

cs.CL

Magnetization Densities as Replica Parameters: The Dilute Ferromagnet

In this paper we compute exactly the ground state energy and entropy of the dilute ferromagnetic Ising model. The two thermodynamic quantities are also computed when a magnetic field with random locations is present. The result is reached in the replica approach frame by a class of replica order parameters introduced by Monasson. The strategy is first illustrated considering the SK model, for which we will show the complete equivalence with the standard replica approach. Then, we apply to the diluted ferromagnetic Ising model with a random located magnetic field, which is mapped into a Potts model. This formalism can be, in principle, applied to all random systems, and we believe that it could be of help in many other contexts.

cond-mat.dis-nn

Exact and Approximate Solutions for the Dilute Ising Model

The ground state energy and entropy of the dilute mean field Ising model is computed exactly by a single order parameter. An analogous exact solution is obtained in presence of a magnetic field with random locations. Results allow for a complete understanding of the geography of the associated random graph. In particular we give the size of the giant component (continent) and the number of isolated clusters of connected spins of all given size (islands). We also compute the average number of bonds per spin in the continent and in the islands. Then, we tackle the problem of computing the free energy of the dilute Ising model at strictly positive temperature. We are able to find out the exact solution in the paramagnetic region and exactly determine the phase transition line. In the ferromagnetic region we provide a solution in terms of an expansion with respect to a second parameter which can be made as accurate as necessary. All results are reached in the replica frame by a strategy which is not based on multi-overlaps.

cond-mat.dis-nn

Measures of lexical distance between languages

The idea of measuring distance between languages seems to have its roots in the work of the French explorer Dumont D'Urville \cite{Urv}. He collected comparative words lists of various languages during his voyages aboard the Astrolabe from 1826 to 1829 and, in his work about the geographical division of the Pacific, he proposed a method to measure the degree of relation among languages. The method used by modern glottochronology, developed by Morris Swadesh in the 1950s, measures distances from the percentage of shared cognates, which are words with a common historical origin. Recently, we proposed a new automated method which uses normalized Levenshtein distance among words with the same meaning and averages on the words contained in a list. Recently another group of scholars \cite{Bak, Hol} proposed a refined of our definition including a second normalization. In this paper we compare the information content of our definition with the refined version in order to decide which of the two can be applied with greater success to resolve relationships among languages.

cs.CL