Searcharxiv⌕ Search

arXiv subjects

Jaroslaw Kwapien

Publications and source records attributed to Jaroslaw Kwapien.

At least 19 recordsLinked to original sources

Universal versus system-specific features of punctuation usage patterns in~major Western~languages

The celebrated proverb that "speech is silver, silence is golden" has a long multinational history and multiple specific meanings. In written texts punctuation can in fact be considered one of its manifestations. Indeed, the virtue of effectively speaking and writing involves - often decisively - the capacity to apply the properly placed breaks. In the present study, based on a large corpus of world-famous and representative literary texts in seven major Western languages, it is shown that the distribution of intervals between consecutive punctuation marks in almost all texts can universally be characterised by only two parameters of the discrete Weibull distribution which can be given an intuitive interpretation in terms of the so-called hazard function. The values of these two parameters tend to be language-specific, however, and even appear to navigate translations. The properties of the computed hazard functions indicate that among the studied languages, English turns out to be the least constrained by the necessity to place a consecutive punctuation mark to partition a sequence of words. This may suggest that when compared to other studied languages, English is more flexible, in the sense of allowing longer uninterrupted sequences of words. Spanish reveals similar tendency to only a bit lesser extent.

cs.CL↗

Hierarchical organization of H. Eugene Stanley scientific collaboration community in weighted network representation

By mapping the most advanced elements of the contemporary social interactions, the world scientific collaboration network develops an extremely involved and heterogeneous organization. Selected characteristics of this heterogeneity are studied here and identified by focusing on the scientific collaboration community of H. Eugene Stanley - one of the most prolific world scholars at the present time. Based on the Web of Science records as of March 28, 2016, several variants of networks are constructed. It is found that the Stanley #1 network - this in analogy to the Erdős # - develops a largely consistent hierarchical organization and Stanley himself obeys rules of the same hierarchy. However, this is seen exclusively in the weighted network representation. When such a weighted network is evolving, an existing relevant model indicates that the spread of weight gets stimulation to the multiplicative bursts over the neighbouring nodes, which leads to a balanced growth of interconnections among them. While not exclusive to Stanley, such a behaviour is not a rule, however. Networks of other outstanding scholars studied here more often develop a star-like form and the central hubs constitute the outliers. This study is complemented by a spectral analysis of the normalised Laplacian matrices derived from the weighted variants of the corresponding networks and, among others, it points to the efficiency of such a procedure for identifying the component communities and relations among them in the complex weighted networks.

physics.soc-ph↗

Minimum spanning tree filtering of correlations for varying time scales and size of fluctuations

Based on a recently proposed $q$-dependent detrended cross-correlation coefficient $ρ_q$, we generalize the concept of minimum spanning tree (MST) by introducing a family of $q$-dependent minimum spanning trees ($q$MST) that are selective to cross-correlations between different fluctuation amplitudes and different time scales. They inherit this ability directly from the coefficients $ρ_q$ that are processed here to construct a distance matrix. Conventional MST with detrending corresponds in this context to $q=2$. We apply the $q$MSTs to sample empirical data from the stock market and discuss the results. We show that the $q$MST graphs can complement $ρ_q$ in disentangling correlations that cannot be observed by the MST graphs based on $ρ_{\rm DCCA}$ and, therefore, they can be useful in many areas where the multivariate cross-correlations are of interest. We apply our method to data from the stock market and obtain more information about correlation structure of the data than by using $q=2$ only. We show that two sets of signals that differ from each other statistically can give comparable trees for $q=2$, while only by using the trees for $q \ne 2$ we become able to distinguish between these sets. We also show that a family of $q$MSTs for a range of $q$ express the diversity of correlations in a manner resembling the multifractal analysis, where one computes a spectrum of the generalized fractal dimensions, the generalized Hurst exponents, or the multifractal singularity spectra: the more diverse the correlations are, the more variable the tree topology is for different $q$s. Our analysis exhibits that the stocks belonging to the same or similar industrial sectors are correlated via the fluctuations of moderate amplitudes, while the largest fluctuations often happen to synchronize in those stocks that do not necessarily belong to the same industry.

q-fin.ST↗

In narrative texts punctuation marks obey the same statistics as words

From a grammar point of view, the role of punctuation marks in a sentence is formally defined and well understood. In semantic analysis punctuation plays also a crucial role as a method of avoiding ambiguity of the meaning. A different situation can be observed in the statistical analyses of language samples, where the decision on whether the punctuation marks should be considered or should be neglected is seen rather as arbitrary and at present it belongs to a researcher's preference. An objective of this work is to shed some light onto this problem by providing us with an answer to the question whether the punctuation marks may be treated as ordinary words and whether they should be included in any analysis of the word co-occurences. We already know from our previous study (S.~Drożdż {\it et al.}, Inf. Sci. 331 (2016) 32-44) that full stops that determine the length of sentences are the main carrier of long-range correlations. Now we extend that study and analyze statistical properties of the most common punctuation marks in a few Indo-European languages, investigate their frequencies, and locate them accordingly in the Zipf rank-frequency plots as well as study their role in the word-adjacency networks. We show that, from a statistical viewpoint, the punctuation marks reveal properties that are qualitatively similar to the properties of the most frequent words like articles, conjunctions, pronouns, and prepositions. This refers to both the Zipfian analysis and the network analysis. By adding the punctuation marks to the Zipf plots, we also show that these plots that are normally described by the Zipf-Mandelbrot distribution largely restore the power-law Zipfian behaviour for the most frequent items.

cs.CL↗

Inferring cultural regions from correlation networks of given baby names

We report investigations on the statistical characteristics of the baby names given between 1910 and 2010 in the United States of America. For each year, the 100 most frequent names in the USA are sorted out. For these names, the correlations between the names profiles are calculated for all pairs of states (minus Hawaii and Alaska). The correlations are used to form a weighted network which is found to vary mildly in time. In fact, the structure of communities in the network remains quite stable till about 1980. The goal is that the calculated structure approximately reproduces the usually accepted geopolitical regions: the North East, the South, and the "Midwest + West" as the third one. Furthermore, the dataset reveals that the name distribution satisfies the Zipf law, separately for each state and each year, i.e. the name frequency $f\propto r^{-α}$, where r is the name rank. Between 1920 and 1980, the exponent alpha is the largest one for the set of states classified as 'the South', but the smallest one for the set of states classified as "Midwest + West". Our interpretation is that the pool of selected names was quite narrow in the Southern states. The data is compared with some related statistics of names in Belgium, a country also with different regions, but having quite a different scale than the USA. There, the Zipf exponent is low for young people and for the Brussels citizens.

physics.soc-ph↗

Detrended fluctuation analysis made flexible to detect range of cross-correlated fluctuations

The detrended cross-correlation coefficient $ρ_{\rm DCCA}$ has recently been proposed to quantify the strength of cross-correlations on different temporal scales in bivariate, non-stationary time series. It is based on the detrended cross-correlation and detrended fluctuation analyses (DCCA and DFA, respectively) and can be viewed as an analogue of the Pearson coefficient in the case of the fluctuation analysis. The coefficient $ρ_{\rm DCCA}$ works well in many practical situations but by construction its applicability is limited to detection of whether two signals are generally cross-correlated, without possibility to obtain information on the amplitude of fluctuations that are responsible for those cross-correlations. In order to introduce some related flexibility, here we propose an extension of $ρ_{\rm DCCA}$ that exploits the multifractal versions of DFA and DCCA: MFDFA and MFCCA, respectively. The resulting new coefficient $ρ_q$ not only is able to quantify the strength of correlations, but also it allows one to identify the range of detrended fluctuation amplitudes that are correlated in two signals under study. We show how the coefficient $ρ_q$ works in practical situations by applying it to stochastic time series representing processes with long memory: autoregressive and multiplicative ones. Such processes are often used to model signals recorded from complex systems and complex physical phenomena like turbulence, so we are convinced that this new measure can successfully be applied in time series analysis. In particular, we present an example of such application to highly complex empirical data from financial markets. The present formulation can straightforwardly be extended to multivariate data in terms of the $q$-dependent counterpart of the correlation matrices and then to the network representation.

physics.data-an↗

Detrended cross-correlations between returns, volatility, trading activity, and volume traded for the stock market companies

We consider a few quantities that characterize trading on a stock market in a fixed time interval: logarithmic returns, volatility, trading activity (i.e., the number of transactions), and volume traded. We search for the power-law cross-correlations among these quantities aggregated over different time units from 1 min to 10 min. Our study is based on empirical data from the American stock market consisting of tick-by-tick recordings of 31 stocks listed in Dow Jones Industrial Average during the years 2008-2011. Since all the considered quantities except the returns show strong daily patterns related to the variable trading activity in different parts of a day, which are the best evident in the autocorrelation function, we remove these patterns by detrending before we proceed further with our study. We apply the multifractal detrended cross-correlation analysis with sign preserving (MFCCA) and show that the strongest power-law cross-correlations exist between trading activity and volume traded, while the weakest ones exist (or even do not exist) between the returns and the remaining quantities. We also show that the strongest cross-correlations are carried by those parts of the signals that are characterized by large and medium variance. Our observation that the most convincing power-law cross-correlations occur between trading activity and volume traded reveals the existence of strong fractal-like coupling between these quantities.

q-fin.ST↗

Modeling the average shortest path length in growth of word-adjacency networks

We investigate properties of evolving linguistic networks defined by the word-adjacency relation. Such networks belong to the category of networks with accelerated growth but their shortest path length appears to reveal the network size dependence of different functional form than the ones known so far. We thus compare the networks created from literary texts with their artificial substitutes based on different variants of the Dorogovtsev-Mendes model and observe that none of them is able to properly simulate the novel asymptotics of the shortest path length. Then, we identify the local chain-like linear growth induced by grammar and style as a missing element in this model and extend it by incorporating such effects. It is in this way that a satisfactory agreement with the empirical result is obtained.

cs.CL↗

Stock returns versus trading volume: is the correspondence more general?

This paper presents a quantitative analysis of the relationship between the stock market returns and corresponding trading volumes using high- frequency data from the Polish stock market. First, for stocks that were traded for suffciently long period of time, we study the return and volume distributions and identify their consistency with the power-law functions. We find that, for majority of stocks, the scaling exponents of both distri- butions are systematically related by about a factor of 2 with the ones for the returns being larger. Second, we study the empirical price impact of trades of a given volume and find that this impact can be well described by a square-root dependence: r(V) V^(1/2). We conclude that the prop- erties of data from the Polish market resemble those reported in literature concerning certain mature markets.

q-fin.ST↗

Complex network analysis of literary and scientific texts

We present results from our quantitative study of statistical and network properties of literary and scientific texts written in two languages: English and Polish. We show that Polish texts are described by the Zipf law with the scaling exponent smaller than the one for the English language. We also show that the scientific texts are typically characterized by the rank-frequency plots with relatively short range of power-law behavior as compared to the literary texts. We then transform the texts into their word-adjacency network representations and find another difference between the languages. For the majority of the literary texts in both languages, the corresponding networks revealed the scale-free structure, while this was not always the case for the scientific texts. However, all the network representations of texts were hierarchical. We do not observe any qualitative and quantitative difference between the languages. However, if we look at other network statistics like the clustering coefficient and the average shortest path length, the English texts occur to possess more clustered structure than do the Polish ones. This result was attributed to differences in grammar of both languages, which was also indicated in the Zipf plots. All the texts, however, show network structure that differs from any of the Watts-Strogatz, the Barabasi-Albert, and the Erdos-Renyi architectures.

physics.soc-ph↗

Asymmetric random matrices: What do we need them for?

Complex systems are typically represented by large ensembles of observations. Correlation matrices provide an efficient formal framework to extract information from such multivariate ensembles and identify in a quantifiable way patterns of activity that are reproducible with statistically significant frequency compared to a reference chance probability, usually provided by random matrices as fundamental reference. The character of the problem and especially the symmetries involved must guide the choice of random matrices to be used for the definition of a baseline reference. For standard correlation matrices this is the Wishart ensemble of symmetric random matrices. The real world complexity however often shows asymmetric information flows and therefore more general correlation matrices are required to adequately capture the asymmetry. Here we first summarize the relevant theoretical concepts. We then present some examples of human brain activity where asymmetric time-lagged correlations are evident and hence highlight the need for further theoretical developments.

physics.data-an↗

The foreign exchange market: return distributions, multifractality, anomalous multifractality and Epps effect

We present a systematic study of various statistical characteristics of high-frequency returns from the foreign exchange market. This study is based on six exchange rates forming two triangles: EUR-GBP-USD and GBP-CHF-JPY. It is shown that the exchange rate return fluctuations for all the pairs considered are well described by the nonextensive statistics in terms of q-Gaussians. There exist some small quantitative variations in the nonextensivity q-parameter values for different exchange rates and this can be related to the importance of a given exchange rate in the world's currency trade. Temporal correlations organize the series of returns such that they develop the multifractal characteristics for all the exchange rates with a varying degree of symmetry of the singularity spectrum f(alpha) however. The most symmetric spectrum is identified for the GBP/USD. We also form time series of triangular residual returns and find that the distributions of their fluctuations develop disproportionately heavier tails as compared to small fluctuations which excludes description in terms of q-Gaussians. The multifractal characteristics for these residual returns reveal such anomalous properties like negative singularity exponents and even negative singularity spectra. Such anomalous multifractal measures have so far been considered in the literature in connection with the diffusion limited aggregation and with turbulence. We find that market inefficiency on short time scales leads to the occurrence of the Epps effect on much longer time scales. Although the currency market is much more liquid than the stock markets and it has much larger transaction frequency, the building-up of correlations takes up to several hours - time that does not differ much from what is observed in the stock markets. This may suggest that non-synchronicity of transactions is not the unique source of the observed effect.

q-fin.ST↗

Linguistic complexity: English vs. Polish, text vs. corpus

We analyze the rank-frequency distributions of words in selected English and Polish texts. We show that for the lemmatized (basic) word forms the scale-invariant regime breaks after about two decades, while it might be consistent for the whole range of ranks for the inflected word forms. We also find that for a corpus consisting of texts written by different authors the basic scale-invariant regime is broken more strongly than in the case of comparable corpus consisting of texts written by the same author. Similarly, for a corpus consisting of texts translated into Polish from other languages the scale-invariant regime is broken more strongly than for a comparable corpus of native Polish texts. Moreover, we find that if the words are tagged with their proper part of speech, only verbs show rank-frequency distribution that is almost scale-invariant.

cs.CL↗

Quantitative features of multifractal subtleties in time series

Based on the Multifractal Detrended Fluctuation Analysis (MFDFA) and on the Wavelet Transform Modulus Maxima (WTMM) methods we investigate the origin of multifractality in the time series. Series fluctuating according to a qGaussian distribution, both uncorrelated and correlated in time, are used. For the uncorrelated series at the border (q=5/3) between the Gaussian and the Levy basins of attraction asymptotically we find a phase-like transition between monofractal and bifractal characteristics. This indicates that these may solely be the specific nonlinear temporal correlations that organize the series into a genuine multifractal hierarchy. For analyzing various features of multifractality due to such correlations, we use the model series generated from the binomial cascade as well as empirical series. Then, within the temporal ranges of well developed power-law correlations we find a fast convergence in all multifractal measures. Besides of its practical significance this fact may reflect another manifestation of a conjectured q-generalized Central Limit Theorem.

physics.data-an↗

Sign and amplitude representation of the forex networks

We decompose the exchange rates returns of 41 currencies (incl. gold) into their sign and amplitude components. Then we group together all exchange rates with a common base currency, construct Minimal Spanning Trees for each group independently, and analyze properties of these trees. We show that both the sign and the amplitude time series have similar correlation properties as far as the core network structure is concerned. There exist however interesting peripheral differences that may open a new perspective to view the Forex dynamics.

q-fin.ST↗

Analysis of a network structure of the foreign currency exchange market

We analyze structure of the world foreign currency exchange (FX) market viewed as a network of interacting currencies. We analyze daily time series of FX data for a set of 63 currencies, including gold, silver and platinum. We group together all the exchange rates with a common base currency and study each group separately. By applying the methods of filtered correlation matrix we identify clusters of closely related currencies. The clusters are formed typically according to the economical and geographical factors. We also study topology of weighted minimal spanning trees for different network representations (i.e., for different base currencies) and find that in a majority of representations the network has a hierarchical scale-free structure. In addition, we analyze the temporal evolution of the network and detect that its structure is not stable over time. A medium-term trend can be identified which affects the USD node by decreasing its centrality. Our analysis shows also an increasing role of euro in the world's currency market.

q-fin.ST↗

Structure and evolution of the foreign exchange networks

We investigate topology and temporal evolution of the foreign currency exchange market viewed from a weighted network perspective. Based on exchange rates for a set of 46 currencies (including precious metals), we construct different representations of the FX network depending on a choice of the base currency. Our results show that the network structure is not stable in time, but there are main clusters of currencies, which persist for a long period of time despite the fact that their size and content are variable. We find a long-term trend in the network's evolution which affects the USD and EUR nodes. In all the network representations, the USD node gradually loses its centrality, while, on contrary, the EUR node has become slightly more central than it used to be in its early years. Despite this directional trend, the overall evolution of the network is noisy.

q-fin.ST↗

Approaching the linguistic complexity

We analyze the rank-frequency distributions of words in selected English and Polish texts. We compare scaling properties of these distributions in both languages. We also study a few small corpora of Polish literary texts and find that for a corpus consisting of texts written by different authors the basic scaling regime is broken more strongly than in the case of comparable corpus consisting of texts written by the same author. Similarly, for a corpus consisting of texts translated into Polish from other languages the scaling regime is broken more strongly than for a comparable corpus of native Polish texts. Moreover, based on the British National Corpus, we consider the rank-frequency distributions of the grammatically basic forms of words (lemmas) tagged with their proper part of speech. We find that these distributions do not scale if each part of speech is analyzed separately. The only part of speech that independently develops a trace of scaling is verbs.

cs.CL↗