SearcharxivSearch

arXiv subjects

Loet Leydesdorff

Publications and source records attributed to Loet Leydesdorff.

At least 19 recordsLinked to original sources

Quantitative Theory of Meaning. Application to Financial Markets. EUR/USD case study

The paper focuses on the link between information, investors' expectations and market price movement. EUR/USD market is examined from communication-theoretical perspective on the dynamics of information and meaning. We build upon the quantitative theory of meaning as a complement to the quantitative theory of information. Different groups of investors entertain different criteria to process information, so that the same information can be supplied with different meanings. Meanings shape investors' expectations which are revealed in market asset price movement. This dynamics can be captured by non-linear evolutionary equation. We use a computationally efficient technique of logistic Continuous Wavelet Transformation (CWT) to analyze EUR/USD market. The results reveal the latent EUR/USD trend structure which coincides with the model predicted time series indicating that proposed model can adequately describe some patterns of investors' behavior. Proposed methodology can be used to better understand and forecast future market assets' price movement.

cs.CY

A discussion of measuring the top-1 percent most-highly cited publications: Quality and impact of Chinese papers

The top 1 percent most highly cited articles are watched closely as the vanguards of the sciences. Using Web of Science data, one can find that China had overtaken the USA in the relative participation in the top 1 percent in 2019, after outcompeting the EU on this indicator in 2015. However, this finding contrasts with repeated reports of Western agencies that the quality of Chinese output in science is lagging other advanced nations, even as it has caught up in numbers of articles. The difference between the results presented here and the previous results depends mainly upon field normalizations, which classify source journals by discipline. Average citation rates of these subsets are commonly used as a baseline so that one can compare among disciplines. However, the expected value of the top 1 percent of a sample of N papers is N 100, ceteris paribus. Using the average citation rates as expected values, errors are introduced by using the mean of highly skewed distributions and a specious precision in the delineations of the subsets. Classifications can be used for the decomposition, but not for the normalization. When the data is thus decomposed, the USA ranks ahead of China in biomedical fields such as virology. Although the number of papers is smaller, China outperforms the US in the field of Business and Finance in the Social Sciences Citation Index when p is less than .05. Using percentile ranks, subsets other than indexing based classifications can be tested for the statistical significance of differences among them.

cs.DL

Are University Rankings Statistically Significant? A Comparison among Chinese Universities and with the USA

Purpose: We address the question of whether differences are statistically significant in the rankings of universities. We propose methods measuring the statistical significance among different universities and illustrate the results by empirical data. Design/methodology/approach: Based on z-testing and overlapping confidence intervals, and using data about 205 Chinese universities included in the Leiden Rankings 2020, we argue that three main groups of Chinese research universities can be distinguished. Findings: When the sample of 205 Chinese universities is merged with the 197 US universities included in Leiden Rankings 2020, the results similarly indicate three main groups: high, middle, low. Using this data (Leiden Rankings and Web-of-Science), the z-scores of the Chinese universities are significantly below those of the US universities albeit with some overlap. Research limitations: We show empirically that differences in ranking may be due to changes in the data, the models, or the modeling effects on the data. The scientometric groupings are not always stable when we use different methods. R&D policy implications: Differences among universities can be tested for their statistical significance. The statistics relativize the values of decimals in the rankings. One can operate with a scheme of low/middle/high in policy debates and leave the more fine-grained rankings of individual universities to operational management and local settings. Originality/value: In the discussion about the rankings of universities, the question of whether differences are statistically significant, is, in our opinion, insufficiently addressed.

cs.DL

Within-Journal Self-citations and the Pinski-Narin Influence Weights

The Journal Impact Factor (JIF) is linearly sensitive to self-citations because each self-citation adds to the numerator, whereas the denominator is not affected. Pinski & Narin (1976) derived the Influence Weight (IW) as an alternative to Garfield's JIF. Whereas the JIF is based on raw citation counts normalized by the number of publications, IWs are based on the eigenvectors in the matrix of aggregated journal-journal citations without a reference to size: the cited and citing sides are combined by a matrix approach. IWs emerge as a vector after recursive iteration of the normalized matrix. Before recursion, IW is a (vector-based) non-network indicator of impact, but after recursion (i.e. repeated improvement by iteration), IWs can be considered a network measure of prestige among the journals in the (sub)graph as a representation of a field of science. As a consequence (not intended by Pinski & Narin in 1976), the self-citations are integrated at the field level and no longer disturb the analysis as outliers. In our opinion, this is a very desirable property of a measure of quality or impact. As illustrations, we use data of journal citation matrices already studied in the literature, and also the complete set of data in the Journal Citation Reports 2017 (n = 11,579 journals). The values of IWs are sometimes counter-intuitive and difficult to interpret. Furthermore, iterations do not always converge. Routines for the computation of IWs are made available at http://www.leydesdorff.net/iw.

cs.DL

Does the $h_α$ index reinforce the Matthew effect in science? Agent-based simulations using Stata and R

Recently, Hirsch (2019a) proposed a new variant of the h index called the $h_α$ index. He formulated as follows: "we define the $h_α$ index of a scientist as the number of papers in the h-core of the scientist (i.e. the set of papers that contribute to the h-index of the scientist) where this scientist is the $α$-author" (p. 673). The $h_α$ index was criticized by Leydesdorff, Bornmann, and Opthof (2019). One of their most important points is that the index reinforces the Matthew effect in science. We address this point in the current study using a recently developed Stata command (h_index) and R package (hindex), which can be used to simulate h index and $h_α$index applications in research evaluation. The user can investigate under which conditions $h_α$ reinforces the Matthew effect. The results of our study confirm what Leydesdorff et al. (2019) expected: the $h_α$ index reinforces the Matthew effect. This effect can be intensified if strategic behavior of the publishing scientists and cumulative advantage effects are additionally considered in the simulation.

cs.DL

The Integrated Impact Indicator (I3) Revisited: A Non-Parametric Alternative to the Journal Impact Factor

We propose the I3* indicator as a non-parametric alternative to the Journal Impact Factor (JIF) and h-index. We apply I3* to more than 10,000 journals. The results can be compared with other journal metrics. I3* is a promising variant within the general scheme of non-parametric indicators I3 introduced previously: it provides a single metric which correlates with both impact in terms of citations (c) and output in terms of publications (p). We argue for weighting using four percentile classes: the top-1% and top-10% as excellence indicators; the top-50% and bottom-50% as output indicators. Like the h-index, which also incorporates both c and p, I3*-values are size-dependent; however, division of I3* by the number of publications (I3*/N) provides a size-independent indicator which correlates strongly with the two- and five-year Journal Impact Factors (JIF2 and JIF5). Unlike the h-index, I3* correlates significantly with both the total number of citations and publications. The values of I3* and I3*/N can be statistically tested against the expectation or against one another using chi-square tests or effect sizes. A template (in Excel) is provided online for relevant tests.

cs.DL

Does the public discuss other topics on climate change than researchers? A comparison of explorative networks based on author keywords and hashtags

Twitter accounts have already been used in many scientometric studies, but the meaningfulness of the data for societal impact measurements in research evaluation has been questioned. Earlier research focused on social media counts and neglected the interactive nature of the data. We explore a new network approach based on Twitter data in which we compare author keywords to hashtags as indicators of topics. We analyze the topics of tweeted publications and compare them with the topics of all publications (tweeted and not tweeted). Our exploratory study is based on a comprehensive publication set of climate change research. We are interested in whether Twitter data are able to reveal topics of public discussions which can be separated from research-focused topics. We find that the most tweeted topics regarding climate change research focus on the consequences of climate change for humans. Twitter users are interested in climate change publications which forecast effects of a changing climate on the environment and to adaptation, mitigation and management issues rather than in the methodology of climate-change research and causes of climate change. Our results indicate that publications using scientific jargon are less likely to be tweeted than publications using more general keywords. Twitter networks seem to be able to visualize public discussions about specific topics.

cs.DL

How well does I3 perform for impact measurement compared to other bibliometric indicators? The convergent validity of several (field-normalized) indicators

Recently, the integrated impact indicator (I3) indicator was introduced where citations are weighted in accordance with the percentile rank class of each publication in a set of publications. I3 can also be used as a field-normalized indicator. Field-normalization is common practice in bibliometrics, especially when institutions and countries are compared. Publication and citation practices are so different among fields that citation impact is normalized for cross-field comparisons. In this study, we test the ability of the indicator to discriminate between quality levels of papers as defined by Faculty members at F1000Prime. F1000Prime is a post-publication peer review system for assessing papers in the biomedical area. Thus, we test the convergent validity of I3 (in this study, we test I3/N - the size-independent variant of I3 where I3 is divided by the number of papers) using assessments by peers as baseline and compare its validity with several other (field-normalized) indicators: the mean-normalized citation score (MNCS), relative-citation ratio (RCR), citation score normalized by cited references (CSNCR), characteristic scores and scales (CSS), source-normalized citation score (SNCS), citation percentile, and proportion of papers which belong to the x% most frequently cited papers (PPtop x%). The results show that the PPtop 1% indicator discriminates best among different quality levels. I3 performs similar as (slightly better than) most of the other field-normalized indicators. Thus, the results point out that the indicator could be a valuable alternative to other indicators in bibliometrics.

cs.DL

Which are the influential publications in the Web of Science subject categories over a long period of time? CRExplorer software used for big-data analyses in bibliometrics

What are the landmark papers in scientific disciplines? On whose shoulders does research in these fields stand? Which papers are indispensable for scientific progress? These are typical questions which are not only of interest for researchers (who frequently know the answers - or guess to know them), but also for the interested general public. Citation counts can be used to identify very useful papers, since they reflect the wisdom of the crowd; in this case, the scientists using the published results for their own research. In this study, we identified with recently developed methods for the program CRExplorer landmark publications in nearly all Web of Science subject categories (WoSSCs). These are publications which belong more frequently than other publications across the citing years to the top-per mill in their subject category. The results for three subject categories "Information Science and Library Science", "Computer Science, Information Systems", and "Computer Science, Software Engineering" are exemplarily discussed in more detail. The results for the other WoSSCs can be found online at http://crexplorer.net.

cs.DL

Synergy in the Knowledge Base of U.S. Innovation Systems at National, State, and Regional Levels: The Contributions of High-Tech Manufacturing and Knowledge-Intensive Services

Using information theory, we measure innovation systemness as synergy among size-classes, zip-codes, and technological classes (NACE-codes) for 8.5 million American companies. The synergy at the national level is decomposed at the level of states, Core-Based Statistical Areas (CBSA), and Combined Statistical Areas (CSA). We zoom in to the state of California and in more detail to Silicon Valley. Our results do not support the assumption of a national system of innovations in the U.S.A. Innovation systems appear to operate at the level of the states; the CBSA are too small, so that systemness spills across their borders. Decomposition of the sample in terms of high-tech manufacturing (HTM), medium-high-tech manufacturing (MHTM), knowledge-intensive services (KIS), and high-tech services (HTKIS) does not change this pattern, but refines it. The East Coast -- New Jersey, Boston, and New York -- and California are the major players, with Texas a third one in the case of HTKIS. Chicago and industrial centers in the Midwest also contribute synergy. Within California, Los Angeles contributes synergy in the sectors of manufacturing, the San Francisco area in KIS. Knowledge-intensive services in Silicon Valley and the Bay area -- a CSA composed of seven CBSA -- spill over to other regions and even globally.

cs.DL

Information, Meaning, and Intellectual Organization in Networks of Inter-Human Communication

The Shannon-Weaver model of linear information transmission is extended with two loops potentially generating redundancies: (i) meaning is provided locally to the information from the perspective of hindsight, and (ii) meanings can be codified differently and then refer to other horizons of meaning. Thus, three layers are distinguished: variations in the communications, historical organization at each moment of time, and evolutionary self-organization of the codes of communication over time. Furthermore, the codes of communication can functionally be different and then the system is both horizontally and vertically differentiated. All these subdynamics operate in parallel and necessarily generate uncertainty. However, meaningful information can be considered as the specific selection of a signal from the noise; the codes of communication are social constructs that can generate redundancy by giving different meanings to the same information. Reflexively, one can translate among codes in more elaborate discourses. The second (instantiating) layer can be operationalized in terms of semantic maps using the vector space model; the third in terms of mutual redundancy among the latent dimensions of the vector space. Using Blaise Cronin's œuvre, the different operations of the three layers are demonstrated empirically.

cs.DL

hα: The Scientist as Chimpanzee or Bonobo

In a recent paper, Hirsch (2018) proposes to attribute the credit for a co-authored paper to the α-author--the author with the highest h-index--regardless of his or her actual contribution, effectively reducing the role of the other co-authors to zero. The indicator hα inherits most of the disadvantages of the h-index from which it is derived, but adds the normative element of reinforcing the Matthew effect in science. Using an example, we show that hα can be extremely unstable. The empirical attribution of credit among co-authors is not captured by abstract models such as h, h_bar , or hα.

cs.DL

Statistical Significance and Effect Sizes of Differences among Research Universities at the Level of Nations and Worldwide based on the Leiden Rankings

The Leiden Rankings can be used for grouping research universities by considering universities which are not statistically significantly different as homogeneous sets. The groups and intergroup relations can be analyzed and visualized using tools from network analysis. Using the so-called "excellence indicator" PPtop-10%--the proportion of the top-10% most-highly-cited papers assigned to a university--we pursue a classification using (i) overlapping stability intervals, (ii) statistical-significance tests, and (iii) effect sizes of differences among 902 universities in 54 countries; we focus on the UK, Germany, Brazil, and the USA as national examples. Although the groupings remain largely the same using different statistical significance levels or overlapping stability intervals, these classifications are uncorrelated with those based on effect sizes. Effect sizes for the differences between universities are small (w <.2). The more detailed analysis of universities at the country level suggests that distinctions beyond three or perhaps four groups of universities (high, middle, low) may not be meaningful. Given similar institutional incentives, isomorphism within each eco-system of universities should not be underestimated. Our results suggest that networks based on overlapping stability intervals can provide a first impression of the relevant groupings among universities. However, the clusters are not well-defined divisions between groups of universities.

cs.DL

Interdisciplinarity as Diversity in Citation Patterns among Journals: Rao-Stirling Diversity, Relative Variety, and the Gini coefficient

Questions of definition and measurement continue to constrain a consensus on the measurement of interdisciplinarity. Using Rao-Stirling (RS) Diversity produces sometimes anomalous results. We argue that these unexpected outcomes can be related to the use of "dual-concept diversity" which combines "variety" and "balance" in the definitions (ex ante). We propose to modify RS Diversity into a new indicator (DIV) which operationalizes variety, balance, and disparity independently and then combines them ex post. "Balance" can be measured using the Gini coefficient. We apply DIV to the aggregated citation patterns of 11,487 journals covered by the Journal Citation Reports 2016 of the Science Citation Index and the Social Sciences Citation Index as an empirical domain and, in more detail, to the citation patterns of 85 journals assigned to the Web-of-Science category "information science & library science" in both the cited and citing directions. We compare the results of the indicators and show that DIV provides improved results in terms of distinguishing between interdisciplinary knowledge integration (citing) versus knowledge diffusion (cited). The new diversity indicator and RS diversity measure different features. A routine for the measurement of the various operationalizations of diversity (in any data matrix) is made available online.

cs.DL

Revisiting Relative Indicators and Provisional Truths

Following discussions in 2010 and 2011, scientometric evaluators have increasingly abandoned relative indicators in favor of comparing observed with expected citation ratios. The latter method provides parameters with error values allowing for the statistical testing of differences in citation scores. A further step would be to proceed to non-parametric statistics (e.g., the top-10%) given the extreme skewness (non-normality) of the citation distributions. In response to a plea for returning to relative indicators in the previous issue of this newsletter, we argue in favor of further progress in the development of citation impact indicators.

cs.DL

The negative effects of citing with a national orientation in terms of recognition: national and international citations in natural-sciences papers from Germany, the Netherlands, and the UK

Nations can be distinguished in terms of whether domestic or international research is cited. We analyzed the research output in natural sciences of three leading European research economies (Germany, the Netherlands, and the UK) and ask where their researchers look for the knowledge that underpins their most highly-cited papers. Is one internationally oriented or is citation limited to national resources? Do the citation patterns reflect a growing differentiation between the domestic and international research enterprise? To evaluate change over time, we include natural-sciences papers published in the countries from three publication years: 2004, 2009, and 2014. The results show that articles co-authored by researchers from Germany or the Netherlands are less likely to be among the globally most highly-cited articles if they also cite "domestic" research (i.e. research authored by authors from the same country). To put this another way, less well-cited research is more likely to stand on domestic shoulders and research that becomes more highly-cited is more likely to stand on international shoulders. A possible reason for the results is that researchers "over-cite" the papers from their own country - lacking the focus on quality in citing. However, these differences between domestic and international shoulders are not visible for the UK.

cs.DL

Diversity and Interdisciplinarity: How Can One Distinguish and Recombine Disparity, Variety, and Balance?

The dilemma which remained unsolved using Rao-Stirling diversity, namely of how variety and balance can be combined into "dual concept diversity" (Stirling, 1998, pp. 48f.) can be clarified by using Nijssen et al.'s (1998) argument that the Gini coefficient is a perfect indicator of balance. However, the Gini coefficient is not an indicator of variety; this latter term can be operationalized independently as relative variety. The three components of diversity--variety, balance, and disparity--can thus be clearly distinguished and independently operationalized as measures varying between zero and one. The new diversity indicator ranges with more resolving power in the empirical case.

cs.DL

Topic Modelling of Empirical Text Corpora: Validity, Reliability, and Reproducibility in Comparison to Semantic Maps

Using the 6,638 case descriptions of societal impact submitted for evaluation in the Research Excellence Framework (REF 2014), we replicate the topic model (Latent Dirichlet Allocation or LDA) made in this context and compare the results with factor-analytic results using a traditional word-document matrix (Principal Component Analysis or PCA). Removing a small fraction of documents from the sample, for example, has on average a much larger impact on LDA than on PCA-based models to the extent that the largest distortion in the case of PCA has less effect than the smallest distortion of LDA-based models. In terms of semantic coherence, however, LDA models outperform PCA-based models. The topic models inform us about the statistical properties of the document sets under study, but the results are statistical and should not be used for a semantic interpretation - for example, in grant selections and micro-decision making, or scholarly work-without follow-up using domain-specific semantic maps.

cs.CL