SearcharxivSearch

arXiv subjects

Marcel Ausloos

Publications and source records attributed to Marcel Ausloos.

At least 19 recordsLinked to original sources

On (Newcomb-)Benford's law: a tale of two papers and of their disproportionate citations. How citation counts can become biased

The first digit (FD) phenomenon i.e., the significant digits of numbers in large data are often distributed according to a logarithmically decreasing function was first reported by S. Newcomb and then many decades later independently by F. Benford. After its century long neglect the last three decades have seen huge growth in the number of relevant publications. However, notwithstanding the rising popularity the two independent proponents of the phenomenon are not equally acknowledged an indication of which is disproportionate number of citations accumulated by Newcomb (1881) and Benford (1938). In the present study we use citation analysis to show that the formalization of the eponym Benford's law, a name questionable itself for overlooking Newcomb's contribution, by Raimi (1976) had a strong adverse effect on the future citations of Newcomb (1881). Furthermore, we identify the papers published over various decades of the developmental history of the FD phenomenon, which latter turned out to be amongst the most cited ones in the field. We find that lack of its consideration, intentional or occasionally out of ignorance for referencing by the prominent papers, is responsible for a far lesser number of citations of Newcomb (1881) in comparison to Benford (1938).

physics.soc-ph

Forsaking your own: unveiling the delayed recognition of Garfield's work on the "delayed recognition" phenomenon

Delayed recognition (DR) implies that the full scholarly potential of certain scientific papers is recognized belatedly many years after their publication. Such papers are initially barely cited (sleep), and then suddenly, sometime in the future, their citation numbers burst (are awakened). After van Raan (2004a) called them "Sleeping Beauties" the DR phenomenon has drawn considerable attention. However, long before van Raan (2004a) Garfield studied the phenomenon in a series of articles from 1970 up to year 2004. In the present study we ask the pertinent question; Has the phenomenon of DR itself suffered the delayed recognition? In search of an answer we study the citation history of the Garfield (1980a) paper in which Garfield addressed DR directly for the first time. We find that the paper hardly received the attention befitting the Garfield's stature as an information scientist. Specifically, the paper received a meager of 10 citations up to the publication year of van Raan (2004a) and was then, in 2007, feebly awakened from its deep sleep of twenty-eight years receiving 20 citations in next four years; up to 2010. Being the undisputed giant of information science that even Garfield's paper on DR can suffer DR is hardly anticipated.

physics.soc-ph

Note on pre-taxation reported data by UK FTSE-listed companies. A search for Benford's laws compatibility

Pre-taxation analysis plays a crucial role in ensuring the fairness of public revenue collection. It can also serve as a tool to reduce the risk of tax avoidance, one of the UK government's concerns. Our report utilises pre-tax income ($PI$) and total assets ($TA$) data from 567 companies listed on the FTSE All-Share index, gathered from the Refinitiv EIKON database, covering 14 years, i.e., the period from 2009 to 2022. We also derive the $PI/TA$ ratio, and distinguish between positive and negative $PI$ cases. We test the conformity of such data to Benford's Laws,- specifically studying the first significant digit ($Fd$), the second significant digit ($Sd$), and the first and second significant digits ($FSd$). We use and justify two pertinent tests, the $\chi^2$ and the Mean Absolute Deviation (MAD). We find that both tests are not leading to conclusions in complete agreement with each other, - in particular the MAD test entirely rejects the Benford's Laws conformity of the reported financial data. From the mere accounting point of view, we conclude that the findings not only cast some doubt on the reported financial data, but also suggest that many more investigations be envisaged on closely related matters. On the other hand, the study of a ratio, like $PI/TA$, of variables which are (or not) Benford's Laws compliant add to the literature debating whether such indirect variables should (or not) be Benford's Laws compliant.

q-fin.ST

Environmental Performance, Financial Constraint and Tax Avoidance Practices: Insights from FTSE All-Share Companies

Through its initiative known as the Climate Change Act (2008), the Government of the United Kingdom encourages corporations to enhance their environmental performance with the significant aim of reducing targeted greenhouse gas emissions by the year 2050. Previous research has predominantly assessed this encouragement favourably, suggesting that improved environmental performance bolsters governmental efforts to protect the environment and fosters commendable corporate governance practices among companies. Studies indicate that organisations exhibiting strong corporate social responsibility (CSR), environmental, social, and governance (ESG) criteria, or high levels of environmental performance often engage in lower occurrences of tax avoidance. However, our findings suggest that an increase in environmental performance may paradoxically lead to a rise in tax avoidance activities. Using a sample of 567 firms listed on the FTSE All Share from 2014 to 2022, our study finds that firms associated with higher environmental performance are more likely to avoid taxation. The study further documents that the effect is more pronounced for firms facing financial constraints. Entropy balancing, propensity score matching analysis, the instrumental variable method, and the Heckman test are employed in our study to address potential endogeneity concerns. Collectively, the findings of our study suggest that better environmental performance helps explain the variation in firms tax avoidance practices.

q-fin.GN

Hybrid Galam--Bass Model for Technology Innovation

This work proposes a hybrid model that combines the Galam model of opinion dynamics with the Bass diffusion model used in technology adoption on Barabasi-Albert complex networks. The main idea is to advance a version of the Bass model that can suitably describe an opinion formation context while introducing irreversible transitions from group B (opponents) to group A (supporters). Moreover, we extend the model to take into account the presence of a charismatic competitor, which fosters conversion back to the old technology. The approach is different from the introduction of a mean field due to the interactions driven by the network structure. Additionally, we introduce the Kolmogorov-Sinai entropy to quantify the system's unpredictability and information loss over time. The results show an increase in the regularity of the trajectories as the preferential attachment parameter increases.

physics.soc-ph

Equity Premium Prediction: Taking into Account the Role of Long, even Asymmetric, Swings in Stock Market Behavior

Through a novel approach, this paper shows that substantial change in stock market behavior has a statistically and economically significant impact on equity risk premium predictability both on in-sample and out-of-sample cases. In line with Auer's ''Bullish ratio'', a ''Bullish index'' is introduced to measure the changes in stock market behavior, which we describe through a ''fluctuation detrending moving average analysis'' (FDMAA) for returns. We consider 28 indicators. We find that a ''positive shock'' of the Bullish Index is closely related to strong equity risk premium predictability for forecasts based on macroeconomic variables for up to six months. In contrast, a ''negative shock'' is associated with strong equity risk premium predictability with adequate forecasts for up to nine months when based on technical indicators.

q-fin.ST

A study about who is interested in stock splitting and why: considering companies, shareholders or managers

There are many misconceptions around stock prices, stock splits, shareholders, investors, and managers behaviour about such informations due to a number of confounding factors. This paper tests hypotheses with a selected database, about the question ''is stock split attractive for companies?'' in another words, ''why companies split their stock?'', ''why managers split their stock?'', sometimes for no benefit, and ''why shareholders agree with such decisions?''. We contribute to the existing knowledge through a discussion of nine events in recent (selectively chosen) years, observing the role of information asymmetries, the returns and traded volumes before and after the event. Therefore, calculating the beta for each sample, it is found that stock splits (i) affect the market and slightly enhance the trading volume in a short-term, (ii) increase the shareholder base for its firm, (iii) have a positive effect on the liquidity of the market. We concur that stock split announcements can reduce the level of information asymmetric. Investors readjust their beliefs in the firm, although most of the firms are mispriced in the stock split year.

q-fin.GN

New inequality indicators for team ranking in multi-stage female professional cyclist races

Cycling competition is highly interesting since the team ranking is based on the best performance of some subset of team members. The paper develops new inequality indicators, a methodology to construct them, and numerical illustrations allowing to provide operative arguments in their favor. The numerical illustrations subsequently deal with hierarchical ranking indicators of (female) cyclist teams, competing in multi-stage races. For the illustration, the 2023 editions of the most famous long races for females are considered: 34th Giro d'Italia Donne, 2nd Tour de France Femmes, 9th Vuelta Femenina. Several classical ranking indicators are recalled and adapted to the study cases. The most usual indicator, $T_L$, is based on the riders arriving time for the various stages, i.e., according to Union Cycliste Internationale (UCI) standard rules. One also uses another indicator, $A_L$, which requires that the riders finish the race, whence each stage, in order to define the race best team. Another contribution of the paper derives from specific developments of these indicators, thereby leading to new measures: the ''leadership gap" based on $A_L-T_L$, and the ''competition temperature", based on entropy. It is argued that the numerical values point to differences in team strategy based on rider skill levels. The ranking of contributions to indicators allow to observe the "crucial core" made of the most competitive teams.

physics.soc-ph

Should one (be allowed to) replace the Cippolini's?

One examines and discusses proposals on whether riders could be replaced in a team during multi-stage races, and how much a team final time at the end of the race would change (be "adjusted") if only the riders having completed the race are taken into account for ranking teams. A few results of the two main multi-stage races, the men Tour de France and the Giro d'Italia, are used as case studies. The impact of disqualification later on, due to doping, much after the end of such a race, is also examined in the case of two Tour de France. The statistical discussion is based on the Kendall-$\tau$ coefficients for comparing team ranks at the end of these multi-stages races cases. One observes that there are significant differences in the results of the discussed measures. It is shown that there is much variety in results significance, whence demonstrating many interests of the "adjusted indicators". Moreover, it is argued that the "adjusted" rank indicator would promote more competitive and more attractive daily stages and lead to more valuable race management.

physics.soc-ph

Eco-Innovation and Earnings Management: Unveiling the Moderating Effects of Financial Constraints and Opacity in FTSE All-Share Firms

Our research investigates the relationship between eco-innovation and earnings management among 567 firms listed on the FTSE All-Share Index from 2014 to 2022. By examining how sustainability-driven innovation influences financial reporting practices, we explore the strategic motivations behind income smoothing in firms engaged in environmental initiatives. The findings reveal a positive association between eco-innovation and earnings management, suggesting that firms may leverage ecoinnovation not only for environmental signalling but also to project financial stability and meet stakeholder expectations. The analysis further uncovers that the propensity for earnings management is amplified in firms facing financial constraints, proxied by low Whited-Wu (WW) scores and weak sales performance, and in those characterised by high financial opacity. We employ a robust multi-method approach to address potential endogeneity and selection bias, including entropy balancing, propensity score matching (PSM), and the Heckman Test correction. Our research contributes to the literature by providing empirical evidence on the dual strategic role of ecoinnovation -balancing sustainability signalling with earnings management, under varying financial conditions. The findings offer actionable insights for regulators, investors, and policymakers navigating the intersection of corporate transparency, financial health, and environmental responsibility.

econ.GN

A theory of best choice selection through objective arguments grounded in Linear Response Theory concepts

In this paper, we propose how to use objective arguments grounded in statistical mechanics concepts in order to obtain a single number, obtained after aggregation, which would allow to rank "agents", "opinions", ..., all defined in a very broad sense. We aim toward any process which should a priori demand or lead to some consensus in order to attain the presumably best choice among many possibilities. In order to precise the framework, we discuss previous attempts, recalling trivial "means of scores", - weighted or not, Condorcet paradox, TOPSIS, etc. We demonstrate through geometrical arguments on a toy example, with 4 criteria, that the pre-selected order of criteria in previous attempts makes a difference on the final result. However, it might be unjustified. Thus, we base our "best choice theory" on the linear response theory in statistical mechanics: we indicate that one should be calculating correlations functions between all possible choice evaluations, thereby avoiding an arbitrarily ordered set of criteria. We justify the point through an example with 6 possible criteria. Applications in many fields are suggested. Beside, two toy models serving as practical examples and illustrative arguments are given in an Appendix.

physics.soc-ph

Hierarchy Selection: New team ranking indicators for cyclist multi-stage races

In this paper, I report some investigation discussing team selection, whence hierarchy, through ranking indicators, for example when measuring professional cyclist team's sportive value, in particular in multistage races. A logical, it seems, constraint is introduced on the riders: they must finish the race. Several new indicators are defined, justified, and compared. These indicators are mainly based on the arriving place of (the best 3) riders instead of their time needed for finishing the stage or the race, - as presently classically used. A case study, serving as an illustration containing the necessary ingredients for a wider discussion, is the 2023 Vuelta de San Juan, but without loss of generality. It is shown that the new indicators offer some new viewpoint for distinguishing the ranking through the cumulative sums of the places of riders rather than their finishing times. On the other hand, the indicators indicate a different team hierarchy if only the finishing riders are considered. Some consideration on the distance between ranking indicators is presented. Moreover, it is argued that these new ranking indicators should hopefully promote more competitive races, not only till the end of the race, but also until the end of each stage. Generalizations and other applications within operational research topics, like in academia, are suggested.

physics.soc-ph

Unleashing the Power of AI. A Systematic Review of Cutting-Edge Techniques in AI-Enhanced Scientometrics, Webometrics, and Bibliometrics

Purpose: The study aims to analyze the synergy of Artificial Intelligence (AI), with scientometrics, webometrics, and bibliometrics to unlock and to emphasize the potential of the applications and benefits of AI algorithms in these fields. Design/methodology/approach: By conducting a systematic literature review, our aim is to explore the potential of AI in revolutionizing the methods used to measure and analyze scholarly communication, identify emerging research trends, and evaluate the impact of scientific publications. To achieve this, we implemented a comprehensive search strategy across reputable databases such as ProQuest, IEEE Explore, EBSCO, Web of Science, and Scopus. Our search encompassed articles published from January 1, 2000, to September 2022, resulting in a thorough review of 61 relevant articles. Findings: (i) Regarding scientometrics, the application of AI yields various distinct advantages, such as conducting analyses of publications, citations, research impact prediction, collaboration, research trend analysis, and knowledge mapping, in a more objective and reliable framework. (ii) In terms of webometrics, AI algorithms are able to enhance web crawling and data collection, web link analysis, web content analysis, social media analysis, web impact analysis, and recommender systems. (iii) Moreover, automation of data collection, analysis of citations, disambiguation of authors, analysis of co-authorship networks, assessment of research impact, text mining, and recommender systems are considered as the potential of AI integration in the field of bibliometrics. Originality/value: This study covers the particularly new benefits and potential of AI-enhanced scientometrics, webometrics, and bibliometrics to highlight the significant prospects of the synergy of this integration through AI.

cs.DL

Identification of the most important external features of highly cited scholarly papers through 3 (i.e., Ridge, Lasso, and Boruta) feature selection data mining methods

Highly cited papers are influenced by external factors that are not directly related to the document's intrinsic quality. In this study, 50 characteristics for measuring the performance of 68 highly cited papers, from the Journal of the American Medical Informatics Association indexed in Web of Sciences (WoS), from 2009 to 2019 were investigated. In the first step, a Pearson correlation analysis is performed to eliminate variables with zero or weak correlation with the target (dependent) variable ([number of citations in WOS]). Consequently, 32 variables are selected for the next step. By applying the Ridge technique, 13 features show a positive effect on the number of citations. Using three different algorithms, i.e., Ridge, Lasso, and Boruta, 6 factors appear to be the most relevant ones. The [Number of citations by international researchers], [Journal self-citations in citing documents], and [Authors' self-citations in citing documents], are recognized as the most important features by all three methods here used. The [First author's scientific age], [Open-access paper], and [Number of first author's citations in WOS] are identified as the important features of highly cited papers by only two methods, Ridge and Lasso. Notice that we use specific machine learning algorithms as feature selection methods (Ridge, Lasso, and Boruta) to identify the most important features of highly cited papers, tools that had not previously been used for this purpose. In conclusion, we re-emphasize the performance resulting from such algorithms. Moreover, we do not advise authors to seek to increase the citations of their articles by manipulating the identified performance features. Indeed, ethical rules regarding these characteristics must be strictly obeyed.

physics.soc-ph

Shannon Entropy and Herfindahl-Hirschman Index as Team's Performance and Competitive Balance Indicators in Cyclist Multi-Stage Races

It seems that one cannot find many papers relating entropy to sport competitions. Thus, in this paper, I use (i) the Shannon intrinsic entropy ($S$) as an indicator of "teams sporting value" (or "competition performance") and (ii) the Herfindahl-Hirschman index (HHi) index as a "teams competitive balance" indicator, in the case of (professional) cyclist multi-stage races. The 2022 Tour de France and 2023 Tour of Oman are used for numerical illustrations and discussion. The numerical values are obtained from classical and and new ranking indices which measure the teams "final time", on one hand, and "final place", on the other hand, based on the "best three" riders in each stage, but also the corresponding times and places throughout the race, for these finishing riders. The analysis data demonstrates that the constraint, "only the finishing riders count", makes much sense for obtaining a more objective measure of "team value" and team performance", at the end of a multi-stage race. A graphical analysis allows to distinguish various team levels, with in each a Feller-Pareto distribution, thereby pointing to self-organized processes. In so doing, one hopefully better relates objective scientific measures to sport team competitions, and, besides, even proposes some paths to elaborate on forecasting through standard probability concepts.

physics.soc-ph

Portfolio Volatility Estimation Relative to Stock Market Cross-Sectional Intrinsic Entropy

Selecting stock portfolios and assessing their relative volatility risk compared to the market as a whole, market indices, or other portfolios is of great importance to professional fund managers and individual investors alike. Our research uses the cross-sectional intrinsic entropy (CSIE) model to estimate the cross-sectional volatility of the stock groups that can be considered together as portfolio constituents. In our study, we benchmark portfolio volatility risks against the volatility of the entire market provided by the CSIE and the volatility of market indices computed using longitudinal data. This article introduces CSIE-based betas to characterise the relative volatility risk of the portfolio against market indices and the market as a whole. We empirically prove that, through CSIE-based betas, multiple sets of symbols that outperform the market indices in terms of rate of return while maintaining the same level of risk or even lower than the one exhibited by the market index can be discovered, for any given time interval. These sets of symbols can be used as constituent stock portfolios and, in connection with the perspective provided by the CSIE volatility estimates, to hierarchically assess their relative volatility risk within the broader context of the overall volatility of the stock market.

q-fin.ST

Markov Chain Monte Carlo for generating ranked textual data

This paper faces a central theme in applied statistics and information science, which is the assessment of the stochastic structure of rank-size laws in text analysis. We consider the words in a corpus by ranking them on the basis of their frequencies in descending order. The starting point is that the ranked data generated in linguistic contexts can be viewed as the realisations of a discrete states Markov chain, whose stationary distribution behaves according to a discretisation of the best fitted rank-size law. The employed methodological toolkit is Markov Chain Monte Carlo, specifically referring to the Metropolis-Hastings algorithm. The theoretical framework is applied to the rank-size analysis of the hapax legomena occurring in the speeches of the US Presidents. We offer a large number of statistical tests leading to the consistency of our methodological proposal. To pursue our scopes, we also offer arguments supporting that hapaxes are rare (``extreme") events resulting from memory-less-like processes. Moreover, we show that the considered sample has the stochastic structure of a Markov chain of order one. Importantly, we discuss the versatility of the method, which is considered suitable for deducing similar outcomes for other applied science contexts.

stat.ME

God ($\equiv Elohim$), the first small world network

In this paper, the approach of network mapping of words in literary texts is extended to ''textual factors'': the network nodes are defined as ''concepts''; the links are ''community connexions''. Thereafter, the text network properties are investigated along modern statistical physics approaches of networks, thereby relating network topology and algebraic properties, to literary texts contents. As a practical illustration, the first chapter of the Genesis in the Bible is mapped into a 10 node network, as in the Kabbalah approach, mentioning God ($\equiv Elohim$). The characteristics of the network are studied starting from its adjacency matrix, and the corresponding Laplacian matrix. Triplets of nodes are particularly examined in order to emphasize the ''textual (community) connexions'' of each agent "emanation", through the so called clustering coefficients and the overlap index, whence measuring the ''semantic flow'' between the different nodes. It is concluded that this graph is a small-world network, weakly dis-assortative, because its average local clustering coefficient is significantly higher than a random graph constructed on the same vertex set.

physics.soc-ph