SearcharxivSearch

arXiv subjects

Angelo Di Iorio

Publications and source records attributed to Angelo Di Iorio.

11 recordsLinked to original sources

Towards a Definition of the Computational Architecture of Open Scholarly Infrastructures

This paper proposes a layered framework for defining the architecture of the computational unit of an Open Scholarly Infrastructure (OSI) and presents a concrete implementation possibility through the introduction of OpenCitations, an OSI dedicated to publishing citation data and bibliographic metadata. Grounded in the Principles of Open Scholarly Infrastructure (POSI), the study focuses on the technical dimensions of openness, sustainability, interoperability, and reproducibility that are required for a robust OSI. The proposed approach first identifies the core principles and technical features that a computational block in an OSI should guarantee, including separation of operational domains, orchestration, scalability, observability, automation, and workflow portability. It then maps these requirements onto a layered architectural model composed of Hardware, Virtualization, Orchestration, and Application layers, complemented by a transversal Meta layer. The OpenCitations case study shows how this framework can be instantiated in practice through an on-premises infrastructure. The case is presented with details regarding the technicalities and actual implementations of the proposed methodology. By combining a conceptual definition with a real-world implementation, this work offers both a practical reference and a methodological basis for designing the computational block of an OSI.

cs.DL

CiteFusion: An Ensemble Framework for Citation Intent Classification Harnessing Dual-Model Binary Couples and SHAP Analyses

Understanding the motivations underlying scholarly citations is essential to evaluate research impact and promote transparent scholarly communication. This study introduces CiteFusion, an ensemble framework designed to address the multi-class Citation Intent Classification task on two benchmark datasets: SciCite and ACL-ARC. The framework employs a one-vs-all decomposition of the multi-class task into class-specific binary subtasks, leveraging complementary pairs of SciBERT and XLNet models, independently tuned, for each citation intent. The outputs of these base models are aggregated through a feedforward neural network meta-classifier to reconstruct the original classification task. To enhance interpretability, SHAP (SHapley Additive exPlanations) is employed to analyze token-level contributions, and interactions among base models, providing transparency into the classification dynamics of CiteFusion, and insights about the kind of misclassifications of the ensemble. In addition, this work investigates the semantic role of structural context by incorporating section titles, as framing devices, into input sentences, assessing their positive impact on classification accuracy. CiteFusion ultimately demonstrates robust performance in imbalanced and data-scarce scenarios: experimental results show that CiteFusion achieves state-of-the-art performance, with Macro-F1 scores of 89.60% on SciCite, and 76.24% on ACL-ARC. Furthermore, to ensure interoperability and reusability, citation intents from both datasets schemas are mapped to Citation Typing Ontology (CiTO) object properties, highlighting some overlaps. Finally, we describe and release a web-based application that classifies citation intents leveraging the CiteFusion models developed on SciCite.

cs.CL

Do open citations give insights on the qualitative peer-review evaluation in research assessments? An analysis of the Italian National Scientific Qualification

In the past, several works have investigated ways for combining quantitative and qualitative methods in research assessment exercises. Indeed, the Italian National Scientific Qualification (NSQ), i.e. the national assessment exercise which aims at deciding whether a scholar can apply to professorial academic positions as Associate Professor and Full Professor, adopts a quantitative and qualitative evaluation process: it makes use of bibliometrics followed by a peer-review process of candidates' CVs. The NSQ divides academic disciplines into two categories, i.e. citation-based disciplines (CDs) and non-citation-based disciplines (NDs), a division that affects the metrics used for assessing the candidates of that discipline in the first part of the process, which is based on bibliometrics. In this work, we aim at exploring whether citation-based metrics, calculated only considering open bibliographic and citation data, can support the human peer-review of NDs and yield insights on how it is conducted. To understand if and what citation-based (and, possibly, other) metrics provide relevant information, we created a series of machine learning models to replicate the decisions of the NSQ committees. As one of the main outcomes of our study, we noticed that the strength of the citational relationship between the candidate and the commission in charge of assessing his/her CV seems to play a role in the peer-review phase of the NSQ of NDs.

cs.DL

Effect of Auditory Stimuli on Electroencephalography-based Authentication

Opposed to standard authentication methods based on credentials, biometric-based authentication has lately emerged as a viable paradigm for attaining rapid and secure authentication of users. Among the numerous categories of biometric traits, electroencephalogram (EEG)-based biometrics is recognized as a promising method owing to its unique characteristics. This paper provides an experimental evaluation of the effect of auditory stimuli (AS) on EEG-based biometrics by studying the following features: i) general change in AS-aided EEG-based biometric authentication in comparison with non-AS-aided EEG-based biometric authentication, ii) role of the language of the AS and ii) influence of the conduction method of the AS. Our results show that the presence of an AS can improve authentication performance by 9.27%. Additionally, the performance achieved with an in-ear AS is better than that obtained using a bone-conducting AS. Finally, we verify that performance is independent of the language of the AS. The results of this work provide a step forward towards designing a robust EEG-based authentication system.

cs.CR

Open bibliographic data and the Italian National Scientific Qualification: measuring coverage of academic fields

The importance of open bibliographic repositories is widely accepted by the scientific community. For evaluation processes, however, there is still some skepticism: even if large repositories of open access articles and free publication indexes exist and are continuously growing, assessment procedures still rely on proprietary databases, mainly due to the richness of the data available in these proprietary databases and the services provided by the companies they are offered by. This paper investigates the status of open bibliographic data of three of the most used open resources, namely Microsoft Academic Graph, Crossref and OpenAIRE, evaluating their potentialities as substitutes of proprietary databases for academic evaluation processes. We focused on the Italian National Scientific Qualification (NSQ), the Italian process for University Professor qualification, which uses data from commercial indexes, and investigated similarities and differences between research areas, disciplines and application roles. The main conclusion is that open datasets are ready to be used for some disciplines, among which mathematics, natural sciences, economics and statistics, even if there is still room for improvement; but there is still a large gap to fill in others - like history, philosophy, pedagogy and psychology - and a stronger effort is required from researchers and institutions.

cs.DL

Academics evaluating academics: a methodology to inform the review process on top of open citations

In the past, several works have investigated ways for combining quantitative and qualitative methods in research assessment exercises. In this work, we aim at introducing a methodology to explore whether citation-based metrics, calculated only considering open bibliographic and citation data, can yield insights on how human peer-review of research assessment exercises is conducted. To understand if and what metrics provide relevant information, we propose to use a series of machine learning models to replicate the decisions of the committees of the research assessment exercises.

cs.DL

Can we assess research using open scientific knowledge graphs? A case study within the Italian National Scientific Qualification

The need for open scientific knowledge graphs is ever increasing. While there are large repositories of open access articles and free publication indexes, there are still few free knowledge graphs exposing citation networks, and often their coverage is partial. Consequently, most evaluation processes based on citation counts rely on commercial citation databases. Things are changing thanks to the Initiative for Open Citations (I4OC, https://i4oc.org) and the Initiative for Open Abstracts (I4OA, https://i4oa.org), whose goal is to campaign for scholarly publishers to open the reference lists and the other metadata of their articles. This paper investigates the growth of the open bibliographic metadata and open citations in two scientific knowledge graphs, OpenCitations' COCI and Crossref, with an experiment on the Italian National Scientific Qualification (NSQ), the National process for University Professor qualification which uses data from commercial indexes. We simulated the procedure by only using such open data and explored similarities and differences with the official results. The outcomes of the experiment show that the amount of open bibliographic metadata and open citation data currently available in the two scientific knowledge graphs adopted is not yet enough for obtaining results similar to those provided using commercial databases.

cs.DL

Open data to evaluate academic researchers: an experiment with the Italian Scientific Habilitation

The need for scholarly open data is ever increasing. While there are large repositories of open access articles and free publication indexes, there are still a few examples of free citation networks and their coverage is partial. One of the results is that most of the evaluation processes based on citation counts rely on commercial citation databases. Things are changing under the pressure of the Initiative for Open Citations (I4OC), whose goal is to campaign for scholarly publishers to make their citations as totally open. This paper investigates the growth of open citations with an experiment on the Italian Scientific Habilitation, the National process for University Professor qualification which instead uses data from commercial indexes. We simulated the procedure by only using open data and explored similarities and differences with the official results. The outcomes of the experiment show that the amount of open citation data currently available is not yet enough for obtaining similar results.

cs.DL

Semantic Publishing Challenge - Assessing the Quality of Scientific Output by Information Extraction and Interlinking

The Semantic Publishing Challenge series aims at investigating novel approaches for improving scholarly publishing using Linked Data technology. In 2014 we had bootstrapped this effort with a focus on extracting information from non-semantic publications - computer science workshop proceedings volumes and their papers - to assess their quality. The objective of this second edition was to improve information extraction but also to interlink the 2014 dataset with related ones in the LOD Cloud, thus paving the way for sophisticated end-user services.

cs.DL

Semantic Publishing Challenge -- Assessing the Quality of Scientific Output

Linked Open Datasets about scholarly publications enable the development and integration of sophisticated end-user services; however, richer datasets are still needed. The first goal of this Challenge was to investigate novel approaches to obtain such semantic data. In particular, we were seeking methods and tools to extract information from scholarly publications, to publish it as LOD, and to use queries over this LOD to assess quality. This year we focused on the quality of workshop proceedings, and of journal articles w.r.t. their citation network. A third, open task, asked to showcase how such semantic data could be exploited and how Semantic Web technologies could help in this emerging context.

cs.DL

Where are your Manners? Sharing Best Community Practices in the Web 2.0

The Web 2.0 fosters the creation of communities by offering users a wide array of social software tools. While the success of these tools is based on their ability to support different interaction patterns among users by imposing as few limitations as possible, the communities they support are not free of rules (just think about the posting rules in a community forum or the editing rules in a thematic wiki). In this paper we propose a framework for the sharing of best community practices in the form of a (potentially rule-based) annotation layer that can be integrated with existing Web 2.0 community tools (with specific focus on wikis). This solution is characterized by minimal intrusiveness and plays nicely within the open spirit of the Web 2.0 by providing users with behavioral hints rather than by enforcing the strict adherence to a set of rules.

cs.CY