SearcharxivSearch

arXiv subjects

Alessio Bucaioni

Publications and source records attributed to Alessio Bucaioni.

6 recordsLinked to original sources

Toward Federated Cognitive Digital Twins over the Edge-to-Cloud Continuum

Digital Twins (DTs) are increasingly adopted to monitor, analyze, and optimize Cyber-Physical Systems (CPSs) through continuous interaction between physical assets and their digital counterparts. However, current DT architectures often rely on centralized and monolithic designs, leading to scalability, latency, and resilience issues in distributed environment such as smart cities. Moreover, they provide limited support for semantic integration and high-level reasoning, reducing the effectiveness of DT-based decision-making. Recent studies on Federated Digital Twins (FDTs) have addressed scalability by decomposing complex systems into interacting twins, but they still largely centralize intelligence in cloud components. In parallel, Cognitive Digital Twins (CDTs) enhance DTs with semantic reasoning, explainability, and AI-driven decision support, yet they are typically difficult to integrate into distributed architectures. This paper proposes a Federated Cognitive Digital Twin (FCDT) architecture that combines federation and cognition within a unified approach. The architecture distributes intelligence across the edge-to-cloud continuum through local twins, which provide real-time monitoring and lightweight cognitive capabilities, and global twins, which perform system-level reasoning, simulation, and coordination. By integrating distributed autonomy with cognitive reasoning, the proposed approach improves scalability, responsiveness, and decision-making in complex distributed CPSs

cs.SE

Bug-Report-Driven Fault Localization: Industrial Benchmarking and Lesson Learned at ABB Robotics

Software quality assurance remains a major challenge in industrial environments, where large-scale and long-lived systems inevitably accumulate defects. Identifying the location of a fault is often time-consuming and costly, particularly during maintenance phases when developers must rely primarily on textual bug reports rather than complete runtime or code-level context. In this study, we investigated if artificial intelligence can support fault localization using only the natural-language content of bug reports. By relying only on textual information, our approach requires no access to source code, execution traces, or static analysis artifacts, making it directly deployable within existing industrial maintenance workflows. We framed fault localization as a supervised text classification problem and evaluated three traditional machine learning models (Logistic Regression, Support Vector Machine, and Random Forest) and two fine-tuned transformer-based language models (RoBERTa-Base and Distil-RoBERTa). Our evaluation used proprietary data from ABB Robotics in Sweden, comprising five years of resolved industrial bug reports, each linked to its verified code fix. This setting allowed us to assess model effectiveness under realistic industrial constraints. Our results showed that traditional models using term frequency-inverse document features consistently outperformed the fine-tuned language models on this dataset, while data augmentation improved Random Forest performance. These findings challenge the assumption that transformer-based models universally outperform classical approaches in industrial contexts with domain-specific data. We demonstrated that historical bug reports can be systematically used for text-based, artificial intelligence-assisted fault localization, providing a scalable, low-cost, and empirically grounded complement to common debugging practices in industry.

cs.SE

When Retriever Meets Generator: A Joint Model for Code Comment Generation

Automatically generating concise, informative comments for source code can lighten documentation effort and accelerate program comprehension. Retrieval-augmented approaches first fetch code snippets with existing comments and then synthesize a new comment, yet retrieval and generation are typically optimized in isolation, allowing irrelevant neighbors topropagate noise downstream. To tackle the issue, we propose a novel approach named RAGSum with the aim of both effectiveness and efficiency in recommendations. RAGSum is built on top offuse retrieval and generation using a single CodeT5 backbone. We report preliminary results on a unified retrieval-generation framework built on CodeT5. A contrastive pre-training phase shapes code embeddings for nearest-neighbor search; these weights then seed end-to-end training with a composite loss that (i) rewards accurate top-k retrieval; and (ii) minimizes comment-generation error. More importantly, a lightweight self-refinement loop is deployed to polish the final output. We evaluated theframework on three cross-language benchmarks (Java, Python, C), and compared it with three well-established baselines. The results show that our approach substantially outperforms thebaselines with respect to BLEU, METEOR, and ROUTE-L. These findings indicate that tightly coupling retrieval and generationcan raise the ceiling for comment automation and motivateforthcoming replications and qualitative developer studies.

cs.SE

ROSE: Transformer-Based Refactoring Recommendation for Architectural Smells

Architectural smells such as God Class, Cyclic Dependency, and Hub-like Dependency degrade software quality and maintainability. Existing tools detect such smells but rarely suggest how to fix them. This paper explores the use of pre-trained transformer models--CodeBERT and CodeT5--for recommending suitable refactorings based on detected smells. We frame the task as a three-class classification problem and fine-tune both models on over 2 million refactoring instances mined from 11,149 open-source Java projects. CodeT5 achieves 96.9% accuracy and 95.2% F1, outperforming CodeBERT and traditional baselines. Our results show that transformer-based models can effectively bridge the gap between smell detection and actionable repair, laying the foundation for future refactoring recommendation systems. We release all code, models, and data under an open license to support reproducibility and further research.

cs.SE

Artificial Intelligence for Software Architecture: Literature Review and the Road Ahead

Artificial intelligence is increasingly applied across software engineering, yet its explicit role in software architecture remains insufficiently understood. Architectural practices rely on complex trade-offs, documentation, and long-term evolution, all of which are traditionally manual, error-prone, and difficult to sustain. To clarify how artificial intelligence can address these challenges, we conducted a systematic literature review of 51 peer-reviewed primary studies and systematically mapped their contributions onto 17 practitioner-reported software architecture challenges derived from empirical interviews. This analysis identifies 14 topical areas where artificial intelligence has been applied to architectural tasks and identifies six artificial intelligence-specific challenges that expose fundamental gaps between current capabilities and practitioner needs. Building on these findings, we chart a research agenda for artificial intelligence-driven software architecture organized around five strategic pillars. By grounding the roadmap in both systematic evidence and practitioner insights, this work provides the first peer-reviewed comprehensive synthesis of artificial intelligence contributions to software architecture, establishes a foundation for future research, and outlines the conditions under which artificial intelligence can become a trustworthy partner in architectural design, evaluation, and evolution.

cs.SE

A Functional Software Reference Architecture for LLM-Integrated Systems

The integration of large language models into software systems is transforming capabilities such as natural language understanding, decision-making, and autonomous task execution. However, the absence of a commonly accepted software reference architecture hinders systematic reasoning about their design and quality attributes. This gap makes it challenging to address critical concerns like privacy, security, modularity, and interoperability, which are increasingly important as these systems grow in complexity and societal impact. In this paper, we describe our \textit{emerging} results for a preliminary functional reference architecture as a conceptual framework to address these challenges and guide the design, evaluation, and evolution of large language model-integrated systems. We identify key architectural concerns for these systems, informed by current research and practice. We then evaluate how the architecture addresses these concerns and validate its applicability using three open-source large language model-integrated systems in computer vision, text processing, and coding.

cs.SE