SearcharxivSearch

arXiv subjects

Danyllo Albuquerque

Publications and source records attributed to Danyllo Albuquerque.

5 recordsLinked to original sources

From Textual Requirements to Microservice Architectures - A Comprehensive Evaluation of LLM-Based Design Synthesis

Microservice architectures have become dominant for modernizing monolithic systems, yet identifying appropriate services remains challenging and largely manual. Existing decomposition approaches are predominantly code-centric, limiting applicability in early design stages where only textual requirements are available. Despite advances in Large Language Models (LLMs), limited empirical evidence exists on their ability to synthesize complete microservice architectures from natural-language requirements, including service definitions and inter-service interactions. This study investigates whether an LLM can bridge requirements engineering and architectural design, generating architectures solely from textual requirements and evaluating structural agreement and perceived quality of results. We conduct a mixed-method study using OpenAI o3 under zero-shot (ZS) and few-shot (FS) prompting across two systems (Bookstore, PetClinic), one execution per system/condition. Architectures are evaluated through (i) comparison with reference architectures using precision, recall, and F1-score for service identification and communication recovery, and (ii) a blinded expert assessment of correctness, completeness, modularity, and plausibility, plus open feedback synthesis. OpenAI o3 identifies services with higher agreement under FS prompting (F1 = 0.79 for ZS versus = 0.97 for FS). Communication recovery is more challenging: ZS produces dense architectures with high recall but low precision (F1 = 0.61), while FS improves agreement, reaching F1 = 0.82 and reducing unsupported dependencies. Expert evaluation corroborates these results, with FS architectures perceived as more modular, coherent, and plausible than ZS outputs. OpenAI o3 shows potential for requirements-driven synthesis when guided by exemplar prompting. Results are model- and context-specific from two small systems, not model-independent proof.

cs.SE

AI-Conducted Interviews in Empirical Software Engineering: An Experience Report

Semi-structured interviews are widely used in empirical software engineering (ESE), but they are resource-intensive and difficult to coordinate across schedules, locations, and natural languages. This experience report examines a customized MyGPT used to conduct short, self-administered interviews in two ESE studies: one on refactoring practices and another on generative AI in Scrum-related activities. Participants accessed the interviewer through shared links, used voice interaction, selected a preferred natural language, and completed the interview without a researcher present. The AI followed a predefined protocol and generated a structured synthesis that participants voluntarily submitted; these artifacts were not treated as verbatim transcripts. We analyzed 66 submissions and questionnaire responses, and audited artifact format, language, length, and protocol consistency. Of the submitted artifacts, 92.4% followed the expected synthesis format, 65 were predominantly in Portuguese and one in English, and two conflicted with the reported protocol. Participants generally rated the experience positively: 90.9% reported a positive overall experience and comfort, 95.5% considered the questions clear, 97.0% rated the pace positively, and 89.4% would participate again. Reported limitations included generic questions, limited sensitivity to answers, insufficient depth, privacy concerns, and missed human interaction. The findings support the operational viability and acceptability of this workflow among analyzed respondents, but do not establish completion rates, time savings, summary fidelity, or equivalence to human-conducted interviews. AI interviewers should therefore be treated as a complementary option for short, focused, low-risk studies, with protocol design, privacy guidance, artifact validation, and human oversight.

cs.SE

Prompting GPT-5 on Scrum Certification Questions: An Empirical Accuracy Study

Large Language Models (LLMs) are increasingly used in Agile Software Development for documentation, coaching, and training. As practitioners adopt these tools to prepare for certifications such as Professional Scrum Master (PSM), a key question is whether LLMs can reliably reason about Scrum, a framework with normative, well-defined rules described in the Scrum Guide (2020). This paper examines how different prompt techniques affect the factual accuracy of LLM responses to Scrum certification-style questions. A dataset of 993 validated PSM-aligned questions was answered by GPT-5 using three techniques: zero-shot, chain-of-thought, and with-source citation. All prompts achieved certification-level accuracy above 85\%, with the citation-based variant performing best (89.1\%) and yielding the lowest error rate. Correct answers concentrated in well-defined topics, such as \emph{Definition of Done}, Events, and Product Backlog Management, and in single-answer multiple-choice items, while multi-select questions and more interpretive areas, such as Scrum Team and Product Value, were less stable. Among questions where at least one prompt failed (16.2\%), errors clustered into misalignment with the Scrum Guide (28\%), content outside its scope (34\%), and outdated or biased interpretations (38\%). Overall, prompt techniques produced modest but consistent improvements, particularly in reducing misinterpretation and version drift, supporting more reliable use of LLMs in Agile learning and certification preparation.

cs.SE

Adoption of Large Language Models in Scrum Management: Insights from Brazilian Practitioners

Scrum is widely adopted in software project management due to its adaptability and collaborative nature. The recent emergence of Large Language Models (LLMs) has created new opportunities to support knowledge-intensive Scrum practices. However, existing research has largely focused on technical activities such as coding and testing, with limited evidence on the use of LLMs in management-related Scrum activities. In this study, we investigate the use of LLMs in Scrum management activities through a survey of 70 Brazilian professionals. Among them, 49 actively use Scrum, and 33 reported using LLM-based assistants in their Scrum practices. The results indicate a high level of proficiency and frequent use of LLMs, with 85% of respondents reporting intermediate or advanced proficiency and 52% using them daily. LLM use concentrates on exploring Scrum practices, with artifacts and events receiving targeted yet uneven support, whereas broader management tasks appear to be adopted more cautiously. The main benefits include increased productivity (78%) and reduced manual effort (75%). However, several critical risks remain, as respondents report 'almost correct' outputs (81%), confidentiality concerns (63%), and hallucinations during use (59%). This work provides one of the first empirical characterizations of LLM use in Scrum management, identifying current practices, quantifying benefits and risks, and outlining directions for responsible adoption and integration in Agile environments.

cs.SE

An Empirical Study of Foundation Models for Variability-Induced Compilation Errors in Configurable C Code

In configurable systems, conditional compilation can hide compilation errors under untested feature combinations. We investigate foundation models for detecting such errors and, in a controlled setting, restoring compilability in configurable C code. Study I evaluates GPT-OSS-20B on 5,000 synthetic snippets generated by ChatGPT-5.2 from 30 curated seeds and exhaustively compiled under all Boolean feature assignments; it also compares TypeChef and evaluates Gemini 3.6 Flash on a stratified sample. GPT-OSS-20B achieved 84.7% micro-precision and 52.1% micro-recall for affected configurations. Coverage depended on reporting style: presence conditions covered 99.4% of failing configurations, whereas explicit enumerations covered 29.5% under a prompt requesting only a minimal justifiable set. GPT-OSS-20B restored compilability for 1,930 of 2,665 faulty snippets (72.4%), while Gemini 3.6 Flash did so for 182 of 190 sampled faulty snippets (95.8%). A paired counterfactual audit found no evidence that an identified label-correlated #define property materially influenced GPT-OSS-20B's predictions. Study II evaluates Codex-GPT5.5 on 100 faulty file-level subjects from five mature configurable systems and reports target-fault-aligned problems in 94 subjects, including four of five historical bugs. Overall, foundation models can support localized detection, explanation, and triage, but should complement compiler-based and variability-aware analyses; compiler acceptance does not establish semantic correctness

cs.SE