SearcharxivSearch

arXiv subjects

Cleyton Magalhaes

Publications and source records attributed to Cleyton Magalhaes.

At least 19 recordsLinked to original sources

Technostress in the Age of AI: A Preliminary Study with Software Professionals

The rapid adoption and evolution of AI are changing software engineering work and requiring professionals to repeatedly adapt their knowledge, practices, and skills. Although technological adaptation has long characterized software development, less is known about how these new and recurring pressures manifest as technostress. This preliminary exploratory study investigates AI related technostress among software professionals. We conducted a survey and performed a thematic analysis of responses from 121 software professionals across 26 countries reporting their recent experiences with AI at work. Our findings suggest that AI related technostress emerges not only from adapting to rapidly changing technologies, but also from having to manage the work, technical responsibilities, and professional changes that accompany their adoption. This characterization shows that AI introduces pressures beyond learning and using new tools, affecting how software professionals perform and remain accountable for technical work and how they prepare for the future of their careers.

cs.SE

Exploring Quantum Software Testing Across Research and Practice: Emerging Results from a Multivocal Literature Review

This paper presents preliminary findings from a multivocal literature review investigating how quantum software testing is characterized across academic and practitioner-oriented sources. Our study integrated peer-reviewed studies with gray literature, including blogs, tutorials, forums, technical reports, documentation pages, and company webpages. Our results indicate a rapidly evolving but fragmented ecosystem involving classical adapted testing approaches, quantum-specific techniques, statistical validation methods, simulators, debugging environments, and verification frameworks. The reviewed material also revealed recurring challenges related to scalability limitations, hardware noise, probabilistic execution, limited observability, and immature tooling ecosystems. These findings provide an initial characterization of how research and practice currently discuss quantum software testing challenges, techniques, and tooling.

cs.SE

Bias Smells in AI Software Development: Recognizing Potential Sources of Fairness Debt

Context: Fairness debt arises from AI development shortcomings that may lead to societal harms. While technical and social debt concern software design decisions and team dynamics, respectively, fairness debt captures the long-term consequences of development decisions that may reinforce bias and inequities in AI-based software systems. Despite growing attention to AI fairness, limited evidence exists on how practitioners recognize potential sources of fairness debt during development. Aim: This study investigates the indicators practitioners recognize as signaling potential sources of fairness debt in AI-based software projects. Method: We conducted an exploratory case study of four AI projects within one organization. Data were collected from 25 professionals through semi-structured interviews and open-ended questionnaires, complemented by observation of internal communication channels and project documentation, and analyzed using iterative qualitative coding, memoing, and constant comparison. Results: We identified six recurring indicators, termed bias smells: Context Oversimplification, Dataset Imbalance, Metrics Inadequacies, Ad hoc Testing, Individual Diversity Unawareness, and Homogeneous Team Composition. These smells span technical and human aspects of software development and signal conditions that may introduce or reinforce bias and contribute to fairness debt. Conclusion: Bias smells extend the software smell paradigm to fairness and provide a foundation for incorporating fairness into software quality assurance through observable indicators.

cs.SE

Exploring Dependence, Overreliance, and Addiction Related Behaviors Associated with Large Language Model Use Among Software Engineers

The widespread adoption of Large Language Models (LLMs) has changed how software engineers perform everyday development activities. While these systems provide substantial support for tasks such as code generation, debugging, and documentation, their increasing integration into professional workflows has also raised questions regarding developers' reliance on these tools and the emergence of dependence, overreliance, and addiction-related behaviors. This study investigates how software engineers experience the use of LLMs during professional software development, with attention to behavioral patterns associated with dependence, overreliance, and addiction-related behaviors. An exploratory survey was conducted with 119 software practitioners. The data were analyzed using descriptive statistics and qualitative thematic analysis of participants' open-ended responses. Participants primarily described functional dependence, with LLMs becoming integrated into routine software engineering activities because of the productivity and efficiency they provide. Responses also suggested patterns consistent with overreliance, particularly through prioritizing LLMs over documentation or peer consultation while continuing to verify generated outputs. Reports associated with addiction-related behaviors were less common and primarily reflected difficulty moderating use or emotional attachment to the technology rather than impaired control. The findings suggest that LLMs are becoming a habitual component of professional software engineering practice. While most reported use appears functional, the results indicate the importance of promoting appropriate reliance by supporting trust calibration, professional judgment, and verification throughout software development.

cs.SE

How Many Interviews Are Enough in a Software Engineering Study? Preliminary Findings on Sample Size and Saturation

Background. Interview based studies are widely used in empirical software engineering to investigate human, organizational, and socio technical phenomena, yet interview sample adequacy and saturation are reported inconsistently across the literature. Aims. This paper investigates how interview sample adequacy and saturation are operationalized in empirical software engineering research. Method. We analyzed papers published between 2016 and 2025 across major software engineering venues, focusing on interview sample sizes, saturation discussions, and sample adequacy justifications. Results. Preliminary findings indicate substantial variation in sample sizes, from highly specialized small sample studies to broader investigations involving large interview datasets. Studies involving fewer than 12 interviewees were common and frequently associated with specialized industrial contexts or constrained organizational access. However, the most recurrent range was 13 to 24 interviewees, suggesting that moderate sized samples represent the most common configuration in empirical software engineering research. Saturation and sample adequacy justification were heterogeneous, with many studies relying on implicit or contextual reasoning rather than explicit methodological discussion. Conclusions. Our findings provide initial empirical insights into methodological reporting practices in interview based software engineering research and contribute to ongoing discussions regarding qualitative rigor and transparency.

cs.SE

The Influence of Fraudulent AI-Generated Responses on Software Engineering Surveys

Background: Large Language Models (LLMs) introduce new concerns regarding fraudulent or AI assisted participation in software engineering surveys. Aims: This study investigates how suspicious or potentially AI assisted responses may affect the validity of software engineering survey findings. Method: We conducted a secondary analysis of four software engineering survey datasets using manual identification of suspicious responses, automated AI generated text detection, descriptive statistical analysis, and thematic analysis. We compared findings obtained from the original and manually cleaned datasets. Results: Quantitative findings generally remained stable after filtering suspicious responses, although some demographic and analytical variables showed moderate variation, affecting the interpretation of specific participant groups and contextual characteristics. In contrast, qualitative findings were more strongly influenced by changes in contextual framing, code prominence, and the nature of the evidence supporting interpretation, shaping how participants' experiences and study contexts were interpreted and characterized. Conclusions: AI assisted participation may influence software engineering survey findings differently depending on the type of analysis being conducted. The findings reinforce the importance of combining multiple validation procedures, particularly in studies relying on open ended responses.

cs.SE

How Software Engineering Students Use LLMs to Write Research Papers: An Experience Report

Large language models are increasingly becoming part of software engineering education, including activities involving empirical software engineering and evidence synthesis. This paper reports an educational experience involving the integration of reflective LLM use into an empirical methods assignment in a third-year software architecture course. Students were asked to develop a short research paper using either a rapid review or a gray literature review methodology and to disclose how LLMs were used throughout the assignment. We analyzed 146 student disclosure statements using a cross-analysis process combining LLM-assisted categorization with manual verification and refinement by the researchers. The reflections describe how students incorporated LLMs during activities such as brainstorming, methodological clarification, organization of findings, and writing refinement, while also reporting concerns regarding inaccuracies and verification of generated content. This experience report discusses lessons learned and educational implications for integrating AI-assisted technologies into empirical software engineering education.

cs.SE

Academic Integrity and Emotional Responses to Inappropriate LLM Use in Software Engineering Education

Academic integrity in higher education is increasingly shaped by complex socio-technical environments marked by automated tools, evolving institutional practices, and heightened performance pressures. Within this context, large language models (LLMs) are becoming prevalent in software engineering education, further blurring boundaries around acceptable assistance and authorship. This study investigates how software engineering students describe their emotional experiences after using LLMs in ways they perceive as academically inappropriate. We conducted a cross-sectional survey with 116 undergraduate students. Results show emotionally heterogeneous responses. Indifference was most frequent, including among students who recognized risks to learning and academic standing. Guilt and anxiety were reported in relation to moral discomfort and concern about penalties. Relief and satisfaction were evident primarily in deadline-driven contexts and situations of unclear guidance.

cs.SE

The role of team diversity in AI systems development

The widespread integration of AI technologies has intensified concerns about fairness and bias, as these systems often perpetuate societal inequalities through flawed data and design choices. While software engineering research has largely concentrated on technical solutions, such as improving datasets and models, the social dynamics that shape AI outcomes remain underexplored. This study investigates the role of team diversity in the development of AI systems. Drawing from the experience of four AI focused teams working in a large software company operating in Brazil and Portugal, and collaborating with global clients, the study explores how diverse teams influence the development of AI systems. Using Grounded Theory, we conducted 25 interviews with software professionals involved in projects spanning domains such as education, energy, accessibility, and facial recognition. Although our study is conducted in an organizational setting, the variety of projects, from regional to multinational, ensures exposure to global development practices and diverse team dynamics, bringing a variety of perspectives into our findings. Our analysis revealed six key roles that team diversity played in AI development: diversifying perspectives for bias identification, bringing empathy to AI development, addressing systemic discrimination, supporting inclusive and participatory decision making, using diversity as a safeguard against bias, and fostering broadened thinking in problem solving. These findings highlight the importance of incorporating diverse perspectives in AI projects and offer practical recommendations for integrating fairness considerations into software development practices.

cs.SE

An Investigation on How AI-Generated Responses Affect SoftwareEngineering Surveys

Survey research is a fundamental empirical method in software engineering, enabling the systematic collection of data on professional practices, perceptions, and experiences. However, recent advances in large language models (LLMs) have introduced new risks to survey integrity, as participants can use generative tools to fabricate or manipulate their responses. This study explores how LLMs are being misused in software engineering surveys and investigates the methodological implications of such behavior for data authenticity, validity, and research integrity. We collected data from two survey deployments conducted in 2025 through the Prolific platform and analyzed the content of participants' answers to identify irregular or falsified responses. A subset of responses suspected of being AI generated was examined through qualitative pattern inspection, narrative characterization, and automated detection using the Scribbr AI Detector. The analysis revealed recurring structural patterns in 49 survey responses indicating synthetic authorship, including repetitive sequencing, uniform phrasing, and superficial personalization. These false narratives mimicked coherent reasoning while concealing fabricated content, undermining construct, internal, and external validity. Our study identifies data authenticity as an emerging dimension of validity in software engineering surveys. We emphasize that reliable evidence now requires combining automated and interpretive verification procedures, transparent reporting, and community standards to detect and prevent AI generated responses, thereby protecting the credibility of surveys in software engineering.

cs.SE

Bita: A Conversational Assistant for Fairness Testing

Bias in AI systems can lead to unfair and discriminatory outcomes, especially when left untested before deployment. Although fairness testing aims to identify and mitigate such bias, existing tools are often difficult to use, requiring advanced expertise and offering limited support for real-world workflows. To address this, we introduce Bita, a conversational assistant designed to help software testers detect potential sources of bias, evaluate test plans through a fairness lens, and generate fairness-oriented exploratory testing charters. Bita integrates a large language model with retrieval-augmented generation, grounding its responses in curated fairness literature. Our validation demonstrates how Bita supports fairness testing tasks on real-world AI systems, providing structured, reproducible evidence of its utility. In summary, our work contributes a practical tool that operationalizes fairness testing in a way that is accessible, systematic, and directly applicable to industrial practice.

cs.SE

Software Testing with Large Language Models: An Interview Study with Practitioners

\textit{Background:} The use of large language models in software testing is growing fast as they support numerous tasks, from test case generation to automation, and documentation. However, their adoption often relies on informal experimentation rather than structured guidance. \textit{Aims:} This study investigates how software testing professionals use LLMs in practice to propose a preliminary, practitioner-informed guideline to support their integration into testing workflows. \textit{Method:} We conducted a qualitative study with 15 software testers from diverse roles and domains. Data were collected through semi-structured interviews and analyzed using grounded theory-based processes focused on thematic analysis. \textit{Results:} Testers described an iterative and reflective process that included defining testing objectives, applying prompt engineering strategies, refining prompts, evaluating outputs, and learning over time. They emphasized the need for human oversight and careful validation, especially due to known limitations of LLMs such as hallucinations and inconsistent reasoning. \textit{Conclusions:} LLM adoption in software testing is growing, but remains shaped by evolving practices and caution around risks. This study offers a starting point for structuring LLM use in testing contexts and invites future research to refine these practices across teams, tools, and tasks.

cs.SE

Model-Assisted and Human-Guided: Perceptions and Practices of Software Professionals Using LLMs for Coding

Large Language Models have quickly become a central component of modern software development workflows, and software practitioners are increasingly integrating LLMs into various stages of the software development lifecycle. Despite the growing presence of LLMs, there is still a limited understanding of how these tools are actually used in practice and how professionals perceive their benefits and limitations. This paper presents preliminary findings from a global survey of 131 software practitioners. Our results reveal how LLMs are utilized for various coding-specific tasks. Software professionals report benefits such as increased productivity, reduced cognitive load, and faster learning, but also raise concerns about LLMs' inaccurate outputs, limited context awareness, and associated ethical risks. Most developers treat LLMs as assistive tools rather than standalone solutions, reflecting a cautious yet practical approach to their integration. Our findings provide an early, practitioner-focused perspective on LLM adoption, highlighting key considerations for future research and responsible use in software engineering.

cs.SE

Tether: A Personalized Support Assistant for Software Engineers with ADHD

Equity, diversity, and inclusion in software engineering often overlook neurodiversity, particularly the experiences of developers with Attention Deficit Hyperactivity Disorder (ADHD). Despite the growing awareness about that population in SE, few tools are designed to support their cognitive challenges (e.g., sustained attention, task initiation, self-regulation) within development workflows. We present Tether, an LLM-powered desktop application designed to support software engineers with ADHD by delivering adaptive, context-aware assistance. Drawing from engineering research methodology, Tether combines local activity monitoring, retrieval-augmented generation (RAG), and gamification to offer real-time focus support and personalized dialogue. The system integrates operating system level system tracking to prompt engagement and its chatbot leverages ADHD-specific resources to offer relevant responses. Preliminary validation through self-use revealed improved contextual accuracy following iterative prompt refinements and RAG enhancements. Tether differentiates itself from generic tools by being adaptable and aligned with software-specific workflows and ADHD-related challenges. While not yet evaluated by target users, this work lays the foundation for future neurodiversity-aware tools in SE and highlights the potential of LLMs as personalized support systems for underrepresented cognitive needs.

cs.SE

Testing the Untestable? An Empirical Study on the Testing Process of LLM-Powered Software Systems

Background: Software systems powered by large language models are becoming a routine part of everyday technologies, supporting applications across a wide range of domains. In software engineering, many studies have focused on how LLMs support tasks such as code generation, debugging, and documentation. However, there has been limited focus on how full systems that integrate LLMs are tested during development. Aims: This study explores how LLM-powered systems are tested in the context of real-world application development. Method: We conducted an exploratory case study using 99 individual reports written by students who built and deployed LLM-powered applications as part of a university course. Each report was independently analyzed using thematic analysis, supported by a structured coding process. Results: Testing strategies combined manual and automated methods to evaluate both system logic and model behavior. Common practices included exploratory testing, unit testing, and prompt iteration. Reported challenges included integration failures, unpredictable outputs, prompt sensitivity, hallucinations, and uncertainty about correctness. Conclusions: Testing LLM-powered systems required adaptations to traditional verification methods, blending source-level reasoning with behavior-aware evaluations. These findings provide evidence on the practical context of testing generative components in software systems.

cs.SE

Software Fairness Testing in Practice

Software testing ensures that a system functions correctly, meets specified requirements, and maintains high quality. As artificial intelligence and machine learning (ML) technologies become integral to software systems, testing has evolved to address their unique complexities. A critical advancement in this space is fairness testing, which identifies and mitigates biases in AI applications to promote ethical and equitable outcomes. Despite extensive academic research on fairness testing, including test input generation, test oracle identification, and component testing, practical adoption remains limited. Industry practitioners often lack clear guidelines and effective tools to integrate fairness testing into real-world AI development. This study investigates how software professionals test AI-powered systems for fairness through interviews with 22 practitioners working on AI and ML projects. Our findings highlight a significant gap between theoretical fairness concepts and industry practice. While fairness definitions continue to evolve, they remain difficult for practitioners to interpret and apply. The absence of industry-aligned fairness testing tools further complicates adoption, necessitating research into practical, accessible solutions. Key challenges include data quality and diversity, time constraints, defining effective metrics, and ensuring model interoperability. These insights emphasize the need to bridge academic advancements with actionable strategies and tools, enabling practitioners to systematically address fairness in AI systems.

cs.SE

From Diverse Origins to a DEI Crisis: The Pushback Against Equity, Diversity, and Inclusion in Software Engineering

Background: Diversity, equity, and inclusion are rooted in the very origins of software engineering, shaped by the contributions from many individuals from underrepresented groups to the field. Yet today, DEI efforts in the industry face growing resistance. As companies retreat from visible commitments, and pushback initiatives started only a few years ago. Aims: This study explores how the DEI backlash is unfolding in the software industry by investigating institutional changes, lived experiences, and the strategies used to sustain DEI practices. Method: We conducted an exploratory case study using 59 publicly available Reddit posts authored by self-identified software professionals. Data were analyzed using reflexive thematic analysis. Results: Our findings show that software companies are responding to the DEI backlash in varied ways, including re-structuring programs, scaling back investments, or quietly continuing efforts under new labels. Professionals reported a wide range of emotional responses, from anxiety and frustration to relief and happiness, shaped by identity, role, and organizational culture. Yet, despite the backlash, multiple forms of resistance and adaptation have emerged to protect inclusive practices in software engineering. Conclusions: The DEI backlash is reshaping DEI in software engineering. While public messaging may soften or disappear, core DEI values persist in adapted forms. This study offers a new perspective into how inclusion is evolving under pressure and highlights the resilience of DEI in software environments.

cs.SE

A Delphi Study on the Adaptation of SCRUM Practices to Remote Work

This study explores how Scrum practices were adjusted for remote and hybrid work during and after the COVID-19 pandemic, using a Delphi study with Scrum Masters to gather expert insights. Preliminary key findings highlight communication as the primary challenge, leading to adjustments in meeting structures, information-sharing practices, and collaboration tools. Teams restructured ceremonies, introduced new meetings, and implemented persistent information-sharing mechanisms to improve their work.

cs.SE