Searcharxiv⌕ Search

arXiv subjects

Md Mustakim Billah

Publications and source records attributed to Md Mustakim Billah.

3 recordsLinked to original sources

Why Do Pull Requests Go Silent? Uncovering the Barriers to Contribution Completion in Open-Source Code Review

Pull requests (PRs) underpin pull-based software development by enabling distributed code review and collaborative contribution in open-source projects. Yet many become inactive before integration and are eventually abandoned or closed, wasting contributor and maintainer effort. Although prior work has examined PR abandonment and review delays, less is known about the contribution types, discussion-level barriers, and post-stalling collaboration patterns associated with inactivity. We investigate which PR types most often stall, why inactivity occurs from authors' and reviewers' perspectives, and how stalling relates to later contributor and reviewer engagement. We analyzed 14,234 stalled PRs and 164,562 review comments from 19 popular GitHub repositories using stale-bot workflows. An LLM-based voting classifier categorized PRs by contribution type, while quantitative analysis was combined with qualitative coding of general and inline review discussions. Feature-enhancement and issue-fixing PRs formed the largest share, together exceeding 77% of classified stalled PRs. General comments linked stalling mainly to communication and coordination breakdowns, including missing interaction, delayed feedback, and unclear follow-up. Inline comments showed that inactivity does not always reflect disengagement: many PRs were blocked by technical or dependency issues, including failing checks, configuration problems, compatibility concerns, and environment mismatches. Only 39.56% of contributors later submitted another PR, and reviewer re-engagement with the same contributors was approximately 21%. PR inactivity is a socio-technical coordination problem involving communication, technical readiness, review ownership, and automation practices. We recommend type-aware triage, clearer review feedback, explicit ownership of next actions, CI blocker management, and cause-aware stale-bot interventions.

cs.SE↗

Do Automatic Comment Generation Techniques Fall Short? Exploring the Influence of Method Dependencies on Code Understanding

Method-level comments are critical for improving code comprehension and supporting software maintenance. With advancements in large language models (LLMs), automated comment generation has become a major research focus. However, existing approaches often overlook method dependencies, where one method relies on or calls others, affecting comment quality and code understandability. This study investigates the prevalence and impact of dependent methods in software projects and introduces a dependency-aware approach for method-level comment generation. Analyzing a dataset of 10 popular Java GitHub projects, we found that dependent methods account for 69.25% of all methods and exhibit higher engagement and change proneness compared to independent methods. Across 448K dependent and 199K independent methods, we observed that state-of-the-art fine-tuned models (e.g., CodeT5+, CodeBERT) struggle to generate comprehensive comments for dependent methods, a trend also reflected in LLM-based approaches like ASAP. To address this, we propose HelpCOM, a novel dependency-aware technique that incorporates helper method information to improve comment clarity, comprehensiveness, and relevance. Experiments show that HelpCOM outperforms baseline methods by 5.6% to 50.4% across syntactic (e.g., BLEU), semantic (e.g., SentenceBERT), and LLM-based evaluation metrics. A survey of 156 software practitioners further confirms that HelpCOM significantly improves the comprehensibility of code involving dependent methods, highlighting its potential to enhance documentation, maintainability, and developer productivity in large-scale systems.

cs.SE↗

Are Large Language Models a Threat to Programming Platforms? An Exploratory Study

Competitive programming platforms like LeetCode, Codeforces, and HackerRank evaluate programming skills, often used by recruiters for screening. With the rise of advanced Large Language Models (LLMs) such as ChatGPT, Gemini, and Meta AI, their problem-solving ability on these platforms needs assessment. This study explores LLMs' ability to tackle diverse programming challenges across platforms with varying difficulty, offering insights into their real-time and offline performance and comparing them with human programmers. We tested 98 problems from LeetCode, 126 from Codeforces, covering 15 categories. Nine online contests from Codeforces and LeetCode were conducted, along with two certification tests on HackerRank, to assess real-time performance. Prompts and feedback mechanisms were used to guide LLMs, and correlations were explored across different scenarios. LLMs, like ChatGPT (71.43% success on LeetCode), excelled in LeetCode and HackerRank certifications but struggled in virtual contests, particularly on Codeforces. They performed better than users in LeetCode archives, excelling in time and memory efficiency but underperforming in harder Codeforces contests. While not immediately threatening, LLMs performance on these platforms is concerning, and future improvements will need addressing.

cs.SE↗