Searcharxiv⌕ Search

arXiv subjects

Md Nazmus Sakib

Publications and source records attributed to Md Nazmus Sakib.

11 recordsLinked to original sources

DeltaSeek: Toward Active Perception in Evolving Construction Environments

Construction environments evolve continuously, causing large geometric changes that degrade static mapping and registration performance. This necessitates active perception, where robots deliberately select sensing configurations to resolve the environment's current state. We present DeltaSeek, an initial framework toward active perception in evolving built environments. While our broader objective is a system that reasons about where, how, and when to observe, this paper addresses a critical prerequisite: how a robot's sensing embodiment constrains the observations it can acquire. We formalize an embodiment's permissible observation set and evaluate with a Husky A300 equipped with a UR5e on an IFC-derived benchmark under chassis-mounted and wrist-mounted RGB-D configurations, scoring observations by geometric visibility and effort by drivable distance. In a room-scale scene with eight controlled changes spanning four observability conditions, exhaustive evaluation over 240 permissible base poses and five arm postures shows that two changes admit no chassis viewpoint whatsoever, while the wrist camera resolves both. For changes observed by both embodiments, the median base travel is $6.0$~m for the wrist camera and $15.2$~m for the chassis camera. These results distinguish sensing limitations from acquisition costs, clarifying whether an observation is impossible or simply requires more travel.

cs.RO↗

Push-Pull Determinants Among Bangladeshi Students Enrolled in NCR Private Universities: A Single-Destination Exploratory Study

International student mobility from Bangladesh is a significant feature of South Asian higher education, yet India's National Capital Region (NCR) remains underexplored as a destination for outbound Bangladeshi students. This exploratory single-destination study examines push-pull factors among Bangladeshi students enrolled at private universities in India's NCR. A structured online survey was administered to students at Sharda University, Noida International University, and Galgotias University (n = 63; n = 56 after quality filtering). Descriptive statistics, K-means clustering, and binary logistic regression were applied within Lee's push-pull framework. Political and administrative disruption in Bangladesh's academic calendar was the leading push factor (M = 3.73). Geographical and cultural proximity was the strongest pull factor (M = 3.80), closely followed by visa accessibility (M = 3.73), with no significant mean difference between them. Australia and Germany were the most frequently considered alternative destinations (33.9% each). Advisory networks influenced 73.2% of respondents under a broad threshold and 66.1% under a stricter threshold, mainly for university and course selection. Satisfaction and recommendation intent were positively associated, and infrastructure satisfaction showed the strongest association with high recommendation intent in a five-predictor logistic model (OR = 2.54, 95% CI [1.09, 5.89]). K-means clustering produced three exploratory decision-profile groups: Comprehensively Motivated, Proximity-Led Enrollers, and Low-Salience Enrollers. The findings suggest that geographic nearness, cultural familiarity, and visa accessibility may operate as a composite accessibility advantage in this short-haul intra-regional corridor.

cs.CY↗

ReflectEd: Evaluating Reflection-Driven Learning in an AI-Assisted System

In collaborative settings, sustaining momentum and engagement between checkpoints (e.g., meetings) can be challenging, often leading to task drift and reduced preparedness. To address this gap, we developed ReflectEd, an AI-assisted system that supports between-checkpoint reflection through theory-driven prompts with progressively structured levels and mechanism-based scaffolding. We evaluated ReflectEd in a mixed-method study comparing two reflection configurations: a regular reflection workflow and a deeper reflection workflow that included an additional transformative reflection activity. Across conditions, participants reported steady engagement early in the week. In the deeper configuration, later reflections tended to exhibit higher actionability and richer forward-looking planning, while also being harder to sustain and more effortful during periods of active work. Partner-visible reflections were frequently described as supporting coordination by surfacing differences in focus and facilitating accountability. Overall, the findings characterize trade-offs between reflection depth, feasibility, and perceived preparedness for subsequent checkpoints. We discuss implications for the design of AI-assisted systems that support collaboration readiness and reflection-oriented regulation in time-constrained collaborative workflows.

cs.HC↗

Expecting Too Much, Getting Too Little: Exploring the Challenges and Design Opportunities of Asynchronous AI Interviewers

Organizations use asynchronous AI interview systems to efficiently manage large applicant pools, enabling quick and uniform evaluations. However, concerns remain about their impact on user agency and the lack of personalization applicants experience with these systems. Although efforts have been made to humanize the interview process, users' expectations are often unmet, especially when compared to the promises made by these systems. To examine how applicants perceive and experience these tools, particularly in the context of their growing familiarity with large language models (LLMs), we conducted a two-phase study. The first phase involved an analysis of 11 subreddit discussions on interview experiences with asynchronous AI interviewers, followed by a semi-structured interview study with 17 participants. Qualitative analysis revealed key issues such as mismatched expectations, amplified by organizational rhetoric and applicant expectations shaped by experiences with LLMs. These factors shaped participants' sense of agency and trust, often leading to workarounds and deceptive practices. In the follow-up study, we designed an interface with two features, response variants and feedback variants, and evaluated it across six groups (N = 180, 30 participants each) to assess whether these features support users' sense of agency, competence, and relatedness. Our analysis suggests that even subtle design changes can enhance user autonomy and that carefully designed feedback can provide meaningful support in high-stakes interview contexts.

cs.HC↗

Competing or Collaborating? The Role of Hackathon Formats in Shaping Team Dynamics and Project Choices

Hackathons have emerged as dynamic platforms for fostering innovation, collaboration, and skill development in the technology sector. Structural differences across hackathon formats raise important questions about how event design can shape student learning experiences and engagement. This study examines two distinct hackathon formats: a gender-specific hackathon (GS) and a regular institutional hackathon (RI). Using a mixed-methods approach, we analyze variations in team dynamics, project themes, role assignments, and environmental settings. Our findings indicate that GS hackathon foster a collaborative and supportive atmosphere, emphasizing personal growth and community learning, with projects often centered on health and well-being. In contrast, RI hackathon tend to promote a competitive, outcome-driven environment, with projects frequently addressing entertainment and environmental sustainability. Based on these insights, we propose a hybrid hackathon model that combines the strengths of both formats to balance competition with inclusivity. This work contributes to the design of more engaging, equitable, and pedagogically effective hackathon experiences.

cs.HC↗

Scene Graph-Guided Generative AI Framework for Synthesizing and Evaluating Industrial Hazard Scenarios

Training vision models to detect workplace hazards accurately requires realistic images of unsafe conditions that could lead to accidents. However, acquiring such datasets is difficult because capturing accident-triggering scenarios as they occur is nearly impossible. To overcome this limitation, this study presents a novel scene graph-guided generative AI framework that synthesizes photorealistic images of hazardous scenarios grounded in historical Occupational Safety and Health Administration (OSHA) accident reports. OSHA narratives are analyzed using GPT-4o to extract structured hazard reasoning, which is converted into object-level scene graphs capturing spatial and contextual relationships essential for understanding risk. These graphs guide a text-to-image diffusion model to generate compositionally accurate hazard scenes. To evaluate the realism and semantic fidelity of the generated data, a visual question answering (VQA) framework is introduced. Across four state-of-the-art generative models, the proposed VQA Graph Score outperforms CLIP and BLIP metrics based on entropy-based validation, confirming its higher discriminative sensitivity.

cs.AI↗

Sketch2BIM: A Multi-Agent Human-AI Collaborative Pipeline to Convert Hand-Drawn Floor Plans to 3D BIM

This study introduces a human-in-the-loop pipeline that converts unscaled, hand-drawn floor plan sketches into semantically consistent 3D BIM models. The workflow leverages multimodal large language models (MLLMs) within a multi-agent framework, combining perceptual extraction, human feedback, schema validation, and automated BIM scripting. Initially, sketches are iteratively refined into a structured JSON layout of walls, doors, and windows. Later, these layouts are transformed into executable scripts that generate 3D BIM models. Experiments on ten diverse floor plans demonstrate strong convergence: openings (doors, windows) are captured with high reliability in the initial pass, while wall detection begins around 83% and achieves near-perfect alignment after a few feedback iterations. Across all categories, precision, recall, and F1 scores remain above 0.83, and geometric errors (RMSE, MAE) progressively decrease to zero through feedback corrections. This study demonstrates how MLLM-driven multi-agent reasoning can make BIM creation accessible to both experts and non-experts using only freehand sketches.

cs.AI↗

An Interdisciplinary Review of Commonsense Reasoning and Intent Detection

This review explores recent advances in commonsense reasoning and intent detection, two key challenges in natural language understanding. We analyze 28 papers from ACL, EMNLP, and CHI (2020-2025), organizing them by methodology and application. Commonsense reasoning is reviewed across zero-shot learning, cultural adaptation, structured evaluation, and interactive contexts. Intent detection is examined through open-set models, generative formulations, clustering, and human-centered systems. By bridging insights from NLP and HCI, we highlight emerging trends toward more adaptive, multilingual, and context-aware models, and identify key gaps in grounding, generalization, and benchmark design.

cs.CL↗

Automatic Pull Request Description Generation Using LLMs: A T5 Model Approach

Developers create pull request (PR) descriptions to provide an overview of their changes and explain the motivations behind them. These descriptions help reviewers and fellow developers quickly understand the updates. Despite their importance, some developers omit these descriptions. To tackle this problem, we propose an automated method for generating PR descriptions based on commit messages and source code comments. This method frames the task as a text summarization problem, for which we utilized the T5 text-to-text transfer model. We fine-tuned a pre-trained T5 model using a dataset containing 33,466 PRs. The model's effectiveness was assessed using ROUGE metrics, which are recognized for their strong alignment with human evaluations. Our findings reveal that the T5 model significantly outperforms LexRank, which served as our baseline for comparison.

cs.LG↗

Risks, Causes, and Mitigations of Widespread Deployments of Large Language Models (LLMs): A Survey

Recent advancements in Large Language Models (LLMs), such as ChatGPT and LLaMA, have significantly transformed Natural Language Processing (NLP) with their outstanding abilities in text generation, summarization, and classification. Nevertheless, their widespread adoption introduces numerous challenges, including issues related to academic integrity, copyright, environmental impacts, and ethical considerations such as data bias, fairness, and privacy. The rapid evolution of LLMs also raises concerns regarding the reliability and generalizability of their evaluations. This paper offers a comprehensive survey of the literature on these subjects, systematically gathered and synthesized from Google Scholar. Our study provides an in-depth analysis of the risks associated with specific LLMs, identifying sub-risks, their causes, and potential solutions. Furthermore, we explore the broader challenges related to LLMs, detailing their causes and proposing mitigation strategies. Through this literature analysis, our survey aims to deepen the understanding of the implications and complexities surrounding these powerful models.

cs.CL↗

Exploring the Influence of Online Videos on Parents or Caregivers of Children with Developmental Delays

Developmental Delays and Disabilities (DDDs) refer to conditions where children are slower or unable to reach developmental milestones compared to typically developing children. This can cause significant stress for parents, leading to social isolation and loneliness. Online videos, particularly those on YouTube, aim to support these parents and caregivers by offering guidance and assistance. Studies show that parents of children with DDDs create videos on YouTube to enhance authenticity and build connections. However, there is limited knowledge about how other parents with children with DDDs perceive and are impacted by these videos. Our study used a mixed-method approach to annotate and analyze more than fifteen hundred YouTube videos on children's DDDs. We found that these videos provide crucial informational content and offer mental and emotional support through shared personal experiences. Comments analysis revealed a strong sense of community among YouTubers and viewers. Interviews with parents of children with DDDs showed that they find these videos relatable and essential for managing their children's diagnosis and treatments. We concluded by discussing platform-centric design implications for supporting parents and other caregivers of children with DDDs.

cs.HC↗