Searcharxiv⌕ Search

arXiv subjects

MD Tamim Hossain

Publications and source records attributed to MD Tamim Hossain.

3 recordsLinked to original sources

Vision-Based Calorie Estimation for Bangladeshi Street Food: A Comparative Study of Detection and Regression Models

With obesity emerging as a major global health concern, accurate calorie estimation systems have become increasingly important for effective dietary management. Current vision-based approaches are inappropriate for Bangladeshi street food, which is widely consumed and culturally significant, because they mostly focus on Western cuisines and often overlook portion size. The purpose of this research is to offer a vision-based calorie estimation methodology that was created especially for street food in Bangladesh. Training, validation and testing splits were created from a proprietary dataset of 3,885 photos from six classes (Singara, Somusa, Puri, Peaju, Beguni, and Coin as a reference). Five detection and segmentation architectures, YOLOv8n, YOLO11n, YOLO12n, YOLO26n and RF-DETR, were methodically compared. Food dimensions were scaled using a Bangladeshi 5 Taka coin as a reference. To predict calories, extracted geometric characteristics were subsequently fed into machine learning regression models such as Random Forest, Gradient Boost and AdaBoost. YOLO11n outperformed other models with the best detection performance, achieving 96.1% mAP@50 and balanced mask metrics. With a mean absolute error (MAE) of 5.68, root mean squared error (RMSE) of 7.23, and an R2 score of 95.0%, Random Forest regression produced the best results for calorie estimation. YOLO11n in conjunction with Random Forest regression offers a precise and effective calorie estimation for street food in Bangladesh, which is useful for dietary monitoring and mobile health apps. The code and dataset are publicly available on https://github.com/hossain-tamim/BD-CAL .

cs.CV↗

Integrating Skeleton Based Representations for Robust Yoga Pose Classification Using Deep Learning Models

Yoga is a popular form of exercise worldwide due to its spiritual and physical health benefits, but incorrect postures can lead to injuries. Automated yoga pose classification has therefore gained importance to reduce reliance on expert practitioners. While human pose keypoint extraction models have shown high potential in action recognition, systematic benchmarking for yoga pose recognition remains limited, as prior works often focus solely on raw images or a single pose extraction model. In this study, we introduce a curated dataset, 'Yoga-16', which addresses limitations of existing datasets, and systematically evaluate three deep learning architectures (VGG16, ResNet50, and Xception), using three input modalities (direct images, MediaPipe Pose skeleton images, and YOLOv8 Pose skeleton images). Our experiments demonstrate that skeleton-based representations outperform raw image inputs, with the highest accuracy of 96.09% achieved by VGG16 with MediaPipe Pose skeleton input. Additionally, we provide interpretability analysis using Grad-CAM, offering insights into model decision-making for yoga pose classification with cross-validation analysis.

cs.CV↗

Evaluating YOLO Architectures: Implications for Real-Time Vehicle Detection in Urban Environments of Bangladesh

Vehicle detection systems trained on Non-Bangladeshi datasets struggle to accurately identify local vehicle types in Bangladesh's unique road environments, creating critical gaps in autonomous driving technology for developing regions. This study evaluates six YOLO model variants on a custom dataset featuring 29 distinct vehicle classes, including region-specific vehicles such as ``Desi Nosimon'', ``Leguna'', ``Battery Rickshaw'', and ``CNG''. The dataset comprises high-resolution images (1920x1080) captured across various Bangladeshi roads using mobile phone cameras and manually annotated using LabelImg with YOLO format bounding boxes. Performance evaluation revealed YOLOv11x as the top performer, achieving 63.7\% mAP@0.5, 43.8\% mAP@0.5:0.95, 61.4\% recall, and 61.6\% F1-score, though requiring 45.8 milliseconds per image for inference. Medium variants (YOLOv8m, YOLOv11m) struck an optimal balance, delivering robust detection performance with mAP@0.5 values of 62.5\% and 61.8\% respectively, while maintaining moderate inference times around 14-15 milliseconds. The study identified significant detection challenges for rare vehicle classes, with Construction Vehicles and Desi Nosimons showing near-zero accuracy due to dataset imbalances and insufficient training samples. Confusion matrices revealed frequent misclassifications between visually similar vehicles, particularly Mini Trucks versus Mini Covered Vans. This research provides a foundation for developing robust object detection systems specifically adapted to Bangladesh traffic conditions, addressing critical needs in autonomous vehicle technology advancement for developing regions where conventional generic-trained models fail to perform adequately.

cs.CV↗