SearcharxivSearch

arXiv subjects

Md Mahmudul Hasan

Publications and source records attributed to Md Mahmudul Hasan.

14 recordsLinked to original sources

A Hybrid Framework for Song Lyric Annotation Based on Human-LLM Alignment

Emotion recognition of song lyrics is a challenging task since lyrics may not necessarily align with the overall emotion of a song. As a result, lyrics annotation remains largely underexplored. Drawing inspiration from research in large language model (LLM) assisted annotation, we examine the alignment between humans and LLMs for annotation of lyrics by creating a new sentence-level dataset of lyrics. Our observations highlight the subjectivity of the task and the inherent challenges. Following this, we present a hybrid annotation framework that optimizes human and LLM annotation by predicting potential misalignment in annotation.

cs.CL

Cognitive Edge Device (CED) for Real-Time Environmental Monitoring in Aquatic Ecosystems

Invasive signal crayfish have a detrimental impact on ecosystems. They spread the fungal-type crayfish plague disease (Aphanomyces astaci) that is lethal to the native white clawed crayfish, the only native crayfish species in Britain. Invasive signal crayfish extensively burrow, causing habitat destruction, erosion of river banks and adverse changes in water quality, while also competing with native species for resources leading to declines in native populations. Moreover, pollution exacerbates the vulnerability of White-clawed crayfish, with their populations declining by over 90%. To safeguard aquatic ecosystems, it is imperative to address the challenges posed by invasive species and pollution in aquatic ecosystem's. This article introduces the Cognitive Edge Device (CED) computing platform for the detection of crayfish and plastic. It also presents two publicly available underwater datasets, annotated with sequences of crayfish and aquatic plastic debris. Four You Only Look Once (YOLO) variants were trained and evaluated for crayfish and plastic object detection. YOLOv5s achieved the highest detection accuracy, with an mAP@0.5 of 0.90, and achieved the best precision

cs.CV

Computational Design of Metal-Free Porphyrin Dyes for Sustainable Dye-Sensitized Solar Cells Informing Energy Informatics and Decision Support

This study aims to evaluate the optoelectronic properties of metal free porphyrin-based D-$π$-A dyes via in-silico performance investigation notifying energy informatics and decision support. To develop novel organic dyes, three acceptor/anchoring groups and five donating groups were introduced to strategic positions of the base porphyrin structure, resulting in a total of fifteen dyes. The singlet ground state geometries of the dyes were optimized utilizing density functional theory (DFT) with B3LYP and the excited state optical properties were explored through time-dependent DFT (TD-DFT) using the PCM model with tetrahydrofuran (THF) as solvent. Both DFT and TD-DFT calculations were carried out using the 6-311G(d,p) basis set. The HOMO energy levels of almost all the modified dyes are lower than the redox potential of I$^-$/I$3^-$ and LUMO energy levels are higher than the conduction band of TiO$2$. The absorption maxima values ranged from 690.64 to 975.55 nm. The dye N1 using triphenylamine group as donor and p-ethynylbenzoic acid group as acceptor, showed optimum optoelectronic properties ($ΔG{reg}=-9.73$ eV, $ΔG{inj}=7.18$ eV, $V_{OC}=1.47$ V and $J_{SC}=15.03$ mA/cm$^2$) with highest PCE 14.37%, making it the best studied dye. This newly modified organic dye with enhanced PCE is remarkably effective for the dye-sensitized solar cells (DSSC) industry. Beyond materials discovery, this study highlights the role of high-performance computing in enabling predictive screening of dye candidates and generating performance indicators (HOMO-LUMO gaps, absorption spectra, charge transfer free energies, photovoltaic metrics). These outputs can serve as key parameters for energy informatics and system modelling.

cond-mat.mtrl-sci

WeedScout: Real-Time Autonomous blackgrass Classification and Mapping using dedicated hardware

Blackgrass (Alopecurus myosuroides) is a competitive weed that has wide-ranging impacts on food security by reducing crop yields and increasing cultivation costs. In addition to the financial burden on agriculture, the application of herbicides as a preventive to blackgrass can negatively affect access to clean water and sanitation. The WeedScout project introduces a Real-Rime Autonomous Black-Grass Classification and Mapping (RT-ABGCM), a cutting-edge solution tailored for real-time detection of blackgrass, for precision weed management practices. Leveraging Artificial Intelligence (AI) algorithms, the system processes live image feeds, infers blackgrass density, and covers two stages of maturation. The research investigates the deployment of You Only Look Once (YOLO) models, specifically the streamlined YOLOv8 and YOLO-NAS, accelerated at the edge with the NVIDIA Jetson Nano (NJN). By optimising inference speed and model performance, the project advances the integration of AI into agricultural practices, offering potential solutions to challenges such as herbicide resistance and environmental impact. Additionally, two datasets and model weights are made available to the research community, facilitating further advancements in weed detection and precision farming technologies.

cs.RO

A Machine Learning Approach to Detect Customer Satisfaction From Multiple Tweet Parameters

Since internet technologies have advanced, one of the primary factors in company development is customer happiness. Online platforms have become prominent places for sharing reviews. Twitter is one of these platforms where customers frequently post their thoughts. Reviews of flights on these platforms have become a concern for the airline business. A positive review can help the company grow, while a negative one can quickly ruin its revenue and reputation. So it's vital for airline businesses to examine the feedback and experiences of their customers and enhance their services to remain competitive. But studying thousands of tweets and analyzing them to find the satisfaction of the customer is quite a difficult task. This tedious process can be made easier by using a machine learning approach to analyze tweets to determine client satisfaction levels. Some work has already been done on this strategy to automate the procedure using machine learning and deep learning techniques. However, they are all purely concerned with assessing the text's sentiment. In addition to the text, the tweet also includes the time, location, username, airline name, and so on. This additional information can be crucial for improving the model's outcome. To provide a machine learning based solution, this work has broadened its perspective to include these qualities. And it has come as no surprise that the additional features beyond text sentiment analysis produce better outcomes in machine learning based models.

cs.LG

IKD+: Reliable Low Complexity Deep Models For Retinopathy Classification

Deep neural network (DNN) models for retinopathy have estimated predictive accuracies in the mid-to-high 90%. However, the following aspects remain unaddressed: State-of-the-art models are complex and require substantial computational infrastructure to train and deploy; The reliability of predictions can vary widely. In this paper, we focus on these aspects and propose a form of iterative knowledge distillation(IKD), called IKD+ that incorporates a tradeoff between size, accuracy and reliability. We investigate the functioning of IKD+ using two widely used techniques for estimating model calibration (Platt-scaling and temperature-scaling), using the best-performing model available, which is an ensemble of EfficientNets with approximately 100M parameters. We demonstrate that IKD+ equipped with temperature-scaling results in models that show up to approximately 500-fold decreases in the number of parameters than the original ensemble without a significant loss in accuracy. In addition, calibration scores (reliability) for the IKD+ models are as good as or better than the base mode

cs.LG

Optimizing Return and Secure Disposal of Prescription Opioids to Reduce the Diversion to Secondary Users and Black Market

Opioid Use Disorder (OUD) has reached an epidemic level in the US. Diversion of unused prescription opioids to secondary users and black market significantly contributes to the abuse and misuse of these highly addictive drugs, leading to the increased risk of OUD and accidental opioid overdose within communities. Hence, it is critical to design effective strategies to reduce the non-medical use of opioids that can occur via diversion at the patient level. In this paper, we aim to address this critical public health problem by designing strategies for the return and safe disposal of unused prescription opioids. We propose a data-driven optimization framework to determine the optimal incentive disbursement plans and locations of easily accessible opioid disposal kiosks to motivate prescription opioid users of diverse profiles in returning their unused opioids. We develop a Mixed-Integer Non-Linear Programming (MINLP) model to solve the decision problem, followed by a reformulation scheme using Benders Decomposition that results in a computationally efficient solution. We present a case study to show the benefits and usability of the model using a dataset created from Massachusetts All Payer Claims Data (MA APCD). Our proposed model allows the policymakers to estimate and include a penalty cost considering the economic and healthcare burden associated with prescription opioid diversion. Our numerical experiments demonstrate the ability of model and usefulness in determining optimal locations of opioid disposal kiosks and incentive disbursement plans for maximizing the disposal of unused opioids. The proposed optimization framework offers various trade-off strategies that can help government agencies design pragmatic policies for reducing the diversion of unused prescription opioids.

math.OC

On the Stability of Explicit Finite Difference Methods for Advection-Diffusion Equations

In this paper we study the stability of explicit finite difference discretizations of linear advection-diffusion equations (ADE) with arbitrary order of accuracy in the context of method of lines. The analysis first focuses on the stability of the system of ordinary differential equations (ODE) that is obtained by discretizing the ADE in space and then extends to fully discretized methods where explicit Runge-Kutta methods are used for integrating the ODE system. In particular, it is proved that all stable semi-discretization of the ADE gives rise to a conditionally stable fully discretized method if the time-integrator is at least first-order accurate, whereas high-order spatial discretization of the advection equation cannot yield a stable method if the temporal order is too low. In the second half of this paper, we extend the analysis to a partially dissipative wave system and obtain the stability results for both semi-discretized and fully-discretized methods. Finally, the major theoretical predictions are verified numerically.

math.NA

A Big Data Analytics Framework to Predict the Risk of Opioid Use Disorder

Overdose related to prescription opioids have reached an epidemic level in the US, creating an unprecedented national crisis. This has been exacerbated partly due to the lack of tools for physicians to help predict the risk of whether a patient will develop opioid use disorder. Little is known about how machine learning can be applied to a big-data platform to ensure an informed, sustained and judicious prescribing of opioids, in particular for commercially insured population. This study explores Massachusetts All Payer Claims Data, a de-identified healthcare dataset, and proposes a machine learning framework to examine how naïve users develop opioid use disorder. We perform several feature selections techniques to identify influential demographic and clinical features associated with opioid use disorder from a class imbalanced analytic sample. We then compare the predictive power of four well-known machine learning algorithms: Logistic Regression, Random Forest, Decision Tree, and Gradient Boosting to predict the risk of opioid use disorder. The study results show that the Random Forest model outperforms the other three algorithms while determining the features, some of which are consistent with prior clinical findings. Moreover, alongside the higher predictive accuracy, the proposed framework is capable of extracting some risk factors that will add significant knowledge to what is already known in the extant literature. We anticipate that this study will help healthcare practitioners improve the current prescribing practice of opioids and contribute to curb the increasing rate of opioid addiction and overdose.

stat.AP

Efficient, Effective and Well Justified Estimation of Active Nodes within a Cluster

Reliable and efficient estimation of the size of a dynamically changing cluster in an IoT network is critical in its nominal operation. Most previous estimation schemes worked with relatively smaller frame size and large number of rounds. Here we propose a new estimator named \textquotedblleft Gaussian Estimator of Active Nodes,\textquotedblright (GEAN), that works with large enough frame size under which testing statistics is well approximated as a Gaussian variable, thereby requiring less number of frames, and thus less total number of channel slots to attain a desired accuracy in estimation. More specifically, the selection of the frame size is done according to Triangular Array Central Limit Theorem which also enables us to quantify the approximation error. Larger frame size helps the statistical average to converge faster to the ensemble mean of the estimator and the quantification of the approximation error helps to determine the number of rounds to keep up with the accuracy requirements. We present the analysis of our scheme under two different channel models i.e. $ \{0,1 \} $ and $ \{0,1,e \} $, whereas all previous schemes worked only under $ \{0,1 \} $ channel model. The overall performance of GEAN is better than the previously proposed schemes considering the number of slots required for estimation to achieve a given level of estimation accuracy.

cs.IT

Latent Factor Analysis of Gaussian Distributions under Graphical Constraints

We explore the algebraic structure of the solution space of convex optimization problem Constrained Minimum Trace Factor Analysis (CMTFA), when the population covariance matrix $Σ_x$ has an additional latent graphical constraint, namely, a latent star topology. In particular, we have shown that CMTFA can have either a rank $ 1 $ or a rank $ n-1 $ solution and nothing in between. The special case of a rank $ 1 $ solution, corresponds to the case where just one latent variable captures all the dependencies among the observables, giving rise to a star topology. We found explicit conditions for both rank $ 1 $ and rank $n- 1$ solutions for CMTFA solution of $Σ_x$. As a basic attempt towards building a more general Gaussian tree, we have found a necessary and a sufficient condition for multiple clusters, each having rank $ 1 $ CMTFA solution, to satisfy a minimum probability to combine together to build a Gaussian tree. To support our analytical findings we have presented some numerical demonstrating the usefulness of the contributions of our work.

cs.IT

Algebraic Properties of Wyner Common Information Solution under Graphical Constraints

The Constrained Minimum Determinant Factor Analysis (CMDFA) setting was motivated by Wyner's common information problem where we seek a latent representation of a given Gaussian vector distribution with the minimum mutual information under certain generative constraints. In this paper, we explore the algebraic structures of the solution space of the CMDFA, when the underlying covariance matrix $Σ_x$ has an additional latent graphical constraint, namely, a latent star topology. In particular, sufficient and necessary conditions in terms of the relationships between edge weights of the star graph have been found. Under such conditions and constraints, we have shown that the CMDFA problem has either a rank one solution or a rank $n-1$ solution where $n$ is the dimension of the observable vector. Numerical results are provided to demonstrate the difference between the optimal mutual information and that derived under a naive star constraint.

cs.IT

Algebraic Properties of Wyner Common Information Solution under Graphical Constraints

The Constrained Minimum Determinant Factor Analysis (CMDFA) setting was motivated by Wyner's common information problem where we seek a latent representation of a given Gaussian vector distribution with the minimum mutual information under certain generative constraints. In this paper, we explore the algebraic structures of the solution space of the CMDFA, when the underlying covariance matrix $Σ_x$ has an additional latent graphical constraint, namely, a latent star topology. In particular, sufficient and necessary conditions in terms of the relationships between edge weights of the star graph have been found. Under such conditions and constraints, we have shown that the CMDFA problem has either a rank one solution or a rank $n-1$ solution where $n$ is the dimension of the observable vector. Further results are given in regards to the solution to the CMDFA with $n-1$ latent factors.

cs.IT

Latent Factor Analysis of Gaussian Distributions under Graphical Constraints

In this paper, we explore the algebraic structures of solution spaces for Gaussian latent factor analysis when the population covariance matrix $Σ_x$ has an additional latent graphical constraint, namely, a latent star topology. In particular, we give sufficient and necessary conditions under which the solutions to constrained minimum trace factor analysis (CMTFA) is still star. We further show that the solution to CMTFA under the star constraint can only have two cases, i.e. the number of latent variable can be only one (star) or $n-1$ where $n$ is the dimension of the observable vector, and characterize the solution for both the cases.

cs.IT