SearcharxivSearch

arXiv subjects

Md Shafiqul Islam

Publications and source records attributed to Md Shafiqul Islam.

9 recordsLinked to original sources

Reasoning under Ambiguity: Uncertainty-Aware Multilingual Emotion Classification under Partial Supervision

Contemporary knowledge-based systems increasingly rely on multilingual emotion identification to support intelligent decision-making, yet they face major challenges due to emotional ambiguity and incomplete supervision. Emotion recognition from text is inherently uncertain because multiple emotional states often co-occur and emotion annotations are frequently missing or heterogeneous. Most existing multi-label emotion classification methods assume fully observed labels and rely on deterministic learning objectives, which can lead to biased learning and unreliable predictions under partial supervision. This paper introduces Reasoning under Ambiguity, an uncertainty-aware framework for multilingual multi-label emotion classification that explicitly aligns learning with annotation uncertainty. The proposed approach uses a shared multilingual encoder with language-specific optimization and an entropy-based ambiguity weighting mechanism that down-weights highly ambiguous training instances rather than treating missing labels as negative evidence. A mask-aware objective with positive-unlabeled regularization is further incorporated to enable robust learning under partial supervision. Experiments on English, Spanish, and Arabic emotion classification benchmarks demonstrate consistent improvements over strong baselines across multiple evaluation metrics, along with improved training stability, robustness to annotation sparsity, and enhanced interpretability.

cs.CL

Improving the Safety and Trustworthiness of Medical AI via Multi-Agent Evaluation Loops

Large Language Models (LLMs) are increasingly applied in healthcare, yet ensuring their ethical integrity and safety compliance remains a major barrier to clinical deployment. This work introduces a multi-agent refinement framework designed to enhance the safety and reliability of medical LLMs through structured, iterative alignment. Our system combines two generative models - DeepSeek R1 and Med-PaLM - with two evaluation agents, LLaMA 3.1 and Phi-4, which assess responses using the American Medical Association's (AMA) Principles of Medical Ethics and a five-tier Safety Risk Assessment (SRA-5) protocol. We evaluate performance across 900 clinically diverse queries spanning nine ethical domains, measuring convergence efficiency, ethical violation reduction, and domain-specific risk behavior. Results demonstrate that DeepSeek R1 achieves faster convergence (mean 2.34 vs. 2.67 iterations), while Med-PaLM shows superior handling of privacy-sensitive scenarios. The iterative multi-agent loop achieved an 89% reduction in ethical violations and a 92% risk downgrade rate, underscoring the effectiveness of our approach. This study presents a scalable, regulator-aligned, and cost-efficient paradigm for governing medical AI safety.

cs.AI

The Kernel Manifold: A Geometric Approach to Gaussian Process Model Selection

Gaussian Process (GP) regression is a powerful nonparametric Bayesian framework, but its performance depends critically on the choice of covariance kernel. Selecting an appropriate kernel is therefore central to model quality, yet remains one of the most challenging and computationally expensive steps in probabilistic modeling. We present a Bayesian optimization framework built on kernel-of-kernels geometry, using expected divergence-based distances between GP priors to explore kernel space efficiently. A multidimensional scaling (MDS) embedding of this distance matrix maps a discrete kernel library into a continuous Euclidean manifold, enabling smooth BO. In this formulation, the input space comprises kernel compositions, the objective is the log marginal likelihood, and featurization is given by the MDS coordinates. When the divergence yields a valid metric, the embedding preserves geometry and produces a stable BO landscape. We demonstrate the approach on synthetic benchmarks, real-world time-series datasets, and an additive manufacturing case study predicting melt-pool geometry, achieving superior predictive accuracy and uncertainty calibration relative to baselines including Large Language Model (LLM)-guided search. This framework establishes a reusable probabilistic geometry for kernel search, with direct relevance to GP modeling and deep kernel learning.

cs.LG

DataScribe: An AI-Native, Policy-Aligned Web Platform for Multi-Objective Materials Design and Discovery

The acceleration of materials discovery requires digital platforms that go beyond data repositories to embed learning, optimization, and decision-making directly into research workflows. We introduce DataScribe, an AI-native, cloud-based materials discovery platform that unifies heterogeneous experimental and computational data through ontology-backed ingestion and machine-actionable knowledge graphs. The platform integrates FAIR-compliant metadata capture, schema and unit harmonization, uncertainty-aware surrogate modeling, and native multi-objective multi-fidelity Bayesian optimization, enabling closed-loop propose-measure-learn workflows across experimental and computational pipelines. DataScribe functions as an application-layer intelligence stack, coupling data governance, optimization, and explainability rather than treating them as downstream add-ons. We validate the platform through case studies in electrochemical materials and high-entropy alloys, demonstrating end-to-end data fusion, real-time optimization, and reproducible exploration of multi-objective trade spaces. By embedding optimization engines, machine learning, and unified access to public and private scientific data directly within the data infrastructure, and by supporting open, free use for academic and non-profit researchers, DataScribe functions as a general-purpose application-layer backbone for laboratories of any scale, including self-driving laboratories and geographically distributed materials acceleration platforms, with built-in support for performance, sustainability, and supply-chain-aware objectives.

cs.LG

Emotion Detection in Speech Using Lightweight and Transformer-Based Models: A Comparative and Ablation Study

Emotion recognition from speech plays a vital role in the development of empathetic human-computer interaction systems. This paper presents a comparative analysis of lightweight transformer-based models, DistilHuBERT and PaSST, by classifying six core emotions from the CREMA-D dataset. We benchmark their performance against a traditional CNN-LSTM baseline model using MFCC features. DistilHuBERT demonstrates superior accuracy (70.64%) and F1 score (70.36%) while maintaining an exceptionally small model size (0.02 MB), outperforming both PaSST and the baseline. Furthermore, we conducted an ablation study on three variants of the PaSST, Linear, MLP, and Attentive Pooling heads, to understand the effect of classification head architecture on model performance. Our results indicate that PaSST with an MLP head yields the best performance among its variants but still falls short of DistilHuBERT. Among the emotion classes, angry is consistently the most accurately detected, while disgust remains the most challenging. These findings suggest that lightweight transformers like DistilHuBERT offer a compelling solution for real-time speech emotion recognition on edge devices. The code is available at: https://github.com/luckymaduabuchi/Emotion-detection-.

cs.SD

Ulam's method for computing stationary densities of invariant measures for piecewise convex maps with countably infinite number of branches

Let $τ: I=[0, 1]\to [0, 1]$ be a piecewise convex map with countably infinite number of branches. In \cite{GIR}, the existence of absolutely continuous invariant measure (ACIM) $μ$ for $τ$ and the exactness of the system $(τ, μ)$ has been proven. In this paper, we develop an Ulam method for approximation of $f^*$, the density of ACIM $μ$. We construct a sequence $\{τ_n\}_{n=1}^\infty$ of maps $τ_n: I\to I$ s. t. $τ_n$ has a finite number of branches and the sequence $τ_n$ converges to $τ$ almost uniformly. Using supremum norms and Lasota-Yorke type inequalities, we prove the existence of ACIMs $μ_n$ for $τ_n$ with the densities $f_n$. For a fixed $n$, we apply Ulam's method with $k$ subintervals to $τ_n$ and compute approximations $f_{n,k}$ of $f_n$. We prove that $f_{n,k}\to f^*$ as $n\to \infty, k\to \infty,$ both a.e. and in $L^1$. We provide examples of piecewise convex maps $τ$ with countably infinite number of branches, their approximations $τ_n$ with finite number of branches and for increasing values of parameter $k$ show the errors $\|f^*-f_{n,k}\|_1$.

math.DS

Frozen Mode Regime in an Optical Waveguide With Distributed Bragg Reflector

We introduce a glide symmetric optical waveguide exhibiting a stationary inflection point (SIP) in the Bloch wavenumber dispersion relation. An SIP is a third order exceptional point of degeneracy (EPD) where three Bloch eigenmodes coalesce to form a so-called frozen mode with vanishing group velocity and diverging amplitude. We show that the incorporation of chirped distributed Bragg reflectors and distributed coupling between waveguides in the periodic structure facilitates the SIP formation and greatly enhances the characteristics of the frozen mode regime. We confirm the existence of an SIP in two ways: by observing the flatness of the dispersion diagram and also by using a coalescence parameter describing the separation of the three eigenvectors collapsing on each other. We find that in the absence of losses, both the quality factor and the group delay at the SIP grow with the cubic power of the cavity length. The frozen mode regime can be very attractive for light amplification and lasing, in optical delay lines, sensors, and modulators.

physics.optics

Design of a Modified Coupled Resonators Optical Waveguide Supporting a Frozen Mode

We design a three-way silicon optical waveguide with the Bloch dispersion relation supporting a stationary inflection point (SIP). The SIP is a third order exceptional point of degeneracy (EPD) where three Bloch modes coalesce forming the frozen mode with greatly enhanced amplitude. The proposed design consists of a coupled resonators optical waveguide (CROW) coupled to a parallel straight waveguide. At any given frequency, this structure supports three pairs of reciprocal Bloch eigenmodes, propagating and/or evanescent. In addition to full-wave simulations, we also employ a so-called ''hybrid model'' that uses transfer matrices obtained from full-wave simulations of sub-blocks of the unit cell. This allows us to account for radiation losses and enables a design procedure based on minimizing the eigenmodes' coalescence parameter. The proposed finite-length CROW displays almost unitary transfer function at the SIP frequency, implying a nearly perfect conversion of the input light into the frozen mode. The group delay and the effective quality factor at the SIP frequency show an $N^{3}$ scaling, where $N$ is the number of unit cells in the cavity. The frozen mode in the CROW can be utilized in various applications like sensors, lasers and optical delay lines.

physics.optics

Vapor Cloud Delayed-DPPM Modulation Technique for nonlinear Optoacoustic Communication

The optoacoustic process can solve the longstanding challenge of wireless information transmission from an airborne unit to an underwater node (UWN). The nonlinear optoacoustic signal generated by proper laser parameters can propagate long distances in water. However, forming such a signal requires a high-power laser, and the buildup of a vapor cloud precludes the subsequent acoustic signal generation. Therefore, pursuing the traditional on-off keying (OOK) modulation technique will limit the data rate and power efficiency. In this paper, we analyze different modulation techniques and propose a vapor cloud delayed-differential pulse position modulation (VCD-DPPM) technique to improve the data rate and achieve high power efficiency for a single stationary laser transmitter. The symbol rate of VCD-DPPM is approximately 6.9 times and 1.69 times higher than OOK in our text communication simulation using a laser repetition rate of 10 kHz and 40 Hz, respectively. Furthermore, VCD-DPPM is 137% more power efficient than the OOK technique for both cases. We have generated different acoustic signal levels in laboratory conditions and simulated the bit error rate (BER) for different depths and positions of the UWN, while considering ambient underwater noises. Our results indicate that VCD-DPPM enables efficient data transmission.

eess.SP