Searcharxiv⌕ Search

arXiv subjects

Deepak Gupta

Publications and source records attributed to Deepak Gupta.

At least 55 records · Page 3Linked to original sources

Towards Answering Health-related Questions from Medical Videos: Datasets and Approaches

The increase in the availability of online videos has transformed the way we access information and knowledge. A growing number of individuals now prefer instructional videos as they offer a series of step-by-step procedures to accomplish particular tasks. The instructional videos from the medical domain may provide the best possible visual answers to first aid, medical emergency, and medical education questions. Toward this, this paper is focused on answering health-related questions asked by the public by providing visual answers from medical videos. The scarcity of large-scale datasets in the medical domain is a key challenge that hinders the development of applications that can help the public with their health-related questions. To address this issue, we first proposed a pipelined approach to create two large-scale datasets: HealthVidQA-CRF and HealthVidQA-Prompt. Later, we proposed monomodal and multimodal approaches that can effectively provide visual answers from medical videos to natural language questions. We conducted a comprehensive analysis of the results, focusing on the impact of the created datasets on model training and the significance of visual features in enhancing the performance of the monomodal and multi-modal approaches. Our findings suggest that these datasets have the potential to enhance the performance of medical visual answer localization tasks and provide a promising future direction to further enhance the performance by using pre-trained language-vision models.

cs.CL↗

Efficient control protocols for an active Ornstein-Uhlenbeck particle

Designing a protocol to efficiently drive a stochastic system is an active field of research. Here we extend such control theory to an active Ornstein-Uhlenbeck particle (AOUP) in a bistable potential, driven by a harmonic trap. We find that protocols designed to minimize the excess work (up to linear-response) perform better than naive protocols with constant velocity for a wide range of protocol durations.

cond-mat.stat-mech↗

Large Scale Generative Multimodal Attribute Extraction for E-commerce Attributes

E-commerce websites (e.g. Amazon) have a plethora of structured and unstructured information (text and images) present on the product pages. Sellers often either don't label or mislabel values of the attributes (e.g. color, size etc.) for their products. Automatically identifying these attribute values from an eCommerce product page that contains both text and images is a challenging task, especially when the attribute value is not explicitly mentioned in the catalog. In this paper, we present a scalable solution for this problem where we pose attribute extraction problem as a question-answering task, which we solve using \textbf{MXT}, consisting of three key components: (i) \textbf{M}AG (Multimodal Adaptation Gate), (ii) \textbf{X}ception network, and (iii) \textbf{T}5 encoder-decoder. Our system consists of a generative model that \emph{generates} attribute-values for a given product by using both textual and visual characteristics (e.g. images) of the product. We show that our system is capable of handling zero-shot attribute prediction (when attribute value is not seen in training data) and value-absent prediction (when attribute value is not mentioned in the text) which are missing in traditional classification-based and NER-based models respectively. We have trained our models using distant supervision, removing dependency on human labeling, thus making them practical for real-world applications. With this framework, we are able to train a single model for 1000s of (product-type, attribute) pairs, thus reducing the overhead of training and maintaining separate models. Extensive experiments on two real world datasets show that our framework improves the absolute recall@90P by 10.16\% and 6.9\% from the existing state of the art models. In a popular e-commerce store, we have deployed our models for 1000s of (product-type, attribute) pairs.

cs.CV↗

Aux-Drop: Handling Haphazard Inputs in Online Learning Using Auxiliary Dropouts

Many real-world applications based on online learning produce streaming data that is haphazard in nature, i.e., contains missing features, features becoming obsolete in time, the appearance of new features at later points in time and a lack of clarity on the total number of input features. These challenges make it hard to build a learnable system for such applications, and almost no work exists in deep learning that addresses this issue. In this paper, we present Aux-Drop, an auxiliary dropout regularization strategy for online learning that handles the haphazard input features in an effective manner. Aux-Drop adapts the conventional dropout regularization scheme for the haphazard input feature space ensuring that the final output is minimally impacted by the chaotic appearance of such features. It helps to prevent the co-adaptation of especially the auxiliary and base features, as well as reduces the strong dependence of the output on any of the auxiliary inputs of the model. This helps in better learning for scenarios where certain features disappear in time or when new features are to be modelled. The efficacy of Aux-Drop has been demonstrated through extensive numerical experiments on SOTA benchmarking datasets that include Italy Power Demand, HIGGS, SUSY and multiple UCI datasets. The code is available at https://github.com/Rohit102497/Aux-Drop.

cs.LG↗

Empowering Language Model with Guided Knowledge Fusion for Biomedical Document Re-ranking

Pre-trained language models (PLMs) have proven to be effective for document re-ranking task. However, they lack the ability to fully interpret the semantics of biomedical and health-care queries and often rely on simplistic patterns for retrieving documents. To address this challenge, we propose an approach that integrates knowledge and the PLMs to guide the model toward effectively capturing information from external sources and retrieving the correct documents. We performed comprehensive experiments on two biomedical and open-domain datasets that show that our approach significantly improves vanilla PLMs and other existing approaches for document re-ranking task.

cs.CL↗

Fluctuations of entropy production of a run-and-tumble particle

Out-of-equilibrium systems continuously generate entropy, with its rate of production being a fingerprint of non-equilibrium conditions. In small-scale dissipative systems subject to thermal noise, fluctuations of entropy production are significant. Hitherto, mean and variance have been abundantly studied, even if higher moments might be important to fully characterize the system of interest. Here, we introduce a graphical method to compute any moment of entropy production for a generic discrete-state system. Then, we focus on a paradigmatic model of active particles, i.e., run-and-tumble dynamics, which resembles the motion observed in several microorganisms. Employing our framework, we compute the first three cumulants of the entropy production for a discrete version of this model. We also compare our analytical results with numerical simulations. We find that as the number of states increases, the distribution of entropy production deviates from a Gaussian. Finally, we extend our framework to a continuous state-space run-and-tumble model, using an appropriate scaling of the transition rates. The approach here presented might help uncover the features of non-equilibrium fluctuations of any current in biological systems operating out-of-equilibrium.

cond-mat.stat-mech↗

Optimal Control of the F${_1}$-ATPase Molecular Motor

F$_{1}$-ATPase is a rotary molecular motor that \emph{in vivo} is subject to strong nonequilibrium driving forces. There is great interest in understanding the operational principles governing its high efficiency of free-energy transduction. Here we use a near-equilibrium framework to design a non-trivial control protocol to minimize dissipation in rotating F$_{1}$ to synthesize ATP. We find that the designed protocol requires much less work than a naive (constant-velocity) protocol across a wide range of protocol durations. Our analysis points to a possible mechanism for energetically efficient driving of F$_{1}$ \emph{in vivo} and provides insight into free-energy transduction for a broader class of biomolecular and synthetic machines.

cond-mat.stat-mech↗

Classifying text using machine learning models and determining conversation drift

Text classification helps analyse texts for semantic meaning and relevance, by mapping the words against this hierarchy. An analysis of various types of texts is invaluable to understanding both their semantic meaning, as well as their relevance. Text classification is a method of categorising documents. It combines computer text classification and natural language processing to analyse text in aggregate. This method provides a descriptive categorization of the text, with features like content type, object field, lexical characteristics, and style traits. In this research, the authors aim to use natural language feature extraction methods in machine learning which are then used to train some of the basic machine learning models like Naive Bayes, Logistic Regression, and Support Vector Machine. These models are used to detect when a teacher must get involved in a discussion when the lines go off-topic.

cs.LG↗

Machine Learning enabled models for YouTube Ranking Mechanism and Views Prediction

With the continuous increase of internet usage in todays time, everyone is influenced by this source of the power of technology. Due to this, the rise of applications and games Is unstoppable. A major percentage of our population uses these applications for multiple purposes. These range from education, communication, news, entertainment, and many more. Out of this, the application that is making sure that the world stays in touch with each other and with current affairs is social media. Social media applications have seen a boom in the last 10 years with the introduction of smartphones and the internet being available at affordable prices. Applications like Twitch and Youtube are some of the best platforms for producing content and expressing their talent as well. It is the goal of every content creator to post the best and most reliable content so that they can gain recognition. It is important to know the methods of achieving popularity easily, which is what this paper proposes to bring to the spotlight. There should be certain parameters based on which the reach of content could be multiplied by a good factor. The proposed research work aims to identify and estimate the reach, popularity, and views of a YouTube video by using certain features using machine learning and AI techniques. A ranking system would also be used keeping the trending videos in consideration. This would eventually help the content creator know how authentic their content is and healthy competition to make better content before uploading the video on the platform will be ensured.

cs.IR↗

Learning to Answer Multilingual and Code-Mixed Questions

Question-answering (QA) that comes naturally to humans is a critical component in seamless human-computer interaction. It has emerged as one of the most convenient and natural methods to interact with the web and is especially desirable in voice-controlled environments. Despite being one of the oldest research areas, the current QA system faces the critical challenge of handling multilingual queries. To build an Artificial Intelligent (AI) agent that can serve multilingual end users, a QA system is required to be language versatile and tailored to suit the multilingual environment. Recent advances in QA models have enabled surpassing human performance primarily due to the availability of a sizable amount of high-quality datasets. However, the majority of such annotated datasets are expensive to create and are only confined to the English language, making it challenging to acknowledge progress in foreign languages. Therefore, to measure a similar improvement in the multilingual QA system, it is necessary to invest in high-quality multilingual evaluation benchmarks. In this dissertation, we focus on advancing QA techniques for handling end-user queries in multilingual environments. This dissertation consists of two parts. In the first part, we explore multilingualism and a new dimension of multilingualism referred to as code-mixing. Second, we propose a technique to solve the task of multi-hop question generation by exploiting multiple documents. Experiments show our models achieve state-of-the-art performance on answer extraction, ranking, and generation tasks on multiple domains of MQA, VQA, and language generation. The proposed techniques are generic and can be widely used in various domains and languages to advance QA systems.

cs.CL↗

Medical Image Retrieval via Nearest Neighbor Search on Pre-trained Image Features

Nearest neighbor search (NNS) aims to locate the points in high-dimensional space that is closest to the query point. The brute-force approach for finding the nearest neighbor becomes computationally infeasible when the number of points is large. The NNS has multiple applications in medicine, such as searching large medical imaging databases, disease classification, diagnosis, etc. With a focus on medical imaging, this paper proposes DenseLinkSearch an effective and efficient algorithm that searches and retrieves the relevant images from heterogeneous sources of medical images. Towards this, given a medical database, the proposed algorithm builds the index that consists of pre-computed links of each point in the database. The search algorithm utilizes the index to efficiently traverse the database in search of the nearest neighbor. We extensively tested the proposed NNS approach and compared the performance with state-of-the-art NNS approaches on benchmark datasets and our created medical image datasets. The proposed approach outperformed the existing approach in terms of retrieving accurate neighbors and retrieval speed. We also explore the role of medical image feature representation in content-based medical image retrieval tasks. We propose a Transformer-based feature representation technique that outperformed the existing pre-trained Transformer approach on CLEF 2011 medical image retrieval task. The source code of our experiments are available at https://github.com/deepaknlp/DLS.

cs.CV↗

CHQ-Summ: A Dataset for Consumer Healthcare Question Summarization

The quest for seeking health information has swamped the web with consumers' health-related questions. Generally, consumers use overly descriptive and peripheral information to express their medical condition or other healthcare needs, contributing to the challenges of natural language understanding. One way to address this challenge is to summarize the questions and distill the key information of the original question. To address this issue, we introduce a new dataset, CHQ-Summ that contains 1507 domain-expert annotated consumer health questions and corresponding summaries. The dataset is derived from the community question-answering forum and therefore provides a valuable resource for understanding consumer health-related posts on social media. We benchmark the dataset on multiple state-of-the-art summarization models to show the effectiveness of the dataset.

cs.CL↗

Work fluctuations for diffusion dynamics submitted to stochastic return

Returning a system to a desired state under a force field involves a thermodynamic cost, i.e., {\it work}. This cost fluctuates for a small-scale system from one experimental realization to another. We introduce a general framework to determine the work distribution for returning a system facilitated by a confining potential with its minimum at the restart location. The general strategy, based on average over {\it resetting pathways}, constitutes a robust method to gain access to the statistical information of observables from resetting systems. We exploit paradigmatic setups, where explicit computations are attainable, to illustrate the theory. Numerical simulations validate our theoretical predictions. For some of these examples, a non-trivial behavior of the work fluctuations opens a door to optimization problems. Specifically, work fluctuations can be minimized by an appropriate tuning of the return rate.

cond-mat.stat-mech↗

A Three-Stage Algorithm for the Large Scale Dynamic Vehicle Routing Problem with an Industry 4.0 Approach

Companies are eager to have a smart supply chain especially when they have a dynamic system. Industry 4.0 is a concept which concentrates on mobility and real-time integration. Thus, it can be considered as a necessary component that has to be implemented for a Dynamic Vehicle Routing Problem. The aim of this research is to solve large-scale DVRP (LSDVRP) in which the delivery vehicles must serve customer demands from a common depot to minimize transit cost while not exceeding the capacity constraint of each vehicle. In LSDVRP, it is difficult to get an exact solution and the computational time complexity grows exponentially. To find near optimal answers for this problem, a hierarchical approach consisting of three stages callled cluster first, route construction second, route improvement third is proposed. The major contribution of this paper is dealing with large-size real-world problems to decrease the computational time complexity. The results confirmed that the proposed methodology is applicable.

cs.AI↗

A Dataset for Medical Instructional Video Classification and Question Answering

This paper introduces a new challenge and datasets to foster research toward designing systems that can understand medical videos and provide visual answers to natural language questions. We believe medical videos may provide the best possible answers to many first aids, medical emergency, and medical education questions. Toward this, we created the MedVidCL and MedVidQA datasets and introduce the tasks of Medical Video Classification (MVC) and Medical Visual Answer Localization (MVAL), two tasks that focus on cross-modal (medical language and medical video) understanding. The proposed tasks and datasets have the potential to support the development of sophisticated downstream applications that can benefit the public and medical practitioners. Our datasets consist of 6,117 annotated videos for the MVC task and 3,010 annotated questions and answers timestamps from 899 videos for the MVAL task. These datasets have been verified and corrected by medical informatics experts. We have also benchmarked each task with the created MedVidCL and MedVidQA datasets and proposed the multimodal learning methods that set competitive baselines for future research.

cs.CV↗

Effective Resource-Competition Model for Species Coexistence

Local coexistence of species in large ecosystems is traditionally explained within the broad framework of niche theory. However, its rationale hardly justifies rich biodiversity observed in nearly homogeneous environments. Here we consider a consumer-resource model in which a coarse-graining procedure accounts for a variety of ecological mechanisms and leads to effective spatial effects which favour species coexistence. Herein, we provide conditions for several species to live in an environment with very few resources. In fact, the model displays two different phases depending on whether the number of surviving species is larger or smaller than the number of resources. We obtain conditions whereby a species can successfully colonize a pool of coexisting species. Finally, we analytically compute the distribution of the population sizes of coexisting species. Numerical simulations as well as empirical distributions of population sizes support our analytical findings.

q-bio.PE↗

Inducing and optimizing Markovian Mpemba effect with stochastic reset

A hot Markovian system can cool down faster than a colder one: this is known as the Mpemba effect. Here, we show that a non-equilibrium driving via stochastic reset can induce this phenomenon, when absent. Moreover, we derive an optimal driving protocol simultaneously optimizing the appearance time of the Mpemba effect, and the total energy dissipation into the environment, revealing the existence of a Pareto front. Building upon previous experimental results, our findings open up the avenue of possible experimental realizations of optimal cooling protocols in Markovian systems.

cond-mat.stat-mech↗

Towards Developing a Multilingual and Code-Mixed Visual Question Answering System by Knowledge Distillation

Pre-trained language-vision models have shown remarkable performance on the visual question answering (VQA) task. However, most pre-trained models are trained by only considering monolingual learning, especially the resource-rich language like English. Training such models for multilingual setups demand high computing resources and multilingual language-vision dataset which hinders their application in practice. To alleviate these challenges, we propose a knowledge distillation approach to extend an English language-vision model (teacher) into an equally effective multilingual and code-mixed model (student). Unlike the existing knowledge distillation methods, which only use the output from the last layer of the teacher network for distillation, our student model learns and imitates the teacher from multiple intermediate layers (language and vision encoders) with appropriately designed distillation objectives for incremental knowledge extraction. We also create the large-scale multilingual and code-mixed VQA dataset in eleven different language setups considering the multiple Indian and European languages. Experimental results and in-depth analysis show the effectiveness of the proposed VQA model over the pre-trained language-vision models on eleven diverse language setups.

cs.CL↗