Searcharxiv⌕ Search

arXiv subjects

Deepak Gupta

Publications and source records attributed to Deepak Gupta.

At least 73 records · Page 4Linked to original sources

Rescaling CNN through Learnable Repetition of Network Parameters

Deeper and wider CNNs are known to provide improved performance for deep learning tasks. However, most such networks have poor performance gain per parameter increase. In this paper, we investigate whether the gain observed in deeper models is purely due to the addition of more optimization parameters or whether the physical size of the network as well plays a role. Further, we present a novel rescaling strategy for CNNs based on learnable repetition of its parameters. Based on this strategy, we rescale CNNs without changing their parameter count, and show that learnable sharing of weights itself can provide significant boost in the performance of any given model without changing its parameter count. We show that small base networks when rescaled, can provide performance comparable to deeper networks with as low as 6% of optimization parameters of the deeper one. The relevance of weight sharing is further highlighted through the example of group-equivariant CNNs. We show that the significant improvements obtained with group-equivariant CNNs over the regular CNNs on classification problems are only partly due to the added equivariance property, and part of it comes from the learnable repetition of network weights. For rot-MNIST dataset, we show that up to 40% of the relative gain reported by state-of-the-art methods for rotation equivariance could actually be due to just the learnt repetition of weights.

cs.CV↗

BloomNet: A Robust Transformer based model for Bloom's Learning Outcome Classification

Bloom taxonomy is a common paradigm for categorizing educational learning objectives into three learning levels: cognitive, affective, and psychomotor. For the optimization of educational programs, it is crucial to design course learning outcomes (CLOs) according to the different cognitive levels of Bloom Taxonomy. Usually, administrators of the institutions manually complete the tedious work of mapping CLOs and examination questions to Bloom taxonomy levels. To address this issue, we propose a transformer-based model named BloomNet that captures linguistic as well semantic information to classify the course learning outcomes (CLOs). We compare BloomNet with a diverse set of basic as well as strong baselines and we observe that our model performs better than all the experimented baselines. Further, we also test the generalization capability of BloomNet by evaluating it on different distributions which our model does not encounter during training and we observe that our model is less susceptible to distribution shift compared to the other considered models. We support our findings by performing extensive result analysis. In ablation study we observe that on explicitly encapsulating the linguistic information along with semantic information improves the model on IID (independent and identically distributed) performance as well as OOD (out-of-distribution) generalization capability.

cs.CL↗

Heat fluctuations in a harmonic chain of active particles

One of the major challenges in stochastic thermodynamics is to compute the distributions of stochastic observables for small-scale systems for which fluctuations play a significant role. Hitherto much theoretical and experimental research has focused on systems composed of passive Brownian particles. In this paper, we study the heat fluctuations in a system of interacting active particles. Specifically we consider a one-dimensional harmonic chain of $N$ active Ornstein-Uhlenbeck particles, with the chain ends connected to heat baths of different temperatures. We compute the moment-generating function for the heat flow in the steady state. We employ our general framework to explicitly compute the moment-generating function for two example single-particle systems. Further, we analytically obtain the scaled cumulants for the heat flow for the chain. Numerical Langevin simulations confirm the long-time analytical expressions for first and second cumulants for the heat flow for a two-particle chain.

cond-mat.stat-mech↗

Reinforcement Learning for Abstractive Question Summarization with Question-aware Semantic Rewards

The growth of online consumer health questions has led to the necessity for reliable and accurate question answering systems. A recent study showed that manual summarization of consumer health questions brings significant improvement in retrieving relevant answers. However, the automatic summarization of long questions is a challenging task due to the lack of training data and the complexity of the related subtasks, such as the question focus and type recognition. In this paper, we introduce a reinforcement learning-based framework for abstractive question summarization. We propose two novel rewards obtained from the downstream tasks of (i) question-type identification and (ii) question-focus recognition to regularize the question generation model. These rewards ensure the generation of semantically valid questions and encourage the inclusion of key medical entities/foci in the question summary. We evaluated our proposed method on two benchmark datasets and achieved higher performance over state-of-the-art models. The manual evaluation of the summaries reveals that the generated questions are more diverse and have fewer factual inconsistencies than the baseline summaries

cs.CL↗

Question-aware Transformer Models for Consumer Health Question Summarization

Searching for health information online is becoming customary for more and more consumers every day, which makes the need for efficient and reliable question answering systems more pressing. An important contributor to the success rates of these systems is their ability to fully understand the consumers' questions. However, these questions are frequently longer than needed and mention peripheral information that is not useful in finding relevant answers. Question summarization is one of the potential solutions to simplifying long and complex consumer questions before attempting to find an answer. In this paper, we study the task of abstractive summarization for real-world consumer health questions. We develop an abstractive question summarization model that leverages the semantic interpretation of a question via recognition of medical entities, which enables the generation of informative summaries. Towards this, we propose multiple Cloze tasks (i.e. the task of filing missing words in a given context) to identify the key medical entities that enforce the model to have better coverage in question-focus recognition. Additionally, we infuse the decoder inputs with question-type information to generate question-type driven summaries. When evaluated on the MeQSum benchmark corpus, our framework outperformed the state-of-the-art method by 10.2 ROUGE-L points. We also conducted a manual evaluation to assess the correctness of the generated summaries.

cs.CL↗

Stochastic resetting with stochastic returns using external trap

In the past few years, stochastic resetting has become a subject of immense interest. Most of the theoretical studies so far focused on instantaneous resetting which is, however, a major impediment to practical realization or experimental verification in the field. This is because in the real world, taking a particle from one place to another requires finite time and thus a generalization of the existing theory to incorporate non-instantaneous resetting is very much in need. In this paper, we propose a method of resetting which involves non-instantaneous returns facilitated by an external confining trap potential $U(x)$ centered at the resetting location. We consider a Brownian particle that starts its random motion from the origin. Upon resetting, the trap is switched on and the particle starts experiencing a force towards the center of the trap which drives it to return to the origin. The return phase ends when the particle makes a first passage to this center. We develop a general framework to study such a set up. Importantly, we observe that the system reaches a non-equilibrium steady state which we analyze in full details for two choices of $U(x)$, namely, (i) linear and (ii) harmonic. Finally, we perform numerical simulations and find an excellent agreement with the theory. The general formalism developed here can be applied to more realistic return protocols opening up a panorama of possibilities for further theoretical and experimental applications.

cond-mat.stat-mech↗

Resetting with stochastic return through linear confining potential

We consider motion of an overdamped Brownian particle subject to stochastic resetting in one dimension. In contrast to the usual setting where the particle is instantaneously reset to a preferred location (say, the origin), here we consider a finite time resetting process facilitated by an external linear potential $V(x)=λ|x|~ (λ>0)$. When resetting occurs, the trap is switched on and the particle experiences a force $-\partial_x V(x)$ which helps the particle to return to the resetting location. The trap is switched off as soon as the particle makes a first passage to the origin. Subsequently, the particle resumes its free diffusion motion and the process keeps repeating. In this set-up, the system attains a non-equilibrium steady state. We study the relaxation to this steady state by analytically computing the position distribution of the particle at all time and then analysing this distribution using the spectral properties of the corresponding Fokker-Planck operator. As seen for the instantaneous resetting problem, we observe a `cone spreading' relaxation with travelling fronts such that there is an inner core region around the resetting point that reaches the steady state, while the region outside the core still grows ballistically with time. In addition to the unusual relaxation phenomena, we compute the large deviation functions associated to the corresponding probability density and find that the large deviation functions describe a dynamical transition similar to what is seen previously in case of instantaneous resetting. Notably, our method, based on spectral properties, complements the existing renewal formalism and reveals the intricate mathematical structure responsible for such relaxation phenomena. We verify our analytical results against extensive numerical simulations.

cond-mat.stat-mech↗

Stochastic efficiency of an isothermal work-to-work converter engine

We investigate the efficiency of an isothermal Brownian work-to-work converter engine, composed of a Brownian particle coupled to a heat bath at a constant temperature. The system is maintained out of equilibrium by using two external time-dependent stochastic Gaussian forces, where one is called load force and the other is called drive force. Work done by these two forces are stochastic quantities. The efficiency of this small engine is defined as the ratio of stochastic work done against load force to stochastic work done by the drive force. The probability density function as well as large deviation function of the stochastic efficiency are studied analytically and verified by numerical simulations.

cond-mat.stat-mech↗

Domain Controlled Title Generation with Human Evaluation

We study automatic title generation and present a method for generating domain-controlled titles for scientific articles. A good title allows you to get the attention that your research deserves. A title can be interpreted as a high-compression description of a document containing information on the implemented process. For domain-controlled titles, we used the pre-trained text-to-text transformer model and the additional token technique. Title tokens are sampled from a local distribution (which is a subset of global vocabulary) of the domain-specific vocabulary and not global vocabulary, thereby generating a catchy title and closely linking it to its corresponding abstract. Generated titles looked realistic, convincing, and very close to the ground truth. We have performed automated evaluation using ROUGE metric and human evaluation using five parameters to make a comparison between human and machine-generated titles. The titles produced were considered acceptable with higher metric ratings in contrast to the original titles. Thus we concluded that our research proposes a promising method for domain-controlled title generation.

cs.CL↗

CovidGAN: Data Augmentation Using Auxiliary Classifier GAN for Improved Covid-19 Detection

Coronavirus (COVID-19) is a viral disease caused by severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). The spread of COVID-19 seems to have a detrimental effect on the global economy and health. A positive chest X-ray of infected patients is a crucial step in the battle against COVID-19. Early results suggest that abnormalities exist in chest X-rays of patients suggestive of COVID-19. This has led to the introduction of a variety of deep learning systems and studies have shown that the accuracy of COVID-19 patient detection through the use of chest X-rays is strongly optimistic. Deep learning networks like convolutional neural networks (CNNs) need a substantial amount of training data. Because the outbreak is recent, it is difficult to gather a significant number of radiographic images in such a short time. Therefore, in this research, we present a method to generate synthetic chest X-ray (CXR) images by developing an Auxiliary Classifier Generative Adversarial Network (ACGAN) based model called CovidGAN. In addition, we demonstrate that the synthetic images produced from CovidGAN can be utilized to enhance the performance of CNN for COVID-19 detection. Classification using CNN alone yielded 85% accuracy. By adding synthetic images produced by CovidGAN, the accuracy increased to 95%. We hope this method will speed up COVID-19 detection and lead to more robust systems of radiology.

eess.IV↗

Ginzburg-Landau amplitude equation for nonlinear nonlocal models

Regular spatial structures emerge in a wide range of different dynamics characterized by local and/or nonlocal coupling terms. In several research fields this has spurred the study of many models, which can explain pattern formation. The modulations of patterns, occurring on long spatial and temporal scales, can not be captured by linear approximation analysis. Here, we show that, starting from a general model with long range couplings displaying patterns, the spatio-temporal evolution of large scale modulations at the onset of instability is ruled by the well-known Ginzburg-Landau equation, independently of the details of the dynamics. Hence, we demonstrate the validity of such equation in the description of the behavior of a wide class of systems. We introduce a novel mathematical framework that is also able to retrieve the analytical expressions of the coefficients appearing in the Ginzburg-Landau equation as functions of the model parameters. Such framework can include higher order nonlocal interactions and has much larger applicability than the model considered here, possibly including pattern formation in models with very different physical features.

cond-mat.stat-mech↗

Can Taxonomy Help? Improving Semantic Question Matching using Question Taxonomy

In this paper, we propose a hybrid technique for semantic question matching. It uses our proposed two-layered taxonomy for English questions by augmenting state-of-the-art deep learning models with question classes obtained from a deep learning based question classifier. Experiments performed on three open-domain datasets demonstrate the effectiveness of our proposed approach. We achieve state-of-the-art results on partial ordering question ranking (POQR) benchmark dataset. Our empirical analysis shows that coupling standard distributional features (provided by the question encoder) with knowledge from taxonomy is more effective than either deep learning (DL) or taxonomy-based knowledge alone.

cs.CL↗

Asymmetric Stochastic Resetting: Modeling Catastrophic Events

In the classical stochastic resetting problem, a particle, moving according to some stochastic dynamics, undergoes random interruptions that bring it to a selected domain, and then, the process recommences. Hitherto, the resetting mechanism has been introduced as a symmetric reset about the preferred location. However, in nature, there are several instances where a system can only reset from certain directions, e.g., catastrophic events. Motivated by this, we consider a continuous stochastic process on the positive real line. The process is interrupted at random times occurring at a constant rate, and then, the former relocates to a value only if the current one exceeds a threshold; otherwise, it follows the trajectory defined by the underlying process without resetting. We present a general framework to obtain the exact non-equilibrium steady state of the system and the mean first passage time for the system to reach the origin. Employing this framework, we obtain the explicit solutions for two different model systems. Some of the classical results found in symmetric resetting such as the existence of an optimal resetting, are strongly modified. Finally, numerical simulations have been performed to verify the analytical findings, showing an excellent agreement.

cond-mat.stat-mech↗

Siamese Tracking with Lingual Object Constraints

Classically, visual object tracking involves following a target object throughout a given video, and it provides us the motion trajectory of the object. However, for many practical applications, this output is often insufficient since additional semantic information is required to act on the video material. Example applications of this are surveillance and target-specific video summarization, where the target needs to be monitored with respect to certain predefined constraints, e.g., 'when standing near a yellow car'. This paper explores, tracking visual objects subjected to additional lingual constraints. Differently from Li et al., we impose additional lingual constraints upon tracking, which enables new applications of tracking. Whereas in their work the goal is to improve and extend upon tracking itself. To perform benchmarks and experiments, we contribute two datasets: c-MOT16 and c-LaSOT, curated through appending additional constraints to the frames of the original LaSOT and MOT16 datasets. We also experiment with two deep models SiamCT-DFG and SiamCT-CA, obtained through extending a recent state-of-the-art Siamese tracking method and adding modules inspired from the fields of natural language processing and visual question answering. Through experimental results, we show that the proposed model SiamCT-CA can significantly outperform its counterparts. Furthermore, our method enables the selective compression of videos, based on the validity of the constraint.

cs.CV↗

Tighter thermodynamic bound on speed limit in systems with unidirectional transitions

We consider a general discrete state-space system with both unidirectional and bidirectional links. In contrast to bidirectional links, there is no reverse transition along the unidirectional links. Herein, we first compute the statistical length and the thermodynamic cost function for transitions in the probability space, highlighting contributions from total, environmental, and resetting (unidirectional) entropy production. Then, we derive the thermodynamic bound on the speed limit to connect two distributions separated by a finite time, showing the effect of the presence of unidirectional transitions. Novel uncertainty relationships can be found for the \textit{temporal} first and second moments of the average resetting entropy production. We derive simple expressions in the limit of slow unidirectional transition rates. Finally, we present a refinement of the thermodynamic bound, by means of an optimization procedure. We numerically investigate these results on systems that stochastically reset with constant and periodic resetting rate.

cond-mat.stat-mech↗

Reinforced Multi-task Approach for Multi-hop Question Generation

Question generation (QG) attempts to solve the inverse of question answering (QA) problem by generating a natural language question given a document and an answer. While sequence to sequence neural models surpass rule-based systems for QG, they are limited in their capacity to focus on more than one supporting fact. For QG, we often require multiple supporting facts to generate high-quality questions. Inspired by recent works on multi-hop reasoning in QA, we take up Multi-hop question generation, which aims at generating relevant questions based on supporting facts in the context. We employ multitask learning with the auxiliary task of answer-aware supporting fact prediction to guide the question generator. In addition, we also proposed a question-aware reward function in a Reinforcement Learning (RL) framework to maximize the utilization of the supporting facts. We demonstrate the effectiveness of our approach through experiments on the multi-hop question answering dataset, HotPotQA. Empirical evaluation shows our model to outperform the single-hop neural question generation models on both automatic evaluation metrics such as BLEU, METEOR, and ROUGE, and human evaluation metrics for quality and coverage of the generated questions.

cs.CL↗

Coarse-grained entropy production with multiple reservoirs: unraveling the role of time-scales and detailed balance in biology-inspired systems

A general framework to describe a vast majority of biology-inspired systems is to model them as stochastic processes in which multiple couplings are in play at the same time. Molecular motors, chemical reaction networks, catalytic enzymes, and particles exchanging heat with different baths, constitute some interesting examples of such a modelization. Moreover, they usually operate out of equilibrium, being characterized by a net production of entropy, which entails a constrained efficiency. Hitherto, in order to investigate multiple processes simultaneously driving a system, all theoretical approaches deal with them independently, at a coarse-grained level, or employing a separation of time-scales. Here, we explicitly take in consideration the interplay among time-scales of different processes, and whether or not their own evolution eventually relaxes toward an equilibrium state in a given sub-space. We propose a general framework for multiple coupling, from which the well-known formulas for the entropy production can be derived, depending on the available information about each single process. Furthermore, when one of the processes does not equilibrate in its sub-space, even if much faster than all the others, it introduces a finite correction to the entropy production. We employ our framework in various simple and pedagogical examples, for which such a corrective term can be related to a typical scaling of physical quantities in play.

cond-mat.stat-mech↗

Hierarchical Deep Multi-modal Network for Medical Visual Question Answering

Visual Question Answering in Medical domain (VQA-Med) plays an important role in providing medical assistance to the end-users. These users are expected to raise either a straightforward question with a Yes/No answer or a challenging question that requires a detailed and descriptive answer. The existing techniques in VQA-Med fail to distinguish between the different question types sometimes complicates the simpler problems, or over-simplifies the complicated ones. It is certainly true that for different question types, several distinct systems can lead to confusion and discomfort for the end-users. To address this issue, we propose a hierarchical deep multi-modal network that analyzes and classifies end-user questions/queries and then incorporates a query-specific approach for answer prediction. We refer our proposed approach as Hierarchical Question Segregation based Visual Question Answering, in short HQS-VQA. Our contributions are three-fold, viz. firstly, we propose a question segregation (QS) technique for VQAMed; secondly, we integrate the QS model to the hierarchical deep multi-modal neural network to generate proper answers to the queries related to medical images; and thirdly, we study the impact of QS in Medical-VQA by comparing the performance of the proposed model with QS and a model without QS. We evaluate the performance of our proposed model on two benchmark datasets, viz. RAD and CLEF18. Experimental results show that our proposed HQS-VQA technique outperforms the baseline models with significant margins. We also conduct a detailed quantitative and qualitative analysis of the obtained results and discover potential causes of errors and their solutions.

cs.CL↗