Searcharxiv⌕ Search

arXiv subjects

Kundan Kumar

Publications and source records attributed to Kundan Kumar.

At least 37 records · Page 2Linked to original sources

High-Fidelity Audio Compression with Improved RVQGAN

Language models have been successfully used to model natural signals, such as images, speech, and music. A key component of these models is a high quality neural compression model that can compress high-dimensional natural signals into lower dimensional discrete tokens. To that end, we introduce a high-fidelity universal neural audio compression algorithm that achieves ~90x compression of 44.1 KHz audio into tokens at just 8kbps bandwidth. We achieve this by combining advances in high-fidelity audio generation with better vector quantization techniques from the image domain, along with improved adversarial and reconstruction losses. We compress all domains (speech, environment, music, etc.) with a single universal model, making it widely applicable to generative modeling of all audio. We compare with competing audio compression algorithms, and find our method outperforms them significantly. We provide thorough ablations for every design choice, as well as open-source code and trained model weights. We hope our work can lay the foundation for the next generation of high-fidelity audio modeling.

cs.SD↗

Coupling of flow, contact mechanics and friction, generating waves in a fractured porous medium

We present a mixed dimensional model for a fractured poro-elasic medium including contact mechanics. The fracture is a lower dimensional surface embedded in a bulk poro-elastic matrix. The flow equation on the fracture is a Darcy type model that follows the cubic law for permeability. The bulk poro-elasticity is governed by fully dynamic Biot equations. The resulting model is a mixed dimensional type where the fracture flow on a surface is coupled to a bulk flow and geomechanics model. The particularity of the work here is in considering fully dynamic Biot equation, that is, including an inertia term, and the contact mechanics including friction for the fracture surface. We prove the well-posedness of the continuous model.

math.AP↗

Tracking an Underwater Target with Unknown Measurement Noise Statistics Using Variational Bayesian Filters

This paper considers a bearings-only tracking problem using noisy measurements of unknown noise statistics from a passive sensor. It is assumed that the process and measurement noise follows the Gaussian distribution where the measurement noise has an unknown non-zero mean and unknown covariance. Here an adaptive nonlinear filtering technique is proposed where the joint distribution of the measurement noise mean and its covariance are considered to be following normal inverse Wishart distribution (NIW). Using the variational Bayesian (VB) method the estimation technique is derived with optimized tuning parameters i.e, the confidence parameter and the initial degree of freedom of the measurement noise mean and the covariance, respectively. The proposed filtering technique is compared with the adaptive filtering techniques based on maximum likelihood and maximum aposteriori in terms of root mean square error in position and velocity, bias norm, average normalized estimation error squared, percentage of track loss, and relative execution time. Both adaptive filtering techniques are implemented using the traditional Gaussian approximate filters and are applied to a bearings-only tracking problem illustrated with moderately nonlinear and highly nonlinear scenarios to track a target following a nearly straight line path. Two cases are considered for each scenario, one when the measurement noise covariance is static and another when the measurement noise covariance is varying linearly with the distance between the target and the ownship. In this work, the proposed adaptive filters using the VB approach are found to be superior to their corresponding adaptive filters based on the maximum aposteriori and the maximum likelihood at the expense of higher computation cost.

eess.SP↗

Application of Top-hat Transformation for Enhanced Blood Vessel Extraction

In the medical domain, different computer-aided diagnosis systems have been proposed to extract blood vessels from retinal fundus images for the clinical treatment of vascular diseases. Accurate extraction of blood vessels from the fundus images using a computer-generated method can help the clinician to produce timely and accurate reports for the patient suffering from these diseases. In this article, we integrate top-hat based preprocessing approach with fine-tuned B-COSFIRE filter to achieve more accurate segregation of blood vessel pixels from the background. The use of top-hat transformation in the preprocessing stage enhances the efficacy of the algorithm to extract blood vessels in presence of structures like fovea, exudates, haemorrhages, etc. Furthermore, to reduce the false positives, small clusters of blood vessel pixels are removed in the postprocessing stage. Further, we find that the proposed algorithm is more efficient as compared to various modern algorithms reported in the literature.

eess.IV↗

Parametric Scaling of Preprocessing assisted U-net Architecture for Improvised Retinal Vessel Segmentation

Extracting blood vessels from retinal fundus images plays a decisive role in diagnosing the progression in pertinent diseases. In medical image analysis, vessel extraction is a semantic binary segmentation problem, where blood vasculature needs to be extracted from the background. Here, we present an image enhancement technique based on the morphological preprocessing coupled with a scaled U-net architecture. Despite a relatively less number of trainable network parameters, the scaled version of U-net architecture provides better performance compare to other methods in the domain. We validated the proposed method on retinal fundus images from the DRIVE database. A significant improvement as compared to the other algorithms in the domain, in terms of the area under ROC curve (>0.9762) and classification accuracy (>95.47%) are evident from the results. Furthermore, the proposed method is resistant to the central vessel reflex while sensitive to detect blood vessels in the presence of background items viz. exudates, optic disc, and fovea.

eess.IV↗

Pattern Based Multivariable Regression using Deep Learning (PBMR-DP)

We propose a deep learning methodology for multivariate regression that is based on pattern recognition that triggers fast learning over sensor data. We used a conversion of sensors-to-image which enables us to take advantage of Computer Vision architectures and training processes. In addition to this data preparation methodology, we explore the use of state-of-the-art architectures to generate regression outputs to predict agricultural crop continuous yield information. Finally, we compare with some of the top models reported in MLCAS2021. We found that using a straightforward training process, we were able to accomplish an MAE of 4.394, RMSE of 5.945, and R^2 of 0.861.

cs.CV↗

Chunked Autoregressive GAN for Conditional Waveform Synthesis

Conditional waveform synthesis models learn a distribution of audio waveforms given conditioning such as text, mel-spectrograms, or MIDI. These systems employ deep generative models that model the waveform via either sequential (autoregressive) or parallel (non-autoregressive) sampling. Generative adversarial networks (GANs) have become a common choice for non-autoregressive waveform synthesis. However, state-of-the-art GAN-based models produce artifacts when performing mel-spectrogram inversion. In this paper, we demonstrate that these artifacts correspond with an inability for the generator to learn accurate pitch and periodicity. We show that simple pitch and periodicity conditioning is insufficient for reducing this error relative to using autoregression. We discuss the inductive bias that autoregression provides for learning the relationship between instantaneous frequency and phase, and show that this inductive bias holds even when autoregressively sampling large chunks of the waveform during each forward pass. Relative to prior state-of-the-art GAN-based models, our proposed model, Chunked Autoregressive GAN (CARGAN) reduces pitch error by 40-60%, reduces training time by 58%, maintains a fast generation speed suitable for real-time or interactive applications, and maintains or improves subjective quality.

eess.AS↗

Wav2CLIP: Learning Robust Audio Representations From CLIP

We propose Wav2CLIP, a robust audio representation learning method by distilling from Contrastive Language-Image Pre-training (CLIP). We systematically evaluate Wav2CLIP on a variety of audio tasks including classification, retrieval, and generation, and show that Wav2CLIP can outperform several publicly available pre-trained audio representation algorithms. Wav2CLIP projects audio into a shared embedding space with images and text, which enables multimodal applications such as zero-shot classification, and cross-modal retrieval. Furthermore, Wav2CLIP needs just ~10% of the data to achieve competitive performance on downstream tasks compared with fully supervised models, and is more efficient to pre-train than competing methods as it does not require learning a visual model in concert with an auditory model. Finally, we demonstrate image generation from Wav2CLIP as qualitative assessment of the shared embedding space. Our code and model weights are open sourced and made available for further applications.

cs.SD↗

Field-scale impacts of long-term wettability alteration in geological CO$_2$ storage

Constitutive functions that govern macroscale capillary pressure and relative permeability are central in constraining both storage efficiency and sealing properties of CO$_2$ storage systems. Constitutive functions for porous systems are in part determined by wettability, which is a pore-scale phenomenon that influences macroscale displacement. While wettability of saline aquifers and caprocks are assumed to remain water-wet when CO$_2$ is injected, there is recent evidence of contact angle change due to long-term CO$_2$ exposure. Weakening of capillary forces alters the saturation functions dynamically over time. Recently, new dynamic models were developed for saturation functions that capture the impact of wettability alteration (WA) due to long-term CO$_2$ exposure. In this paper, these functions are implemented into a two-phase two-component simulator to study long-term WA dynamics for field-scale CO$_2$ storage. We simulate WA effects on horizontal migration patterns under injection and buoyancy-driven migration in the caprock. We characterize the behavior of each scenario for different flow regimes. Our results show the impact on storage efficiency can be described by the capillary number, while vertical leakage can be scaled by caprock sealing parameters. Scaling models for CO$_2$ migration into the caprock show that long-term WA poses little risk to CO$_2$ containment over relevant timescales.

physics.geo-ph↗

Radially excited (n=3)charm mesons in heavy quark effective theory

By exploring heavy quark effective theory (HQET), we use theoretical available data for bottom mesons to analysis the masses and decays for n = 3 charm mesons. From the predicted masses, we studied ground state strong decay modes in terms of couplings. Comparing the decays with available total decay widths, we provide upper bounds on the associated couplings. We also plot Regge trajectories for our predicted data in planes (J, $M^2$ ) and ($n_r$, $M^2$ ) and estimated higher masses (n = 4) by fixing Regge slopes and intercepts. These Regge trajectories are used to clarify $D_2^*(3000)$ state's $J^P$ as 1F ($2^+$) state. The presented results may further get confirmation through upcoming experimental information.

hep-ph↗

Numerical studies of long-term wettability alteration effects in CO$_2$ storage applications

The wettability of the rock surface in porous media has an effect on the constitutive saturation functions that govern capillary pressure and relative permeability. The term wettability alteration refers to the change of this property over time by processes such as CO$_2$ interactions with the rock. In this work, we perform numerical simulations considering a two-phase two-component flow model including time-dependent wettability alteration in a two-dimensional aquifer-caprock system using the open porous media framework. Particularly, we study the spatial distribution over time of injected CO$_2$ into the aquifer neglecting and including wettability alteration effects. The numerical simulations show that wettability alteration on the caprock results in a loss of containment; however, the CO$_2$ front into the caprock advances very slow since the unexposed caprock along the vertical migration path also needs to be changed by the slow wettability alteration process. The simulations also show that wettability alteration on the aquifer results in an enhancement of storage efficiency; this since the CO$_2$ front migrates more slowly and the capillary entry pressure decreases after wettability alteration.

physics.geo-ph↗

Numerical studies of CO$_2$ leakage remediation by micp-based plugging technology

Microbially induced calcite precipitation (MICP) is a technology for sealing leakage paths to ensure the safe storage of CO$_2$ in geological formations. In this work we introduce a numerical simulator of MICP for field-scale studies. This simulator is implemented in the open porous media (OPM) framework. We compare the numerical results to simulations using an upgraded implementation of the mathematical model in the MATLAB reservoir simulation toolbox (MRST). Finally, we consider a 3D system consisting of two aquifers separated by caprock with a leakage path across the width of the reservoir. We study a strategy where microbial solution is injected only at the beginning of the treatment and subsequently either growth solution or cementation solution is injected for biofilm development or calcite precipitation. By applying this strategy, the numerical results show that the MICP technology could be used to seal these leakage paths.

cs.CE↗

Practical approaches to study microbially induced calcite precipitation at the field scale

Microbially induced calcite precipitation (MICP) is a new and sustainable technology which utilizes biochemical processes to create barriers by calcium carbonate cementation; therefore, this technology has a potential to be used for sealing leakage zones in geological formations. The complexity of current MICP models and present computer power limit the size of numerical simulations. We describe a mathematical model for MICP suitable for field-scale studies. The main mechanisms in the conceptual model are as follow: suspended microbes attach themselves to the pore walls to form biofilm, growth solution is added to stimulate the biofilm development, the biofilm uses cementation solution for production of calcite, and the calcite reduces the pore space which in turn decreases the rock permeability. We apply the model to study the MICP technology in two sets of reservoir properties including a well-established field-scale benchmark system for CO$_2$ leakage. A two-phase flow model for CO$_2$ and water is used to assess the leakage prior to and with MICP treatment. Based on the numerical results, this study confirms the potential for this technology to seal leakage paths in reservoir-caprock systems.

physics.geo-ph↗

NU-GAN: High resolution neural upsampling with GAN

In this paper, we propose NU-GAN, a new method for resampling audio from lower to higher sampling rates (upsampling). Audio upsampling is an important problem since productionizing generative speech technology requires operating at high sampling rates. Such applications use audio at a resolution of 44.1 kHz or 48 kHz, whereas current speech synthesis methods are equipped to handle a maximum of 24 kHz resolution. NU-GAN takes a leap towards solving audio upsampling as a separate component in the text-to-speech (TTS) pipeline by leveraging techniques for audio generation using GANs. ABX preference tests indicate that our NU-GAN resampler is capable of resampling 22 kHz to 44.1 kHz audio that is distinguishable from original audio only 7.4% higher than random chance for single speaker dataset, and 10.8% higher than chance for multi-speaker dataset.

cs.SD↗

Antimony Chalcogenide-based Solid State Sensitizers for Solar Cells: A Forgotten Hero or Low Potential Candidate

The use of stibnite (Sb2S3) as sensitizers in the solid-state sensitized solar cells received considerable research interest during the transition of the millennium. However, the use of perovskite diminished the research in the field and the potential of antimony chalcogenide (Sb2(S,Se)3) was not explored thoroughly. Although these materials also provide bandgap tuning like perovskite by varying the composition of S and Se, it is not as popular as perovskite mainly because of the low efficiency of the solar cells based on it. In this paper, we present a landscape of the functional role of various device parameters on the performance of Sb2(S,Se)3 based solar cells. For the purpose, we first calibrate the optoelectronic model used for the simulation with the experimental results from the literature. The model is then subjected to parametric variations to explore the performance metrics for this class of solar cells. Our results show that despite the belief that open circuit voltage is independent of contact layers doping in proper band aligned sensitized solar cells, here we observe otherwise and the open circuit voltage is indeed dependent on the doping density of the contact layers. Using the detailed numerical simulation and analytical model we further identify the performance optimization map of Sb2(S,Se)3 based sensitized solar cells.

physics.app-ph↗

Guaranteed and computable error bounds for approximations constructed by an iterative decoupling of the Biot problem

The paper is concerned with guaranteed a posteriori error estimates for a class of evolutionary problems related to poroelastic media governed by the quasi-static linear Biot equations. The system is decoupled employing the fixed-stress split scheme, which leads to a semi-discrete system solved iteratively. The error bounds are derived by combining a posteriori estimates for contractive mappings with those of the functional type for elliptic partial differential equations. The estimates are applicable for any approximation in the admissible functional space and are independent of the discretization method. They are fully computable, do not contain mesh dependent constants, and provide reliable global estimates of the error measured in the energy norm. Moreover, they suggest efficient error indicators for the distribution of local errors, which can be used in adaptive procedures.

math.NA↗

Mathematical Modeling, Laboratory Experiments, and Sensitivity Analysis of Bioplug Technology at Darcy Scale

In this paper we study a Darcy-scale mathematical model for biofilm formation in porous media. The pores in the core are divided into three phases: water, oil, and biofilm. The water and oil flow are modeled by an extended version of Darcy's law and the substrate is transported by diffusion and convection in the water phase. Initially there is biofilm on the pore walls. The biofilm consumes substrate for production of biomass and modifies the pore space which changes the rock permeability. The model includes detachment of biomass due to water flux and death of bacteria, and is implemented in MRST. We discuss the capability of the numerical simulator to capture results from laboratory experiments. We perform a novel sensitivity analysis based on sparse-grid interpolation and multi-wavelet expansion to identify the critical model parameters. Numerical experiments using diverse injection strategies are performed to study the impact of different porosity-permeability relations in a core saturated with water and oil.

physics.app-ph↗

MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis

Previous works (Donahue et al., 2018a; Engel et al., 2019a) have found that generating coherent raw audio waveforms with GANs is challenging. In this paper, we show that it is possible to train GANs reliably to generate high quality coherent waveforms by introducing a set of architectural changes and simple training techniques. Subjective evaluation metric (Mean Opinion Score, or MOS) shows the effectiveness of the proposed approach for high quality mel-spectrogram inversion. To establish the generality of the proposed techniques, we show qualitative results of our model in speech synthesis, music domain translation and unconditional music synthesis. We evaluate the various components of the model through ablation studies and suggest a set of guidelines to design general purpose discriminators and generators for conditional sequence synthesis tasks. Our model is non-autoregressive, fully convolutional, with significantly fewer parameters than competing models and generalizes to unseen speakers for mel-spectrogram inversion. Our pytorch implementation runs at more than 100x faster than realtime on GTX 1080Ti GPU and more than 2x faster than real-time on CPU, without any hardware specific optimization tricks.

eess.AS↗