SearcharxivSearch

arXiv subjects

Gautam Bhattacharya

Publications and source records attributed to Gautam Bhattacharya.

12 recordsLinked to original sources

Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning

Safety aligned Large Language Models (LLMs) are vulnerable to harmful fine-tuning attacks -- a few harmful data mixed in the fine-tuning dataset can break the LLMs's safety alignment. While several defenses have been proposed, our evaluation shows that existing defenses fail \textit{when some specific training hyper-parameters are chosen} -- a large learning rate or a large number of training epochs in the fine-tuning stage can easily invalidate the defense. To this end, we propose Antidote, a post-fine-tuning stage solution, which remains \textbf{\textit{agnostic to the training hyper-parameters in the fine-tuning stage}}. Antidote relies on the philosophy that by removing the harmful parameters, the harmful model can be recovered from the harmful behaviors, regardless of how those harmful parameters are formed in the fine-tuning stage. With this philosophy, we introduce a one-shot pruning stage after harmful fine-tuning to remove the harmful weights that are responsible for the generation of harmful content. Despite its embarrassing simplicity, empirical results show that Antidote can reduce harmful score while maintaining accuracy on downstream tasks. Code is available at https://github.com/git-disl/Antidote.

cs.AI

Room Impulse Response Generation Conditioned on Acoustic Parameters

The generation of room impulse responses (RIRs) using deep neural networks has attracted growing research interest due to its applications in virtual and augmented reality, audio postproduction, and related fields. Most existing approaches condition generative models on physical descriptions of a room, such as its size, shape, and surface materials. However, this reliance on geometric information limits their usability in scenarios where the room layout is unknown or when perceptual realism (how a space sounds to a listener) is more important than strict physical accuracy. In this study, we propose an alternative strategy: conditioning RIR generation directly on a set of RIR acoustic parameters. These parameters include various measures of reverberation time and direct sound to reverberation ratio, both broadband and bandwise. By specifying how the space should sound instead of how it should look, our method enables more flexible and perceptually driven RIR generation. We explore both autoregressive and non-autoregressive generative models operating in the Descript Audio Codec domain, using either discrete token sequences or continuous embeddings. Specifically, we have selected four models to evaluate: an autoregressive transformer, the MaskGIT model, a flow matching model, and a classifier-based approach. Objective and subjective evaluations are performed to compare these methods with state-of-the-art alternatives. Results show that the proposed models match or outperform state-of-the-art alternatives, with the MaskGIT model achieving the best performance.

cs.SD

Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation

Audio token modeling has become a powerful framework for speech synthesis, with two-stage approaches employing semantic tokens remaining prevalent. In this paper, we aim to simplify this process by introducing a semantic knowledge distillation method that enables high-quality speech generation in a single stage. Our proposed model improves speech quality, intelligibility, and speaker similarity compared to a single-stage baseline. Although two-stage systems still lead in intelligibility, our model significantly narrows the gap while delivering comparable speech quality. These findings showcase the potential of single-stage models to achieve efficient, high-quality TTS with a more compact and streamlined architecture.

cs.SD

CLIPSonic: Text-to-Audio Synthesis with Unlabeled Videos and Pretrained Language-Vision Models

Recent work has studied text-to-audio synthesis using large amounts of paired text-audio data. However, audio recordings with high-quality text annotations can be difficult to acquire. In this work, we approach text-to-audio synthesis using unlabeled videos and pretrained language-vision models. We propose to learn the desired text-audio correspondence by leveraging the visual modality as a bridge. We train a conditional diffusion model to generate the audio track of a video, given a video frame encoded by a pretrained contrastive language-image pretraining (CLIP) model. At test time, we first explore performing a zero-shot modality transfer and condition the diffusion model with a CLIP-encoded text query. However, we observe a noticeable performance drop with respect to image queries. To close this gap, we further adopt a pretrained diffusion prior model to generate a CLIP image embedding given a CLIP text embedding. Our results show the effectiveness of the proposed method, and that the pretrained diffusion prior can reduce the modality transfer gap. While we focus on text-to-audio synthesis, the proposed model can also generate audio from image queries, and it shows competitive performance against a state-of-the-art image-to-audio synthesis model in a subjective listening test. This study offers a new direction of approaching text-to-audio synthesis that leverages the naturally-occurring audio-visual correspondence in videos and the power of pretrained language-vision models.

cs.SD

Full-band General Audio Synthesis with Score-based Diffusion

Recent works have shown the capability of deep generative models to tackle general audio synthesis from a single label, producing a variety of impulsive, tonal, and environmental sounds. Such models operate on band-limited signals and, as a result of an autoregressive approach, they are typically conformed by pre-trained latent encoders and/or several cascaded modules. In this work, we propose a diffusion-based generative model for general audio synthesis, named DAG, which deals with full-band signals end-to-end in the waveform domain. Results show the superiority of DAG over existing label-conditioned generators in terms of both quality and diversity. More specifically, when compared to the state of the art, the band-limited and full-band versions of DAG achieve relative improvements that go up to 40 and 65%, respectively. We believe DAG is flexible enough to accommodate different conditioning schemas while providing good quality synthesis.

cs.SD

Adapting End-to-End Neural Speaker Verification to New Languages and Recording Conditions with Adversarial Training

In this article we propose a novel approach for adapting speaker embeddings to new domains based on adversarial training of neural networks. We apply our embeddings to the task of text-independent speaker verification, a challenging, real-world problem in biometric security. We further the development of end-to-end speaker embedding models by combing a novel 1-dimensional, self-attentive residual network, an angular margin loss function and adversarial training strategy. Our model is able to learn extremely compact, 64-dimensional speaker embeddings that deliver competitive performance on a number of popular datasets using simple cosine distance scoring. One the NIST-SRE 2016 task we are able to beat a strong i-vector baseline, while on the Speakers in the Wild task our model was able to outperform both i-vector and x-vector baselines, showing an absolute improvement of 2.19% over the latter. Additionally, we show that the integration of adversarial training consistently leads to a significant improvement over an unadapted model.

eess.AS

Generative Adversarial Speaker Embedding Networks for Domain Robust End-to-End Speaker Verification

This article presents a novel approach for learning domain-invariant speaker embeddings using Generative Adversarial Networks. The main idea is to confuse a domain discriminator so that is can't tell if embeddings are from the source or target domains. We train several GAN variants using our proposed framework and apply them to the speaker verification task. On the challenging NIST-SRE 2016 dataset, we are able to match the performance of a strong baseline x-vector system. In contrast to the the baseline systems which are dependent on dimensionality reduction (LDA) and an external classifier (PLDA), our proposed speaker embeddings can be scored using simple cosine distance. This is achieved by optimizing our models end-to-end, using an angular margin loss function. Furthermore, we are able to significantly boost verification performance by averaging our different GAN models at the score level, achieving a relative improvement of 7.2% over the baseline.

eess.AS

U(1) Problem Revisited

In the anomaly equation for the singlet axial current the chiral limit of the quark mass term does not vanish but comprises contribution from fermion zero modes whose integral exactly cancels the topological charge arising from the Adler-Bell-Jackiw anomaly. This signals chiral symmetry and opens a window for restoring the status of Goldstone boson for the singlet $η^\prime$ without having to invoke the large $N_c$ limit in the underlying QCD. We construct the anomaly term in the effective action that incorporates the chiral symmetry property and yet accounts for the excess mass of $η^\prime$ only when chiral symmetry is broken explicitly by the quark masses. The anomaly term in the present scenario thus plays the role of a catalytic agent that enhances the mass of $η^\prime$ so that the singlet axial current obeys the popular PCAC condition.

hep-ph

A Critical String Theory in 3+1 Dimensions

Redefining the vacuum state of a free twofold N=1 covariant supersymmetric string action as the one with all the world sheet fermionic excited states occupied, makes the theory anomaly free in D=4 with Minkowski signature. The theory thus describes a critical string in 3+1 dimensions as opposed to earlier N=2 supersymmetric theories describing a 2+2 dimensional target space. While in the NS sector the spectrum basically resembles the same for the standard N=1 superstring theory with one of the $N$ species in the background, in the R sector both the species of fermions and superconformal ghosts are required to describe the relevant spin operators to describe the fermion spectra. A crucial difference from D=10 case is that the fermion states are Dirac particles instead of Majorana-Weyl. Even though the full spectrum of the theory contains both bosons and fermions of various spin, there is no space-time supersymmetry due to obvious lack of triality.

hep-th

Cosmological perturbations in multiple-field inflation

We analyze cosmological perturbations to the linear order in the context of inflation with an arbitrary number of scalar fields. The fields take values on a non-trivial manifold with a positive-definite metric and are non-minimally coupled to Einstein gravity. The perturbations are decomposed into three different types. The scalar-type perturbations are presented in a gauge-ready form without fixing the temporal gauge condition, as well as in terms of gauge-invariant variables. The gauge-ready method enables us to impose different gauge conditions which are most suitable to the problem at hand. We quantize the scalar perturbations and obtain the solutions.

astro-ph

Perturbative analysis of multiple-field cosmological inflation

We develop a general formalism for analyzing linear perturbations in multiple-field cosmological inflation based on the gauge-ready approach. Our inflationary model consists of an arbitrary number of scalar fields with non-minimal kinetic terms. We solve the equations for scalar- and tensor-type perturbations during inflation to the first order in slow-roll, and then obtain the super-horizon solutions for adiabatic and isocurvature perturbations after inflation. Analytic expressions for power-spectra and spectral indices arising from multiple-field inflation are presented.

astro-ph

Criteria for Exact Solubility of Relativistic Field Theories by Scattering Transform

Scattering transform is a well known powerful tool for quantisation of field theories in (1+1) dimensions. Conventionally only those models whose classical counterparts admit a Lax pair (origin of which is always mysterious) have been quantised in this way. In relativistic quantum field theories we show that the scattering transforms can be constructed ab initio from its invariance under Lorentz transformation (both proper and improper), irreducible transformation nature of scalar and Dirac fields, the existence of a momentum scale associated with asymptotic nature of the scattering transform and the closure of short distance operator product algebra. For single fields it turns out that theories quantisable by scattering transforms are restricted to sine-Gordon type for spin-0 and Massive Thirring type for spin-1/2 if the target space of the scattering transform matrix is assumed to be parity invariant. There are interesting unexplored extensions if the target space is given chirality.

hep-th