Searcharxiv⌕ Search

arXiv subjects

Sachin Kumar

Publications and source records attributed to Sachin Kumar.

At least 73 records · Page 4Linked to original sources

On the Blind Spots of Model-Based Evaluation Metrics for Text Generation

In this work, we explore a useful but often neglected methodology for robustness analysis of text generation evaluation metrics: stress tests with synthetic data. Basically, we design and synthesize a wide range of potential errors and check whether they result in a commensurate drop in the metric scores. We examine a range of recently proposed evaluation metrics based on pretrained language models, for the tasks of open-ended generation, translation, and summarization. Our experiments reveal interesting insensitivities, biases, or even loopholes in existing metrics. For example, we find that BERTScore is confused by truncation errors in summarization, and MAUVE (built on top of GPT-2) is insensitive to errors at the beginning or middle of generations. Further, we investigate the reasons behind these blind spots and suggest practical workarounds for a more reliable evaluation of text generation. We have released our code and data at https://github.com/cloudygoose/blindspot_nlg.

cs.CL↗

Assessing Language Model Deployment with Risk Cards

This paper introduces RiskCards, a framework for structured assessment and documentation of risks associated with an application of language models. As with all language, text generated by language models can be harmful, or used to bring about harm. Automating language generation adds both an element of scale and also more subtle or emergent undesirable tendencies to the generated text. Prior work establishes a wide variety of language model harms to many different actors: existing taxonomies identify categories of harms posed by language models; benchmarks establish automated tests of these harms; and documentation standards for models, tasks and datasets encourage transparent reporting. However, there is no risk-centric framework for documenting the complexity of a landscape in which some risks are shared across models and contexts, while others are specific, and where certain conditions may be required for risks to manifest as harms. RiskCards address this methodological gap by providing a generic framework for assessing the use of a given language model in a given scenario. Each RiskCard makes clear the routes for the risk to manifest harm, their placement in harm taxonomies, and example prompt-output pairs. While RiskCards are designed to be open-source, dynamic and participatory, we present a "starter set" of RiskCards taken from a broad literature survey, each of which details a concrete risk presentation. Language model RiskCards initiate a community knowledge base which permits the mapping of risks and harms to a specific model or its application scenario, ultimately contributing to a better, safer and shared understanding of the risk landscape.

cs.CL↗

Language Generation Models Can Cause Harm: So What Can We Do About It? An Actionable Survey

Recent advances in the capacity of large language models to generate human-like text have resulted in their increased adoption in user-facing settings. In parallel, these improvements have prompted a heated discourse around the risks of societal harms they introduce, whether inadvertent or malicious. Several studies have explored these harms and called for their mitigation via development of safer, fairer models. Going beyond enumerating the risks of harms, this work provides a survey of practical methods for addressing potential threats and societal harms from language generation models. We draw on several prior works' taxonomies of language model risks to present a structured overview of strategies for detecting and ameliorating different kinds of risks/harms of language generators. Bridging diverse strands of research, this survey aims to serve as a practical guide for both LM researchers and practitioners, with explanations of different mitigation strategies' motivations, their limitations, and open problems for future research.

cs.CL↗

Gradient-Based Constrained Sampling from Language Models

Large pretrained language models generate fluent text but are notoriously hard to controllably sample from. In this work, we study constrained sampling from such language models: generating text that satisfies user-defined constraints, while maintaining fluency and the model's performance in a downstream task. We propose MuCoLa -- a sampling procedure that combines the log-likelihood of the language model with arbitrary (differentiable) constraints in a single energy function, and then generates samples in a non-autoregressive manner. Specifically, it initializes the entire output sequence with noise and follows a Markov chain defined by Langevin Dynamics using the gradients of the energy function. We evaluate MuCoLa on text generation with soft and hard constraints as well as their combinations obtaining significant improvements over competitive baselines for toxicity avoidance, sentiment control, and keyword-guided generation.

cs.CL↗

Referee: Reference-Free Sentence Summarization with Sharper Controllability through Symbolic Knowledge Distillation

We present Referee, a novel framework for sentence summarization that can be trained reference-free (i.e., requiring no gold summaries for supervision), while allowing direct control for compression ratio. Our work is the first to demonstrate that reference-free, controlled sentence summarization is feasible via the conceptual framework of Symbolic Knowledge Distillation (West et al., 2022), where latent knowledge in pre-trained language models is distilled via explicit examples sampled from the teacher models, further purified with three types of filters: length, fidelity, and Information Bottleneck. Moreover, we uniquely propose iterative distillation of knowledge, where student models from the previous iteration of distillation serve as teacher models in the next iteration. Starting off from a relatively modest set of GPT3-generated summaries, we demonstrate how iterative knowledge distillation can lead to considerably smaller, but better summarizers with sharper controllability. A useful by-product of this iterative distillation process is a high-quality dataset of sentence-summary pairs with varying degrees of compression ratios. Empirical results demonstrate that the final student models vastly outperform the much larger GPT3-Instruct model in terms of the controllability of compression ratios, without compromising the quality of resulting summarization.

cs.CL↗

Nonlinear Dynamic analysis of vector-host model for Zika infection with predatory fish Gambusia Affinis

In the present paper, we study the dynamics of a nine compartmental vector-host model for Zika virus infection where the predatory fish Gambusia Affinis is introduced into the system to control the zika infection by preying on the vector. The system has six practically feasible equilibrium points where four of them are disease-free, and the rest are endemic. We discuss the existence and stability conditions for the equilibria. We find that when sexual transmission of zika comes to a halt then in absence of mosquitoes infection cannot persist. Hence, one needs to eradicate mosquitoes to eradicate infection. Moreover, we deduce that in the case of zika infection pushing the basic reproduction number below unity is next to impossible. Therefore, O_0, the mosquito survival threshold parameter, and O, the mosquito survival threshold parameter with predation play a crucial role in getting rid of the infection in respective cases since mosquitoes cannot survive when these are less than unity. Sensitivity analysis shows the importance of reducing mosquito biting rate and mutual contact rates between vector and host. It exhibits the importance of increasing the natural mortality rate of vectors to reduce the basic reproduction number. Numerical simulation shows that when the basic reproduction number is close but greater than unity, the introduction of a small amount of predatory fish Gambusia Affinis can completely swipe off the infection. In case of high transmission or high basic reproduction number, this fish increases the susceptible human population and keeps the infection under control, hence, prohibiting it from becoming an epidemic.

math.DS↗

Improving the Diversity of Unsupervised Paraphrasing with Embedding Outputs

We present a novel technique for zero-shot paraphrase generation. The key contribution is an end-to-end multilingual paraphrasing model that is trained using translated parallel corpora to generate paraphrases into "meaning spaces" -- replacing the final softmax layer with word embeddings. This architectural modification, plus a training procedure that incorporates an autoencoding objective, enables effective parameter sharing across languages for more fluent monolingual rewriting, and facilitates fluency and diversity in generation. Our continuous-output paraphrase generation models outperform zero-shot paraphrasing baselines when evaluated on two languages using a battery of computational metrics as well as in human assessment.

cs.CL↗

Machine Translation into Low-resource Language Varieties

State-of-the-art machine translation (MT) systems are typically trained to generate the "standard" target language; however, many languages have multiple varieties (regional varieties, dialects, sociolects, non-native varieties) that are different from the standard language. Such varieties are often low-resource, and hence do not benefit from contemporary NLP solutions, MT included. We propose a general framework to rapidly adapt MT systems to generate language varieties that are close to, but different from, the standard target language, using no parallel (source--variety) data. This also includes adaptation of MT systems to low-resource typologically-related target languages. We experiment with adapting an English--Russian MT system to generate Ukrainian and Belarusian, an English--Norwegian Bokmål system to generate Nynorsk, and an English--Arabic system to generate four Arabic dialects, obtaining significant improvements over competitive baselines.

cs.CL↗

Controlled Text Generation as Continuous Optimization with Multiple Constraints

As large-scale language model pretraining pushes the state-of-the-art in text generation, recent work has turned to controlling attributes of the text such models generate. While modifying the pretrained models via fine-tuning remains the popular approach, it incurs a significant computational cost and can be infeasible due to lack of appropriate data. As an alternative, we propose MuCoCO -- a flexible and modular algorithm for controllable inference from pretrained models. We formulate the decoding process as an optimization problem which allows for multiple attributes we aim to control to be easily incorporated as differentiable constraints to the optimization. By relaxing this discrete optimization to a continuous one, we make use of Lagrangian multipliers and gradient-descent based techniques to generate the desired text. We evaluate our approach on controllable machine translation and style transfer with multiple sentence-level attributes and observe significant improvements over baselines.

cs.CL↗

Topics to Avoid: Demoting Latent Confounds in Text Classification

Despite impressive performance on many text classification tasks, deep neural networks tend to learn frequent superficial patterns that are specific to the training data and do not always generalize well. In this work, we observe this limitation with respect to the task of native language identification. We find that standard text classifiers which perform well on the test set end up learning topical features which are confounds of the prediction task (e.g., if the input text mentions Sweden, the classifier predicts that the author's native language is Swedish). We propose a method that represents the latent topical confounds and a model which "unlearns" confounding features by predicting both the label of the input text and the confound; but we train the two predictors adversarially in an alternating fashion to learn a text representation that predicts the correct label but is less prone to using information about the confound. We show that this model generalizes better and learns features that are indicative of the writing style rather than the content.

cs.LG↗

An Exploration of Data Augmentation Techniques for Improving English to Tigrinya Translation

It has been shown that the performance of neural machine translation (NMT) drops starkly in low-resource conditions, often requiring large amounts of auxiliary data to achieve competitive results. An effective method of generating auxiliary data is back-translation of target language sentences. In this work, we present a case study of Tigrinya where we investigate several back-translation methods to generate synthetic source sentences. We find that in low-resource conditions, back-translation by pivoting through a higher-resource language related to the target language proves most effective resulting in substantial improvements over baselines.

cs.CL↗

CAMTA: Causal Attention Model for Multi-touch Attribution

Advertising channels have evolved from conventional print media, billboards and radio advertising to online digital advertising (ad), where the users are exposed to a sequence of ad campaigns via social networks, display ads, search etc. While advertisers revisit the design of ad campaigns to concurrently serve the requirements emerging out of new ad channels, it is also critical for advertisers to estimate the contribution from touch-points (view, clicks, converts) on different channels, based on the sequence of customer actions. This process of contribution measurement is often referred to as multi-touch attribution (MTA). In this work, we propose CAMTA, a novel deep recurrent neural network architecture which is a casual attribution mechanism for user-personalised MTA in the context of observational data. CAMTA minimizes the selection bias in channel assignment across time-steps and touchpoints. Furthermore, it utilizes the users' pre-conversion actions in a principled way in order to predict pre-channel attribution. To quantitatively benchmark the proposed MTA model, we employ the real world Criteo dataset and demonstrate the superior performance of CAMTA with respect to prediction accuracy as compared to several baselines. In addition, we provide results for budget allocation and user-behaviour modelling on the predicted channel attribution.

cs.LG↗

An update on coherent scattering from complex non-PT-symmetric Scarf II potential with new analytic forms

The versatile and exactly solvable Scarf II has been predicting, confirming and demonstrating interesting phenomena in complex PT-symmetric sector, most impressively. However, for the non-PT-symmetric sector it has gone underutilized. Here, we present most simple analytic forms for the scattering coefficients $(T(k),R(k),|\det S(k)|)$. On one hand, these forms demonstrate earlier effects and confirm the recent ones. On the other hand they make new predictions - all simply and analytically. We show the possibilities of both self-dual and non-self-dual spectral singularities (NSDSS) in two non-PT sectors (potentials). The former one is not accompanied by time-reversed coherent perfect absorption (CPA) and gives rise to the parametrically controlled splitting of SS in to a finite number of complex conjugate pairs of eigenvalues (CCPEs). The latter ones (NSDSS) behave just oppositely: CPA but no splitting of SS. We demonstrate a one-sided reflectionlessness without invisibility. Most importantly, we bring out a surprising co-existence of both real discrete spectrum and a single SS in a fixed potential. Nevertheless, the complex Scarf II is not known to be pseudo-Hermitian ($η^{-1} Hη=H^\dagger$) under a metric of the type $η(x)$, so far.

quant-ph↗

Neural Abstractive Summarization with Structural Attention

Attentional, RNN-based encoder-decoder architectures have achieved impressive performance on abstractive summarization of news articles. However, these methods fail to account for long term dependencies within the sentences of a document. This problem is exacerbated in multi-document summarization tasks such as summarizing the popular opinion in threads present in community question answering (CQA) websites such as Yahoo! Answers and Quora. These threads contain answers which often overlap or contradict each other. In this work, we present a hierarchical encoder based on structural attention to model such inter-sentence and inter-document dependencies. We set the popular pointer-generator architecture and some of the architectures derived from it as our baselines and show that they fail to generate good summaries in a multi-document setting. We further illustrate that our proposed model achieves significant improvement over the baselines in both single and multi-document summarization settings -- in the former setting, it beats the best baseline by 1.31 and 7.8 ROUGE-1 points on CNN and CQA datasets, respectively; in the latter setting, the performance is further improved by 1.6 ROUGE-1 points on the CQA dataset.

cs.CL↗

A Deep Reinforced Model for Zero-Shot Cross-Lingual Summarization with Bilingual Semantic Similarity Rewards

Cross-lingual text summarization aims at generating a document summary in one language given input in another language. It is a practically important but under-explored task, primarily due to the dearth of available data. Existing methods resort to machine translation to synthesize training data, but such pipeline approaches suffer from error propagation. In this work, we propose an end-to-end cross-lingual text summarization model. The model uses reinforcement learning to directly optimize a bilingual semantic similarity metric between the summaries generated in a target language and gold summaries in a source language. We also introduce techniques to pre-train the model leveraging monolingual summarization and machine translation objectives. Experimental results in both English--Chinese and English--German cross-lingual summarization settings demonstrate the effectiveness of our methods. In addition, we find that reinforcement learning models with bilingual semantic similarity as rewards generate more fluent sentences than strong baselines.

cs.CL↗

On exact solutions, conservation laws and invariant analysis of the generalized Rosenau-Hyman equation

In this paper, the nonlinear Rosenau-Hyman equation with time dependent variable coefficients is considered for investigating its invariant properties, exact solutions and conservation laws. Using Lie classical method, we derive symmetries admitted by considered equation. Symmetry reductions are performed for each components of optimal set. Also nonclassical approach is employed on considered equation to find some additional supplementary symmetries and corresponding symmetry reductions are performed. Later three kinds of exact solutions of considered equation are presented graphically for different parameters. In addition, local conservation laws are constructed for considered equation by multiplier approach.

nlin.SI↗

PT-symmetric potentials with imaginary asymptotic saturation

We point out that PT-symmetric potentials $V_{PT}(x)$ having imaginary asymptotic saturation: $V_{PT}(x=\pm \infty) =\pm i V_1, V_1 \in \Re$ are devoid of scattering states and spectral singularity. We show the existence of real (positive and negative) discrete spectrum both with and without complex conjugate pair(s) of eigenvalues (CCPEs). If the states are arranged in the ascending order or real part of discrete eigenvalues, the initial states have few nodes but latter ones oscillate fast. Both real and imaginary parts of $ψ(x)$ vanish asymptotically, $|ψ(x)|$ for the CCPEs are asymmetric and for real energies these are symmetric about origin. For CCPEs $E_{\pm}$ the eigenstates $ψ_{\pm}$ follow an interesting property that $|ψ_+(x)|= N |ψ_-(-x)|, N \in \Re^+$. We remark that, the fast oscillating real discrete energy states discussed are likely to be confused with: reflectionless states, one dimensional version of von Neumann states of Hermitian and spectral singularity state of complex PT-symmetric potentials.

quant-ph↗

Low reflection at zero or low-energies in the well-barrier scattering potentials

Probability of reflection $R(E)$ off a finite attractive scattering potential at zero or low energies is ordinarily supposed to be 1. However, a fully attractive potential presents a paradoxical result that $R(0)=0$ or $R(0)<1$, when an effective parameter $q$ of the potential admits special discrete values. Here, we report another class of finite potentials which are well-barrier (attractive-repulsive) type and which can be made to possess much less reflection at zero and low energies for a band of low values of $q$. These well-barrier potentials have only two real turning points for $E \in(V_{min}, V_{max})$, excepting $E=0$. We present two exactly solvable and two numerically solved models to confirm this phenomenon.

quant-ph↗