SearcharxivSearch

arXiv subjects

Anuj Goyal

Publications and source records attributed to Anuj Goyal.

15 recordsLinked to original sources

Parallel-SFT: Improving Zero-Shot Cross-Programming-Language Transfer for Code RL

Modern language models demonstrate impressive coding capabilities in common programming languages (PLs), such as C++ and Python, but their performance in lower-resource PLs is often limited by training data availability. In principle, however, most programming skills are universal across PLs, so the capability acquired in one PL should transfer to others. In this work, we propose the task of zero-shot cross-programming-language transfer for code RL. We find that, for Llama-3.1, RL training for code generation in a source PL fails to improve, and sometimes even degrades, the performance on other target PLs. To address this, we hypothesize that effective RL transfer requires a generalizable SFT initialization before RL. We thus propose **Parallel-SFT**, an SFT strategy that incorporates "parallel programs" -- functionally equivalent code implemented in multiple PLs -- into the data mixture. We demonstrate that this improves transferability: when we subsequently perform RL on our Parallel-SFT model, we observe better generalization to unseen PLs. Analysis of the model internal representations reveals that Parallel-SFT leads to a more functionality-centric latent space, where equivalent programs across PLs are more tightly clustered, which we hypothesize to contribute to the improved transferability.

cs.CL

Toward More Accurate and Generalizable Evaluation Metrics for Task-Oriented Dialogs

Measurement of interaction quality is a critical task for the improvement of spoken dialog systems. Existing approaches to dialog quality estimation either focus on evaluating the quality of individual turns, or collect dialog-level quality measurements from end users immediately following an interaction. In contrast to these approaches, we introduce a new dialog-level annotation workflow called Dialog Quality Annotation (DQA). DQA expert annotators evaluate the quality of dialogs as a whole, and also label dialogs for attributes such as goal completion and user sentiment. In this contribution, we show that: (i) while dialog quality cannot be completely decomposed into dialog-level attributes, there is a strong relationship between some objective dialog attributes and judgments of dialog quality; (ii) for the task of dialog-level quality estimation, a supervised model trained on dialog-level annotations outperforms methods based purely on aggregating turn-level features; and (iii) the proposed evaluation model shows better domain generalization ability compared to the baselines. On the basis of these results, we argue that having high-quality human-annotated data is an important component of evaluating interaction quality for large industrial-scale voice assistant platforms.

cs.CL

Investigating the Electronic Structure of Prospective Water-splitting Oxide BaCe$_{0.25}$Mn$_{0.75}$O$_{3-\delta}$ Before and After Thermal Reduction

BaCe$_{0.25}$Mn$_{0.75}$O$_{3-\delta}$ (BCM), a non-stoichiometric oxide closely resembling a perovskite crystal structure, has recently emerged as a prospective contender for application in renewable energy harvesting by solar thermochemical hydrogen generation. Using solar energy, oxygen-vacancies can be created in BCM and the reduced crystal so obtained can, in turn, produce H2 by stripping oxygen from H2O. Therefore, a first step toward understanding the working mechanism and optimizing the performance of BCM, is a thorough and comparative analysis of the electronic structure of the pristine and the reduced material. In this paper, we probe the electronic structure of BCM using the combined effort of first-principles calculations and experimental O K-edge x-ray absorption spectroscopy (XAS). The computed projected density-of-states (PDOS) and orbital-plots are used to propose a simplified model for orbital-mixing between the oxygen and the ligand atoms. With the help of state-of-the-art simulations, we are able to find the origins of the XAS peaks and to categorize them on the basis of contribution from Ce and Mn. For the reduced crystal, the calculations show that, as a consequence of dielectric screening, the change in electron-density resulting from the reduction is strongly localized around the oxygen vacancy. Our experimental studies reveal a marked lowering of the first O K-edge peak in the reduced crystal which is shown to result from a diminished O-2p contribution to the frontier unoccupied orbitals, in accordance with the tight-binding scheme. Our study paves the way for investigation of the working-mechanism of BCM and for computational and experimental efforts aimed at design and discovery of efficient water-splitting oxides.

cond-mat.mtrl-sci

Alexa Conversations: An Extensible Data-driven Approach for Building Task-oriented Dialogue Systems

Traditional goal-oriented dialogue systems rely on various components such as natural language understanding, dialogue state tracking, policy learning and response generation. Training each component requires annotations which are hard to obtain for every new domain, limiting scalability of such systems. Similarly, rule-based dialogue systems require extensive writing and maintenance of rules and do not scale either. End-to-End dialogue systems, on the other hand, do not require module-specific annotations but need a large amount of data for training. To overcome these problems, in this demo, we present Alexa Conversations, a new approach for building goal-oriented dialogue systems that is scalable, extensible as well as data efficient. The components of this system are trained in a data-driven manner, but instead of collecting annotated conversations for training, we generate them using a novel dialogue simulator based on a few seed dialogues and specifications of APIs and entities provided by the developer. Our approach provides out-of-the-box support for natural conversational phenomena like entity sharing across turns or users changing their mind during conversation without requiring developers to provide any such dialogue flows. We exemplify our approach using a simple pizza ordering task and showcase its value in reducing the developer burden for creating a robust experience. Finally, we evaluate our system using a typical movie ticket booking task and show that the dialogue simulator is an essential component of the system that leads to over $50\%$ improvement in turn-level action signature prediction accuracy.

cs.CL

OodGAN: Generative Adversarial Network for Out-of-Domain Data Generation

Detecting an Out-of-Domain (OOD) utterance is crucial for a robust dialog system. Most dialog systems are trained on a pool of annotated OOD data to achieve this goal. However, collecting the annotated OOD data for a given domain is an expensive process. To mitigate this issue, previous works have proposed generative adversarial networks (GAN) based models to generate OOD data for a given domain automatically. However, these proposed models do not work directly with the text. They work with the text's latent space instead, enforcing these models to include components responsible for encoding text into latent space and decoding it back, such as auto-encoder. These components increase the model complexity, making it difficult to train. We propose OodGAN, a sequential generative adversarial network (SeqGAN) based model for OOD data generation. Our proposed model works directly on the text and hence eliminates the need to include an auto-encoder. OOD data generated using OodGAN model outperforms state-of-the-art in OOD detection metrics for ROSTD (67% relative improvement in FPR 0.95) and OSQ datasets (28% relative improvement in FPR 0.95) (Zheng et al., 2020).

cs.CL

Computational Fermi level engineering and doping-type conversion of Ga2O3 via three-step synthesis process

Ga2O3 is being actively explored for high-power and high-temperature electronics, deep-ultraviolet optoelectronics, and other applications. Efficient n-type doping of Ga2O3 has been achieved, but p-type doping faces fundamental obstacles due to compensation, deep acceptor levels, and the polaron transport mechanism of free holes. However, aside from achieving p-type conductivity, plenty of opportunity exists to engineer the position of the Fermi level for improved design of Ga2O3 based devices. We use first-principles defect theory and defect equilibrium calculations to simulate a 3-step growth-annealing-quench synthesis protocol for hydrogen assisted Mg doping in beta-Ga2O3, taking into account the gas phase equilibrium between H2, O2 and H2O, which determines the H chemical potential. We predict Ga2O3 doping-type conversion to a net p-type regime after growth under reducing conditions in the presence of H2 followed by O-rich annealing, which is a similar process to the Mg acceptor activation by H removal in GaN. For equilibrium annealing there is an optimal temperature that maximizes the Ga2O3 net acceptor density for a given Mg doping level, which is further increased in the non-equilibrium annealing scenario without re-equilibration. After quenching to operating temperature, the Ga2O3 Fermi level drops below mid-gap down to about +1.5 eV above the valence band maximum, creating a significant number of uncompensated neutral MgGa0 acceptors. The Fermi level reduction down to +1.5 eV and suppression of free electron density in this doping type converted (NA > ND) Ga2O3 material is of significance and impact for the design of Ga2O3 power electronics devices.

cond-mat.mtrl-sci

The influence of alloying on the stacking fault energy of gold from density functional theory calculations

The generalized stacking fault (SFE) energy curves of pure gold (Au) and its binary alloys with transition metals are determined from density functional theory (DFT). Alloy elements Ag, Al, Cu, Ni, Ti, Zr, Zn, In, Ga, Sn, Mn, Cd, Sn, Ta and Cr are substituted into Au at concentrations up to 4%. A comparison of various proposed methodologies to calculate SFEs is given. The intrinsic SFE decreases for all alloying elements from its value for pure Au, but SFE energies (both stable and unstable) vary strongly with the distance of the alloying element from the stacking fault region, and with alloy concentration. The compositional dependence of the SFE on the volume change associated with alloying element is determined. This work demonstrates that the SFE is strongly influenced by misfit strain caused by the alloying elements. Moreover, the computed generalized SFE curves provide information valuable to developing an understanding of the deformation behavior of Au and Au-alloys.

cond-mat.mtrl-sci

On the Dopability of Semiconductors and Governing Material Properties

To be practical, semiconductors need to be doped. Sometimes, to nearly degenerate levels, e.g. in applications such as thermoelectric, transparent electronics or power electronics. However, many materials with finite band gaps are not dopable at all, while many others exhibit strong preference toward allowing either p- or n-type doping, but not both. In this work, we develop a model description of semiconductor dopability and formulate design principles in terms of governing materials properties. Our approach, which builds upon the semiconductor defect theory applied to a suitably devised (tight-binding) model system, reveals analytic relationships between intrinsic materials properties and the semiconductor dopability, and elucidates the role and the insufficiency of previously suggested descriptors such as the absolute band edge positions. We validate our model against a number of classic binary semiconductors and discuss its extension to more complex chemistries and the utility in large-scale material searches.

cond-mat.mtrl-sci

Controlled Text Generation for Data Augmentation in Intelligent Artificial Agents

Data availability is a bottleneck during early stages of development of new capabilities for intelligent artificial agents. We investigate the use of text generation techniques to augment the training data of a popular commercial artificial agent across categories of functionality, with the goal of faster development of new functionality. We explore a variety of encoder-decoder generative models for synthetic training data generation and propose using conditional variational auto-encoders. Our approach requires only direct optimization, works well with limited data and significantly outperforms the previous controlled text generation techniques. Further, the generated data are used as additional training samples in an extrinsic intent classification task, leading to improved performance by up to 5\% absolute f-score in low-resource cases, validating the usefulness of our approach.

cs.CL

Simple Question Answering with Subgraph Ranking and Joint-Scoring

Knowledge graph based simple question answering (KBSQA) is a major area of research within question answering. Although only dealing with simple questions, i.e., questions that can be answered through a single knowledge base (KB) fact, this task is neither simple nor close to being solved. Targeting on the two main steps, subgraph selection and fact selection, the research community has developed sophisticated approaches. However, the importance of subgraph ranking and leveraging the subject--relation dependency of a KB fact have not been sufficiently explored. Motivated by this, we present a unified framework to describe and analyze existing approaches. Using this framework as a starting point, we focus on two aspects: improving subgraph selection through a novel ranking method and leveraging the subject--relation dependency by proposing a joint scoring CNN model with a novel loss function that enforces the well-order of scores. Our methods achieve a new state of the art (85.44% in accuracy) on the SimpleQuestions dataset.

cs.CL

Unsupervised Transfer Learning for Spoken Language Understanding in Intelligent Agents

User interaction with voice-powered agents generates large amounts of unlabeled utterances. In this paper, we explore techniques to efficiently transfer the knowledge from these unlabeled utterances to improve model performance on Spoken Language Understanding (SLU) tasks. We use Embeddings from Language Model (ELMo) to take advantage of unlabeled data by learning contextualized word representations. Additionally, we propose ELMo-Light (ELMoL), a faster and simpler unsupervised pre-training method for SLU. Our findings suggest unsupervised pre-training on a large corpora of unlabeled utterances leads to significantly better SLU performance compared to training from scratch and it can even outperform conventional supervised transfer. Additionally, we show that the gains from unsupervised transfer techniques can be further improved by supervised transfer. The improvements are more pronounced in low resource settings and when using only 1000 labeled in-domain samples, our techniques match the performance of training from scratch on 10-15x more labeled in-domain data.

cs.CL

Fast and Scalable Expansion of Natural Language Understanding Functionality for Intelligent Agents

Fast expansion of natural language functionality of intelligent virtual agents is critical for achieving engaging and informative interactions. However, developing accurate models for new natural language domains is a time and data intensive process. We propose efficient deep neural network architectures that maximally re-use available resources through transfer learning. Our methods are applied for expanding the understanding capabilities of a popular commercial agent and are evaluated on hundreds of new domains, designed by internal or external developers. We demonstrate that our proposed methods significantly increase accuracy in low resource settings and enable rapid development of accurate models with less data.

cs.CL

Metastable rocksalt ZnO is $p$-type dopable

Despite decades of efforts, achieving $p$-type conductivity in the wide band gap ZnO in its ground-state wurtzite structure continues to be a challenge. Here we detail how $p$-type ZnO can be realized in the metastable, high-pressure rocksalt phase (also wide-gap) with Li as an external dopant. Using modern first-principles defect theory, we predict Li to dope the rocksalt phase $p$-type by preferentially substituting for Zn and introducing shallow acceptor levels. Formation of compensating donors like interstitial Li and/or hydrogen, ubiqutous in the wurtzite phase, is inhibited by the close-packed nature of the rocksalt structure, which also exhibits relatively high absolute valence band edge that promotes low hole effective mass and hole delocalization. Resulting concentrations of free holes are predicted to exceed $\sim10^{19}$ cm$^{-3}$ under O-rich synthesis conditions while under O-poor conditions the system remains $n$-type dopable. In addition to revealing compelling opportunities offered by the metastable rocksalt structure in realizing a long-sought $p$-type ZnO our results present polymorphism as a promising route to overcoming strong doping asymmetry of wide-band gap oxides.

cond-mat.mtrl-sci

The conundrum of relaxation volumes in first-principles calculations of charge defects

The defect relaxation volumes obtained from density-functional theory (DFT) calculations of charged vacancies and interstitials are much larger than their neutral counterparts, seemingly unphysically large. In this work, we investigate the possible reasons for this and revisit the methods that address the calculation of charged defect structures in periodic DFT. We probe the dependence of the proposed energy corrections to charged defect formation energies on relaxation volumes and find that corrections such as the image charge correction and the alignment correction, which can lead to sizable changes in defect formation energies, have an almost negligible effect on the charged defect relaxation volume. We also investigate the volume for the net neutral defect reactions comprised of individual charged defects, and find that the aggregate formation volumes have reasonable magnitudes. This work highlights an important issue that, as for defect formation energies, the defect formation volumes depend on the choice of reservoir. We show that considering the change in volume of the electron reservoir in the formation reaction of the charged defects, analogous to how volumes of atoms are accounted for in defect formation volumes, can renormalize the formation volumes of charged defects such that they are comparable to neutral defects. This approach enables the description of the elastic properties of isolated charged defects within the overall neutral material, beyond the context of the overall defect reactions that produce the charged defect.

cond-mat.mtrl-sci

A Computational Framework for Automation of Point Defect Calculations

A complete and rigorously validated open-source Python framework to automate point defect calculations using density functional theory has been developed. The framework provides an effective and efficient method for defect structure generation, and creation of simple yet customizable workflows to analyze defect calculations. The package provides the capability to compute widely-accepted correction schemes to overcome finite-size effects, including (1) potential alignment, (2) image-charge correction, and (3) band filling correction to shallow defects. Using Si, ZnO and In$_2$O$_3$ as test examples, we demonstrate the package capabilities and validate the methodology.

cond-mat.mtrl-sci