Searcharxiv⌕ Search

arXiv subjects

Thai-Son Nguyen

Publications and source records attributed to Thai-Son Nguyen.

18 recordsLinked to original sources

Second-Order Optical Nonlinearity of AlScN Films Grown By Molecular Beam Epitaxy

Alloys of AlN have rapidly emerged as a material platform for nonlinear optics. In this paper, we measure the second-order optical nonlinearity of AlScN films grown directly on nitrided c-plane sapphire by molecular beam epitaxy. This direct growth approach, which bypasses a thick AlN buffer layer, allows us to isolate the true nonlinear response of the AlScN film. Our results show a large enhancement of d31, but a suppression of d33 in AlScN films compared to AlN. We observe that d31 can be as high as 4.92 pm/V , which is 60 times larger than that of AlN. The development of AlScN-based photonic devices can enable energy-efficient nonlinear optical operations that can be epitaxially integrated with electronic and photonic devices based on Si, GaN and AlN.

cond-mat.mtrl-sci↗

Shubnikov-de Haas oscillations of two-dimensional electron gases in AlYN/GaN and AlScN/GaN heterostructures

AlYN and AlScN have recently emerged as promising nitride materials that can be integrated with GaN to form two-dimensional electron gases (2DEGs) at heterojunctions. Electron transport properties in these heterostructures have been enhanced through careful design and optimization of epitaxial growth conditions. In this work, we report for the first time Shubnikov-de Haas (SdH) oscillations of 2DEGs in AlYN/GaN and AlScN/GaN heterostructures, grown by metal-organic chemical vapor deposition. SdH oscillations provide direct access to key 2DEG parameters at the Fermi level: (1) carrier density, (2) electron effective mass (m* ~ 0.24 me for AlYN/GaN and m* ~ 0.25 me for AlScN/GaN), and (3) quantum scattering time (~ 68 fs for AlYN/GaN and ~ 70 fs for AlScN/GaN). These measurements of fundamental transport properties provide critical insights for advancing emerging nitride semiconductors for future high-frequency and power electronics.

cond-mat.mes-hall↗

High Mobility Multiple-Channel AlScN/GaN Heterostructures

Aluminum scandium nitride (AlScN) is a promising barrier material for gallium nitride (GaN)-based transistors for the next generation of radio-frequency electronic devices. In this work, we examine the transport properties of two dimensional electron gases (2DEGs) in single- and multi-channel AlScN/GaN heterostructures grown by molecular beam epitaxy, and demonstrate the lowest sheet resistance among AlScN-based systems reported to date. Assorted schemes of GaN/AlN interlayers are first introduced in single-channel structures between AlScN and GaN to improve conductivity, increasing electron mobility up to $1370$ cm$^{2}$/V$\cdot$s at 300 K and $4160$ cm$^{2}$/V$\cdot$s at 77 K, reducing the sheet resistance down to 170 $Ω/\square$ and 70 $Ω/\square$ respectively. These improvements are then leveraged in multi-channel heterostructures, reaching sheet resistances of 65 $Ω/\square$ for three channels and 45 $Ω/\square$ for five channels at 300 K, further reduced to 21 $Ω/\square$ and 13 $Ω/\square$ at 2 K, respectively, confirming the presence of multiple 2DEGs. Structural characterization indicates pseudomorphic growth with smooth surfaces, while partial barrier relaxation and surface roughening are observed at high scandium content, with no impact on mobility. This first demonstration of ultra-low sheet resistance multi-channel AlScN/GaN heterostructures places AlScN on par with state-of-the-art multi-channel Al(In)N/GaN systems, showcasing its capacity to advance existing and enable new high-speed, high-power electronic devices.

cond-mat.mtrl-sci↗

A Compact Model for Polar Multiple-Channel Field Effect Transistors: A Case Study in III-V Nitride Semiconductors

A compact analytical model is developed for the mobile charge density of polar multiple channel field effect transistors. Two dimensional electron and hole gases can be potentially induced by spontaneous and piezoelectric polarization in polar heterostructures. Focusing on the active region of devices that employ a multiple quantum-well layout, the total electron and hole populations are estimated from fundamental electrostatic and quantum mechanical principles. Hole gas depletion techniques, revolving around intentional donor doping, are modeled and evaluated, culminating in a generalized closed-form equation for the mobile carrier density across the doping schemes examined. The utility of this model is illustrated for the III-Nitride material system, exploring AlGaN/GaN, AlInN/GaN and AlScN/GaN heterostructures. The compact framework provided herein considerably elucidates and enhances the efficiency of multi-layered transistor design.

physics.app-ph↗

Epitaxial high-K AlBN barrier GaN HEMTs

We report a polarization-induced 2D electron gas (2DEG) at an epitaxial AlBN/GaN heterojunction grown on a SiC substrate. Using this 2DEG in a long conducting channel, we realize ultra-thin barrier AlBN/GaN high electron mobility transistors that exhibit current densities of more than 0.25 A/mm, clean current saturation, a low pinch-off voltage of -0.43 V, and a peak transconductance of 0.14 S/mm. Transistor performance in this preliminary realization is limited by the contact resistance. Capacitance-voltage measurements reveal that introducing 7 % B in the epitaxial AlBN barrier on GaN boosts the relative dielectric constant of AlBN to 16, higher than the AlN dielectric constant of 9. Epitaxial high-K barrier AlBN/GaN HEMTs can thus extend performance beyond the capabilities of current GaN transistors.

physics.app-ph↗

Lattice-Matched Multiple Channel AlScN/GaN Heterostructures

AlScN is a new wide bandgap, high-k, ferroelectric material for RF, memory, and power applications. Successful integration of high quality AlScN with GaN in epitaxial layer stacks depends strongly on the ability to control lattice parameters and surface or interface through growth. This study investigates the molecular beam epitaxy growth and transport properties of AlScN/GaN multilayer heterostructures. Single layer Al$_{1-x}$Sc$_x$N/GaN heterostructures exhibited lattice-matched composition within $x$ = 0.09 -- 0.11 using substrate (thermocouple) growth temperatures between 330 $ ^\circ$C and 630 $ ^\circ$C. By targeting the lattice-matched Sc composition, pseudomorphic AlScN/GaN multilayer structures with ten and twenty periods were achieved, exhibiting excellent structural and interface properties as confirmed by X-ray diffraction (XRD) and scanning transmission electron microscopy (STEM). These multilayer heterostructures exhibited substantial polarization-induced net mobile charge densities of up to 8.24 $\times$ 10$^{14}$/cm$^2$ for twenty channels. The sheet density scales with the number of AlScN/GaN periods. By identifying lattice-matched growth condition and using it to generate multiple conductive channels, this work enhances our understanding of the AlScN/GaN material platform.

cond-mat.mtrl-sci↗

Ferroelectric AlBN Films by Molecular Beam Epitaxy

We report the properties of molecular beam epitaxy deposited AlBN thin films on a recently developed epitaxial nitride metal electrode Nb2N. While a control AlN thin film exhibits standard capacitive behavior, distinct ferroelectric switching is observed in the AlBN films with increasing Boron mole fraction. The measured remnant polarization Pr of 15 uC/cm2 and coercive field Ec of 1.45 MV/cm in these films are smaller than those recently reported on films deposited by sputtering, due to incomplete wake-up, limited by current leakage. Because AlBN preserves the ultrawide energy bandgap of AlN compared to other nitride hi-K dielectrics and ferroelectrics, and it can be epitaxially integrated with GaN and AlN semiconductors, its development will enable several opportunities for unique electronic, photonic, and memory devices.

cond-mat.mtrl-sci↗

Multi-stage Large Language Model Correction for Speech Recognition

In this paper, we investigate the usage of large language models (LLMs) to improve the performance of competitive speech recognition systems. Different from previous LLM-based ASR error correction methods, we propose a novel multi-stage approach that utilizes uncertainty estimation of ASR outputs and reasoning capability of LLMs. Specifically, the proposed approach has two stages: the first stage is about ASR uncertainty estimation and exploits N-best list hypotheses to identify less reliable transcriptions; The second stage works on these identified transcriptions and performs LLM-based corrections. This correction task is formulated as a multi-step rule-based LLM reasoning process, which uses explicitly written rules in prompts to decompose the task into concrete reasoning steps. Our experimental results demonstrate the effectiveness of the proposed method by showing 10% ~ 20% relative improvement in WER over competitive ASR systems -- across multiple test domains and in zero-shot settings.

cs.CL↗

Epitaxial lattice-matched Al$_{0.89}$Sc$_{0.11}$N/GaN distributed Bragg reflectors

We demonstrate epitaxial lattice-matched Al$_{0.89}$Sc$_{0.11}$N/GaN ten and twenty period distributed Bragg reflectors (DBRs) grown on c-plane bulk n-type GaN substrates by plasma-enhanced molecular beam epitaxy (PA-MBE). Resulting from a rapid increase of in-plane lattice coefficient as scandium is incorporated into AlScN, we measure a lattice-matched condition to $c$-plane GaN for a Sc content of just 11\%, resulting in a large refractive index mismatch $\mathrm{Δn}$ greater than 0.3 corresponding to an index contrast of $\mathrm{Δn/n_{GaN}}$ = 0.12 with GaN. The DBRs demonstrated here are designed for a peak reflectivity at a wavelength of 400 nm reaching a reflectivity of 0.98 for twenty periods. It is highlighted that AlScN/GaN multilayers require fewer periods for a desired reflectivity than other lattice-matched Bragg reflectors such as those based on AlInN/GaN multilayers.

cond-mat.mtrl-sci↗

Super-Human Performance in Online Low-latency Recognition of Conversational Speech

Achieving super-human performance in recognizing human speech has been a goal for several decades, as researchers have worked on increasingly challenging tasks. In the 1990's it was discovered, that conversational speech between two humans turns out to be considerably more difficult than read speech as hesitations, disfluencies, false starts and sloppy articulation complicate acoustic processing and require robust handling of acoustic, lexical and language context, jointly. Early attempts with statistical models could only reach error rates over 50% and far from human performance (WER of around 5.5%). Neural hybrid models and recent attention-based encoder-decoder models have considerably improved performance as such contexts can now be learned in an integral fashion. However, processing such contexts requires an entire utterance presentation and thus introduces unwanted delays before a recognition result can be output. In this paper, we address performance as well as latency. We present results for a system that can achieve super-human performance (at a WER of 5.0%, over the Switchboard conversational benchmark) at a word based latency of only 1 second behind a speaker's speech. The system uses multiple attention-based encoder-decoder networks integrated within a novel low latency incremental inference approach.

cs.CV↗

High Performance Sequence-to-Sequence Model for Streaming Speech Recognition

Recently sequence-to-sequence models have started to achieve state-of-the-art performance on standard speech recognition tasks when processing audio data in batch mode, i.e., the complete audio data is available when starting processing. However, when it comes to performing run-on recognition on an input stream of audio data while producing recognition results in real-time and with low word-based latency, these models face several challenges. For many techniques, the whole audio sequence to be decoded needs to be available at the start of the processing, e.g., for the attention mechanism or the bidirectional LSTM (BLSTM). In this paper, we propose several techniques to mitigate these problems. We introduce an additional loss function controlling the uncertainty of the attention mechanism, a modified beam search identifying partial, stable hypotheses, ways of working with BLSTM in the encoder, and the use of chunked BLSTM. Our experiments show that with the right combination of these techniques, it is possible to perform run-on speech recognition with low word-based latency without sacrificing in word error rate performance.

eess.AS↗

ELITR Non-Native Speech Translation at IWSLT 2020

This paper is an ELITR system submission for the non-native speech translation task at IWSLT 2020. We describe systems for offline ASR, real-time ASR, and our cascaded approach to offline SLT and real-time SLT. We select our primary candidates from a pool of pre-existing systems, develop a new end-to-end general ASR system, and a hybrid ASR trained on non-native speech. The provided small validation set prevents us from carrying out a complex validation, but we submit all the unselected candidates for contrastive evaluation on the test set.

cs.CL↗

Relative Positional Encoding for Speech Recognition and Direct Translation

Transformer models are powerful sequence-to-sequence architectures that are capable of directly mapping speech inputs to transcriptions or translations. However, the mechanism for modeling positions in this model was tailored for text modeling, and thus is less ideal for acoustic inputs. In this work, we adapt the relative position encoding scheme to the Speech Transformer, where the key addition is relative distance between input states in the self-attention network. As a result, the network can better adapt to the variable distributions present in speech data. Our experiments show that our resulting model achieves the best recognition result on the Switchboard benchmark in the non-augmentation condition, and the best published result in the MuST-C speech translation benchmark. We also show that this model is able to better utilize synthetic data than the Transformer, and adapts better to variable sentence segmentation quality for speech translation.

eess.AS↗

Toward Cross-Domain Speech Recognition with End-to-End Models

In the area of multi-domain speech recognition, research in the past focused on hybrid acoustic models to build cross-domain and domain-invariant speech recognition systems. In this paper, we empirically examine the difference in behavior between hybrid acoustic models and neural end-to-end systems when mixing acoustic training data from several domains. For these experiments we composed a multi-domain dataset from public sources, with the different domains in the corpus covering a wide variety of topics and acoustic conditions such as telephone conversations, lectures, read speech and broadcast news. We show that for the hybrid models, supplying additional training data from other domains with mismatched acoustic conditions does not increase the performance on specific domains. However, our end-to-end models optimized with sequence-based criterion generalize better than the hybrid models on diverse domains. In term of word-error-rate performance, our experimental acoustic-to-word and attention-based models trained on multi-domain dataset reach the performance of domain-specific long short-term memory (LSTM) hybrid models, thus resulting in multi-domain speech recognition systems that do not suffer in performance over domain specific ones. Moreover, the use of neural end-to-end models eliminates the need of domain-adapted language models during recognition, which is a great advantage when the input domain is unknown.

eess.AS↗

Improving sequence-to-sequence speech recognition training with on-the-fly data augmentation

Sequence-to-Sequence (S2S) models recently started to show state-of-the-art performance for automatic speech recognition (ASR). With these large and deep models overfitting remains the largest problem, outweighing performance improvements that can be obtained from better architectures. One solution to the overfitting problem is increasing the amount of available training data and the variety exhibited by the training data with the help of data augmentation. In this paper we examine the influence of three data augmentation methods on the performance of two S2S model architectures. One of the data augmentation method comes from literature, while two other methods are our own development - a time perturbation in the frequency domain and sub-sequence sampling. Our experiments on Switchboard and Fisher data show state-of-the-art performance for S2S models that are trained solely on the speech training data and do not use additional text data.

eess.AS↗

Using multi-task learning to improve the performance of acoustic-to-word and conventional hybrid models

Acoustic-to-word (A2W) models that allow direct mapping from acoustic signals to word sequences are an appealing approach to end-to-end automatic speech recognition due to their simplicity. However, prior works have shown that modelling A2W typically encounters issues of data sparsity that prevent training such a model directly. So far, pre-training initialization is the only approach proposed to deal with this issue. In this work, we propose to build a shared neural network and optimize A2W and conventional hybrid models in a multi-task manner. Our results show that training an A2W model is much more stable with our multi-task model without pre-training initialization, and results in a significant improvement compared to a baseline model. Experiments also reveal that the performance of a hybrid acoustic model can be further improved when jointly training with a sequence-level optimization criterion such as acoustic-to-word.

eess.AS↗

Very Deep Self-Attention Networks for End-to-End Speech Recognition

Recently, end-to-end sequence-to-sequence models for speech recognition have gained significant interest in the research community. While previous architecture choices revolve around time-delay neural networks (TDNN) and long short-term memory (LSTM) recurrent neural networks, we propose to use self-attention via the Transformer architecture as an alternative. Our analysis shows that deep Transformer networks with high learning capacity are able to exceed performance from previous end-to-end approaches and even match the conventional hybrid systems. Moreover, we trained very deep models with up to 48 Transformer layers for both encoder and decoders combined with stochastic residual connections, which greatly improve generalizability and training efficiency. The resulting models outperform all previous end-to-end ASR approaches on the Switchboard benchmark. An ensemble of these models achieve 9.9% and 17.7% WER on Switchboard and CallHome test sets respectively. This finding brings our end-to-end models to competitive levels with previous hybrid systems. Further, with model ensembling the Transformers can outperform certain hybrid systems, which are more complicated in terms of both structure and training procedure.

cs.CL↗

Learning Shared Encoding Representation for End-to-End Speech Recognition Models

In this work, we learn a shared encoding representation for a multi-task neural network model optimized with connectionist temporal classification (CTC) and conventional framewise cross-entropy training criteria. Our experiments show that the multi-task training not only tackles the complexity of optimizing CTC models such as acoustic-to-word but also results in significant improvement compared to the plain-task training with an optimal setup. Furthermore, we propose to use the encoding representation learned by the multi-task network to initialize the encoder of attention-based models. Thereby, we train a deep attention-based end-to-end model with 10 long short-term memory (LSTM) layers of encoder which produces 12.2\% and 22.6\% word-error-rate on Switchboard and CallHome subsets of the Hub5 2000 evaluation.

eess.AS↗