Searcharxiv⌕ Search

arXiv subjects

Wei Han

Publications and source records attributed to Wei Han.

At least 109 records · Page 6Linked to original sources

Noise2Music: Text-conditioned Music Generation with Diffusion Models

We introduce Noise2Music, where a series of diffusion models is trained to generate high-quality 30-second music clips from text prompts. Two types of diffusion models, a generator model, which generates an intermediate representation conditioned on text, and a cascader model, which generates high-fidelity audio conditioned on the intermediate representation and possibly the text, are trained and utilized in succession to generate high-fidelity music. We explore two options for the intermediate representation, one using a spectrogram and the other using audio with lower fidelity. We find that the generated audio is not only able to faithfully reflect key elements of the text prompt such as genre, tempo, instruments, mood, and era, but goes beyond to ground fine-grained semantics of the prompt. Pretrained large language models play a key role in this story -- they are used to generate paired text for the audio of the training set and to extract embeddings of the text prompts ingested by the diffusion models. Generated examples: https://google-research.github.io/noise2music

cs.SD↗

Absence of localized $5d^1$ electrons in KTaO$_3$ interface superconductors

Recently, an exciting discovery of orientation-dependent superconductivity was made in two-dimensional electron gas (2DEG) at the interfaces of LaAlO$_3$/KTaO$_3$ (LAO/KTO) or EuO/KTaO$_3$ (EuO/KTO). The superconducting transition temperature can reach a $T_c$ of up to $\sim$ 2.2 K, which is significantly higher than its 3$d$ counterpart LaAlO$_3$/SrTiO$_3$ (LAO/STO) with a $T_c$ of $\sim$ 0.2 K. However, the underlying origin remains to be understood. To uncover the nature of electrons in KTO-based interfaces, we employ x-ray absorption spectroscopy (XAS) and resonant inelastic x-ray spectroscopy (RIXS) to study LAO/KTO and EuO/KTO with different orientations. We reveal the absence of $dd$ orbital excitations in all the measured samples. Our RIXS results are well reproduced by calculations that considered itinerant $5d$ electrons hybridized with O $2p$ electrons. This suggests that there is a lack of localized Ta $5d^1$ electrons in KTO interface superconductors, which is consistent with the absence of magnetic hysteresis observed in magneto-resistance (MR) measurements. These findings offer new insights into our understanding of superconductivity in Ta $5d$ interface superconductors and their potential applications.

cond-mat.supr-con↗

Rashba spin-orbit coupling enhanced magnetoresistance in junctions with one ferromagnet

We explain how Rashba spin-orbit coupling (SOC) in a two-dimensional electron gas (2DEG), or in a conventional $s$-wave superconductor, can lead to a large magnetoresistance even with one ferromagnet. However, such enhanced magnetoresistance is not generic and can be nonmonotonic and change its sign with Rashba SOC. For an in-plane rotation of magnetization, it is typically negligibly small for a 2DEG and depends on the perfect transmission which emerges from a spin-parity-time symmetry of the scattering states, while this symmetry is generally absent from the Hamiltonian of the system. The key difference from considering the normal-state magnetoresistance is the presence of the spin-dependent Andreev reflection at superconducting interfaces. In the fabricated junctions of quasi-2D van der Waals ferromagnets with conventional $s$-wave superconductors (Fe$_{0.29}$TaS$_2$/NbN) we find another example of enhanced magnetoresistance where the presence of Rashba SOC reduces the effective interfacial strength and is responsible for an equal-spin Andreev reflection. The observed nonmonotonic trend in the out-of-plane magnetoresistance with the interfacial barrier is an evidence for the proximity-induced equal-spin-triplet superconductivity.

cond-mat.mes-hall↗

Efficient Domain Adaptation for Speech Foundation Models

Foundation models (FMs), that are trained on broad data at scale and are adaptable to a wide range of downstream tasks, have brought large interest in the research community. Benefiting from the diverse data sources such as different modalities, languages and application domains, foundation models have demonstrated strong generalization and knowledge transfer capabilities. In this paper, we present a pioneering study towards building an efficient solution for FM-based speech recognition systems. We adopt the recently developed self-supervised BEST-RQ for pretraining, and propose the joint finetuning with both source and unsupervised target domain data using JUST Hydra. The FM encoder adapter and decoder are then finetuned to the target domain with a small amount of supervised in-domain data. On a large-scale YouTube and Voice Search task, our method is shown to be both data and model parameter efficient. It achieves the same quality with only 21.6M supervised in-domain data and 130.8M finetuned parameters, compared to the 731.1M model trained from scratch on additional 300M supervised in-domain data.

cs.CL↗

Trapped mode control in metasurfaces composed of particles with the form birefringence property

Progress in developing advanced photonic devices relies on introducing new materials, discovered physical principles, and optimal designs when constructing their components. Optical systems operating on the principles of excitation of extremely high-quality factor trapped modes (also known as the bound states in the continuum, BICs) are of great interest since they allow the implementation of laser and sensor devices with outstanding characteristics. In this paper, we discuss how one can utilize the anisotropic properties of novel materials (transition metal dichalcogenides, TMDs), particularly, the bulk molybdenum disulfide (MoS2), to realize the excitation of trapped modes in dielectric metasurfaces. The bulk MoS2 is a thin-film structure in which the light wave behaves the same way as that in the uniaxial anisotropic material with the form birefringence property. Our metasurface is composed of an array of disk-shaped nanoparticles (resonators) made of the MoS2 material under the assumption that the anisotropy axis of MoS2 can be tilted to the rotation axis of the disks. We perform a detailed analysis of eigenwaves and scattering properties of such anisotropic resonators as well as the spectral features of the metasurface revealing dependence of the excitation conditions of the trapped mode on the anisotropy axis orientation of the MoS2 material used.

physics.optics↗

Electromagnetic-Compliant Channel Modeling and Performance Evaluation for Holographic MIMO

Recently, the concept of holographic multiple-input multiple-output (MIMO) is emerging as one of the promising technologies beyond massive MIMO. Many challenges need to be addressed to bring this novel idea into practice, including electromagnetic (EM)-compliant channel modeling and accurate performance evaluation. In this paper, an EM-compliant channel model is proposed for the holographic MIMO systems, which is able to model both the characteristics of the propagation channel and the non-ideal factors caused by mutual coupling at the transceivers, including the antenna pattern distortion and the decrease of antenna efficiency. Based on the proposed channel model, a more realistic performance evaluation is conducted to show the performance of the holographic MIMO system in both the single-user and the multi-user scenarios. Key challenges and future research directions are further provided based on the theoretical analyses and numerical results.

cs.IT↗

Speech Aware Dialog System Technology Challenge (DSTC11)

Most research on task oriented dialog modeling is based on written text input. However, users interact with practical dialog systems often using speech as input. Typically, systems convert speech into text using an Automatic Speech Recognition (ASR) system, introducing errors. Furthermore, these systems do not address the differences in written and spoken language. The research on this topic is stymied by the lack of a public corpus. Motivated by these considerations, our goal in hosting the speech-aware dialog state tracking challenge was to create a public corpus or task which can be used to investigate the performance gap between the written and spoken forms of input, develop models that could alleviate this gap, and establish whether Text-to-Speech-based (TTS) systems is a reasonable surrogate to the more-labor intensive human data collection. We created three spoken versions of the popular written-domain MultiWoz task -- (a) TTS-Verbatim: written user inputs were converted into speech waveforms using a TTS system, (b) Human-Verbatim: humans spoke the user inputs verbatim, and (c) Human-paraphrased: humans paraphrased the user inputs. Additionally, we provided different forms of ASR output to encourage wider participation from teams that may not have access to state-of-the-art ASR systems. These included ASR transcripts, word time stamps, and latent representations of the audio (audio encoder outputs). In this paper, we describe the corpus, report results from participating teams, provide preliminary analyses of their results, and summarize the current state-of-the-art in this domain.

cs.AI↗

Universal Paralinguistic Speech Representations Using Self-Supervised Conformers

Many speech applications require understanding aspects beyond the words being spoken, such as recognizing emotion, detecting whether the speaker is wearing a mask, or distinguishing real from synthetic speech. In this work, we introduce a new state-of-the-art paralinguistic representation derived from large-scale, fully self-supervised training of a 600M+ parameter Conformer-based architecture. We benchmark on a diverse set of speech tasks and demonstrate that simple linear classifiers trained on top of our time-averaged representation outperform nearly all previous results, in some cases by large margins. Our analyses of context-window size demonstrate that, surprisingly, 2 second context-windows achieve 96\% the performance of the Conformers that use the full long-term context on 7 out of 9 tasks. Furthermore, while the best per-task representations are extracted internally in the network, stable performance across several layers allows a single universal representation to reach near optimal performance on all tasks.

cs.SD↗

Superconductor/Ferromagnet Heterostructures: A Platform for Superconducting Spintronics and Quantum Computation

The interplay between superconductivity and ferromagnetism in the superconductor/ferromagnet (SC/FM) heterostructures generates many interesting physical phenomena, including spin-triplet superconductivity, superconducting order parameter oscillation, and topological superconductivity. The unique physical properties make the SC/FM heterostructures as promising platforms for future superconducting spintronics and quantum computation applications. In this article, we review important research progress of SC/FM heterostructures from superconducting spintronics to quantum computation, and it is organized as follows. Firstly, we discuss the progress of spin current carriers in SC/FM heterostructures including Bogoliubov quasiparticles, superconducting vortex, and spin-triplet Cooper pairs which might be used for long-range spin transport. Then, we will describe the π Josephson junctions and its application for constructing π qubits. Finally, we will briefly review experimental signatures of Majorana states in the SC/FM heterostructures and the theoretically proposed manipulation, which could be useful to realize fault-tolerant topological quantum computing.

cond-mat.mes-hall↗

An Interpretable Neuron Embedding for Static Knowledge Distillation

Although deep neural networks have shown well-performance in various tasks, the poor interpretability of the models is always criticized. In the paper, we propose a new interpretable neural network method, by embedding neurons into the semantic space to extract their intrinsic global semantics. In contrast to previous methods that probe latent knowledge inside the model, the proposed semantic vector externalizes the latent knowledge to static knowledge, which is easy to exploit. Specifically, we assume that neurons with similar activation are of similar semantic information. Afterwards, semantic vectors are optimized by continuously aligning activation similarity and semantic vector similarity during the training of the neural network. The visualization of semantic vectors allows for a qualitative explanation of the neural network. Moreover, we assess the static knowledge quantitatively by knowledge distillation tasks. Empirical experiments of visualization show that semantic vectors describe neuron activation semantics well. Without the sample-by-sample guidance from the teacher model, static knowledge distillation exhibit comparable or even superior performance with existing relation-based knowledge distillation methods.

cs.LG↗

TOSE: A Fast Capacity Determination Algorithm Based on Random Matrix Theory

Wireless network capacity is one of the most important performance metrics for wireless communication networks. Future wireless networks will be composed of extremely large number of base stations (BSs) and users, and organized in the form of multiple clusters. Unfortunately, the determination of average cluster capacity for such future wireless networks is difficult, and lacks of both analytical expressions and fast algorithms. In this paper, we propose a fast algorithm TOSE to estimate the average cluster capacity based on the random matrix theory (RMT). It can avoid the exact eigenvalue derivations of large dimensional matrices, which are complicated and inevitable in conventional capacity determination methods. Instead, fast eigenvalue estimations can be realized based on RMT in our TOSE algorithm. In addition, we derive the analytical upper and lower bounds of the average cluster capacity. Our numerical experiments show that TOSE is faster than the conventional Cholesky decomposition method, by at least three orders of magnitude. Besides, TOSE has superior generality, since it is independent of the distributions of BSs and users, and the shape of network areas.

cs.IT↗

Accelerating RNN-T Training and Inference Using CTC guidance

We propose a novel method to accelerate training and inference process of recurrent neural network transducer (RNN-T) based on the guidance from a co-trained connectionist temporal classification (CTC) model. We made a key assumption that if an encoder embedding frame is classified as a blank frame by the CTC model, it is likely that this frame will be aligned to blank for all the partial alignments or hypotheses in RNN-T and it can be discarded from the decoder input. We also show that this frame reduction operation can be applied in the middle of the encoder, which result in significant speed up for the training and inference in RNN-T. We further show that the CTC alignment, a by-product of the CTC decoder, can also be used to perform lattice reduction for RNN-T during training. Our method is evaluated on the Librispeech and SpeechStew tasks. We demonstrate that the proposed method is able to accelerate the RNN-T inference by 2.2 times with similar or slightly better word error rates (WER).

eess.AS↗

SAT: Improving Semi-Supervised Text Classification with Simple Instance-Adaptive Self-Training

Self-training methods have been explored in recent years and have exhibited great performance in improving semi-supervised learning. This work presents a Simple instance-Adaptive self-Training method (SAT) for semi-supervised text classification. SAT first generates two augmented views for each unlabeled data and then trains a meta-learner to automatically identify the relative strength of augmentations based on the similarity between the original view and the augmented views. The weakly-augmented view is fed to the model to produce a pseudo-label and the strongly-augmented view is used to train the model to predict the same pseudo-label. We conducted extensive experiments and analyses on three text classification datasets and found that with varying sizes of labeled training data, SAT consistently shows competitive performance compared to existing semi-supervised learning methods. Our code can be found at \url{https://github.com/declare-lab/SAT.git}.

cs.CL↗

MM-Align: Learning Optimal Transport-based Alignment Dynamics for Fast and Accurate Inference on Missing Modality Sequences

Existing multimodal tasks mostly target at the complete input modality setting, i.e., each modality is either complete or completely missing in both training and test sets. However, the randomly missing situations have still been underexplored. In this paper, we present a novel approach named MM-Align to address the missing-modality inference problem. Concretely, we propose 1) an alignment dynamics learning module based on the theory of optimal transport (OT) for indirect missing data imputation; 2) a denoising training algorithm to simultaneously enhance the imputation results and backbone network performance. Compared with previous methods which devote to reconstructing the missing inputs, MM-Align learns to capture and imitate the alignment dynamics between modality sequences. Results of comprehensive experiments on three datasets covering two multimodal tasks empirically demonstrate that our method can perform more accurate and faster inference and relieve overfitting under various missing conditions.

cs.CL↗

SANCL: Multimodal Review Helpfulness Prediction with Selective Attention and Natural Contrastive Learning

With the boom of e-commerce, Multimodal Review Helpfulness Prediction (MRHP), which aims to sort product reviews according to the predicted helpfulness scores has become a research hotspot. Previous work on this task focuses on attention-based modality fusion, information integration, and relation modeling, which primarily exposes the following drawbacks: 1) the model may fail to capture the really essential information due to its indiscriminate attention formulation; 2) lack appropriate modeling methods that take full advantage of correlation among provided data. In this paper, we propose SANCL: Selective Attention and Natural Contrastive Learning for MRHP. SANCL adopts a probe-based strategy to enforce high attention weights on the regions of greater significance. It also constructs a contrastive learning framework based on natural matching properties in the dataset. Experimental results on two benchmark datasets with three categories show that SANCL achieves state-of-the-art baseline performance with lower memory consumption.

cs.CL↗

DoubleMix: Simple Interpolation-Based Data Augmentation for Text Classification

This paper proposes a simple yet effective interpolation-based data augmentation approach termed DoubleMix, to improve the robustness of models in text classification. DoubleMix first leverages a couple of simple augmentation operations to generate several perturbed samples for each training data, and then uses the perturbed data and original data to carry out a two-step interpolation in the hidden space of neural models. Concretely, it first mixes up the perturbed data to a synthetic sample and then mixes up the original data and the synthetic perturbed data. DoubleMix enhances models' robustness by learning the "shifted" features in hidden space. On six text classification benchmark datasets, our approach outperforms several popular text augmentation methods including token-level, sentence-level, and hidden-level data augmentation techniques. Also, experiments in low-resource settings show our approach consistently improves models' performance when the training data is scarce. Extensive ablation studies and case studies confirm that each component of our approach contributes to the final performance and show that our approach exhibits superior performance on challenging counterexamples. Additionally, visual analysis shows that text features generated by our approach are highly interpretable. Our code for this paper can be found at https://github.com/declare-lab/DoubleMix.git.

cs.CL↗

An Efficient Two-Stage SPARC Decoder for Massive MIMO Unsourced Random Access

In this paper, we study a concatenate coding scheme based on sparse regression code (SPARC) and tree code for unsourced random access in massive multiple-input and multiple-output systems. Our focus is concentrated on efficient decoding for the inner SPARC with practical concerns. A two-stage method is proposed to achieve near-optimal performance while maintaining low computational complexity. Specifically, a one-step thresholding-based algorithm is first used for reducing large dimensions of the SPARC decoding, after which a relaxed maximum-likelihood estimator is employed for refinement. Adequate simulation results are provided to validate the near-optimal performance and the low computational complexity. Besides, for covariance-based sparse recovery method, theoretical analyses are given to characterize the upper bound of the number of active users supported when convex relaxation is considered, and the probability of successful dimension reduction by the one-step thresholding-based algorithm.

cs.IT↗

Semantic Compression with Side Information: A Rate-Distortion Perspective

We consider the semantic rate-distortion problem motivated by task-oriented video compression. The semantic information corresponding to the task, which is not observable to the encoder, shows impacts on the observations through a joint probability distribution. The similarities among intra-frame segments and inter-frames in video compression are formulated as side information available at both the encoder and the decoder. The decoder is interested in recovering the observation and making an inference of the semantic information under certain distortion constraints. We establish the information-theoretic limits for the tradeoff between compression rates and distortions by fully characterizing the rate-distortion function. We further evaluate the rate-distortion function under specific Markov conditions for three scenarios: i) both the task and the observation are binary sources; ii) the task is a binary classification of an integer observation as even and odd; iii) Gaussian correlated task and observation. We also illustrate through numerical results that recovering only the semantic information can reduce the coding rate comparing to recovering the source observation.

cs.IT↗