SearcharxivSearch

arXiv subjects

Mu Yang

Publications and source records attributed to Mu Yang.

At least 37 records · Page 2Linked to original sources

Filtering one-way Einstein-Podolsky-Rosen steering

Einstein-Podolsky-Rosen (EPR) steering, a fundamental concept of quantum nonlocality, describes one observer's capability to remotely affect another distant observer's state by local measurements. Unlike quantum entanglement and Bell nonlocality, both associated with the symmetric quantum correlation, EPR steering depicts the unique asymmetric property of quantum nonlocality. With the local filter operation in which some system components are discarded, quantum nonlocality can be distilled to enhance the nonlocal correlation, and even the hidden nonlocality can be activated. However, asymmetric quantum nonlocality in the filter operation still lacks a well-rounded investigation, especially considering the discarded parts where quantum nonlocal correlations may still exist with probabilities. Here, in both theory and experiment, we investigate the effect of reusing the discarded particles from local filter. We observe all configurations of EPR steering simultaneously and other intriguing evolution of asymmetric quantum nonlocality, such as reversing the direction of one-way EPR steering. This work provides a perspective to answer "What is the essential role of utilizing quantum steering as a resource?", and demonstrates a practical toolbox for manipulating asymmetric quantum systems with significant potential applications in quantum information tasks.

quant-ph

Realization of edge states along a synthetic orbital angular momentum dimension

The synthetic dimension is a rising method to study topological physics, which enables us to implement high-dimensional physics in low-dimensional geometries. Photonic orbital angular momentum (OAM), a degree of freedom characterized by discrete yet unbounded, serves as a suitable synthetic dimension. However, a sharp boundary along a synthetic OAM dimension has not been demonstrated, dramatically limiting the investigation of topological edge effects in an open boundary lattice system. In this work, we make a sharp boundary along a Floquet Su-Schrieffer-Heeger OAM lattice and form approximate semi-infinite lattices by drilling a pinhole on the optical elements in a cavity. The band structures with zero ($\pmπ$) energy boundary states are measured directly, benefiting from the spectra detection of the cavity. Moreover, we obtain the edge modes moving from the gap to the bulk by dynamically changing the boundary phase, and we reveal that interference near the surface leads to spectrum discretization. Our work provides a new perspective to observe edge effects and explore practical photonics tools.

physics.optics

Learning ASR pathways: A sparse multilingual ASR model

Neural network pruning compresses automatic speech recognition (ASR) models effectively. However, in multilingual ASR, language-agnostic pruning may lead to severe performance drops on some languages because language-agnostic pruning masks may not fit all languages and discard important language-specific parameters. In this work, we present ASR pathways, a sparse multilingual ASR model that activates language-specific sub-networks ("pathways"), such that the parameters for each language are learned explicitly. With the overlapping sub-networks, the shared parameters can also enable knowledge transfer for lower-resource languages via joint multilingual training. We propose a novel algorithm to learn ASR pathways, and evaluate the proposed method on 4 languages with a streaming RNN-T model. Our proposed ASR pathways outperform both dense models and a language-agnostically pruned model, and provide better performance on low-resource languages compared to the monolingual sparse models.

eess.AS

What Can an Accent Identifier Learn? Probing Phonetic and Prosodic Information in a Wav2vec2-based Accent Identification Model

This study is focused on understanding and quantifying the change in phoneme and prosody information encoded in the Self-Supervised Learning (SSL) model, brought by an accent identification (AID) fine-tuning task. This problem is addressed based on model probing. Specifically, we conduct a systematic layer-wise analysis of the representations of the Transformer layers on a phoneme correlation task, and a novel word-level prosody prediction task. We compare the probing performance of the pre-trained and fine-tuned SSL models. Results show that the AID fine-tuning task steers the top 2 layers to learn richer phoneme and prosody representation. These changes share some similarities with the effects of fine-tuning with an Automatic Speech Recognition task. In addition, we observe strong accent-specific phoneme representations in layer 9. To sum up, this study provides insights into the understanding of SSL features and their interactions with fine-tuning tasks.

eess.AS

Reconstructing the multiphoton spatial wave function with coincidence wavefront sensing

The quantum wave function of multiple particles provides additional information which is inaccessible to detectors working alone. Here, we introduce the coincidence wavefront sensing (CWS) method to reconstruct the phase of the multiphoton transverse spatial wave function. The spatially resolved coincidence photon counting is involved. Numerical simulations of two-photon cases using the weak measurement wavefront sensor are performed to test its correctness, and the phase information hidden in the correlation are revealed. Our work provides a direct spatial way to characterize multipartite quantum systems, and leads to fundamental studies like experimental Bohmian mechanics and applications in quantum optical technologies.

quant-ph

Simulating topological materials with photonic synthetic dimensions in cavities

Photons play essential roles in fundamental physics and practical technologies. They have become one of the attractive informaiton carriers for quantum computation and quantum simulation. Recently, various photonic degrees of freedom supported by optical resonant cavities form photonic synthetic dimensions, which contribute to all-optical platforms for simulating novel topological materials. The photonic discrete or continuous degrees of freedom are mapped to the lattices or momenta of the simulated topological matter, and the couplings between optical modes are equivalent to the interactions among quasi-particles. Mature optical modulations enable flexible engineering of the simulated Hamiltonian. Meanwhile, the resonant detection methods provide direct approaches to obtaining the corresponding energy band structures, particle distributions and dynamical evolutions. In this Review, we give an overview of the synthetic dimensions in optical cavities, including frequency, orbital angular momentum, time-multiplexed lattice, and independent parameters. Abundant higher-dimensional topological models have been demonstrated in lower dimensional synthetic systems. We further discuss the potential development of photonic synthetic dimensions in the future.

physics.optics

Realization of exceptional points along a synthetic orbital angular momentum dimension

Exceptional points (EPs), at which more than one eigenvalue and eigenvector coalesce, are unique spectral features of Non-Hermiticity (NH) systems. They exist widely in open systems with complex energy spectra. We experimentally demonstrate the appearance of paired EPs in a periodical driven degenerate optical cavity along the synthetic orbital angular momentum (OAM) dimension with a tunable parameter. The complex-energy band structures and the key features of EPs, i.e. their Fermi arcs, parity-time symmetry breaking transition, energy swapping, and half-integer band windings are directly observed by detecting the cavity's transmission spectrum. Our results advance the fundamental understanding of NH physics and demonstrate the flexibility of using the photonic synthetic dimensions to implement NH systems.

physics.optics

Improving Mispronunciation Detection with Wav2vec2-based Momentum Pseudo-Labeling for Accentedness and Intelligibility Assessment

Current leading mispronunciation detection and diagnosis (MDD) systems achieve promising performance via end-to-end phoneme recognition. One challenge of such end-to-end solutions is the scarcity of human-annotated phonemes on natural L2 speech. In this work, we leverage unlabeled L2 speech via a pseudo-labeling (PL) procedure and extend the fine-tuning approach based on pre-trained self-supervised learning (SSL) models. Specifically, we use Wav2vec 2.0 as our SSL model, and fine-tune it using original labeled L2 speech samples plus the created pseudo-labeled L2 speech samples. Our pseudo labels are dynamic and are produced by an ensemble of the online model on-the-fly, which ensures that our model is robust to pseudo label noise. We show that fine-tuning with pseudo labels achieves a 5.35% phoneme error rate reduction and 2.48% MDD F1 score improvement over a labeled-samples-only fine-tuning baseline. The proposed PL method is also shown to outperform conventional offline PL methods. Compared to the state-of-the-art MDD systems, our MDD solution produces a more accurate and consistent phonetic error diagnosis. In addition, we conduct an open test on a separate UTD-4Accents dataset, where our system recognition outputs show a strong correlation with human perception, based on accentedness and intelligibility.

eess.AS

Toward practical weak measurement wavefront sensing: spatial resolution and achromatism

The weak measurement wavefront sensor detects the phase gradient of light like the Shack-Hartmann sensor does. However, the use of one thin birefringent crystal to displace light beams results in a wavelength-dependent phase difference between the two polarization components, which limits the practical application. Using a Savart plate which consists of two such crystals can compensate for the phase difference and realize achromatic wavefront sensing when combined with an achromatic retarder. We discuss the spatial resolution of the sensor and experimentally reconstruct a wavefront modulated by a pattern. Then we obtain the Zernike coefficients with three different wavelengths before and after modulation. Our work makes this new wavefront sensor more applicable to actual tasks like biomedical imaging.

physics.optics

Towards Lifelong Learning of Multilingual Text-To-Speech Synthesis

This work presents a lifelong learning approach to train a multilingual Text-To-Speech (TTS) system, where each language was seen as an individual task and was learned sequentially and continually. It does not require pooled data from all languages altogether, and thus alleviates the storage and computation burden. One of the challenges of lifelong learning methods is "catastrophic forgetting": in TTS scenario it means that model performance quickly degrades on previous languages when adapted to a new language. We approach this problem via a data-replay-based lifelong learning method. We formulate the replay process as a supervised learning problem, and propose a simple yet effective dual-sampler framework to tackle the heavily language-imbalanced training samples. Through objective and subjective evaluations, we show that this supervised learning formulation outperforms other gradient-based and regularization-based lifelong learning methods, achieving 43% Mel-Cepstral Distortion reduction compared to a fine-tuning baseline.

eess.AS

Detecting momentum weak value: Shack-Hartmann versus a weak measurement wavefront sensor

The task of wavefront sensing is to measure the phase of the optical field. Here, we demonstrate that the widely used Shack-Hartmann wavefront sensor detects the weak value of transverse momentum, usually achieved by the method of quantum weak measurement. We extend its input states to partially coherent states and compare it with the weak measurement wavefront sensor, which has a higher spatial resolution but a smaller dynamic range. Since weak values are commonly used in investigating fundamental quantum physics and quantum metrology, our work would find essential applications in these fields.

physics.optics

Demonstrating shareability of multipartite Einstein-Podolsky-Rosen steering

Einstein-Podolsky-Rosen (EPR) steering, a category of quantum nonlocal correlations describing the ability of one observer to influence another party's state via local measurements, is different from both entanglement and Bell nonlocality by possessing an asymmetric property. For multipartite EPR steering, the monogamous situation, where two observers cannot simultaneously steer the state of the third party, has been investigated rigorously both in theory and experiment. In contrast to the monogamous situation, the shareability of EPR steering in reduced subsystems allows the state of one party to be steered by two or more observers and thus reveals more configurations of multipartite EPR steering. However, the experimental implementation of such a kind of shareability has still been absent until now. Here, in an optical experiment, we provide a proof-of-principle demonstration of the shareability of EPR steering without the constraint of monogamy in a three-qubit system. Moreover, based on the reduced bipartite EPR steering detection results, we verify the genuine three-qubit entanglement results. This work provides a complementary viewpoint for understanding multipartite EPR steering and has potential applications in many quantum information protocols, such as multipartite entanglement detection, quantum cryptography, and the construction of quantum networks.

quant-ph

Topological band structure via twisted photons in a degenerate cavity

Synthetic dimensions based on particles' internal degrees of freedom, such as frequency, spatial modes and arrival time, have attracted significant attention. They offer ideal large-scale lattices to simulate nontrivial topological phenomena. Exploring more synthetic dimensions is one of the paths toward higher dimensional physics. In this work, we design and experimentally control the coupling among synthetic dimensions consisting of the intrinsic photonic orbital angular momentum and spin angular momentum degrees of freedom in a degenerate optical resonant cavity, which generates a periodically driven spin-orbital coupling system. We directly characterize the system's properties, including the density of states, energy band structures and topological windings, through the transmission intensity measurements. Our work demonstrates a novel mechanism for exploring the spatial modes of twisted photons as the synthetic dimension, which paves the way to design rich topological physics in a highly compact platform.

physics.optics

InverseMV: Composing Piano Scores with a Convolutional Video-Music Transformer

Many social media users prefer consuming content in the form of videos rather than text. However, in order for content creators to produce videos with a high click-through rate, much editing is needed to match the footage to the music. This posts additional challenges for more amateur video makers. Therefore, we propose a novel attention-based model VMT (Video-Music Transformer) that automatically generates piano scores from video frames. Using music generated from models also prevent potential copyright infringements that often come with using existing music. To the best of our knowledge, there is no work besides the proposed VMT that aims to compose music for video. Additionally, there lacks a dataset with aligned video and symbolic music. We release a new dataset composed of over 7 hours of piano scores with fine alignment between pop music videos and MIDI files. We conduct experiments with human evaluation on VMT, SeqSeq model (our baseline), and the original piano version soundtrack. VMT achieves consistent improvements over the baseline on music smoothness and video relevance. In particular, with the relevance scores and our case study, our model has shown the capability of multimodality on frame-level actors' movement for music generation. Our VMT model, along with the new dataset, presents a promising research direction toward composing the matching soundtrack for videos. We have released our code at https://github.com/linchintung/VMT

cs.LG

Topological contextuality and anyonic statistics of photonic-encoded parafermions

Quasiparticle poisoning, expected to arise during the measurement of Majorana zero mode state, poses a fundamental problem towards the realization of Majorana-based quantum computation. Parafermions, a natural generalization of Majorana fermions, can encode topological qudits immune to quasiparticle poisoning. While parafermions are expected to emerge in superconducting fractional quantum Hall systems, they are not yet attainable with current technology. To bypass this problem, we employ a photonic quantum simulator to experimentally demonstrate the key components of parafermion-based universal quantum computation. Our contributions in this article are twofold. First, by manipulating the photonic states, we realize Clifford operator Berry phases that correspond to braiding statistics of parafermions. Second, we investigate the quantum contextuality in a topological system for the first time by demonstrating the contextuality of parafermion encoded qudit states. Importantly, we find that the topologically-encoded contextuality opens the way to magic state distillation, while both the contextuality and the braiding-induced Clifford gates are resilient against local noise. By introducing contextuality, our photonic quantum simulation provides the first step towards a physically robust methodology for realizing topological quantum computation.

quant-ph

EventPlus: A Temporal Event Understanding Pipeline

We present EventPlus, a temporal event understanding pipeline that integrates various state-of-the-art event understanding components including event trigger and type detection, event argument detection, event duration and temporal relation extraction. Event information, especially event temporal knowledge, is a type of common sense knowledge that helps people understand how stories evolve and provides predictive hints for future events. EventPlus as the first comprehensive temporal event understanding pipeline provides a convenient tool for users to quickly obtain annotations about events and their temporal information for any user-provided document. Furthermore, we show EventPlus can be easily adapted to other domains (e.g., biomedical domain). We make EventPlus publicly available to facilitate event-related information extraction and downstream applications.

cs.CL

Photonic implementation of quantum information masking

Masking of quantum information spreads it over nonlocal correlations and hides it from the subsystems. It is known that no operation can simultaneously mask all pure states [Phys. Rev. Lett. 120, 230501 (2018)], so in what sense is quantum information masking useful? Here, we extend the definition of quantum information masking to general mixed states, and show that the resource of maskable quantum states are far more abundant than the no-go theorem seemingly suggests. Geometrically, the simultaneously maskable states lays on hyperdisks in the state hypersphere, and strictly contain the broadcastable states. We devise a photonic quantum information masking machine using time-correlated photons to experimentally investigate the properties of qubit masking, and demonstrate the transfer of quantum information into bipartite correlations and its faithful retrieval. The versatile masking machine has decent extensibility, and may be applicable to quantum secret sharing and fault-tolerant quantum communication. Our results provide some insights on the comprehension and potential application of quantum information masking.

quant-ph

Biomedical Event Extraction with Hierarchical Knowledge Graphs

Biomedical event extraction is critical in understanding biomolecular interactions described in scientific corpus. One of the main challenges is to identify nested structured events that are associated with non-indicative trigger words. We propose to incorporate domain knowledge from Unified Medical Language System (UMLS) to a pre-trained language model via Graph Edge-conditioned Attention Networks (GEANet) and hierarchical graph representation. To better recognize the trigger words, each sentence is first grounded to a sentence graph based on a jointly modeled hierarchical knowledge graph from UMLS. The grounded graphs are then propagated by GEANet, a novel graph neural networks for enhanced capabilities in inferring complex events. On BioNLP 2011 GENIA Event Extraction task, our approach achieved 1.41% F1 and 3.19% F1 improvements on all events and complex events, respectively. Ablation studies confirm the importance of GEANet and hierarchical KG.

cs.CL