SearcharxivSearch

arXiv subjects

Ngoc Duong

Publications and source records attributed to Ngoc Duong.

6 recordsLinked to original sources

Design and Optimization of Spin Dynamics in Ge Quantum Dots: g-Factor Modulation, Geometry-Induced Dephasing Sweet Spots, and Phonon-Induced Relaxation

Gate geometry and bias asymmetry can be used to engineer spin dynamics in gate-defined Ge hole quantum dots by reshaping the confinement potential and driving transitions between distinct confinement regimes. In this work, we show that these transitions strongly modify wavefunction localization, heavy-hole/light-hole mixing, and the effective vertical electric field, leading to pronounced g-factor modulation and geometry-induced dephasing sweet spots where the qubit becomes first-order insensitive to vertical electric-field fluctuations. We further find that phonon-induced spin relaxation exhibits a strong dependence on device size and bias, with T1 following a magnetic-field scaling close to B-9, consistent with Rashba-dominated heavy-hole spin dynamics. These results are obtained using a comprehensive three-dimensional simulation framework for strained Si0.2Ge0.8/Ge gate-defined hole spin qubits, combining realistic electrostatics with a four-band Luttinger-Kohn Hamiltonian. Unlike simplified symmetric confinement models, this approach captures asymmetric wavefunction redistribution, g-tensor anisotropy, and the coupled electrostatic and spin response of realistic devices. Our results establish gate pattern and bias design as practical tools for optimizing spin coherence in Ge hole-spin qubits.

cond-mat.mes-hall

Inplace knowledge distillation with teacher assistant for improved training of flexible deep neural networks

Deep neural networks (DNNs) have achieved great success in various machine learning tasks. However, most existing powerful DNN models are computationally expensive and memory demanding, hindering their deployment in devices with low memory and computational resources or in applications with strict latency requirements. Thus, several resource-adaptable or flexible approaches were recently proposed that train at the same time a big model and several resource-specific sub-models. Inplace knowledge distillation (IPKD) became a popular method to train those models and consists in distilling the knowledge from a larger model (teacher) to all other sub-models (students). In this work a novel generic training method called IPKD with teacher assistant (IPKD-TA) is introduced, where sub-models themselves become teacher assistants teaching smaller sub-models. We evaluated the proposed IPKD-TA training method using two state-of-the-art flexible models (MSDNet and Slimmable MobileNet-V1) with two popular image classification benchmarks (CIFAR-10 and CIFAR-100). Our results demonstrate that the IPKD-TA is on par with the existing state of the art while improving it in most cases.

eess.SP

Identify, locate and separate: Audio-visual object extraction in large video collections using weak supervision

We tackle the problem of audiovisual scene analysis for weakly-labeled data. To this end, we build upon our previous audiovisual representation learning framework to perform object classification in noisy acoustic environments and integrate audio source enhancement capability. This is made possible by a novel use of non-negative matrix factorization for the audio modality. Our approach is founded on the multiple instance learning paradigm. Its effectiveness is established through experiments over a challenging dataset of music instrument performance videos. We also show encouraging visual object localization results.

cs.CV

Audio style transfer

'Style transfer' among images has recently emerged as a very active research topic, fuelled by the power of convolution neural networks (CNNs), and has become fast a very popular technology in social media. This paper investigates the analogous problem in the audio domain: How to transfer the style of a reference audio signal to a target audio content? We propose a flexible framework for the task, which uses a sound texture model to extract statistics characterizing the reference audio style, followed by an optimization-based audio texture synthesis to modify the target content. In contrast to mainstream optimization-based visual transfer method, the proposed process is initialized by the target content instead of random noise and the optimized loss is only about texture, not structure. These differences proved key for audio style transfer in our experiments. In order to extract features of interest, we investigate different architectures, whether pre-trained on other tasks, as done in image style transfer, or engineered based on the human auditory system. Experimental results on different types of audio signal confirm the potential of the proposed approach.

cs.SD

MediaEval 2018: Predicting Media Memorability Task

In this paper, we present the Predicting Media Memorability task, which is proposed as part of the MediaEval 2018 Benchmarking Initiative for Multimedia Evaluation. Participants are expected to design systems that automatically predict memorability scores for videos, which reflect the probability of a video being remembered. In contrast to previous work in image memorability prediction, where memorability was measured a few minutes after memorization, the proposed dataset comes with short-term and long-term memorability annotations. All task characteristics are described, namely: the task's challenges and breakthrough, the released data set and ground truth, the required participant runs and the evaluation metrics.

cs.CV

Under-determined reverberant audio source separation using a full-rank spatial covariance model

This article addresses the modeling of reverberant recording environments in the context of under-determined convolutive blind source separation. We model the contribution of each source to all mixture channels in the time-frequency domain as a zero-mean Gaussian random variable whose covariance encodes the spatial characteristics of the source. We then consider four specific covariance models, including a full-rank unconstrained model. We derive a family of iterative expectationmaximization (EM) algorithms to estimate the parameters of each model and propose suitable procedures to initialize the parameters and to align the order of the estimated sources across all frequency bins based on their estimated directions of arrival (DOA). Experimental results over reverberant synthetic mixtures and live recordings of speech data show the effectiveness of the proposed approach.

stat.ML