SearcharxivSearch

arXiv subjects

Shilong Wu

Publications and source records attributed to Shilong Wu.

13 recordsLinked to original sources

M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset

In the field of speaker diarization, the development of technology is constrained by two problems: insufficient data resources and poor generalization ability of deep learning models. To address these two problems, firstly, we propose an automated method for constructing speaker diarization datasets, which generates more accurate pseudo-labels for massive data through the combination of audio and video. Relying on this method, we have released Multi-modal, Multi-scenario and Multi-language Speaker Diarization (M3SD) datasets. This dataset is derived from real network videos and is highly diverse. Our dataset and code have been open-sourced at https://huggingface.co/spaces/OldDragon/m3sd.

eess.AS

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Meetings are a valuable yet challenging scenario for speech applications due to complex acoustic conditions. This paper summarizes the outcomes of the MISP 2025 Challenge, hosted at Interspeech 2025, which focuses on multi-modal, multi-device meeting transcription by incorporating video modality alongside audio. The tasks include Audio-Visual Speaker Diarization (AVSD), Audio-Visual Speech Recognition (AVSR), and Audio-Visual Diarization and Recognition (AVDR). We present the challenge's objectives, tasks, dataset, baseline systems, and solutions proposed by participants. The best-performing systems achieved significant improvements over the baseline: the top AVSD model achieved a Diarization Error Rate (DER) of 8.09%, improving by 7.43%; the top AVSR system achieved a Character Error Rate (CER) of 9.48%, improving by 10.62%; and the best AVDR system achieved a concatenated minimum-permutation Character Error Rate (cpCER) of 11.56%, improving by 72.49%.

cs.SD

Neural Speaker Diarization Using Memory-Aware Multi-Speaker Embedding with Sequence-to-Sequence Architecture

We propose a novel neural speaker diarization system using memory-aware multi-speaker embedding with sequence-to-sequence architecture (NSD-MS2S), which integrates the strengths of memory-aware multi-speaker embedding (MA-MSE) and sequence-to-sequence (Seq2Seq) architecture, leading to improvement in both efficiency and performance. Next, we further decrease the memory occupation of decoding by incorporating input features fusion and then employ a multi-head attention mechanism to capture features at different levels. NSD-MS2S achieved a macro diarization error rate (DER) of 15.9% on the CHiME-7 EVAL set, which signifies a relative improvement of 49% over the official baseline system, and is the key technique for us to achieve the best performance for the main track of CHiME-7 DASR Challenge. Additionally, we introduce a deep interactive module (DIM) in MA-MSE module to better retrieve a cleaner and more discriminative multi-speaker embedding, enabling the current model to outperform the system we used in the CHiME-7 DASR Challenge. Our code will be available at https://github.com/liyunlongaaa/NSD-MS2S.

eess.AS

The USTC-NERCSLIP Systems for the CHiME-7 DASR Challenge

This technical report details our submission system to the CHiME-7 DASR Challenge, which focuses on speaker diarization and speech recognition under complex multi-speaker scenarios. Additionally, it also evaluates the efficiency of systems in handling diverse array devices. To address these issues, we implemented an end-to-end speaker diarization system and introduced a rectification strategy based on multi-channel spatial information. This approach significantly diminished the word error rates (WER). In terms of recognition, we utilized publicly available pre-trained models as the foundational models to train our end-to-end speech recognition models. Our system attained a Macro-averaged diarization-attributed WER (DA-WER) of 21.01% on the CHiME-7 evaluation set, which signifies a relative improvement of 62.04% over the official baseline system.

eess.AS

The Multimodal Information Based Speech Processing (MISP) 2023 Challenge: Audio-Visual Target Speaker Extraction

Previous Multimodal Information based Speech Processing (MISP) challenges mainly focused on audio-visual speech recognition (AVSR) with commendable success. However, the most advanced back-end recognition systems often hit performance limits due to the complex acoustic environments. This has prompted a shift in focus towards the Audio-Visual Target Speaker Extraction (AVTSE) task for the MISP 2023 challenge in ICASSP 2024 Signal Processing Grand Challenges. Unlike existing audio-visual speech enhance-ment challenges primarily focused on simulation data, the MISP 2023 challenge uniquely explores how front-end speech processing, combined with visual clues, impacts back-end tasks in real-world scenarios. This pioneering effort aims to set the first benchmark for the AVTSE task, offering fresh insights into enhancing the ac-curacy of back-end speech recognition systems through AVTSE in challenging and real acoustic environments. This paper delivers a thorough overview of the task setting, dataset, and baseline system of the MISP 2023 challenge. It also includes an in-depth analysis of the challenges participants may encounter. The experimental results highlight the demanding nature of this task, and we look forward to the innovative solutions participants will bring forward.

eess.AS

Semi-supervised multi-channel speaker diarization with cross-channel attention

Most neural speaker diarization systems rely on sufficient manual training data labels, which are hard to collect under real-world scenarios. This paper proposes a semi-supervised speaker diarization system to utilize large-scale multi-channel training data by generating pseudo-labels for unlabeled data. Furthermore, we introduce cross-channel attention into the Neural Speaker Diarization Using Memory-Aware Multi-Speaker Embedding (NSD-MA-MSE) to learn channel contextual information of speaker embeddings better. Experimental results on the CHiME-7 Mixer6 dataset which only contains partial speakers' labels of the training set, show that our system achieved 57.01% relative DER reduction compared to the clustering-based model on the development set. We further conducted experiments on the CHiME-6 dataset to simulate the scenario of missing partial training set labels. When using 80% and 50% labeled training data, our system performs comparably to the results obtained using 100% labeled data for training.

eess.AS

The Multimodal Information based Speech Processing (MISP) 2022 Challenge: Audio-Visual Diarization and Recognition

The Multi-modal Information based Speech Processing (MISP) challenge aims to extend the application of signal processing technology in specific scenarios by promoting the research into wake-up words, speaker diarization, speech recognition, and other technologies. The MISP2022 challenge has two tracks: 1) audio-visual speaker diarization (AVSD), aiming to solve ``who spoken when'' using both audio and visual data; 2) a novel audio-visual diarization and recognition (AVDR) task that focuses on addressing ``who spoken what when'' with audio-visual speaker diarization results. Both tracks focus on the Chinese language, and use far-field audio and video in real home-tv scenarios: 2-6 people communicating each other with TV noise in the background. This paper introduces the dataset, track settings, and baselines of the MISP2022 challenge. Our analyses of experiments and examples indicate the good performance of AVDR baseline system, and the potential difficulties in this challenge due to, e.g., the far-field video quality, the presence of TV noise in the background, and the indistinguishable speakers.

cs.MM

Electronic Nature of Charge Density Wave and Electron-Phonon Coupling in Kagome Superconductor KV$_3$Sb$_5$

The Kagome superconductors AV3Sb5 (A=K, Rb, Cs) have received enormous attention due to their nontrivial topological electronic structure, anomalous physical properties and superconductivity. Unconventional charge density wave (CDW) has been detected in AV3Sb5. High-precision electronic structure determination is essential to understand its origin. Here we unveil electronic nature of the CDW phase in our high-resolution angle-resolved photoemission measurements on KV3Sb5. We have observed CDW-induced Fermi surface reconstruction and the associated band folding. The CDW-induced band splitting and the associated gap opening have been revealed at the boundary of the pristine and reconstructed Brillouin zones. The Fermi surface- and momentum-dependent CDW gap is measured and the strongly anisotropic CDW gap is observed for all the V-derived Fermi surface. In particular, we have observed signatures of the electron-phonon coupling in KV3Sb5. These results provide key insights in understanding the nature of the CDW state and its interplay with superconductivity in AV3Sb5 superconductors.

cond-mat.supr-con

Multiple topological states in iron-based superconductors

Topological insulators and semimetals as well as unconventional iron-based superconductors have attracted major recent attention in condensed matter physics. Previously, however, little overlap has been identified between these two vibrant fields, even though the principal combination of topological bands and superconductivity promises exotic unprecedented avenues of superconducting states and Majorana bound states (MBSs), the central building block for topological quantum computation. Along with progressing laser-based spin-resolved and angle-resolved photoemission spectroscopy (ARPES) towards high energy and momentum resolution, we have resolved topological insulator (TI) and topological Dirac semimetal (TDS) bands near the Fermi level ($E_{\text{F}}$) in the iron-based superconductors Li(Fe,Co)As and Fe(Te,Se), respectively. The TI and TDS bands can be individually tuned to locate close to $E_{\text{F}}$ by carrier doping, allowing to potentially access a plethora of different superconducting topological states in the same material. Our results reveal the generic coexistence of superconductivity and multiple topological states in iron-based superconductors, rendering these materials a promising platform for high-$T_{\text{c}}$ topological superconductivity.

cond-mat.supr-con

Topological Dirac semimetal phase in the iron-based superconductor Fe(Te,Se)

Topological Dirac semimetals (TDSs) exhibit bulk Dirac cones protected by time reversal and crystal symmetry, as well as surface states originating from non-trivial topology. While there is a manifold possible onset of superconducting order in such systems, few observations of intrinsic superconductivity have so far been reported for TDSs. We observe evidence for a TDS phase in FeTe$_{1-x}$Se$_x$ ($x$ = 0.45), one of the high transition temperature ($T_c$) iron-based superconductors. In angle-resolved photoelectron spectroscopy (ARPES) and transport experiments, we find spin-polarized states overlapping with the bulk states on the (001) surface, and linear magnetoresistance (MR) starting from 6 T. Combined, this strongly suggests the existence of a TDS phase, which is confirmed by theoretical calculations. In total, the topological electronic states in Fe(Te,Se) provide a promising high $T_c$ platform to realize multiple topological superconducting phases.

cond-mat.supr-con

Experimental observation of node-line-like surface states in LaBi

In a Dirac nodal line semimetal, the bulk conduction and valence bands touch at extended lines in the Brillouin zone. To date, most of the theoretically predicted and experimentally discovered nodal lines derive from the bulk bands of two- and three-dimensional materials. Here, based on combined angle-resolved photoemission spectroscopy measurements and first-principles calculations, we report the discovery of node-line-like surface states on the (001) surface of LaBi. These bands derive from the topological surface states of LaBi and bridge the band gap opened by spin-orbit coupling and band inversion. Our first-principles calculations reveal that these "nodal lines" have a tiny gap, which is beyond typical experimental resolution. These results may provide important information to understand the extraordinary physical properties of LaBi, such as the extremely large magnetoresistance and resistivity plateau.

cond-mat.mtrl-sci

Discovery of two-dimensional Dirac nodal line fermions in monolayer Cu2Si

Topological nodal line semimetals, a novel quantum state of materials, possess topologically nontrivial valence and conduction bands that touch at a line near the Fermi level. The exotic band structure can lead to various novel properties, such as long-range Coulomb interaction and flat Landau levels. Recently, topological nodal lines have been observed in several bulk materials, such as PtSn4, ZrSiS, TlTaSe2 and PbTaSe2. However, in two-dimensional materials, experimental research on nodal line fermions is still lacking. Here, we report the discovery of two-dimensional Dirac nodal line fermions in monolayer Cu2Si based on combined theoretical calculations and angle-resolved photoemission spectroscopy measurements. The Dirac nodal lines in Cu2Si form two concentric loops centred around the Γ point and are protected by mirror reflection symmetry. Our results establish Cu2Si as a new platform to study the novel physical properties in two-dimensional Dirac materials and provide new opportunities to realize high-speed low-dissipation devices.

cond-mat.mtrl-sci

High quality atomically thin PtSe2 films grown by molecular beam epitaxy

Atomically thin PtSe2 films have attracted extensive research interests for potential applications in high-speed electronics, spintronics and photodetectors. Obtaining high quality, single crystalline thin films with large size is critical. Here we report the first successful layer-by-layer growth of high quality PtSe2 films by molecular beam epitaxy. Atomically thin films from 1 ML to 22 ML have been grown and characterized by low-energy electron diffraction, Raman spectroscopy and X-ray photoemission spectroscopy. Moreover, a systematic thickness dependent study of the electronic structure is revealed by angle-resolved photoemission spectroscopy (ARPES), and helical spin texture is revealed by spin-ARPES. Our work provides new opportunities for growing large size single crystalline films for investigating the physical properties and potential applications of PtSe2.

cond-mat.mtrl-sci