Searcharxiv⌕ Search

arXiv subjects

Bing Han

Publications and source records attributed to Bing Han.

At least 73 records · Page 4Linked to original sources

SegRap2023: A Benchmark of Organs-at-Risk and Gross Tumor Volume Segmentation for Radiotherapy Planning of Nasopharyngeal Carcinoma

Radiation therapy is a primary and effective NasoPharyngeal Carcinoma (NPC) treatment strategy. The precise delineation of Gross Tumor Volumes (GTVs) and Organs-At-Risk (OARs) is crucial in radiation treatment, directly impacting patient prognosis. Previously, the delineation of GTVs and OARs was performed by experienced radiation oncologists. Recently, deep learning has achieved promising results in many medical image segmentation tasks. However, for NPC OARs and GTVs segmentation, few public datasets are available for model development and evaluation. To alleviate this problem, the SegRap2023 challenge was organized in conjunction with MICCAI2023 and presented a large-scale benchmark for OAR and GTV segmentation with 400 Computed Tomography (CT) scans from 200 NPC patients, each with a pair of pre-aligned non-contrast and contrast-enhanced CT scans. The challenge's goal was to segment 45 OARs and 2 GTVs from the paired CT scans. In this paper, we detail the challenge and analyze the solutions of all participants. The average Dice similarity coefficient scores for all submissions ranged from 76.68\% to 86.70\%, and 70.42\% to 73.44\% for OARs and GTVs, respectively. We conclude that the segmentation of large-size OARs is well-addressed, and more efforts are needed for GTVs and small-size or thin-structure OARs. The benchmark will remain publicly available here: https://segrap2023.grand-challenge.org

eess.IV↗

InstructME: An Instruction Guided Music Edit And Remix Framework with Latent Diffusion Models

Music editing primarily entails the modification of instrument tracks or remixing in the whole, which offers a novel reinterpretation of the original piece through a series of operations. These music processing methods hold immense potential across various applications but demand substantial expertise. Prior methodologies, although effective for image and audio modifications, falter when directly applied to music. This is attributed to music's distinctive data nature, where such methods can inadvertently compromise the intrinsic harmony and coherence of music. In this paper, we develop InstructME, an Instruction guided Music Editing and remixing framework based on latent diffusion models. Our framework fortifies the U-Net with multi-scale aggregation in order to maintain consistency before and after editing. In addition, we introduce chord progression matrix as condition information and incorporate it in the semantic space to improve melodic harmony while editing. For accommodating extended musical pieces, InstructME employs a chunk transformer, enabling it to discern long-term temporal dependencies within music sequences. We tested InstructME in instrument-editing, remixing, and multi-round editing. Both subjective and objective evaluations indicate that our proposed method significantly surpasses preceding systems in music quality, text relevance and harmony. Demo samples are available at https://musicedit.github.io/

cs.SD↗

Ultra-high-speed coherent anti-Stokes Raman spectroscopy with a hybrid dual-comb source

Coherent anti-Stokes Raman scattering (CARS) spectroscopy with time-delayed ultrashort pulses and a single-pixel photodetector has shown great potential for spectroscopic imaging and transient studies in chemistry and biological research. However, those systems rely on mechanical delay lines or two asynchronous optical combs with inflexible repetition frequencies, technically limiting their acquisition speeds. Here, we demonstrate a hybrid dual-comb CARS system involving a broadband fiber laser and a highly-flexible, frequency-modulated electro-optic comb. We achieve multiplex CARS spectra (2800-3200 cm-1), with a moderate resolution (22 cm-1), at a maximum refresh rate of 1 MHz, limited by the radio-frequency synthesizer we use. Fast spectroscopic CARS imaging is demonstrated for liquid mixtures. Our system enables spectral measurements in the high-wavenumber C-H stretching region at a record speed that is an order of magnitude higher than state-of-the-art systems, which may open up new opportunities for fast chemical sensing and imaging.

physics.optics↗

Market Crowds' Trading Behaviors, Agreement Prices, and the Implications of Trading Volume

It has been long that literature in financial academics focuses mainly on price and return but much less on trading volume. In the past twenty years, it has already linked both price and trading volume to economic fundamentals, and explored the behavioral implications of trading volume such as investor's attitude toward risks, overconfidence, disagreement, and attention etc. However, what is surprising is how little we really know about trading volume. Here we show that trading volume probability represents the frequency of market crowd's trading action in terms of behavior analysis, and test two adaptive hypotheses relevant to the volume uncertainty associated with price in China stock market. The empirical work reveals that market crowd trade a stock in efficient adaptation except for simple heuristics, gradually tend to achieve agreement on an outcome or an asset price widely on a trading day, and generate such a stationary equilibrium price very often in interaction and competition among themselves no matter whether it is highly overestimated or underestimated. This suggests that asset prices include not only a fundamental value but also private information, speculative, sentiment, attention, gamble, and entertainment values etc. Moreover, market crowd adapt to gain and loss by trading volume increase or decrease significantly in interaction with environment in any two consecutive trading days. Our results demonstrate how interaction between information and news, the trading action, and return outcomes in the three-term feedback loop produces excessive trading volume which includes various internal and external causes.

q-fin.GN↗

Elucidating Dynamic Conductive State Changes in Amorphous Lithium Lanthanum Titanate for Resistive Switching Devices

Exploration of novel resistive switching materials attracts attention to replace conventional Si-based transistors and to achieve neuromorphic computing that can surpass the limit of the current Von-Neumann computing for the time of Internet of Things (IoT). Materials priorly used to serve in batteries have demonstrated metal-insulator transitions upon an electrical biasing due to resulting compositional change. This property is desirable for future resistive switching devices. Amorphous lithium lanthanum titanate (a-LLTO) was originally developed as a solid-state electrolyte with relatively high lithium ionic conductivity and low electronic conductivity among oxide-type solid electrolytes. However, it has been suggested that electric conductivity of a-LLTO changes depending on oxygen content. In this work, the investigation of switching behavior of a-LLTO was conducted by employing a range of voltage sweep techniques, ultimately establishing a stable and optimal operating condition within the voltage window of -3.5 V to 3.5 V. This voltage range effectively balances the desirable trait of a substantial resistance change by three orders of magnitude with the imperative avoidance of LLTO decomposition. This switching behavior is also confirmed at nanodevice of Ni/LLTO/Ni through in-situ biasing inside focused-ion beam/scanning electron microscope (FIB-SEM). Experiment and computation with different LLTO composition shows that LLTO has two distinct conductivity states due to Ti reduction. The distribution of these two states is discussed using simplified binary model, implying the conductive filament growth during low resistance state. Consequently, our study deepens understanding of LLTO electronic properties and encourages the interdisciplinary application of battery materials for resistive switching devices.

physics.app-ph↗

Leveraging In-the-Wild Data for Effective Self-Supervised Pretraining in Speaker Recognition

Current speaker recognition systems primarily rely on supervised approaches, constrained by the scale of labeled datasets. To boost the system performance, researchers leverage large pretrained models such as WavLM to transfer learned high-level features to the downstream speaker recognition task. However, this approach introduces extra parameters as the pretrained model remains in the inference stage. Another group of researchers directly apply self-supervised methods such as DINO to speaker embedding learning, yet they have not explored its potential on large-scale in-the-wild datasets. In this paper, we present the effectiveness of DINO training on the large-scale WenetSpeech dataset and its transferability in enhancing the supervised system performance on the CNCeleb dataset. Additionally, we introduce a confidence-based data filtering algorithm to remove unreliable data from the pretraining dataset, leading to better performance with less training data. The associated pretrained models, confidence files, pretraining and finetuning scripts will be made available in the Wespeaker toolkit.

eess.AS↗

Quantitative Analysis of Sodium Metal Deposition and Interphase in Na Metal Batteries

Sodium-ion batteries exhibit significant promise as a viable alternative to current lithium-ion technologies owing to their sustainability, low cost per energy density, reliability, and safety. Despite recent advancements in cathode materials for this category of energy storage systems, the primary challenge in realizing practical applications of sodium-ion systems is the absence of an anode system with high energy density and durability. Although Na metal is the ultimate anode that can facilitate high-energy sodium-ion batteries, its use remains limited due to safety concerns and the high-capacity loss associated with the high reactivity of Na metal. In this study, titration gas chromatography is employed to accurately quantify the sodium inventory loss in ether- and carbonate-based electrolytes. Uniaxial pressure is developed as a powerful tool to control the deposition of sodium metal with dense morphology, thereby enabling high initial coulombic efficiencies. In ether-based electrolytes, the Na metal surface exhibits the presence of a uniform solid electrolyte interphase layer, primarily characterized by favorable inorganic chemical components with close-packed structures. The full cell, utilizing a controlled electroplated sodium metal in ether-based electrolyte, provides capacity retention of 91.84% after 500 cycles at 2C current rate and delivers 86 mAh/g discharge capacity at 45C current rate, suggesting the potential to enable Na metal in the next generation of sodium-ion technologies with specifications close to practical requirements.

physics.app-ph↗

Attention-based Encoder-Decoder End-to-End Neural Diarization with Embedding Enhancer

Deep neural network-based systems have significantly improved the performance of speaker diarization tasks. However, end-to-end neural diarization (EEND) systems often struggle to generalize to scenarios with an unseen number of speakers, while target speaker voice activity detection (TS-VAD) systems tend to be overly complex. In this paper, we propose a simple attention-based encoder-decoder network for end-to-end neural diarization (AED-EEND). In our training process, we introduce a teacher-forcing strategy to address the speaker permutation problem, leading to faster model convergence. For evaluation, we propose an iterative decoding method that outputs diarization results for each speaker sequentially. Additionally, we propose an Enhancer module to enhance the frame-level speaker embeddings, enabling the model to handle scenarios with an unseen number of speakers. We also explore replacing the transformer encoder with a Conformer architecture, which better models local information. Furthermore, we discovered that commonly used simulation datasets for speaker diarization have a much higher overlap ratio compared to real data. We found that using simulated training data that is more consistent with real data can achieve an improvement in consistency. Extensive experimental validation demonstrates the effectiveness of our proposed methodologies. Our best system achieved a new state-of-the-art diarization error rate (DER) performance on all the CALLHOME (10.08%), DIHARD II (24.64%), and AMI (13.00%) evaluation benchmarks, when no oracle voice activity detection (VAD) is used. Beyond speaker diarization, our AED-EEND system also shows remarkable competitiveness as a speech type detection model.

cs.SD↗

Attention-based Encoder-Decoder Network for End-to-End Neural Speaker Diarization with Target Speaker Attractor

This paper proposes a novel Attention-based Encoder-Decoder network for End-to-End Neural speaker Diarization (AED-EEND). In AED-EEND system, we incorporate the target speaker enrollment information used in target speaker voice activity detection (TS-VAD) to calculate the attractor, which can mitigate the speaker permutation problem and facilitate easier model convergence. In the training process, we propose a teacher-forcing strategy to obtain the enrollment information using the ground-truth label. Furthermore, we propose three heuristic decoding methods to identify the enrollment area for each speaker during the evaluation process. Additionally, we enhance the attractor calculation network LSTM used in the end-to-end encoder-decoder based attractor calculation (EEND-EDA) system by incorporating an attention-based model. By utilizing such an attention-based attractor decoder, our proposed AED-EEND system outperforms both the EEND-EDA and TS-VAD systems with only 0.5s of enrollment data.

cs.SD↗

Enhancing Efficient Continual Learning with Dynamic Structure Development of Spiking Neural Networks

Children possess the ability to learn multiple cognitive tasks sequentially, which is a major challenge toward the long-term goal of artificial general intelligence. Existing continual learning frameworks are usually applicable to Deep Neural Networks (DNNs) and lack the exploration on more brain-inspired, energy-efficient Spiking Neural Networks (SNNs). Drawing on continual learning mechanisms during child growth and development, we propose Dynamic Structure Development of Spiking Neural Networks (DSD-SNN) for efficient and adaptive continual learning. When learning a sequence of tasks, the DSD-SNN dynamically assigns and grows new neurons to new tasks and prunes redundant neurons, thereby increasing memory capacity and reducing computational overhead. In addition, the overlapping shared structure helps to quickly leverage all acquired knowledge to new tasks, empowering a single network capable of supporting multiple incremental tasks (without the separate sub-network mask for each task). We validate the effectiveness of the proposed model on multiple class incremental learning and task incremental learning benchmarks. Extensive experiments demonstrated that our model could significantly improve performance, learning speed and memory capacity, and reduce computational overhead. Besides, our DSD-SNN model achieves comparable performance with the DNNs-based methods, and significantly outperforms the state-of-the-art (SOTA) performance for existing SNNs-based continual learning methods.

cs.AI↗

Exploring Binary Classification Loss For Speaker Verification

The mismatch between close-set training and open-set testing usually leads to significant performance degradation for speaker verification task. For existing loss functions, metric learning-based objectives depend strongly on searching effective pairs which might hinder further improvements. And popular multi-classification methods are usually observed with degradation when evaluated on unseen speakers. In this work, we introduce SphereFace2 framework which uses several binary classifiers to train the speaker model in a pair-wise manner instead of performing multi-classification. Benefiting from this learning paradigm, it can efficiently alleviate the gap between training and evaluation. Experiments conducted on Voxceleb show that the SphereFace2 outperforms other existing loss functions, especially on hard trials. Besides, large margin fine-tuning strategy is proven to be compatible with it for further improvements. Finally, SphereFace2 also shows its strong robustness to class-wise noisy labels which has the potential to be applied in the semi-supervised training scenario with inaccurate estimated pseudo labels. Codes are available in https://github.com/Hunterhuan/sphereface2_speaker_verification

eess.AS↗

Wespeaker baselines for VoxSRC2023

This report showcases the results achieved using the wespeaker toolkit for the VoxSRC2023 Challenge. Our aim is to provide participants, especially those with limited experience, with clear and straightforward guidelines to develop their initial systems. Via well-structured recipes and strong results, we hope to offer an accessible and good enough start point for all interested individuals. In this report, we describe the results achieved on the VoxSRC2023 dev set using the pretrained models, you can check the CodaLab evaluation server for the results on the evaluation set.

eess.AS↗

Build a SRE Challenge System: Lessons from VoxSRC 2022 and CNSRC 2022

Many speaker recognition challenges have been held to assess the speaker verification system in the wild and probe the performance limit. Voxceleb Speaker Recognition Challenge (VoxSRC), based on the voxceleb, is the most popular. Besides, another challenge called CN-Celeb Speaker Recognition Challenge (CNSRC) is also held this year, which is based on the Chinese celebrity multi-genre dataset CN-Celeb. This year, our team participated in both speaker verification closed tracks in CNSRC 2022 and VoxSRC 2022, and achieved the 1st place and 3rd place respectively. In most system reports, the authors usually only provide a description of their systems but lack an effective analysis of their methods. In this paper, we will outline how to build a strong speaker verification challenge system and give a detailed analysis of each method compared with some other popular technical means.

cs.SD↗

The SJTU X-LANCE Lab System for CNSRC 2022

This technical report describes the SJTU X-LANCE Lab system for the three tracks in CNSRC 2022. In this challenge, we explored the speaker embedding modeling ability of deep ResNet (Deeper r-vector). All the systems are only trained on the Cnceleb training set and we use the same systems for the three tracks in CNSRC 2022. In this challenge, our system ranks the first place in the fixed track of speaker verification task. Our best single system and fusion system achieve 0.3164 and 0.2975 minDCF respectively. Besides, we submit the result of ResNet221 to the speaker retrieval track and achieve 0.4626 mAP. More importantly, we have helped the wespeaker [1] toolkit reproduce our result: https://github.com/wenet-e2e/wespeaker.

cs.SD↗

Elucidating the Role of Prelithiation in Si-based Anodes for Interface Stabilization

Prelithiation as a facile and effective method to compensate the lithium inventory loss in the initial cycle has progressed considerably both on anode and cathode sides. However, much less research has been devoted to the prelithiation effect on the interface stabilization for long-term cycling of Si-based anodes. An in-depth quantitative analysis of the interface that form during the prelithiation of SiO$_x$ is presented here and the results are compared with prelithiaton of Si anodes. Local structure probe combined with detailed electrochemical analysis reveals that a characteristic mosaic interface is formed on both prelithiated SiO$_x$ and Si anodes. This mosaic interface containing multiple lithium silicates phases, is fundamentally different from the solid electrolyte interface (SEI) formed without prelithiation. The ideal conductivity and mechanical properties of lithium silicates enable improved cycling stability of both prelithiated anodes. With a higher ratio of lithium silicates due to the oxygen participation, prelithiated SiO$_{1.3}$ anode improves the initial coulombic efficiency to 94% in full cell and delivers good cycling retention after hundreds cycles under lean electrolyte conditions. The insights provided in this work could be used to further optimize high Si loading based anode in future high energy density batteries.

physics.chem-ph↗

Freestanding LiPON: from Fundamental Study to Uniformly Dense Li Metal Deposition Under Zero External Pressure

Lithium phosphorus oxynitride (LiPON) is a well-known amorphous thin film solid electrolyte that has been extensively studied in the last three decades. Despite the promises to pair with Li metal anode and various cathode materials, the presence of rigid substrate and LiPONs unique amorphous, air-sensitive nature set limitations to comprehensively understand its intrinsic properties for future development and applications. This work demonstrates a methodology to synthesize LiPON in a freestanding form that exhibits remarkable flexibility and a Young s modulus of ~33 GPa. Solid-state nuclear magnetic resonance (ss-NMR) and differential scanning calorimetry (DSC) results with unprecedented high signal-to-noise ratio could be obtained with such freestanding LiPON (FS-LiPON), revealing the Li-LiPON interface bonding environments quantitatively and a well-defined glass transition temperature for LiPON. Combining interfacial stress and a seeding layer, FS-LiPON demonstrates a uniform and fully dense Li metal deposition without the aid of external pressure. Such a FS-LiPON film offers new opportunities for fundamental study of LiPON material and associated interfaces, and provides perspectives for interface engineering in bulk solid-state battery.

cond-mat.mtrl-sci↗

Self-Supervised Learning with Cluster-Aware-DINO for High-Performance Robust Speaker Verification

Automatic speaker verification task has made great achievements using deep learning approaches with the large-scale manually annotated dataset. However, it's very difficult and expensive to collect a large amount of well-labeled data for system building. In this paper, we propose a novel and advanced self-supervised learning framework which can construct a high performance speaker verification system without using any labeled data. To avoid the impact of false negative pairs, we adopt the self-distillation with no labels (DINO) framework as the initial model, which can be trained without exploiting negative pairs. Then, we introduce a cluster-aware training strategy for DINO to improve the diversity of data. In the iteration learning stage, due to a mass of unreliable labels from clustering, the quality of pseudo labels is important for the system training. This motivates us to propose dynamic loss-gate and label correction (DLG-LC) methods to alleviate the performance degradation caused by unreliable labels. More specifically, we model the loss distribution with GMM and obtain the loss-gate threshold dynamically to distinguish the reliable and unreliable labels. Besides, we adopt the model predictions to correct the unreliable label, for better utilizing the unreliable data rather than dropping them directly. Moreover, we extend the DLG-LC to multi-modality to further improve the performance. The experiments are performed on the commonly used Voxceleb dataset. Compared to the best-known self-supervised speaker verification system, our proposed method obtain 22.17%, 27.94% and 25.56% relative EER improvement on Vox-O, Vox-E and Vox-H test sets, even with fewer iterations, smaller models, and simpler clustering methods. More importantly, the newly proposed system even achieves comparable results with the fully supervised system, but without using any human labeled data.

cs.SD↗

Adaptive structure evolution and biologically plausible synaptic plasticity for recurrent spiking neural networks

The architecture design and multi-scale learning principles of the human brain that evolved over hundreds of millions of years are crucial to realizing human-like intelligence. Spiking Neural Network (SNN) based Liquid State Machine (LSM) serves as a suitable architecture to study brain-inspired intelligence because of its brain-inspired structure and the potential for integrating multiple biological principles. Existing researches on LSM focus on different certain perspectives, including high-dimensional encoding or optimization of the liquid layer, network architecture search, and application to hardware devices. There is still a lack of in-depth inspiration from the learning and structural evolution mechanism of the brain. Considering these limitations, this paper presents a novel LSM learning model that integrates adaptive structural evolution and multi-scale biological learning rules. For structural evolution, an adaptive evolvable LSM model is developed to optimize the neural architecture design of liquid layer with separation property. For brain-inspired learning of LSM, we propose a dopamine-modulated Bienenstock-Cooper-Munros (DA-BCM) method that incorporates global long-term dopamine regulation and local trace-based BCM synaptic plasticity. Comparative experimental results on different decision-making tasks show that introducing structural evolution of the liquid layer, and the DA-BCM regulation of the liquid layer and the readout layer could improve the decision-making ability of LSM and flexibly adapt to rule reversal. This work is committed to exploring how evolution can help to design more appropriate network architectures and how multi-scale neuroplasticity principles coordinated to enable the optimization and learning of LSMs for relatively complex decision-making tasks.

cs.NE↗