SearcharxivSearch

arXiv subjects

Shen Hu

Publications and source records attributed to Shen Hu.

3 recordsLinked to original sources

Wafer-scale Demonstration of High-voltage beta-Ga2O3 MOSFETs with Excellent Uniformity and over 3kV Breakdown Voltages

This study demonstrates a wafer-scale growth of a 2-inch Si-doped $\beta$-Ga2O3 (100) epitaxial wafer and the realization of uniform, high-voltage lateral $\beta$-Ga2O3 MOSFET arrays. The 2-inch homoepitaxial $\beta$-Ga2O3 (100) film grown by MOCVD exhibit excellent crystalline uniformity with an average rocking curve FWHM of ~27.0 arcsec and a low surface roughness less than 1 nm, alongside a uniform net doping concentration on the value of 4.60 $\times$ 1E17 cm-3. The fabricated MOSFETs deliver a threshold voltage of -31.75 V, a drain-current on/off ratio over 1E9, a specific on-resistance of 126.52 mohm$\cdot$cm2 and breakdown voltage exceeding 3 kV. Statistical analysis across the entire wafer presents good device uniformity, with threshold voltages ranging from -28 V to -36 V, output current densities of 60-75 mA/mm, and a breakdown voltage over 3 kV. These results provide the demonstration using the 2-inch $\beta$-Ga2O3 epitaxial wafer to realize high-voltage $\beta$-Ga2O3 MOSFETs with wafer-scale performance uniformity for next-generation power device application.

cond-mat.mtrl-sci

Balanced Multimodal Learning: An Unidirectional Dynamic Interaction Perspective

Multimodal learning typically utilizes multimodal joint loss to integrate different modalities and enhance model performance. However, this joint learning strategy can induce modality imbalance, where strong modalities overwhelm weaker ones and limit exploitation of individual information from each modality and the inter-modality interaction information. Existing strategies such as dynamic loss weighting, auxiliary objectives and gradient modulation mitigate modality imbalance based on joint loss. These methods remain fundamentally reactive, detecting and correcting imbalance after it arises, while leaving the competitive nature of the joint loss untouched. This limitation drives us to explore a new strategy for multimodal imbalance learning that does not rely on the joint loss, enabling more effective interactions between modalities and better utilization of information from individual modalities and their interactions. In this paper, we introduce Unidirectional Dynamic Interaction (UDI), a novel strategy that abandons the conventional joint loss in favor of a proactive, sequential training scheme. UDI first trains the anchor modality to convergence, then uses its learned representations to guide the other modality via unsupervised loss. Furthermore, the dynamic adjustment of modality interactions allows the model to adapt to the task at hand, ensuring that each modality contributes optimally. By decoupling modality optimization and enabling directed information flow, UDI prevents domination by any single modality and fosters effective cross-modal feature learning. Our experimental results demonstrate that UDI outperforms existing methods in handling modality imbalance, leading to performance improvement in multimodal learning tasks.

cs.LG

A New Approach to Voice Authenticity

Voice faking, driven primarily by recent advances in text-to-speech (TTS) synthesis technology, poses significant societal challenges. Currently, the prevailing assumption is that unaltered human speech can be considered genuine, while fake speech comes from TTS synthesis. We argue that this binary distinction is oversimplified. For instance, altered playback speeds can be used for malicious purposes, like in the 'Drunken Nancy Pelosi' incident. Similarly, editing of audio clips can be done ethically, e.g., for brevity or summarization in news reporting or podcasts, but editing can also create misleading narratives. In this paper, we propose a conceptual shift away from the binary paradigm of audio being either 'fake' or 'real'. Instead, our focus is on pinpointing 'voice edits', which encompass traditional modifications like filters and cuts, as well as TTS synthesis and VC systems. We delineate 6 categories and curate a new challenge dataset rooted in the M-AILABS corpus, for which we present baseline detection systems. And most importantly, we argue that merely categorizing audio as fake or real is a dangerous over-simplification that will fail to move the field of speech technology forward.

cs.SD