Searcharxiv⌕ Search

arXiv subjects

Min Tang

Publications and source records attributed to Min Tang.

At least 55 records · Page 3Linked to original sources

gDist: Efficient Distance Computation between 3D Meshes on GPU

Computing maximum/minimum distances between 3D meshes is crucial for various applications, i.e., robotics, CAD, VR/AR, etc. In this work, we introduce a highly parallel algorithm (gDist) optimized for Graphics Processing Units (GPUs), which is capable of computing the distance between two meshes with over 15 million triangles in less than 0.4 milliseconds (Fig. 1). By testing on benchmarks with varying characteristics, the algorithm achieves remarkable speedups over prior CPU-based and GPU-based algorithms on a commodity GPU (NVIDIA GeForce RTX 4090). Notably, the algorithm consistently maintains high-speed performance, even in challenging scenarios that pose difficulties for prior algorithms.

cs.GR↗

E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS

This paper introduces Embarrassingly Easy Text-to-Speech (E2 TTS), a fully non-autoregressive zero-shot text-to-speech system that offers human-level naturalness and state-of-the-art speaker similarity and intelligibility. In the E2 TTS framework, the text input is converted into a character sequence with filler tokens. The flow-matching-based mel spectrogram generator is then trained based on the audio infilling task. Unlike many previous works, it does not require additional components (e.g., duration model, grapheme-to-phoneme) or complex techniques (e.g., monotonic alignment search). Despite its simplicity, E2 TTS achieves state-of-the-art zero-shot TTS capabilities that are comparable to or surpass previous works, including Voicebox and NaturalSpeech 3. The simplicity of E2 TTS also allows for flexibility in the input representation. We propose several variants of E2 TTS to improve usability during inference. See https://aka.ms/e2tts/ for demo samples.

eess.AS↗

An Adaptive Angular Domain Compression Scheme For Solving Multiscale Radiative Transfer Equation

When dealing with the steady-state multiscale radiative transfer equation (RTE) with heterogeneous coefficients, spatially localized low-rank structures are present in the angular space. This paper introduces an adaptive tailored finite point scheme (TFPS) for RTEs in heterogeneous media, which can adaptively compress the angular space. It does so by selecting reduced TFPS basis functions based on the local optical properties of the background media. These reduced basis functions capture the important local modes in the velocity domain. A detailed a posteriori error analysis is performed to quantify the discrepancy between the reduced and full TFPS solutions. Additionally, numerical experiments demonstrate the efficiency and accuracy of the adaptive TFPS in solving multiscale RTEs, especially in scenarios involving boundary and interface layers.

math.NA↗

Restricted sumsets in $\mathbb{Z}$

Let $k\geqslant 3$ and let $A=\{0=a_{0}<a_{1}<\cdots<a_{k-1}\}$ with $\gcd(A)=1$. Freiman-Lev conjecture [V.F. Lev, Restricted set addition in groups, I. The classical setting, J. London Math. Soc. 62(2000), 27-40] is a well-known conjecture which related to restricted sumsets. Up to now, Freiman-Lev conjecture is open for all $a_{k-1}\geqslant 2k-2$. In this paper, we prove the Freiman-Lev conjecture is true for $a_{k-1}\geqslant 2k-2$ and $a_{k-2}<2k-4$. That is, Freiman-Lev conjecture is still open for the case $a_{k-1}\geqslant 2k-2$ and $a_{k-2}\geq 2k-4$.

math.NT↗

Efficient Sparse Attention needs Adaptive Token Release

In recent years, Large Language Models (LLMs) have demonstrated remarkable capabilities across a wide array of text-centric tasks. However, their `large' scale introduces significant computational and storage challenges, particularly in managing the key-value states of the transformer, which limits their wider applicability. Therefore, we propose to adaptively release resources from caches and rebuild the necessary key-value states. Particularly, we accomplish this by a lightweight controller module to approximate an ideal top-$K$ sparse attention. This module retains the tokens with the highest top-$K$ attention weights and simultaneously rebuilds the discarded but necessary tokens, which may become essential for future decoding. Comprehensive experiments in natural language generation and modeling reveal that our method is not only competitive with full attention in terms of performance but also achieves a significant throughput improvement of up to 221.8%. The code for replication is available on the https://github.com/WHUIR/ADORE.

cs.CL↗

SpeechX: Neural Codec Language Model as a Versatile Speech Transformer

Recent advancements in generative speech models based on audio-text prompts have enabled remarkable innovations like high-quality zero-shot text-to-speech. However, existing models still face limitations in handling diverse audio-text speech generation tasks involving transforming input speech and processing audio captured in adverse acoustic conditions. This paper introduces SpeechX, a versatile speech generation model capable of zero-shot TTS and various speech transformation tasks, dealing with both clean and noisy signals. SpeechX combines neural codec language modeling with multi-task learning using task-dependent prompting, enabling unified and extensible modeling and providing a consistent way for leveraging textual input in speech enhancement and transformation tasks. Experimental results show SpeechX's efficacy in various tasks, including zero-shot TTS, noise suppression, target speaker extraction, speech removal, and speech editing with or without background noise, achieving comparable or superior performance to specialized models across tasks. See https://aka.ms/speechx for demo samples.

eess.AS↗

An Investigation of Noise Robustness for Flow-Matching-Based Zero-Shot TTS

Recently, zero-shot text-to-speech (TTS) systems, capable of synthesizing any speaker's voice from a short audio prompt, have made rapid advancements. However, the quality of the generated speech significantly deteriorates when the audio prompt contains noise, and limited research has been conducted to address this issue. In this paper, we explored various strategies to enhance the quality of audio generated from noisy audio prompts within the context of flow-matching-based zero-shot TTS. Our investigation includes comprehensive training strategies: unsupervised pre-training with masked speech denoising, multi-speaker detection and DNSMOS-based data filtering on the pre-training data, and fine-tuning with random noise mixing. The results of our experiments demonstrate significant improvements in intelligibility, speaker similarity, and overall audio quality compared to the approach of applying speech enhancement to the audio prompt.

eess.AS↗

Total-Duration-Aware Duration Modeling for Text-to-Speech Systems

Accurate control of the total duration of generated speech by adjusting the speech rate is crucial for various text-to-speech (TTS) applications. However, the impact of adjusting the speech rate on speech quality, such as intelligibility and speaker characteristics, has been underexplored. In this work, we propose a novel total-duration-aware (TDA) duration model for TTS, where phoneme durations are predicted not only from the text input but also from an additional input of the total target duration. We also propose a MaskGIT-based duration model that enhances the diversity and quality of the predicted phoneme durations. Our results demonstrate that the proposed TDA duration models achieve better intelligibility and speaker similarity for various speech rate configurations compared to the baseline models. We also show that the proposed MaskGIT-based model can generate phoneme durations with higher quality and diversity compared to its regression or flow-matching counterparts.

eess.AS↗

Dynamic Phase Enabled Topological Mode Steering in Composite Su-Schrieffer-Heeger Waveguide Arrays

Topological boundary states localize at interfaces whenever the interface implies a change of the associated topological invariant encoded in the geometric phase. The generically present dynamic phase, however, which is energy and time dependent, has been known to be non-universal, and hence not to intertwine with any topological geometric phase. Using the example of topological zero modes in composite Su-Schrieffer-Heeger (c-SSH) waveguide arrays with a central defect, we report on the selective excitation and transition of topological boundary mode based on dynamic phase-steered interferences. Our work thus provides a new knob for the control and manipulation of topological states in composite photonic devices, indicating promising applications where topological modes and their bandwidth can be jointly controlled by the dynamic phase, geometric phase, and wavelength in on-chip topological devices.

physics.optics↗

Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like

Laughter is one of the most expressive and natural aspects of human speech, conveying emotions, social cues, and humor. However, most text-to-speech (TTS) systems lack the ability to produce realistic and appropriate laughter sounds, limiting their applications and user experience. While there have been prior works to generate natural laughter, they fell short in terms of controlling the timing and variety of the laughter to be generated. In this work, we propose ELaTE, a zero-shot TTS that can generate natural laughing speech of any speaker based on a short audio prompt with precise control of laughter timing and expression. Specifically, ELaTE works on the audio prompt to mimic the voice characteristic, the text prompt to indicate the contents of the generated speech, and the input to control the laughter expression, which can be either the start and end times of laughter, or the additional audio prompt that contains laughter to be mimicked. We develop our model based on the foundation of conditional flow-matching-based zero-shot TTS, and fine-tune it with frame-level representation from a laughter detector as additional conditioning. With a simple scheme to mix small-scale laughter-conditioned data with large-scale pre-training data, we demonstrate that a pre-trained zero-shot TTS model can be readily fine-tuned to generate natural laughter with precise controllability, without losing any quality of the pre-trained zero-shot TTS model. Through objective and subjective evaluations, we show that ELaTE can generate laughing speech with significantly higher quality and controllability compared to conventional models. See https://aka.ms/elate/ for demo samples.

eess.AS↗

An asymptotic-preserving method for the three-temperature radiative transfer model

We present an asymptotic-preserving (AP) numerical method for solving the three-temperature radiative transfer model, which holds significant importance in inertial confinement fusion. A carefully designedsplitting method is developed that can provide a general framework of extending AP schemes for the gray radiative transport equation to the more complex three-temperature radiative transfer model. The proposed scheme captures two important limiting models: the three-temperature radiation diffusion equation (3TRDE) when opacity approaches infinity and the two-temperature limit when the ion-electron coupling coefficient goes to infinity. We have rigorously demonstrated the AP property and energy conservation characteristics of the proposed scheme and its efficiency has been validated through a series of benchmark tests in the numerical part.

math.NA↗

High order conservative LDG-IMEX methods for the degenerate nonlinear non-equilibrium radiation diffusion problems

In this paper, we develop a class of high-order conservative methods for simulating non-equilibrium radiation diffusion problems. Numerically, this system poses significant challenges due to strong nonlinearity within the stiff source terms and the degeneracy of nonlinear diffusion terms. Explicit methods require impractically small time steps, while implicit methods, which offer stability, come with the challenge to guarantee the convergence of nonlinear iterative solvers. To overcome these challenges, we propose a predictor-corrector approach and design proper implicit-explicit time discretizations. In the predictor step, the system is reformulated into a nonconservative form and linear diffusion terms are introduced as a penalization to mitigate strong nonlinearities. We then employ a Picard iteration to secure convergence in handling the nonlinear aspects. The corrector step guarantees the conservation of total energy, which is vital for accurately simulating the speeds of propagating sharp fronts in this system. For spatial approximations, we utilize local discontinuous Galerkin finite element methods, coupled with positive-preserving and TVB limiters. We validate the orders of accuracy, conservation properties, and suitability of using large time steps for our proposed methods, through numerical experiments conducted on one- and two-dimensional spatial problems. In both homogeneous and heterogeneous non-equilibrium radiation diffusion problems, we attain a time stability condition comparable to that of a fully implicit time discretization. Such an approach is also applicable to many other reaction-diffusion systems.

math.NA↗

NOTSOFAR-1 Challenge: New Datasets, Baseline, and Tasks for Distant Meeting Transcription

We introduce the first Natural Office Talkers in Settings of Far-field Audio Recordings (``NOTSOFAR-1'') Challenge alongside datasets and baseline system. The challenge focuses on distant speaker diarization and automatic speech recognition (DASR) in far-field meeting scenarios, with single-channel and known-geometry multi-channel tracks, and serves as a launch platform for two new datasets: First, a benchmarking dataset of 315 meetings, averaging 6 minutes each, capturing a broad spectrum of real-world acoustic conditions and conversational dynamics. It is recorded across 30 conference rooms, featuring 4-8 attendees and a total of 35 unique speakers. Second, a 1000-hour simulated training dataset, synthesized with enhanced authenticity for real-world generalization, incorporating 15,000 real acoustic transfer functions. The tasks focus on single-device DASR, where multi-channel devices always share the same known geometry. This is aligned with common setups in actual conference rooms, and avoids technical complexities associated with multi-device tasks. It also allows for the development of geometry-specific solutions. The NOTSOFAR-1 Challenge aims to advance research in the field of distant conversational speech recognition, providing key resources to unlock the potential of data-driven methods, which we believe are currently constrained by the absence of comprehensive high-quality training and benchmarking datasets.

cs.SD↗

A fast offline/online forward solver for stationary transport equation with multiple inflow boundary conditions and varying coefficients

It is of great interest to solve the inverse problem of stationary radiative transport equation (RTE) in optical tomography. The standard way is to formulate the inverse problem into an optimization problem, but the bottleneck is that one has to solve the forward problem repeatedly, which is time-consuming. Due to the optical property of biological tissue, in real applications, optical thin and thick regions coexist and are adjacent to each other, and the geometry can be complex. To use coarse meshes and save the computational cost, the forward solver has to be asymptotic preserving across the interface (APAL). In this paper, we propose an offline/online solver for RTE. The cost at the offline stage is comparable to classical methods, while the cost at the online stage is much lower. Two cases are considered. One is to solve the RTE with fixed scattering and absorption cross sections while the boundary conditions vary; the other is when cross sections vary in a small domain and the boundary conditions change many times. The solver can be decomposed into offline/online stages in these two cases. One only needs to calculate the offline stage once and update the online stage when the parameters vary. Our proposed solver is much cheaper when one needs to solve RTE with multiple right-hand sides or when the cross sections vary in a small domain, thus can accelerate the speed of solving inverse RTE problems. We illustrate the online/offline decomposition based on the Tailored Finite Point Method (TFPM), which is APAL on general quadrilateral meshes.

math.NA↗

Kinetic chemotaxis tumbling kernel determined from macroscopic quantities

Chemotaxis is the physical phenomenon that bacteria adjust their motions according to chemical stimulus. A classical model for this phenomenon is a kinetic equation that describes the velocity jump process whose tumbling/transition kernel uniquely determines the effect of chemical stimulus on bacteria. The model has been shown to be an accurate model that matches with bacteria motion qualitatively. For a quantitative modeling, biophysicists and practitioners are also highly interested in determining the explicit value of the tumbling kernel. Due to the experimental limitations, measurements are typically macroscopic in nature. Do macroscopic quantities contain enough information to recover microscopic behavior? In this paper, we give a positive answer. We show that when given a special design of initial data, the population density, one specific macroscopic quantity as a function of time, contains sufficient information to recover the tumbling kernel and its associated damping coefficient. Moreover, we can read off the chemotaxis tumbling kernel using the values of population density directly from this specific experimental design. This theoretical result using kinetic theory sheds light on how practitioners may conduct experiments in laboratories.

math.AP↗

Symmetry induced selective excitation of topological states in SSH waveguide arrays

The investigation of topological state transition in carefully designed photonic lattices is of high interest for fundamental research, as well as for applied studies such as manipulating light flow in on-chip photonic systems. Here, we report on topological phase transition between symmetric topological zero modes (TZM) and antisymmetric TZMs in Su-Schrieffer-Heeger (SSH) mirror symmetric waveguides. The transition of TZMs is realized by adjusting the coupling ratio between neighboring waveguide pairs, which is enabled by selective modulation of the refractive index in the waveguide gaps. Bi-directional topological transitions between symmetric and antisymmetric TZMs can be achieved with our proposed switching strategy. Selective excitation of topological edge mode is demonstrated owing to the symmetry characteristics of the TZMs. The flexible manipulation of topological states is promising for on-chip light flow control and may spark further investigations on symmetric/antisymmetric TZM transitions in other photonic topological frameworks.

physics.optics↗

CTSN: Predicting Cloth Deformation for Skeleton-based Characters with a Two-stream Skinning Network

We present a novel learning method to predict the cloth deformation for skeleton-based characters with a two-stream network. The characters processed in our approach are not limited to humans, and can be other skeletal-based representations of non-human targets such as fish or pets. We use a novel network architecture which consists of skeleton-based and mesh-based residual networks to learn the coarse and wrinkle features as the overall residual from the template cloth mesh. Our network is used to predict the deformation for loose or tight-fitting clothing or dresses. We ensure that the memory footprint of our network is low, and thereby result in reduced storage and computational requirements. In practice, our prediction for a single cloth mesh for the skeleton-based character takes about 7 milliseconds on an NVIDIA GeForce RTX 3090 GPU. Compared with prior methods, our network can generate fine deformation results with details and wrinkles.

cs.GR↗

Real-Time Joint Personalized Speech Enhancement and Acoustic Echo Cancellation

Personalized speech enhancement (PSE) is a real-time SE approach utilizing a speaker embedding of a target person to remove background noise, reverberation, and interfering voices. To deploy a PSE model for full duplex communications, the model must be combined with acoustic echo cancellation (AEC), although such a combination has been less explored. This paper proposes a series of methods that are applicable to various model architectures to develop efficient causal models that can handle the tasks of PSE, AEC, and joint PSE-AEC. We present extensive evaluation results using both simulated data and real recordings, covering various acoustic conditions and evaluation metrics. The results show the effectiveness of the proposed methods for two different model architectures. Our best joint PSE-AEC model comes close to the expert models optimized for individual tasks of PSE and AEC in their respective scenarios and significantly outperforms the expert models for the combined PSE-AEC task.

eess.AS↗