Searcharxiv⌕ Search

arXiv subjects

Min Tang

Publications and source records attributed to Min Tang.

At least 73 records · Page 4Linked to original sources

Transition behavior of the waiting time distribution in a jumping model with the internal state

It has been noticed that when the waiting time distribution exhibits a transition from an intermediate time power law decay to a long-time exponential decay in the continuous time random walk model, a transition from anomalous diffusion to normal diffusion can be observed at the population level. However, the mechanism behind the transition of waiting time distribution is rarely studied. In this paper, we provide one possible mechanism to explain the origin of such transition. A jump model terminated by a state-dependent Poisson clock is studied by a formal asymptotic analysis for the time evolutionary equation of its probability density function. The waiting time behavior under a more relaxed setting can be rigorously characterized by probability tools. Both approaches show the transition phenomenon of the waiting time $T$, which is further verified numerically by particle simulations. Our results indicate that a small drift and strong noise in the state equation and a stiff response in the Poisson rate are crucial to the transitional phenomena.

math.AP↗

ICASSP 2023 Deep Noise Suppression Challenge

Deep Speech Enhancement Challenge is the 5th edition of deep noise suppression (DNS) challenges organized at ICASSP 2023 Signal Processing Grand Challenges. DNS challenges were organized during 2019-2023 to stimulate research in deep speech enhancement (DSE). Previous DNS challenges were organized at INTERSPEECH 2020, ICASSP 2021, INTERSPEECH 2021, and ICASSP 2022. From prior editions, we learnt that improving signal quality (SIG) is challenging particularly in presence of simultaneously active interfering talkers and noise. This challenge aims to develop models for joint denosing, dereverberation and suppression of interfering talkers. When primary talker wears a headphone, certain acoustic properties of their speech such as direct-to-reverberation (DRR), signal to noise ratio (SNR) etc. make it possible to suppress neighboring talkers even without enrollment data for primary talker. This motivated us to create two tracks for this challenge: (i) Track-1 Headset; (ii) Track-2 Speakerphone. Both tracks has fullband (48kHz) training data and testset, and each testclips has a corresponding enrollment data (10-30s duration) for primary talker. Each track invited submissions of personalized and non-personalized models all of which are evaluated through same subjective evaluation. Most models submitted to challenge were personalized models, same team is winner in both tracks where the best models has improvement of 0.145 and 0.141 in challenge's Score as compared to noisy blind testset.

cs.SD↗

Pattern formation of a pathway-based diffusion model: linear stability analysis and an asymptotic preserving method

We investigate the linear stability analysis of a pathway-based diffusion model (PBDM), which characterizes the dynamics of the engineered Escherichia coli populations [X. Xue and C. Xue and M. Tang, P LoS Computational Biology, 14 (2018), pp. e1006178]. This stability analysis considers small perturbations of the density and chemical concentration around two non-trivial steady states, and the linearized equations are transformed into a generalized eigenvalue problem. By formal analysis, when the internal variable responds to the outside signal fast enough, the PBDM converges to an anisotropic diffusion model, for which the probability density distribution in the internal variable becomes a delta function. We introduce an asymptotic preserving (AP) scheme for the PBDM that converges to a stable limit scheme consistent with the anisotropic diffusion model. Further numerical simulations demonstrate the theoretical results of linear stability analysis, i.e., the pattern formation, and the convergence of the AP scheme.

math.NA↗

Real-Time Audio-Visual End-to-End Speech Enhancement

Audio-visual speech enhancement (AV-SE) methods utilize auxiliary visual cues to enhance speakers' voices. Therefore, technically they should be able to outperform the audio-only speech enhancement (SE) methods. However, there are few works in the literature on an AV-SE system that can work in real time on a CPU. In this paper, we propose a low-latency real-time audio-visual end-to-end enhancement (AV-E3Net) model based on the recently proposed end-to-end enhancement network (E3Net). Our main contribution includes two aspects: 1) We employ a dense connection module to solve the performance degradation caused by the deep model structure. This module significantly improves the model's performance on the AV-SE task. 2) We propose a multi-stage gating-and-summation (GS) fusion module to merge audio and visual cues. Our results show that the proposed model provides better perceptual quality and intelligibility than the baseline E3net model with a negligible computational cost increase.

eess.AS↗

A fully asymptotic preserving decomposed multi-group method for the frequency-dependent radiative transfer equations

The opacity of FRTE depends on not only the material temperature but also the frequency, whose values may vary several orders of magnitude for different frequencies. The gray radiation diffusion and frequency-dependent diffusion equations are two simplified models that can approximate the solution to FRTE in the thick opacity regime. The frequency discretization for the two limit models highly affects the numerical accuracy. However, classical frequency discretization for FRTE considers only the absorbing coefficient. In this paper, we propose a new decomposed multi-group method for frequency discretization that is not only AP in both gray radiation diffusion and frequency-dependent diffusion limits, but also the frequency discretization of the limiting models can be tuned. Based on the decomposed multi-group method, a full AP scheme in frequency, time, and space is proposed. Several numerical examples are used to verify the performance of the proposed scheme.

math.NA↗

Exploring WavLM on Speech Enhancement

There is a surge in interest in self-supervised learning approaches for end-to-end speech encoding in recent years as they have achieved great success. Especially, WavLM showed state-of-the-art performance on various speech processing tasks. To better understand the efficacy of self-supervised learning models for speech enhancement, in this work, we design and conduct a series of experiments with three resource conditions by combining WavLM and two high-quality speech enhancement systems. Also, we propose a regression-based WavLM training objective and a noise-mixing data configuration to further boost the downstream enhancement performance. The experiments on the DNS challenge dataset and a simulation dataset show that the WavLM benefits the speech enhancement task in terms of both speech quality and speech recognition accuracy, especially for low fine-tuning resources. For the high fine-tuning resource condition, only the word error rate is substantially improved.

eess.AS↗

Tumor boundary instability induced by nutrient consumption and supply

We investigate the tumor boundary instability induced by nutrient consumption and supply based on a Hele-Shaw model derived from taking the incompressible limit of a cell density model. We analyze the boundary stability/instability in two scenarios: 1) the front of the traveling wave; 2) the radially symmetric boundary. In each scenario, we investigate the boundary behaviors under two different nutrient supply regimes, in vitro, and in vivo. Our main conclusion is that for either scenario, the in vitro regime always stabilizes the tumor's boundary regardless of the nutrient consumption rate. However, boundary instability may occur when the tumor cells aggressively consume nutrients, and the nutrient supply is governed by the in vivo regime.

math.AP↗

N-Cloth: Predicting 3D Cloth Deformation with Mesh-Based Networks

We present a novel mesh-based learning approach (N-Cloth) for plausible 3D cloth deformation prediction. Our approach is general and can handle cloth or obstacles represented by triangle meshes with arbitrary topologies. We use graph convolution to transform the cloth and object meshes into a latent space to reduce the non-linearity in the mesh space. Our network can predict the target 3D cloth mesh deformation based on the initial state of the cloth mesh template and the target obstacle mesh. Our approach can handle complex cloth meshes with up to 100K triangles and scenes with various objects corresponding to SMPL humans, non-SMPL humans or rigid bodies. In practice, our approach can be used to generate plausible cloth simulation at 30-45 fps on an NVIDIA GeForce RTX 3090 GPU. We highlight its benefits over prior learning-based methods and physically-based cloth simulators.

cs.GR↗

Reconstructing Recognizable 3D Face Shapes based on 3D Morphable Models

Many recent works have reconstructed distinctive 3D face shapes by aggregating shape parameters of the same identity and separating those of different people based on parametric models (e.g., 3D morphable models (3DMMs)). However, despite the high accuracy in the face recognition task using these shape parameters, the visual discrimination of face shapes reconstructed from those parameters is unsatisfactory. The following research question has not been answered in previous works: Do discriminative shape parameters guarantee visual discrimination in represented 3D face shapes? This paper analyzes the relationship between shape parameters and reconstructed shape geometry and proposes a novel shape identity-aware regularization(SIR) loss for shape parameters, aiming at increasing discriminability in both the shape parameter and shape geometry domains. Moreover, to cope with the lack of training data containing both landmark and identity annotations, we propose a network structure and an associated training strategy to leverage mixed data containing either identity or landmark labels. We compare our method with existing methods in terms of the reconstruction error, visual distinguishability, and face recognition accuracy of the shape parameters. Experimental results show that our method outperforms the state-of-the-art methods.

cs.CV↗

A Spatial-Temporal asymptotic preserving scheme for radiation magnetohydrodynamics in the equilibrium and non-equilibrium diffusion limit

The radiation magnetohydrodynamics (RMHD) system couples the ideal magnetohydrodynamics equations with a gray radiation transfer equation. The main challenge is that the radiation travels at the speed of light while the magnetohydrodynamics changes with the time scale of the fluid. The time scales of these two processes can vary dramatically. In order to use mesh sizes and time steps that are independent of the speed of light, asymptotic preserving (AP) schemes in both space and time are desired. In this paper, we develop an AP scheme in both space and time for the RMHD system. Two different scalings are considered. One results in an equilibrium diffusion limit system, while the other results in a non-equilibrium system. The main idea is to decompose the radiative intensity into three parts, each part is treated differently with suitable combinations of explicit and implicit discretizations guaranteeing the favorable stability conditionand computational efficiency. The performance of the AP method is presented, for both optically thin and thick regions, as well as for the radiative shock problem.

math.NA↗

Multiscale convergence of the inverse problem for chemotaxis in the Bayesian setting

Chemotaxis describes the movement of an organism, such as single or multi-cellular organisms and bacteria, in response to a chemical stimulus. Two widely used models to describe the phenomenon are the celebrated Keller-Segel equation and a chemotaxis kinetic equation. These two equations describe the organism movement at the macro- and mesoscopic level respectively, and are asymptotically equivalent in the parabolic regime. How the organism responds to a chemical stimulus is embedded in the diffusion/advection coefficients of the Keller-Segel equation or the turning kernel of the chemotaxis kinetic equation. Experiments are conducted to measure the time dynamics of the organisms' population level movement when reacting to certain stimulation. From this one infers the chemotaxis response, which constitutes an inverse problem. \\ In this paper we discuss the relation between both the macro- and mesoscopic inverse problems, each of which is associated to two different forward models. The discussion is presented in the Bayesian framework, where the posterior distribution of the turning kernel of the organism population is sought after. We prove the asymptotic equivalence of the two posterior distributions.

math.AP↗

Sphere Face Model:A 3D Morphable Model with Hypersphere Manifold Latent Space

3D Morphable Models (3DMMs) are generative models for face shape and appearance. However, the shape parameters of traditional 3DMMs satisfy the multivariate Gaussian distribution while the identity embeddings satisfy the hypersphere distribution, and this conflict makes it challenging for face reconstruction models to preserve the faithfulness and the shape consistency simultaneously. To address this issue, we propose the Sphere Face Model(SFM), a novel 3DMM for monocular face reconstruction, which can preserve both shape fidelity and identity consistency. The core of our SFM is the basis matrix which can be used to reconstruct 3D face shapes, and the basic matrix is learned by adopting a two-stage training approach where 3D and 2D training data are used in the first and second stages, respectively. To resolve the distribution mismatch, we design a novel loss to make the shape parameters have a hyperspherical latent space. Extensive experiments show that SFM has high representation ability and shape parameter space's clustering performance. Moreover, it produces fidelity face shapes, and the shapes are consistent in challenging conditions in monocular face reconstruction.

cs.CV↗

VarArray: Array-Geometry-Agnostic Continuous Speech Separation

Continuous speech separation using a microphone array was shown to be promising in dealing with the speech overlap problem in natural conversation transcription. This paper proposes VarArray, an array-geometry-agnostic speech separation neural network model. The proposed model is applicable to any number of microphones without retraining while leveraging the nonlinear correlation between the input channels. The proposed method adapts different elements that were proposed before separately, including transform-average-concatenate, conformer speech separation, and inter-channel phase differences, and combines them in an efficient and cohesive way. Large-scale evaluation was performed with two real meeting transcription tasks by using a fully developed transcription system requiring no prior knowledge such as reference segmentations, which allowed us to measure the impact that the continuous speech separation system could have in realistic settings. The proposed model outperformed a previous approach to array-geometry-agnostic modeling for all of the geometry configurations considered, achieving asclite-based speaker-agnostic word error rates of 17.5% and 20.4% for the AMI development and evaluation sets, respectively, in the end-to-end setting using no ground-truth segmentations.

eess.AS↗

On kinetic and macroscopic models for the stripe formation in engineered bacterial populations

We study the well-posedness of the biological models with AHL-dependent cell mobility on engineered Escherichia coli populations. For the kinetic model proposed by Xue-Xue-Tang recently, the local existence for large initial data is proved first. Furthermore, the positivity and local conservation laws for density $ρ(t,x,z)$ and nutrient $n(t,x)$ with initial assumptions are justified. Based on these properties, it can be extended globally in time near the equilibrium $(0,0,0)$. Considering the asymptotic behaviors of faster response CheZ turnover rate (i.e.,$\varepsilon\rightarrow 0$), one formally derives an anisotropic diffusion engineered Escherichia coli populations model (in short, AD-EECP) for which we find a key extra a priori estimate to overcome the difficulties coming from the nonlinearity of the diffusion structure. The local well-posedness and the positivity and local conservation laws for density and nutrient of the AD-EECP are justified. Furthermore, the global existence around the steady state $(\varrho_a, h_a, 0)$ with $\varrho_a \in [0, Λ_b)$ is obtained.

math.AP↗

Human Listening and Live Captioning: Multi-Task Training for Speech Enhancement

With the surge of online meetings, it has become more critical than ever to provide high-quality speech audio and live captioning under various noise conditions. However, most monaural speech enhancement (SE) models introduce processing artifacts and thus degrade the performance of downstream tasks, including automatic speech recognition (ASR). This paper proposes a multi-task training framework to make the SE models unharmful to ASR. Because most ASR training samples do not have corresponding clean signal references, we alternately perform two model update steps called SE-step and ASR-step. The SE-step uses clean and noisy signal pairs and a signal-based loss function. The ASR-step applies a pre-trained ASR model to training signals enhanced with the SE model. A cross-entropy loss between the ASR output and reference transcriptions is calculated to update the SE model parameters. Experimental results with realistic large-scale settings using ASR models trained on 75,000-hour data show that the proposed framework improves the word error rate for the SE output by 11.82% with little compromise in the SE quality. Performance analysis is also carried out by changing the ASR model, the data used for the ASR-step, and the schedule of the two update steps.

eess.AS↗

Multi-scale GCN-assisted two-stage network for joint segmentation of retinal layers and disc in peripapillary OCT images

An accurate and automated tissue segmentation algorithm for retinal optical coherence tomography (OCT) images is crucial for the diagnosis of glaucoma. However, due to the presence of the optic disc, the anatomical structure of the peripapillary region of the retina is complicated and is challenging for segmentation. To address this issue, we developed a novel graph convolutional network (GCN)-assisted two-stage framework to simultaneously label the nine retinal layers and the optic disc. Specifically, a multi-scale global reasoning module is inserted between the encoder and decoder of a U-shape neural network to exploit anatomical prior knowledge and perform spatial reasoning. We conducted experiments on human peripapillary retinal OCT images. The Dice score of the proposed segmentation network is 0.820$\pm$0.001 and the pixel accuracy is 0.830$\pm$0.002, both of which outperform those from other state-of-the-art techniques.

eess.IV↗

P-Cloth: Interactive Complex Cloth Simulation on Multi-GPU Systems using Dynamic Matrix Assembly and Pipelined Implicit Integrators

We present a novel parallel algorithm for cloth simulation that exploits multiple GPUs for fast computation and the handling of very high resolution meshes. To accelerate implicit integration, we describe new parallel algorithms for sparse matrix-vector multiplication (SpMV) and for dynamic matrix assembly on a multi-GPU workstation. Our algorithms use a novel work queue generation scheme for a fat-tree GPU interconnect topology. Furthermore, we present a novel collision handling scheme that uses spatial hashing for discrete and continuous collision detection along with a non-linear impact zone solver. Our parallel schemes can distribute the computation and storage overhead among multiple GPUs and enable us to perform almost interactive simulation on complex cloth meshes, which can hardly be handled on a single GPU due to memory limitations. We have evaluated the performance with two multi-GPU workstations (with 4 and 8 GPUs, respectively) on cloth meshes with 0.5-1.65M triangles. Our approach can reliably handle the collisions and generate vivid wrinkles and folds at 2-5 fps, which is significantly faster than prior cloth simulation systems. We observe almost linear speedups with respect to the number of GPUs.

cs.GR↗

On an inverse problem in additive number theory

For a set $A$, let $P(A)$ be the set of all finite subset sums of $A$. In this paper, for a sequence of integers $B=\{1<b_1<b_2<\cdots\}$ and $3b_1+5\leq b_2\leq 6b_1+10$, we determine the critical value for $b_3$ such that there exists an infinite sequence $A$ of positive integers for which $P(A)=\mathbb{N}\setminus B$. This result shows that we partially solve the problem of Fang and Fang [`On an inverse problem in additive number theory', Acta Math. Hungar. 158(2019), 36-39].

math.NT↗