SearcharxivSearch

arXiv subjects

Zhaofeng Lin

Publications and source records attributed to Zhaofeng Lin.

15 recordsLinked to original sources

Candor-LR: A Dyadic Conversational Dataset for Audio-Visual Speech Recognition

Current audio-visual speech recognition (AVSR) benchmarks, like LRS3, rely heavily on clean, scripted and rehearsed speech. They fail to reflect the complexity of natural conversation, which involves overlapping speech, spontaneous turn-taking, unscripted vocabulary and variable acoustic conditions. To shift the field toward realistic dialogue, we introduce Candor-LR, a conversational benchmark derived from the CANDOR corpus of 1,656 natural dyadic videoconferences. Our custom data preparation pipeline yields 713.5, 10.1, and 60.1 hours of training, validation, and test data, respectively. Evaluating pretrained AVSR models on Candor-LR reveals that audio-only accuracy drops sharply compared to LRS3, but visual cues compensate effectively, driving much larger performance gains on Candor-LR than on LRS3. Furthermore, training on this corpus significantly improves cross-domain robustness under both clean and noisy conditions, as its realistic conversational data captures broader audio-video features. We open-source our pipeline to ensure reproducibility, establishing Candor-LR as a challenging benchmark for conversational AVSR.

eess.AS

Assessing True Generalisability of Audio-Visual Speech Recognisers

Current Audio-Visual Speech Recognition (AVSR) models achieve near-perfect performance on the standard LRS3 benchmark, raising concerns of adaptive overfitting. To systematically assess true generalisability, we construct a highly controlled, unseen evaluation set subsampled from the massive MultiVSR dataset. Unlike standard out-of-distribution benchmarks, our subset strictly matches the acoustic, visual, and demographic distributions of the LRS3 test set. Evaluating five state-of-the-art architectures reveals a universal performance collapse, proving that current systems fail to generalise even under strictly aligned conditions. Through a fine-grained attribute analysis across seven factors, we isolate the specific drivers of this degradation. Furthermore, we uncover a profound lexical bias, expose distinct error patterns, and surprisingly reveal that audio-visual performance even lags behind audio-only settings. We release our matched test set for future benchmarking.

eess.AS

Darboux-type formula for Jacobi biorthogonal polynomials

In this paper, we study the asymptotic behavior of Jacobi biorthogonal polynomials. A Darboux-type formula is established using the method of steepest descent. In the proof, we construct an appropriate contour to apply the Rodrigues formula. Our result reduces to the classical Darboux formula in the orthogonal case.

math.CA

Normal approximation for iterated inner functions

A Berry--Ess\'{e}en theorem for linear combinations of iterates of an inner function is obtained. Our proof, which is based an elementary transfer argument and classical results in martingale theory, also leads to a simple proof of Nicolau and Soler i Gibert's central limit theorem for inner functions.

math.PR

Exact values of Fourier dimensions of Gaussian multiplicative chaos on high dimensional torus

We determine the exact values of the Fourier dimensions for Gaussian Multiplicative Chaos measures on the $d$-dimensional torus $\mathbb{T}^d$ for all integers $d \ge 1$. This resolves a problem left open in previous works [LQT24,LQT25] for high dimensions $d\ge 3$. The proof relies on a new construction of log-correlated Gaussian fields admitting specific decompositions into smooth processes with high regularity. This construction enables a multi-resolution analysis to obtain sharp local estimates on the measure's Fourier decay. These local estimates are then integrated into a global bound using Pisier's martingale type inequality for vector-valued martingales.

math.PR

Harmonic analysis of multiplicative chaos Part II: a unified approach to Fourier dimensions

We introduce a unified approach for studying the polynomial Fourier decay of classical multiplicative chaos measures. As consequences, we obtain the precise Fourier dimensions for multiplicative chaos measures arising from the following key models: the sub-critical 1D and 2D GMC (which in particular resolves the Garban-Vargas conjecture); the sub-critical $d$-dimensional GMC with $d \ge 3$ when the parameter $\gamma$ is near the critical value; the canonical Mandelbrot random coverings; the canonical Mandelbrot cascades. For various other models, we establish the non-trivial lower bounds of the Fourier dimensions and in various cases we conjecture that they are all optimal and provide the exact values of Fourier dimensions.

math.PR

Uncovering the Visual Contribution in Audio-Visual Speech Recognition

Audio-Visual Speech Recognition (AVSR) combines auditory and visual speech cues to enhance the accuracy and robustness of speech recognition systems. Recent advancements in AVSR have improved performance in noisy environments compared to audio-only counterparts. However, the true extent of the visual contribution, and whether AVSR systems fully exploit the available cues in the visual domain, remains unclear. This paper assesses AVSR systems from a different perspective, by considering human speech perception. We use three systems: Auto-AVSR, AVEC and AV-RelScore. We first quantify the visual contribution using effective SNR gains at 0 dB and then investigate the use of visual information in terms of its temporal distribution and word-level informativeness. We show that low WER does not guarantee high SNR gains. Our results suggest that current methods do not fully exploit visual information, and we recommend future research to report effective SNR gains alongside WERs.

eess.AS

Harmonic analysis of multiplicative chaos Part I: the proof of Garban-Vargas conjecture for 1D GMC

In this paper, we establish the exact Fourier dimensions of all standard sub-critical Gaussian multiplicative chaos on the unit interval, thereby confirming the Garban-Vargas conjecture. The proof relies on a significant improvement of the vector-valued martingale method, initially developed by Chen-Han-Qiu-Wang in the studies of the Fourier dimensions of Mandelbrot cascade random measures.

math.PR

Number rigid determinantal point processes induced by generalized Cantor sets

We consider the Ghosh-Peres number rigidity of translation-invariant determinantal point processes on the real line $\mathbb{R}$, whose correlation kernels are induced by the Fourier transform of the indicators of generalized Cantor sets in the unit interval. Our main results show that for any given $\theta\in(0,1)$, there exists a generalized Cantor set with Lebesgue measure $\theta$, such that the corresponding determinantal point process is Ghosh-Peres number rigid.

math.PR

A Law of large numbers for vector-valued linear statistics of Bergman DPP

We establish a law of large numbers for a certain class of vector-valued linear statistics for the Bergman determinantal point process on the unit disk. Our result seems to be the first LLN for vector-valued linear statistics in the setting of determinantal point processes. As an application, we prove that, for almost all configurations $X$ with respect to with respect to the Bergman determinantal point process, the weighted Poincar\'e series (we denote by $d_{h}(\cdot,\cdot)$ the hyperbolic distance on $\mathbb{D}$) \begin{align*} \sum_{k=0}^\infty\sum_{x\in X\atop k\le d_{h}(z,x)<k+1}e^{-sd_{\mathrm{h}}(z,x)}f(x) \end{align*} cannot be simultaneously convergent for all Bergman functions $f\in A^2(\mathbb{D})$ whenever $1<s<3/2$. This confirms a result announced without proof in Bufetov-Qiu's work.

math.PR

Improving Whispered Speech Recognition Performance using Pseudo-whispered based Data Augmentation

Whispering is a distinct form of speech known for its soft, breathy, and hushed characteristics, often used for private communication. The acoustic characteristics of whispered speech differ substantially from normally phonated speech and the scarcity of adequate training data leads to low automatic speech recognition (ASR) performance. To address the data scarcity issue, we use a signal processing-based technique that transforms the spectral characteristics of normal speech to those of pseudo-whispered speech. We augment an End-to-End ASR with pseudo-whispered speech and achieve an 18.2% relative reduction in word error rate for whispered speech compared to the baseline. Results for the individual speaker groups in the wTIMIT database show the best results for US English. Further investigation showed that the lack of glottal information in whispered speech has the largest impact on whispered speech ASR performance.

eess.AS

Truncations of random unitary matrices drawn from Hua-Pickrell distribution

Let $U$ be a random unitary matrix drawn from the Hua-Pickrell distribution $\mu_{\mathrm{U}(n+m)}^{(\delta)}$ on the unitary group $\mathrm{U}(n+m)$. We show that the eigenvalues of the truncated unitary matrix $[U_{i,j}]_{1\leq i,j\leq n}$ form a determinantal point process $\mathscr{X}_n^{(m,\delta)}$ on the unit disc $\mathbb{D}$ for any $\delta\in\mathbb{C}$ satisfying $\mathrm{Re}\,\delta>-1/2$. We also prove that the limiting point process taken by $n\to\infty$ of the determinantal point process $\mathscr{X}_n^{(m,\delta)}$ is always $\mathscr{X}^{[m]}$, independent of $\delta$. Here $\mathscr{X}^{[m]}$ is the determinantal point process on $\mathbb{D}$ with weighted Bergman kernel \begin{equation*} \begin{split} K^{[m]}(z,w)=\frac{1}{(1-z\overline w)^{m+1}} \end{split} \end{equation*} with respect to the reference measure $d\mu^{[m]}(z)=\frac{m}{\pi}(1-|z|)^{m-1}d\sigma(z)$, where $d\sigma(z)$ is the Lebesgue measure on $\mathbb{D}$.

math.PR

Fluctuations of the process of moduli for the Ginibre and hyperbolic ensembles

We investigate the point process of moduli of the Ginibre and hyperbolic ensembles. We show that far from the origin and at an appropriate scale, these processes exhibit Gaussian and Poisson fluctuations. Among the possible Gaussian fluctuations, we can find white noise but also fluctuations with non-trivial covariance at a particular scale.

math.PR

Gaussian limit for determinantal point processes with $J$-Hermitian kernels

We show that the central limit theorem for linear statistics over determinantal point processes with $J$-Hermitian kernels holds under fairly general conditions. In particular, We establish Gaussian limit for linear statistics over determinantal point processes on union of two copies of $\mathbb{R}^d$ when the correlation kernels are $J$-Hermitian translation-invariant.

math.PR