Searcharxiv⌕ Search

arXiv subjects

Kong Aik Lee

Publications and source records attributed to Kong Aik Lee.

At least 19 recordsLinked to original sources

Source-Directed Trajectory Perturbation at First-Order Cost for Domain Generalization in Speech Deepfake Detection

Speech deepfake detectors often lose accuracy when the distribution of the test data differs from that of the training data. Meta-learning for domain generalization (MLDG) shows promise by simulating domain shifts with episodic meta-train and meta-test splits. However, the MLDG meta-objective considers only a single clean adaptation trajectory and does not account for the local loss landscape around its endpoints. To address this limitation, we first propose an explicit worst-case MLDG variant, dubbed WC-MLDG-4P, which perturbs both the meta-train and meta-test states but requires four gradient evaluations per episode. We then introduce WC-MLDG-2P, a source-directed alternative to the explicit robust objective. It shifts the clean MLDG endpoint toward a locally higher source-loss state and evaluates the meta-test gradient there, retaining the two-evaluation cost of the first-order MLDG. A first-order expansion relates the resulting gradient change to meta-test directional curvature along the source gradient without explicitly computing a Hessian. Relative to MLDG, WC-MLDG-2P achieves relative mean-EER reductions of 23.4% with XLSR-AASIST and 7.6% with XLSR-Conformer-TCM, respectively.

eess.AS↗

DuRe-ST: Dual-Relation Spectro-Temporal Modeling for Speech Deepfake Detection

Previous speech deepfake detectors can adaptively capture spectro-temporal dependencies through graph attention, yet they largely overlook the co-variation between spectral and temporal representations. To address this gap, we construct a normalized affinity graph from their joint covariance and apply polynomial graph filtering to capture higher-order covariance-induced dependencies. We first develop Cov-ST to isolate the contribution of covariance-based relational modeling. Although it improves detection performance, its sensitivity to the polynomial order suggests limited robustness when covariance relations are modeled alone. We therefore propose DuRe-ST, which jointly exploits covariance-induced and graph-attention-induced relations to capture complementary second-order co-variation and adaptive spectro-temporal dependencies. Experiments show that DuRe-ST achieves an average relative EER reduction of 25.9% over XLSR-AASIST on the ASVspoof benchmarks and 28.4% across four cross-dataset benchmarks with only 4-8k additional trainable back-end parameters. It further outperforms the strongest publicly available comparison models by 2.2-13.2% in relative EER on four benchmarks, while remaining smaller than the publicly available models considered.

eess.AS↗

TS-SP: Learning Speaker-Preserving Representations in Audio Large Language Models

Audio large language models (ALLMs) can understand speech content, yet their ability to use speaker identity for verification remains limited. We propose TS-SP (Two-Stage Speaker Preservation), a parameter-efficient framework for learning speaker-preserving representations and making them accessible to an ALLM's language-model component. We instantiate and evaluate TS-SP on Qwen2.5-Omni-7B. First, we adapt the audio encoder with speaker identity supervision. We then freeze the adapted encoder and train the language model to compare speakers. Both stages use low-rank adaptation (LoRA), keeping the pretrained base weights fixed. On Vox1-O, TS-SP reduces the equal error rate (EER) from 7.01\% for the Paired Loss Adaptation Baseline to 4.37\%. EER remains within 4.31--4.79\% under unseen prompts. Cross-domain evaluation on CN-Celeb yields a similar EER to the baseline, but lower accuracy at the native decision threshold. These findings support two-stage adaptation for improving speaker verification on the evaluated backbone.

cs.SD↗

Learning as Deepfakes Evolve: RF-Prompt for Continual Audio Deepfake Detection

Continual audio deepfake detection requires learning newly emerging deepfake methods while retaining discrimination of previously encountered speech. Existing dataset-incremental evaluation changes both real-speech domains and deepfake mechanisms, making their effects difficult to distinguish. We construct five task organizations over identical training, development, and evaluation pools to study these factors under a controlled sample budget. Our proposed Real-Anchored Mechanism-Incremental (RAMI) protocol reflects the practical setting in which available real speech provides a recurring mixed-domain reference while new deepfake mechanisms arrive incrementally. We further propose RF-Prompt, an asymmetric continual prompt-learning method that preserves reusable real-speech knowledge through a shared real prompt and expands mechanism-specific knowledge through inherited fake experts with orthogonal residuals. Input-adaptive soft fusion combines the accumulated experts into a fixed number of injected tokens without requiring task identity at inference. On RAMI, RF-Prompt achieves 10.110% average EER and 10.370% pooled EER, outperforming all evaluated continual-learning baselines. Across the five controlled protocols, RAMI yields the lowest common-average and pooled EER. Component ablations, limited-data experiments, and cross-backbone evaluations further validate the proposed design.

cs.SD↗

Cross-modal Translation via Conditional Latent Denoising for Video Deepfake Detection

The growing threat of video deepfakes necessitates multimodal detection. Beyond serving as independent indicators of authenticity, audio and visual signals have intrinsic dependencies that also provide an essential criterion for detection. Previous methods often overlook the cross-modal correspondences, hindering information transfer between domains and leaving crucial detection cues unexplored. To address this challenge, we propose a framework called Cross-modal Translation via Conditional Latent Denoising (CTCLD) for video deepfake detection. It connects the distinct distributions of heterogeneous modalities in latent spaces, enabling smooth cross-domain information transfer to improve detection performance. We first establish a Bayesian foundation by decomposing the audio-visual joint distribution. Subsequently, CTCLD translates both modalities via bidirectional latent denoising conditioned on each other, effectively capturing subtle inconsistencies in the manipulated signals. Experimental results demonstrate that the proposed CTCLD enables comprehensive domain alignment, resulting in a robust video deepfake detection approach with competitive performance.

cs.AI↗

DGS-MLDG: Domain Gradient Surgery Guided Meta-Learning for Domain Generalization in Speech Deepfake Detection

Speech deepfake detection faces significant challenges due to domain shifts. Domain generalization (DG), particularly meta-learning for domain generalization (MLDG), offers a promising solution by simulating and mitigating domain shifts. However, MLDG is often hindered by conflicting gradients between its meta-train and meta-test objectives, leading to suboptimal performance. To address this problem, we propose domain gradient surgery (DGS), a meta-learning method that resolves conflicts through an asymmetric projection strategy. DGS removes the destructive component from the meta-test gradient, ensuring a conflict-free optimization trajectory versus the meta-train gradient. Furthermore, we introduce layer-wise DGS (LW-DGS), an efficient variant of DGS that dynamically identifies and intervenes only conflict-prone layers. Extensive experiments on challenging benchmarks demonstrate that DGS-MLDG and LW-DGS-MLDG achieve an average relative EER reduction of 5.29% and 4.04%, respectively.

eess.AS↗

Domain-Adaptive Dual-Gating Mixture of Experts for Generalizable Speech Deepfake Detection

Recent advances in speech deepfake detection (SDD) have leveraged the Mixture of Experts (MoE) to enhance generalization capacity. However, existing gating networks often overlook the acoustic and temporal cues of deepfakes. In this work, we propose a novel domain-adaptive dual-gating MoE (DADGMoE) framework for SDD under unseen attack types and acoustic conditions. Our innovative dual-gating mechanism leverages Sinc-layer-based filters to process both low-level acoustic signals (raw waveforms) and high-level speech representations from a large self-supervised learning (SSL) model. It further incorporates domain prototypes to guide expert routing based on implicit deepfake patterns. The lightweight affine experts process the routed inputs. Experiments show that our DADGMoE significantly outperforms the baseline, achieving up to a 40.8% relative EER reduction on challenging out-of-dataset benchmarks. This framework demonstrates superior generalization capabilities and efficient design.

eess.AS↗

Boundary and Intra-Segment Learning for Partial Audio Deepfake Localization

Partial audio deepfakes manipulate only selected speech regions, making them difficult to be localized. Existing methods exploit boundary cues for partial deepfake localization, but primarily focus on identifying boundary positions rather than modeling the feature changes that characterize authenticity transitions. Meanwhile, the internal characteristics of continuous bona fide and spoofed segments remain underexplored. In this paper, we propose Boundary and Intra-Segment Learning (BISL), which introduces boundary learning to model feature differences between adjacent frames and distinguish authenticity transitions from general acoustic variations. In addition, intra-segment learning captures the overall characteristics of continuous bona fide and spoofed segments while enhancing feature consistency within each segment. By jointly learning frame, boundary, and segment information, BISL enables more effective fine-grained partial audio deepfake localization. Experiments on multiple localization benchmarks show that BISL achieves an EER of 2.52\% and an F1-score of 97.40\% on PartialSpoof, outperforming the compared methods, while maintaining competitive performance on HAD and improved cross-dataset performance on LPS. The code will be made publicly available upon acceptance.

cs.SD↗

Reducing Speaker Residual by Considering Pinhole Effect in Voice Anonymization

Voice anonymization aims to protect privacy by suppressing speaker identity while preserving linguistic content and prosody. However, residual speaker attributes in non-identity representations may still increase linkability and weaken privacy protection. To this end, this paper proposes a fine-tuning strategy with a pinhole loss for well-trained voice anonymization frameworks to further reduce residual speaker attributes. Inspired by the pinhole effect, the pinhole loss measures the linkability of anonymized utterances from the same source speaker. By minimizing this loss, linkability is reduced, thereby improving privacy protection. Experiments on multiple anonymization frameworks, pseudo-speaker generation methods, and datasets show improved privacy protection while maintaining utility. Audio samples can be found in https://anonymous.4open.science/r/Pinhole-loss-fine-tunning-4628.

eess.AS↗

A Unified Uncertainty-Aware Back-End for Speaker Verification: Scoring, Normalization, and Calibration

Speaker verification back-ends commonly combine similarity scoring, score normalization, and calibration. However, speaker embeddings extracted from real-world utterances have trial-dependent reliability because of factors such as duration, noise, and channel variation. Existing uncertainty-aware methods primarily improve the speaker encoder or the initial similarity score, while the estimated uncertainty is typically not propagated through subsequent normalization and calibration. We represent each utterance by a speaker embedding, interpreted as a posterior mean, together with its covariance as an uncertainty estimate. We present a unified uncertainty-aware back-end comprising uncertainty-aware cosine scoring, uncertainty-aware AS-Norm (UAS-Norm), and uncertainty-aware Quality Measure Function calibration (UQMF). Covariance information is incorporated throughout this pipeline to adjust score scaling, cohort statistics, normalized-score combination, and calibration features. Experiments with ECAPA-TDNN and ResNet show consistent EER reductions and improved target--non-target separation across both architectures.

cs.SD↗

Not All Attacks Are Learned Equally in Speech Deepfake Detection

Speech deepfake detection (SDD) models are trained on multi-attack datasets containing diverse spoofing systems, such as text-to-speech (TTS) and voice conversion (VC). In standard classifier training on multi-attack datasets, all attacks are treated as one spoofed class, and performance is reported using overall Equal Error Rate (EER). This aggregate view obscures how individual attacks shape learning and generalization. To better understand this attack-level behavior, we first balance TTS and VC exposure using sample and attack omission. We then measure attack-wise EER at inference and analyze attack-wise training loss and predictive entropy to characterize optimization. Results show that attacks contribute unequally: some attacks have high EER sensitivity and concentrated entropy with low loss, indicating strong influence on the decision boundary. We define these as high-impact attacks. To reduce uneven generalization across attacks, we propose a replay-regularized, attack-aware curriculum that steps exposure based on measured attack influence. Experiments on ASVspoof 2019, 2021, ASVspoof 5, and Fake-or-Real show improved overall robustness and reduced attack-level imbalance compared with standard multi-attack training.

eess.AS↗

BiMTokenizer: Preserving Semantic-Acoustic Balance in Low-Bitrate Speech Tokenization via Bidirectional State-Space Modeling

Speech codecs serve as bridges between continuous speech signals and large language models, yet face an inherent conflict between acoustic fidelity and semantic preservation. To mitigate this conflict, recent works increasingly adopt dual-tower architectures to decouple semantic and acoustic modeling with separate encoders. However, these dual-tower designs incur substantial architectural overhead. To avoid such complexity, we revisit the single-tower paradigm and propose BiMTokenizer, a low-bitrate speech codec (around 1.1 kbps) combining a bidirectional state-space backbone with Residual Spherical Leech Quantization (RSLQ). The bidirectional backbone strengthens temporal modeling, while RSLQ offers a fixed, well-separated lattice bottleneck for robust semantic and acoustic tokenization without learned-codebook collapse. Experiments show that BiMTokenizer achieves superior acoustic reconstruction and the lowest WER among low-bitrate codec baselines across both clean and noisy environments, while using less than half the parameters of recent dual-tower baselines. Furthermore, its robust semantic representations yield strong performance on downstream speech understanding tasks, confirming that a well-designed single-tower codec can preserve the semantic-acoustic balance at low bitrates. The code and model weights are available at https://github.com/ZhangXinWhut/BiMTokenizer.

cs.SD↗

Learning to Prefer Reliably: Error-Augmented Emotion Preference Optimization with Calibrated Fusion

Emotion preference learning uses pairwise comparisons between candidate descriptions to align multimodal large language models (MLLMs) with human judgments of open-ended emotion descriptions and to train reward models that capture human emotional preferences. However, conventional pairwise supervision is often sparse, typically providing only a single negative description for each positive description, and therefore offers limited coverage of the diverse ways in which an emotion description can be incorrect. In particular, models may be insufficiently exposed to semantically fluent but emotionally inconsistent descriptions. Beyond this data-level limitation, relying on a single MLLM judge introduces a distinct model-level concern: its judgments can be affected by model-specific biases when interpreting fine-grained or ambiguous multimodal emotional cues. To address these limitations, we propose Error-Augmented Preference Optimization (EAPO), a framework for improving the reliability of MLLM-based emotion preference judgment at both the data and model levels. First, we construct an error-augmented dataset by generating multiple controlled and emotion-aware negative descriptions from each preferred description. We then adapt multiple independent MLLM judges to this richer supervision and aggregate their preference margins using margin-calibrated soft fusion, which maps heterogeneous margins to a common scale before aggregation. Experiments on the MER2026-EmoPrefer Challenge dataset and our error-augmented dataset demonstrate that EAPO improves emotion preference prediction and enhances the robustness of MLLM judges when evaluating fluent descriptions that conflict with the video's multimodal emotional evidence. Our code is available at https://github.com/slash1028/EAPO-EmoPrefer.

cs.MM↗

MeMo: Attentional Momentum for Real-Time Audio-Visual Target Speaker Extraction Under Impaired Visual Conditions

Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate a target speaker's voice from multi-speaker environments by leveraging visual cues as guidance. However, the performance of AV-TSE systems heavily relies on the quality of these visual cues. In extreme scenarios where visual cues are missing or severely degraded, the system may fail to accurately extract the target speaker. In contrast, humans can maintain attention on a target speaker even in the absence of explicit auxiliary information. Motivated by such human cognitive ability, we propose a novel framework called MeMo, which incorporates two adaptive memory banks to store attention-related information. MeMo is specifically designed for real-time scenarios: once initial attention is established, the system maintains attentional momentum over time, even when visual cues become unavailable. We conduct comprehensive experiments to verify the effectiveness of MeMo. Experimental results demonstrate that our proposed framework achieves SI-SNR improvements of at least 2 dB over the corresponding baseline.

cs.SD↗

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks

Speech deepfake detection (SDD) systems achieve strong performance on conventional benchmarks; however, existing datasets provide limited coverage of emotionally expressive and recent large audio-language model (LALM)-based attacks. Existing emotional spoofing datasets are also limited in scale and attack diversity, typically covering only voice conversion (VC) or text-to-speech (TTS) attacks. We introduce AffectDF, the most comprehensive benchmark for emotionally expressive speech deepfakes, spanning TTS, VC, emotional VC, and LALM-based spoofing attacks across both acted and spontaneous emotional speech. AffectDF contains approximately 260 hours of speech generated using 21 spoofing attacks across five emotional states. We benchmark state-of-the-art SDD systems under conventional and emotional spoofing conditions, including LALM-based detectors evaluated with both inference-only prompting and supervised fine-tuning. Our experiments reveal severe robustness degradation when models trained on conventional benchmarks are evaluated on AffectDF, with several systems approaching near-random performance. Surprisingly, even large-scale emotional training does not consistently improve cross-domain robustness, indicating that current SDD systems fail to learn generalized spoof representations under emotional and prosodic variability. Robustness further varies substantially across emotional states, attack families, and acted vs spontaneous emotional speech conditions. These findings expose fundamental limitations of current SDD systems and establish AffectDF as a benchmark for developing more robust spoof detection models.

eess.AS↗

Beyond Residual Connections: Manifold-Constrained Hyper-Connections for Robust Speaker Representation Learning

Residual connections are fundamental to deep speaker recogni- tion models, such as ECAPA-TDNN and ResNet. However, standard identity mapping limits information flow to a sin- gle path, constraining representation capacity. We introduce Manifold-Constrained Hyper-Connections (mHC), reformulat- ing residual paths as a multi-stream evolution where informa- tion is mixed through a doubly stochastic matrix. By employing Sinkhorn-Knopp iterations, mHC ensures energy conservation by preserving signal intensity and feature mean, which stabi- lizes gradients and mitigates signal degradation in complex net- works. We evaluate mHC by replacing standard residual con- nections in backbones including ECAPA-TDNN, ResNet-34, Res2Net, and E-Res2Net. Extensive experiments on VoxCeleb1 demonstrate that mHC connections consistently enhance per- formance across all architectures, highlighting its effectiveness for robust speaker representation learning.

cs.SD↗

EII-SCL: Harnessing Emotional Inertia for Multimodal Emotion Recognition in Conversation

Multimodal emotion recognition in conversation (MERC) achieves accurate predictions by integrating multimodal and contextual information in dialogues. While current MERC approaches focus on modeling complex contextual dependencies in conversation, they often overlook the impact of contextual emotional inertia in emotion shift, leading to sub-optimal performance. To address this issue, we propose a novel Emotional Inertia-Informed Supervised Contrastive Learning module (EII-SCL) that informs the contrastive objective by constructing inertia-affected samples within temporal windows, effectively leveraging emotional inertia as a prior while enabling seamless integration with existing MERC models without requiring additional data. Extensive experiments on IEMOCAP and MELD show that our approach consistently outperforms state-of-the-art methods.

cs.MM↗

EmoEUS: Uncertainty Supervision for Multimodal Emotion Recognition in Conversation

Multimodal emotion recognition in conversation (MERC) can leverage multimodal and contextual cues to boost recognition performance. However, existing fusion approaches in MERC often ignore modality-specific uncertainty across utterances caused by conflicting cues, varying noise, and missing modality-specific signals. We propose EmoEUS, an explicit uncertainty supervision framework for MERC. EmoEUS performs uncertainty-aware multimodal fusion by dynamically weighting modalities using learned variance estimates. We also introduce an explicitly supervised loss that aligns each utterance's predicted variance with the distance between the utterance's distributional representation and its emotion- and modality-specific cluster center. Experiments on IEMOCAP and MELD show that EmoEUS consistently outperforms state-of-the-art methods.

cs.MM↗