SearcharxivSearch

arXiv subjects

Yuhong Yang

Publications and source records attributed to Yuhong Yang.

At least 19 recordsLinked to original sources

ProLombard: Structured Multi-Scale Modeling for Normal-to-Lombard Speech Conversion

Normal-to-Lombard (N2L) speech conversion aims to improve speech intelligibility in noisy environments by transforming normal speech into Lombard-style speech while preserving linguistic content, speaker identity, and speech quality. Despite recent progress, existing methods typically model the Lombard effect at the utterance level or the frame level, overlooking its hierarchical nature and its entanglement with both speaker identity and phoneme-level content. This limitation leads to Lombard leakage in speaker representations and incomplete separation between Lombard characteristics and linguistic content. In this work, we propose ProLombard, a structured multi-scale N2L framework that explicitly models the Lombard effect across utterance-, phoneme-, and frame-level representations. To address Lombard-speaker entanglement, we introduce an aligned speaker encoder (ASE) that suppresses Lombard leakage by aligning Lombard-speech speaker embeddings with their normal-speech counterparts. To achieve more complete Lombard-content disentanglement, we develop a phoneme-aware disentanglement and injection mechanism that extends conventional frame-level modeling to the phoneme level. Furthermore, we design a vector quantization (VQ)-median module that provides robust phoneme-level representations through VQ-based segmentation and median-frame-based aggregation. Extensive experiments on Mandarin and English Lombard datasets demonstrate that the proposed approach consistently improves speech intelligibility, Lombard similarity, and perceptual quality over baselines while maintaining speaker identity. These results highlight the importance of structured multi-scale modeling for effective N2L speech conversion.

cs.SD

Towards a Statistical Understanding of Mixture-of-Experts

Mixture-of-experts (MoE) architectures increase model capacity by combining a collection of expert predictors through input-dependent routing, while often activating only a small subset of experts for each input. Despite their growing importance in modern large-scale models, the statistical roles of their design choices, especially routing, sparse activation, and shared experts, remain only partially understood, as existing theory has largely focused on parametric or correctly specified MoE models. In this paper, we view MoE as a form of localized aggregation and show how this localization reshapes the approximation-estimation-computation tradeoff. We derive oracle risk bounds for learning dense and sparse routing with evolving experts, separating approximation, expert-learning, and router-estimation errors, and characterize how sparse Top-K routing can retain the benefits of localized aggregation while controlling per-input computation. We also interpret gating through the geometry of input space, relating routing performance to regions of local expert advantage, and show how shared experts, as adopted in architectures such as DeepSeekMoE, can extract common predictive structure so that routed experts focus on residual local variation. Together, these results provide a unified statistical framework for understanding MoE through input-dependent expert aggregation, in which expert specialization and computational tradeoffs are governed by local predictive structure.

stat.ML

Power and sample size calculations for causal mediation analysis with a binary mediator in randomized trials

Mediation analyses are increasingly conducted in randomized trials, but a sample size adequate for the total treatment effect may leave the natural indirect effect (NIE) or natural direct effect (NDE) substantially underpowered. Randomization does not extend to the mediator, so precision depends on the conditional mediator distribution and the mediator-outcome association, neither of which enters a total-effect calculation. Planning outside linear structural equation models is largely based on simulation under a fully specified data-generating mechanism rarely available at the design stage. This paper develops analytic power and sample size formulas for the NIE and NDE with a binary mediator and a continuous or binary outcome. Under standard identification assumptions, we focus on the ratio-of-mediator-probability weighting (RMPW) estimator that does not require an outcome model for effect estimation. We decompose the oracle variances of the RMPW estimators into components capturing mediator-probability-ratio variability, outcome variation, and their association, with an additional shared-arm covariance term for the NIE. Under a probit latent-index mediator and a working outcome model, these components are determined by a small number of interpretable design inputs rather than by the full joint distribution of covariates, mediator, and outcome. Simulations show that the analytic sample sizes closely match simulation-based benchmarks, attain the target power, and maintain type I error near the nominal level. An ACTG175 illustration shows how pilot data can calibrate the inputs.

stat.ME

A Unified Descriptive-Complexity Framework for Model Selection under Correlated Designs

Model selection becomes particularly challenging under strong predictor dependence and model-class uncertainty, especially when there are exponentially many models. We propose a Descriptive-Complexity Information Criterion (DCIC) that regularizes large candidate model collections through Kraft-admissible code lengths. Under sub-Weibull noise, we establish selection consistency through approximation-error separation without relying on RIP-type conditions, together with nonasymptotic oracle risk bounds that remain valid under model misspecification. The same coding principle places heterogeneous classes on a common complexity scale at a small additional class-identification cost. This extension yields class--model recovery under suitable identifiability conditions and risk adaptation across classes. We further develop a complexity-guided search path that makes the computation--statistics trade-off explicit. Large penalties yield polynomial-size retained search regions with high probability, whereas smaller penalties sharpen the oracle risk benchmark. Numerical experiments illustrate stable support recovery and favorable estimation performance under strong dependence and model-class uncertainty.

stat.ML

Transporting Trial Evidence Under Posterior Drift and Possible Hidden Confounding

Randomized trials provide internally valid treatment-effect evidence, but trial participants may not represent the target population. In contrast, observational studies are often closer to the target population, but their treatment assignment may be affected by possible hidden confounding. We develop a robust posterior-drift framework for estimating the average treatment effect in an observational target population when exact conditional-effect transportability may fail. The framework represents observational conditional potential-outcome regressions as their randomized-trial counterparts plus source-specific drifts. The randomized trial serves as an internally valid anchor, while the observational study supplies the target covariate distribution and partial information about the target causal contrast. To account for possible hidden confounding, we consider a Rosenbaum-type uncertainty set induced by a sensitivity parameter on the generalized propensity score and estimate the drift through a minimax worst-case risk criterion. We derive efficiency results in auxiliary regimes, establish uniform concentration and near-optimality guarantees for the minimax estimator, and handle general parametric and smooth nonparametric drift classes. Simulations and an ACTG 175--WIHS application show that the proposed analysis yields more cautious and interpretable target-population effect estimates than exact-transportability analyses.

stat.ME

Local spectral clustering for heterogeneous clustering structures

Classical clustering methods typically assume that all informative features support a single latent partition of the observations. This assumption can be overly restrictive for modern high-dimensional data, where different subsets of features may encode distinct notions of similarity and induce heterogeneous sample partitions, while some features may contain no meaningful clustering information. We develop a frequentist framework for local clustering that simultaneously identifies feature groups and estimates the sample clustering structure associated with each group. Our approach represents each sample partition by a label-invariant clustering matrix and groups features according to their shared clustering structures, thereby reformulating local clustering as a feature-grouping, or clustering-of-clusterings, problem. Under a heterogeneous sub-Gaussian mixture model, we construct feature-specific Gaussian-kernel similarity matrices and propose a local spectral clustering procedure based on a clustering-matrix optimization criterion. The proposed method avoids explicit likelihood specification and Bayesian posterior computation, accommodates heterogeneous feature distributions, and permits the presence of non-informative features. Extensive simulations and applications further demonstrate the practical utility and superiority of the proposed approach.

stat.ME

Sign Hacking with Auxiliary Variable Exploration in the Age of Big Data

In linear regression, the signs of coefficients convey the direction of covariate effects and are central to empirical interpretation. In high-dimensional settings, however, the abundance of candidate covariates introduces substantial model selection uncertainty. We study the deliberate manipulation of coefficient signs through the inclusion of a carefully chosen auxiliary variable, a practice we term SHAVE (\textit{Sign Hacking with Auxiliary Variable Exploration}). We show that, conditional on the outcome and variables of interest, there exists a set of auxiliary-variable realizations with positive Lebesgue measure that lead to sign reversals upon inclusion. Moreover, with high probability, such variables can be found when many auxiliary candidates are available, leading simultaneously to reversed signs, inflated $t$- and $F$-statistics. Simulation studies and an empirical application corroborate these theoretical findings. We further propose detection strategies for SHAVE when augmented or independent datasets are available, as SHAVE has important implications for reproducibility, $p$-hacking, and research integrity.

stat.ME

Noisy Environment Adaptation of Neural Speech Codec via Focal Mask and Noise Feature Separation

Neural speech codec has attracted extensive attention for high-quality reconstruction at low-bitrate. However, real-world noise severely degrades its performance and hinders high-quality clean speech reconstruction. To tackle this problem, we propose FocalSE, a novel speech enhancement method that performs feature denoising, noise feature separation and noise recognition in the continuous embedding space of neural speech codecs. Specifically, we develop focal modulation-based compression and decompression to capture global context and local mutual information, and generate focal masks to recover clean feature embeddings. We then separate noise embeddings from noisy embeddings to improve denoising performance. Finally, we use ResNet1D-18 to recognize noise categories for better separation effectiveness. Extensive experiments on two standard datasets, LibriTTS and ESC50, demonstrate that our method outperforms state-of-the-art approaches under low-bitrate and low-SNR conditions.

eess.AS

Multi-Armed Bandits with Arriving Arms: Sequential Screening, Dynamic Regret, and Sublinear Guarantees

We study a stochastic multi-armed bandit problem in which the set of available arms expands over time. This setting arises in sequential experimentation when new actions or treatments become available during an ongoing study, making regret against a single best arm in hindsight inappropriate. We instead evaluate performance relative to the best arm currently available, leading to a dynamic-regret criterion for arriving-arm environments. To address the resulting challenges of arrival information discrepancy (AID) and a drifting benchmark (DB), we propose UCB for Arriving Arms (UCB-AA), an elimination-based procedure with an aiding preliminary screening step for newly arrived arms before full competition with incumbent arms. We show that UCB-AA attains regret bounds that depend explicitly on the arrival process, achieves sublinear dynamic regret under regularity conditions on gap evolution, and admits an online extension for unknown horizons. Simulation results show that UCB-AA reduces wasted pulls and maintains a smaller active arm set while preserving competitive regret performance.

stat.ML

Assessing Estimate of CATE from Observational Data via an RCT Study

Conditional average treatment effects (CATEs) are increasingly estimated from observational data and used to guide policy and individualized treatment decisions. Before such estimates can be trusted in practice, their predictive fitness needs to be assessed, yet observational data alone offer limited opportunities for doing so. We propose CATE Assessment via Fitness Evaluation (CAFE), a formal framework for directly assessing the goodness-of-fit of a CATE estimate learned from observational data, rather than the full underlying outcome model, using evidence from a randomized trial. CAFE partitions the trial covariate space according to estimated propensity scores (or the like) and compares observationally derived conditional treatment effects with group-level experimental averages. The framework accommodates a broad class of CATE learners, including parametric models and flexible machine learning methods such as causal forest and boosting. We establish theoretical guarantees under both the null and alternative hypotheses, and introduce a maximum-type extension to improve sensitivity to localized lack of fit. When both randomized trial and observational data are available, we further develop a two-stage procedure to detect the existence of unobserved confounders. Extensive numerical studies show the utility of the CAFE approach when assessing observational-derived CATE estimates.

stat.ME

Combining pre-trained models via localized model averaging

Many pre-trained models (PTMs) are available in modern applications. Because different PTMs are often trained on different datasets, their performances can vary substantially for different new tasks, and the ranking of the candidates may depend heavily on the input. Motivated by this, we propose a localized model averaging method with weights modeled as functions of the covariates, making it substantially more versatile than existing model averaging methods. This formulation allows the model averaging procedure to adaptively capture the varying relative advantages of different PTMs across heterogeneous contexts. Specifically, we learn flexible local weights under a general loss framework that accommodates a broad class of prediction tasks. We further establish the asymptotic optimality of the proposed method for both in-sample and out-of-sample risks, as well as the consistency of the estimated weights. Extensive numerical experiments further demonstrate the effectiveness of the proposed method.

stat.ME

Impossibility of Distribution-Free Predictive Inference for Individual Treatment Effects

Uncertainty quantification for individual treatment effects (ITEs) is a daunting challenge in causal inference. Motivated by recent advances in conformal prediction, several works aim to construct distribution-free prediction sets for ITEs with desired coverage under standard assumptions such as strong ignorability and overlap. In this paper, we show that such goals are fundamentally unattainable in the presence of continuous covariates. Specifically, we establish finite-sample and asymptotic impossibility results demonstrating that any distribution-free prediction set achieving desired coverage for ITEs must be trivial, in the sense that it has infinite expected length. Our analysis relies on a connection between ITE inference and the hardness of conditional independence testing, and highlights the intrinsic limitations imposed by the missing data nature of causal inference. These results provide a new perspective on existing methods, clarifying that their apparent success necessarily relies on additional structural assumptions beyond standard causal assumptions.

stat.ME

Conformal Prediction Assessment: A Framework for Conditional Coverage Evaluation and Selection

Conformal prediction provides rigorous distribution-free finite-sample guarantees for marginal coverage under the assumption of exchangeability, but may exhibit systematic undercoverage or overcoverage for specific subpopulations. Assessing conditional validity is challenging, as standard stratification methods suffer from the curse of dimensionality. We propose Conformal Prediction Assessment (CPA), a framework that reframes the evaluation of conditional coverage as a supervised learning task by training a reliability estimator that predicts instance-level coverage probabilities. Building on this estimator, we introduce the Conditional Validity Index (CVI), which decomposes reliability into safety (undercoverage risk) and efficiency (overcoverage cost). We establish convergence rates for the reliability estimator and prove the consistency of CVI-based model selection. Extensive experiments on synthetic and real-world datasets demonstrate that CPA effectively diagnoses local failure modes and that CC-Select, our CVI-based model selection algorithm, consistently identifies predictors with superior conditional coverage performance.

stat.ME

Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake Detection

Standard supervised training for deepfake detection treats all samples with uniform importance, which can be suboptimal for learning robust and generalizable features. In this work, we propose a novel Tutor-Student Reinforcement Learning (TSRL) framework to dynamically optimize the training curriculum. Our method models the training process as a Markov Decision Process where a ``Tutor'' agent learns to guide a ``Student'' (the deepfake detector). The Tutor, implemented as a Proximal Policy Optimization (PPO) agent, observes a rich state representation for each training sample, encapsulating not only its visual features but also its historical learning dynamics, such as EMA loss and forgetting counts. Based on this state, the Tutor takes an action by assigning a continuous weight (0-1) to the sample's loss, thereby dynamically re-weighting the training batch. The Tutor is rewarded based on the Student's immediate performance change, specifically rewarding transitions from incorrect to correct predictions. This strategy encourages the Tutor to learn a curriculum that prioritizes high-value samples, such as hard-but-learnable examples, leading to a more efficient and effective training process. We demonstrate that this adaptive curriculum improves the Student's generalization capabilities against unseen manipulation techniques compared to traditional training methods. Code is available at https://github.com/wannac1/TSRL.

cs.CV

Cross-Validation in Bipartite Networks

Bipartite networks, which encode interactions between two distinct types of entities, arise widely in applications and exhibit inherent asymmetry across node sets. Despite a growing literature on bipartite community detection, estimating community numbers $(K_1, K_2)$, a critical issue for bipartite network analysis, remains theoretically underdeveloped without any model selection consistency established, to our knowledge. Indeed, the inherent asymmetry and the two-dimensional parameter space with possibly drastically different $K_1$ and $K_2$ pose unique challenges that differ from unipartite cases. In particular, the candidate models may simultaneously overfit one node set while underfitting the other. To address these challenges, we propose Bipartite Cross-Validation (BCV), a penalized cross-validation framework that jointly selects $(K_1,K_2)$ in a fully data-driven manner. We establish the first model selection consistency for bipartite networks, notably accommodating the regime where the numbers of communities scale with the network size, revealing the intricate interplay between sparsity and model complexity. Simulations and real-data applications demonstrate strong finite-sample performance of BCV.

stat.ME

DeformTrace: A Deformable State Space Model with Relay Tokens for Temporal Forgery Localization

Temporal Forgery Localization (TFL) aims to precisely identify manipulated segments in video and audio, offering strong interpretability for security and forensics. While recent State Space Models (SSMs) show promise in precise temporal reasoning, their use in TFL is hindered by ambiguous boundaries, sparse forgeries, and limited long-range modeling. We propose DeformTrace, which enhances SSMs with deformable dynamics and relay mechanisms to address these challenges. Specifically, Deformable Self-SSM (DS-SSM) introduces dynamic receptive fields into SSMs for precise temporal localization. To further enhance its capacity for temporal reasoning and mitigate long-range decay, a Relay Token Mechanism is integrated into DS-SSM. Besides, Deformable Cross-SSM (DC-SSM) partitions the global state space into query-specific subspaces, reducing non-forgery information accumulation and boosting sensitivity to sparse forgeries. These components are integrated into a hybrid architecture that combines the global modeling of Transformers with the efficiency of SSMs. Extensive experiments show that DeformTrace achieves state-of-the-art performance with fewer parameters, faster inference, and stronger robustness.

cs.CV

GEM-TFL: Bridging Weak and Full Supervision for Forgery Localization through EM-Guided Decomposition and Temporal Refinement

Temporal Forgery Localization (TFL) aims to precisely identify manipulated segments within videos or audio streams, providing interpretable evidence for multimedia forensics and security. While most existing TFL methods rely on dense frame-level labels in a fully supervised manner, Weakly Supervised TFL (WS-TFL) reduces labeling cost by learning only from binary video-level labels. However, current WS-TFL approaches suffer from mismatched training and inference objectives, limited supervision from binary labels, gradient blockage caused by non-differentiable top-k aggregation, and the absence of explicit modeling of inter-proposal relationships. To address these issues, we propose GEM-TFL (Graph-based EM-powered Temporal Forgery Localization), a two-phase classification-regression framework that effectively bridges the supervision gap between training and inference. Built upon this foundation, (1) we enhance weak supervision by reformulating binary labels into multi-dimensional latent attributes through an EM-based optimization process; (2) we introduce a training-free temporal consistency refinement that realigns frame-level predictions for smoother temporal dynamics; and (3) we design a graph-based proposal refinement module that models temporal-semantic relationships among proposals for globally consistent confidence estimation. Extensive experiments on benchmark datasets demonstrate that GEM-TFL achieves more accurate and robust temporal forgery localization, substantially narrowing the gap with fully supervised methods.

cs.CV

On damage of interpolation to adversarial robustness in regression

Deep neural networks (DNNs) typically involve a large number of parameters and are trained to achieve zero or near-zero training error. Despite such interpolation, they often exhibit strong generalization performance on unseen data, a phenomenon that has motivated extensive theoretical investigations. Comforting results show that interpolation indeed may not affect the minimax rate of convergence under the squared error loss. In the mean time, DNNs are well known to be highly vulnerable to adversarial perturbations in future inputs. A natural question then arises: Can interpolation also escape from suboptimal performance under a future $X$-attack? In this paper, we investigate the adversarial robustness of interpolating estimators in a framework of nonparametric regression. A finding is that interpolating estimators must be suboptimal even under a subtle future $X$-attack, and achieving perfect fitting can substantially damage their robustness. An interesting phenomenon in the high interpolation regime, which we term the curse of simple size, is also revealed and discussed. Numerical experiments support our theoretical findings.

stat.ML