Searcharxiv⌕ Search

arXiv subjects

Yuke Li

Publications and source records attributed to Yuke Li.

At least 55 records · Page 3Linked to original sources

CityGo: Lightweight Urban Modeling and Rendering with Proxy Buildings and Residual Gaussians

Accurate and efficient modeling of large-scale urban scenes is critical for applications such as AR navigation, UAV based inspection, and smart city digital twins. While aerial imagery offers broad coverage and complements limitations of ground-based data, reconstructing city-scale environments from such views remains challenging due to occlusions, incomplete geometry, and high memory demands. Recent advances like 3D Gaussian Splatting (3DGS) improve scalability and visual quality but remain limited by dense primitive usage, long training times, and poor suit ability for edge devices. We propose CityGo, a hybrid framework that combines textured proxy geometry with residual and surrounding 3D Gaussians for lightweight, photorealistic rendering of urban scenes from aerial perspectives. Our approach first extracts compact building proxy meshes from MVS point clouds, then uses zero order SH Gaussians to generate occlusion-free textures via image-based rendering and back-projection. To capture high-frequency details, we introduce residual Gaussians placed based on proxy-photo discrepancies and guided by depth priors. Broader urban context is represented by surrounding Gaussians, with importance-aware downsampling applied to non-critical regions to reduce redundancy. A tailored optimization strategy jointly refines proxy textures and Gaussian parameters, enabling real-time rendering of complex urban scenes on mobile GPUs with significantly reduced training and memory requirements. Extensive experiments on real-world aerial datasets demonstrate that our hybrid representation significantly reduces training time, achieving on average 1.4x speedup, while delivering comparable visual fidelity to pure 3D Gaussian Splatting approaches. Furthermore, CityGo enables real-time rendering of large-scale urban scenes on mobile consumer GPUs, with substantially reduced memory usage and energy consumption.

cs.GR↗

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts

Controllable speech generation methods typically rely on single or fixed prompts, hindering creativity and flexibility. These limitations make it difficult to meet specific user needs in certain scenarios, such as adjusting the style while preserving a selected speaker's timbre, or choosing a style and generating a voice that matches a character's visual appearance. To overcome these challenges, we propose \textit{FleSpeech}, a novel multi-stage speech generation framework that allows for more flexible manipulation of speech attributes by integrating various forms of control. FleSpeech employs a multimodal prompt encoder that processes and unifies different text, audio, and visual prompts into a cohesive representation. This approach enhances the adaptability of speech synthesis and supports creative and precise control over the generated speech. Additionally, we develop a data collection pipeline for multimodal datasets to facilitate further research and applications in this field. Comprehensive subjective and objective experiments demonstrate the effectiveness of FleSpeech. Audio samples are available at https://kkksuper.github.io/FleSpeech/

eess.AS↗

Cross-Modal Few-Shot Learning: a Generative Transfer Learning Framework

Most existing studies on few-shot learning focus on unimodal settings, where models are trained to generalize to unseen data using a limited amount of labeled examples from a single modality. However, real-world data are inherently multi-modal, and such unimodal approaches limit the practical applications of few-shot learning. To bridge this gap, this paper introduces the Cross-modal Few-Shot Learning (CFSL) task, which aims to recognize instances across multiple modalities while relying on scarce labeled data. This task presents unique challenges compared to classical few-shot learning arising from the distinct visual attributes and structural disparities inherent to each modality. To tackle these challenges, we propose a Generative Transfer Learning (GTL) framework by simulating how humans abstract and generalize concepts. Specifically, the GTL jointly estimates the latent shared concept across modalities and the in-modality disturbance through a generative structure. Establishing the relationship between latent concepts and visual content among abundant unimodal data enables GTL to effectively transfer knowledge from unimodal to novel multimodal data, as humans did. Comprehensive experiments demonstrate that the GTL achieves state-of-the-art performance across seven multi-modal datasets across RGB-Sketch, RGB-Infrared, and RGB-Depth.

cs.CV↗

Anomalous Reynolds stress and dynamic mechanisms in two-dimensional elasto-inertial turbulence of viscoelastic channel flow

Elasto-inertial turbulence (EIT) has been demonstrated to be able to sustain in two-dimensional (2D) channel flow; however the systematic investigations on 2D EIT remain scare. This study addresses this gap by examining the statistical characteristics and dynamic mechanisms of 2D EIT, while exploring its similarities to and differences from three-dimensional (3D) EIT. We demonstrate that the influence of elasticity on the statistical properties of 2D EIT follows distinct trends compared to those observed in 3D EIT and drag-reducing turbulence (DRT). These differences can be attributed to variations in the underlying dynamical processes. As nonlinear elasticity increases, the dominant dynamic evolution in 3D flows involves the gradual suppression of inertial turbulence (IT). In contrast, 2D flows exhibit a progressive enhancement of EIT. More strikingly, we identify an anomalous Reynolds stress in 2D EIT that contributes negatively to flow resistance, a behavior opposite to that of IT. Quadrant analysis of velocity fluctuations reveals the predominance of motions in the first and third quadrants. These motions are closely associated with polymer sheet-like extension structures, which are inclined from the near-wall region toward the channel center along the streamwise direction. Finally, we present the dynamical budget of 2D EIT, which shows significant similarities to that of 3D EIT, thereby providing compelling evidence for the objective existence of the 2D nature of EIT.

physics.flu-dyn↗

BLR-MoE: Boosted Language-Routing Mixture of Experts for Domain-Robust Multilingual E2E ASR

Recently, the Mixture of Expert (MoE) architecture, such as LR-MoE, is often used to alleviate the impact of language confusion on the multilingual ASR (MASR) task. However, it still faces language confusion issues, especially in mismatched domain scenarios. In this paper, we decouple language confusion in LR-MoE into confusion in self-attention and router. To alleviate the language confusion in self-attention, based on LR-MoE, we propose to apply attention-MoE architecture for MASR. In our new architecture, MoE is utilized not only on feed-forward network (FFN) but also on self-attention. In addition, to improve the robustness of the LID-based router on language confusion, we propose expert pruning and router augmentation methods. Combining the above, we get the boosted language-routing MoE (BLR-MoE) architecture. We verify the effectiveness of the proposed BLR-MoE in a 10,000-hour MASR dataset.

cs.CL↗

Anisotropic transport properties and topological Hall effect in the annealed kagome antiferromagnet FeGe

Electron correlation often gives birth to various orders in quantum materials. Recently, a strongly correlated kagome antiferromagnet FeGe is discovered to undergo a charge density wave transition inside the A-type antiferromagnetic state, providing an opportunity to explore the interplay between charge order and magnetism. Here, we reported the observation of anisotropic resistivity and Hall effect, along with a topological Hall effect, in the annealed FeGe crystals. As the current flows along the \emph{ab}-plane, the temperature dependence of $ρ_{ab}$ exhibits a distinct resistivity loop related to a first-order transition at $T_{cdw}$. The applied magnetic fields do not alter $T_{cdw}$ but can induce a spin-flop transition at $H_{sf}$. Consequently, a field-induced large topological Hall effect is observed in the canting antiferromagnetic (CAFM) state below $T_{cant}$, which is possibly attributed to the non-trivial spin texture during the spin-flop process. Whereas, as current is parallel to \emph{c}-axis, both the field-induced transitions in $ρ_{c}$ and $χ_{c}$ disappear. Instead, the Hall resistivity in the annealed FeGe significantly exhibits a deviation from the linear field-dependent. These findings provide valuable insight into revealing the interplay among magnetism, charge order and topology in the kagome magnets.

cond-mat.str-el↗

Universal Scaling Behavior of Transport Properties in Non-Magnetic RuO$_{2}$

As a prototypical altermagnet, RuO$_{2}$ has been subject to many controversial reports regarding its magnetic ground state and the existence of crystal Hall effects. We obtained high-quality RuO$_{2}$ single crystal with a residual resistivity ratio (RRR = 152), and carefully measured its magnetization, longitudinal resistivity ($ρ_{xx}$) and Hall resistivity ($ρ_{yx}$) up to 35 T magnetic field. We also calculated its electronic band, Fermi surface, and conducted numerical simulations for its transport properties. It was found that no magnetic transition occurs below 400 K, and that all the transport properties are consistent with the numerical simulations results, indicating that the magnetotransport properties originate from the intrinsic electronic structures and are dominated by the Lorentz force. Particularly, no crystal Hall effects were observed in our RuO$_{2}$ samples and both magnetoresistance and Hall resistivity follow scaling behavior. These results demonstrate that RuO$_{2}$ is a typical semimetal, rather than an altermagnet.

cond-mat.mtrl-sci↗

CoDiff-VC: A Codec-Assisted Diffusion Model for Zero-shot Voice Conversion

Zero-shot voice conversion (VC) aims to convert the original speaker's timbre to any target speaker while keeping the linguistic content. Current mainstream zero-shot voice conversion approaches depend on pre-trained recognition models to disentangle linguistic content and speaker representation. This results in a timbre residue within the decoupled linguistic content and inadequacies in speaker representation modeling. In this study, we propose CoDiff-VC, an end-to-end framework for zero-shot voice conversion that integrates a speech codec and a diffusion model to produce high-fidelity waveforms. Our approach involves employing a single-codebook codec to separate linguistic content from the source speech. To enhance content disentanglement, we introduce Mix-Style layer normalization (MSLN) to perturb the original timbre. Additionally, we incorporate a multi-scale speaker timbre modeling approach to ensure timbre consistency and improve voice detail similarity. To improve speech quality and speaker similarity, we introduce dual classifier-free guidance, providing both content and timbre guidance during the generation process. Objective and subjective experiments affirm that CoDiff-VC significantly improves speaker similarity, generating natural and higher-quality speech.

cs.SD↗

Weak antilocalization in the transition metal telluride Ta$_2$Pd$_3$Te$_5$

We report transport studies on the layered van der Waals topological crystalline insulator Ta$_2$Pd$_3$Te$_5$. The temperature-dependent resistance at high temperature is dominated by a bulk insulating gap and tend to saturate at low temperatures. Low temperature magnetotransport shows that Ta$_2$Pd$_3$Te$_5$ exhibits weak antilocatization (WAL) effect in both perpendicular orientation and parallel orientation, suggesting an contribution of the WAL effect from both topological edge states and bulk states. By measuring the anisotropic magnetoconductance and then subtracting the contribution of bulk states, the WAL effect associated with topological edge states can be revealed and analyzed quantitatively based on the two-dimensional Hikami-Larkin-Nagaoka model. Our results have important implications in understanding the WAL phenomena in Ta$_2$Pd$_3$Te$_5$.

cond-mat.mtrl-sci↗

Enhancing Traffic Object Detection in Variable Illumination with RGB-Event Fusion

Traffic object detection under variable illumination is challenging due to the information loss caused by the limited dynamic range of conventional frame-based cameras. To address this issue, we introduce bio-inspired event cameras and propose a novel Structure-aware Fusion Network (SFNet) that extracts sharp and complete object structures from the event stream to compensate for the lost information in images through cross-modality fusion, enabling the network to obtain illumination-robust representations for traffic object detection. Specifically, to mitigate the sparsity or blurriness issues arising from diverse motion states of traffic objects in fixed-interval event sampling methods, we propose the Reliable Structure Generation Network (RSGNet) to generate Speed Invariant Frames (SIF), ensuring the integrity and sharpness of object structures. Next, we design a novel Adaptive Feature Complement Module (AFCM) which guides the adaptive fusion of two modality features to compensate for the information loss in the images by perceiving the global lightness distribution of the images, thereby generating illumination-robust representations. Finally, considering the lack of large-scale and high-quality annotations in the existing event-based object detection datasets, we build a DSEC-Det dataset, which consists of 53 sequences with 63,931 images and more than 208,000 labels for 8 classes. Extensive experimental results demonstrate that our proposed SFNet can overcome the perceptual boundaries of conventional cameras and outperform the frame-based method by 8.0% in mAP50 and 5.9% in mAP50:95. Our code and dataset will be available at https://github.com/YN-Yang/SFNet.

cs.CV↗

Coarse-to-fine Alignment Makes Better Speech-image Retrieval

In this paper, we propose a novel framework for speech-image retrieval. We utilize speech-image contrastive (SIC) learning tasks to align speech and image representations at a coarse level and speech-image matching (SIM) learning tasks to further refine the fine-grained cross-modal alignment. SIC and SIM learning tasks are jointly trained in a unified manner. To optimize the learning process, we utilize an embedding queue that facilitates efficient sampling of high-quality and diverse negative representations during SIC learning. Additionally, it enhances the learning of SIM tasks by effectively mining hard negatives based on contrastive similarities calculated in SIC tasks. To further optimize learning under noisy supervision, we incorporate momentum distillation into the training process. Experimental results show that our framework outperforms the state-of-the-art method by more than 4% in R@1 on two benchmark datasets for the speech-image retrieval tasks. Moreover, as observed in zero-shot experiments, our framework demonstrates excellent generalization capabilities.

cs.CL↗

Cross-Modal Denoising: A Novel Training Paradigm for Enhancing Speech-Image Retrieval

The success of speech-image retrieval relies on establishing an effective alignment between speech and image. Existing methods often model cross-modal interaction through simple cosine similarity of the global feature of each modality, which fall short in capturing fine-grained details within modalities. To address this issue, we introduce an effective framework and a novel learning task named cross-modal denoising (CMD) to enhance cross-modal interaction to achieve finer-level cross-modal alignment. Specifically, CMD is a denoising task designed to reconstruct semantic features from noisy features within one modality by interacting features from another modality. Notably, CMD operates exclusively during model training and can be removed during inference without adding extra inference time. The experimental results demonstrate that our framework outperforms the state-of-the-art method by 2.0% in mean R@1 on the Flickr8k dataset and by 1.7% in mean R@1 on the SpokenCOCO dataset for the speech-image retrieval tasks, respectively. These experimental results validate the efficiency and effectiveness of our framework.

cs.CL↗

Learning Causal Domain-Invariant Temporal Dynamics for Few-Shot Action Recognition

Few-shot action recognition aims at quickly adapting a pre-trained model to the novel data with a distribution shift using only a limited number of samples. Key challenges include how to identify and leverage the transferable knowledge learned by the pre-trained model. We therefore propose CDTD, or Causal Domain-Invariant Temporal Dynamics for knowledge transfer. To identify the temporally invariant and variant representations, we employ the causal representation learning methods for unsupervised pertaining, and then tune the classifier with supervisions in next stage. Specifically, we assume the domain information can be well estimated and the pre-trained image decoder and transition models can be well transferred. During adaptation, we fix the transferable temporal dynamics and update the image encoder and domain estimator. The efficacy of our approach is revealed by the superior accuracy of CDTD over leading alternatives across standard few-shot action recognition datasets.

cs.CV↗

Mechanism of stochastic resonance in viscoelastic channel flow

We have recently discovered stochastic resonance (SR) in chaotic inertia-less viscoelastic channel flow. SR appears just above a pure elastic instability at a critical Weissenberg number, $Wi_c=150$, of a transition regime. In this lower sub-region up to $Wi\sim 300$, only the streamwise velocity, $u$, exhibits a chaotic spectrum, $E_u$, while the spanwise velocity continues to exhibit white noise, verified by its flat spectrum, $E_w$, accompanied by weak intensity elastic waves. However, SR vanishes at the upper limit at $Wi\sim 300$, when $E_w$ becomes chaotic, indicating $Wi$ as the control parameter. Here we clarify the mechanism of SR emergence by validating the control parameters, namely $Wi$ and the rms velocity fluctuations, $u_{rms}$, measured at multiple channel locations, which determine the range of SR existence. Our experiments verify three key ingredients of the SR mechanism: chaotic $E_u$, white noise $E_w$, and weak elastic waves, which are consistent with three constituents of autonomous dynamical systems exhibiting SR.

physics.flu-dyn↗

Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning

This paper aims to build a multi-speaker expressive TTS system, synthesizing a target speaker's speech with multiple styles and emotions. To this end, we propose a novel contrastive learning-based TTS approach to transfer style and emotion across speakers. Specifically, contrastive learning from different levels, i.e. utterance and category level, is leveraged to extract the disentangled style, emotion, and speaker representations from speech for style and emotion transfer. Furthermore, a semi-supervised training strategy is introduced to improve the data utilization efficiency by involving multi-domain data, including style-labeled data, emotion-labeled data, and abundant unlabeled data. To achieve expressive speech with diverse styles and emotions for a target speaker, the learned disentangled representations are integrated into an improved VITS model. Experiments on multi-domain data demonstrate the effectiveness of the proposed method.

eess.AS↗

From laminar to chaotic flow via stochastic resonance in viscoelastic channel flow

Recent research indicates that low-inertia viscoelastic channel flow experiences supercritical non-normal mode elastic instability from laminar to sustained chaotic flow due to finite-size perturbations. The challenge of this study is to elucidate a realization of such a pathway when the intensity of the elastic wave is too low to amplify velocity fluctuations above the instability onset. The study identifies two subregions in the transition flow regime at Weissenberg number $Wi>Wi_c$, the instability onset. In the lower subregion at $Wi_c\leq Wi\leq 300$, we discover periodic spikes in the streamwise velocity time series $u(t)$ that appear in the chaotic power spectrum as low-frequency, high-intensity peaks resembling stochastic resonance (SR). In contrast, the spanwise velocity power spectrum, $E_w$, remains flat with low-intensity, noisy, and broad elastic wave peaks. The spikes significantly distort the probability density function of $u$, initiating and amplifying random streaks and wall-normal vorticity fluctuations. The SR appearance is similar to dynamical systems where chaotic attractor and limit cycle interact with external white noise. This similarity is confirmed by presenting a phase portrait in two subregions of the transition regime. In the upper subregion at $Wi>400$ the periodic spikes disappear and $E_w$ becomes chaotic with a large intensity elastic wave sufficient to self-organize and synchronize the streaks into cycles and to amplify the wall normal vorticity according to a recently proposed mechanism.

physics.flu-dyn↗

Split to Merge: Unifying Separated Modalities for Unsupervised Domain Adaptation

Large vision-language models (VLMs) like CLIP have demonstrated good zero-shot learning performance in the unsupervised domain adaptation task. Yet, most transfer approaches for VLMs focus on either the language or visual branches, overlooking the nuanced interplay between both modalities. In this work, we introduce a Unified Modality Separation (UniMoS) framework for unsupervised domain adaptation. Leveraging insights from modality gap studies, we craft a nimble modality separation network that distinctly disentangles CLIP's features into language-associated and vision-associated components. Our proposed Modality-Ensemble Training (MET) method fosters the exchange of modality-agnostic information while maintaining modality-specific nuances. We align features across domains using a modality discriminator. Comprehensive evaluations on three benchmarks reveal our approach sets a new state-of-the-art with minimal computational costs. Code: https://github.com/TL-UESTC/UniMoS

cs.CV↗

HaltingVT: Adaptive Token Halting Transformer for Efficient Video Recognition

Action recognition in videos poses a challenge due to its high computational cost, especially for Joint Space-Time video transformers (Joint VT). Despite their effectiveness, the excessive number of tokens in such architectures significantly limits their efficiency. In this paper, we propose HaltingVT, an efficient video transformer adaptively removing redundant video patch tokens, which is primarily composed of a Joint VT and a Glimpser module. Specifically, HaltingVT applies data-adaptive token reduction at each layer, resulting in a significant reduction in the overall computational cost. Besides, the Glimpser module quickly removes redundant tokens in shallow transformer layers, which may even be misleading for video recognition tasks based on our observations. To further encourage HaltingVT to focus on the key motion-related information in videos, we design an effective Motion Loss during training. HaltingVT acquires video analysis capabilities and token halting compression strategies simultaneously in a unified training process, without requiring additional training procedures or sub-networks. On the Mini-Kinetics dataset, we achieved 75.0% top-1 ACC with 24.2 GFLOPs, as well as 67.2% top-1 ACC with an extremely low 9.9 GFLOPs. The code is available at https://github.com/dun-research/HaltingVT.

cs.CV↗