SearcharxivSearch

arXiv subjects

Hang Yu

Publications and source records attributed to Hang Yu.

At least 109 records · Page 6Linked to original sources

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

We introduce MiniMax-Speech, an autoregressive Transformer-based Text-to-Speech (TTS) model that generates high-quality speech. A key innovation is our learnable speaker encoder, which extracts timbre features from a reference audio without requiring its transcription. This enables MiniMax-Speech to produce highly expressive speech with timbre consistent with the reference in a zero-shot manner, while also supporting one-shot voice cloning with exceptionally high similarity to the reference voice. In addition, the overall quality of the synthesized audio is enhanced through the proposed Flow-VAE. Our model supports 32 languages and demonstrates excellent performance across multiple objective and subjective evaluations metrics. Notably, it achieves state-of-the-art (SOTA) results on objective voice cloning metrics (Word Error Rate and Speaker Similarity) and has secured the top position on the public TTS Arena leaderboard. Another key strength of MiniMax-Speech, granted by the robust and disentangled representations from the speaker encoder, is its extensibility without modifying the base model, enabling various applications such as: arbitrary voice emotion control via LoRA; text to voice (T2V) by synthesizing timbre features directly from text description; and professional voice cloning (PVC) by fine-tuning timbre features with additional data. We encourage readers to visit https://minimax-ai.github.io/tts_tech_report for more examples.

eess.AS

CAMeL: Cross-modality Adaptive Meta-Learning for Text-based Person Retrieval

Text-based person retrieval aims to identify specific individuals within an image database using textual descriptions. Due to the high cost of annotation and privacy protection, researchers resort to synthesized data for the paradigm of pretraining and fine-tuning. However, these generated data often exhibit domain biases in both images and textual annotations, which largely compromise the scalability of the pre-trained model. Therefore, we introduce a domain-agnostic pretraining framework based on Cross-modality Adaptive Meta-Learning (CAMeL) to enhance the model generalization capability during pretraining to facilitate the subsequent downstream tasks. In particular, we develop a series of tasks that reflect the diversity and complexity of real-world scenarios, and introduce a dynamic error sample memory unit to memorize the history for errors encountered within multiple tasks. To further ensure multi-task adaptation, we also adopt an adaptive dual-speed update strategy, balancing fast adaptation to new tasks and slow weight updates for historical tasks. Albeit simple, our proposed model not only surpasses existing state-of-the-art methods on real-world benchmarks, including CUHK-PEDES, ICFG-PEDES, and RSTPReid, but also showcases robustness and scalability in handling biased synthetic images and noisy text annotations. Our code is available at https://github.com/Jahawn-Wen/CAMeL-reID.

cs.CV

Contextual Preference Collaborative Measure Framework Based on Belief System

To reduce the human intervention in the preference measure process,this article proposes a preference collaborative measure framework based on an updated belief system,which is also capable of improving the accuracy and efficiency of preferen-ce measure algorithms.Firstly,the distance of rules and the average internal distance of rulesets are proposed for specifying the relationship between the rules.For discovering the most representative preferences that are common in all users,namely common preference,a algorithm based on average internal distance of ruleset,PRA algorithm,is proposed,which aims to finish the discoveryprocess with minimum information loss rate.Furthermore,the concept of Common belief is proposed to update the belief system,and the common preferences are the evidences of updated belief system.Then,under the belief system,the proposed belief degree and deviation degree are used to determine whether a rule confirms the belief system or not and classify the preference rules into two kinds(generalized or personalized),and eventually filters out Top-K interesting rules relying on belief degree and deviation degree.Based on above,a scalable interestingness calculation framework that can apply various formulas is proposed for accurately calculating interestingness in different conditions.At last,IMCos algorithm and IMCov algorithm are proposed as exemplars to verify the accuracy and efficiency of the framework by using weighted cosine similarity and correlation coefficients as belief degree.In experiments,the proposed algorithms are compared to two state-of-the-art algorithms and the results show that IMCos and IMCov outperform than the other two in most aspects.

cs.AI

Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM

Recent advancements in code large language models (LLMs) have demonstrated remarkable capabilities in code generation and understanding. It is still challenging to build a code LLM with comprehensive performance yet ultimate efficiency. Many attempts have been released in the open source community to break the trade-off between performance and efficiency, such as the Qwen Coder series and the DeepSeek Coder series. This paper introduces yet another attempt in this area, namely Ling-Coder-Lite. We leverage the efficient Mixture-of-Experts (MoE) architecture along with a set of high-quality data curation methods (especially those based on program analytics) to build an efficient yet powerful code LLM. Ling-Coder-Lite exhibits on-par performance on 12 representative coding benchmarks compared to state-of-the-art models of similar size, such as Qwen2.5-Coder-7B and DeepSeek-Coder-V2-Lite, while offering competitive latency and throughput. In practice, we achieve a 50\% reduction in deployment resources compared to the similar-sized dense model without performance loss. To facilitate further research and development in this area, we open-source our models as well as a substantial portion of high-quality data for the annealing and post-training stages. The models and data can be accessed at~\url{https://huggingface.co/inclusionAI/Ling-Coder-lite}.

cs.LG

Resonance Locking of Anharmonic $g$-Modes in Coalescing Neutron Star Binaries

Neutron stars in coalescing binaries deform due to the tidal gravitational fields generated by their companions. During the inspiral phase, the tidal deformation is dominated by the fundamental oscillation ($f$-) mode of the stars. The tide also has sub-dominant gravity ($g$-) modes that are resonantly excited when the linear tidal forcing sweeps through their eigenfrequencies. Beyond the linear order in perturbed fluid displacement, the $g$-modes are anharmonic, i.e., their oscillation frequencies depend on the mode energy. For the lowest-order $g$-mode, we show that when the tidal forcing reaches its linear eigenfrequency, the mode starts to dynamically adjust its energy so that its nonlinearly shifted oscillation frequency always matches that of the driving field. This phenomenon, which we term `resonance locking', persists through the rest of the inspiral, and hence, the mode grows to substantially larger energies than in the linear theory. Using a $1.4$--$1.4\, M_{\odot}$ binary neutron star system with the SLy4 equation of state, we find this results in an extra correction to the frequency-domain gravitational wave (GW) phase of $|ΔΨ|\approx 3\,{\rm rad}$ accumulated from the onset of resonance locking at the GW frequency of $94\,{\rm Hz}$ to the merger at $1.05\,{\rm kHz}$. This effect probes details of the internal structure of merging neutron stars beyond their bulk properties such as tidal deformability.

gr-qc

Resonance locking: radian-level phase shifts due to nonlinear hydrodynamics of $g$-modes in merging neutron star binaries

A neutron star (NS) in a binary system deforms due to the companion's tidal gravitational field. As the binary inspirals due to gravitational wave (GW) emission, the NS's deformation evolves; this evolution is typically modeled as the star's linear response to the companion's time-evolving tidal potential. In principle, the fluid elements' displacements can be excited and evolve nonlinearly since the equations of hydrodynamics and the tidal forcing have nonlinear terms. Recently, Kwon, Yu, and Venumadhav (KYV I [arXiv:2410.03831]) showed that nonlinear terms in the hydrodynamic equations of motion make the low-frequency response of NSs, characterized by gravity ($g$-) modes, behave in an anharmonic manner. The anharmonicity is dominantly generated by the mutual coupling of the four lowest-order ($n=1$, $l=|m|=2$) $g$-modes, and allows them to stay locked in a resonant state that oscillates phase-coherently with the orbit throughout the inspiral. As a result, the $g$-modes grow to larger amplitudes than the linear response suggests, leading to an extra phase correction to the frequency-domain GW signal $|ΔΨ|\approx 3\,{\rm rad}$ at a GW frequency of $1.05\,{\rm kHz}$. This effect is part of the truly dynamical tide, in the sense that the amplitude depends not just on the binary's instantaneous frequency but the entire history of the inspiral. In this paper, we explain the phenomenology of resonance locking in detail and analytically validate the numerical dephasing calculations in KYV I. We also demonstrate that the effect is only significant for the lowest-order $g$-modes.

gr-qc

Effective-one-body model for coalescing binary neutron stars: Incorporating tidal spin and enhanced radiation from dynamical tides

Tidal interactions in a coalescing binary neutron star (BNS) or neutron star-black hole (NSBH) system driven by gravitational wave (GW) radiation contain precious information about physics both at extreme density and in the highly relativistic regime. In the late inspiral stage, where the tidal effects are the strongest, dynamical corrections to the tidal response become significant. Previous analyses model the finite-frequency correction through the effective Love number approach, which only accounts for the correction in the radial interaction but ignores the lag in the tidal bulge behind the companion due to the continuous orbital shrinkage. The lag provides a torque, causing the star's spin to change over time. We dub the evolving component of the spin the tidal spin, whose dimensionless value can reach 0.03-0.4 depending on how rapidly the background star rotates. We present an effective-one-body (EOB) waveform model for BNSs and NSBHs incorporating the tidal spin, particularly its back reaction to the orbit due to the Newtonian tidal torque and the relativistic orbital hang-up. Beyond the conservative dynamics, we also derive the corrections to the dissipative radiation due to finite-frequency effects to the first post-Newtonian order. Depending on the star's background spin, the phase error in the time-domain waveform due to ignoring the tidal spin ranges from 0.3 to 4 radians at the waveform's peak amplitude. The difference in the waveforms with and without the tidal spin remarkably resembles the difference between previous effective Love number models and numerical relativity simulations, underscoring the significance of tidal spin in the construction of faithful models. Our model further extends the description of dynamics in the high-background spin regions of the parameter space that are yet to be covered by numerical simulations.

gr-qc

Impact of on-site potentials on $q$-breathers in nonlinear chains

On-site potentials are ubiquitous in physical systems and strongly influence their heat transport and energy localization. These potentials will inevitably affect the dynamical properties of $q$-breathers (QBs), defined as periodic orbits exponentially localized in normal mode space. By integrating on-site terms into the Fermi-Pasta-Ulam-Tsingou-$β$ system, this work utilizes numerical simulations and Floquet analysis to systematically explore the influence of on-site potentials on QB stability. For most QBs, except those at the phonon band edges, the instability is primarily governed by parametric resonance, and effectively described by coupled Mathieu equations. This approach provides a theoretical expression for the instability thresholds, which aligns well with numerical results. We demonstrate that the instability thresholds can be controlled through the strength of on-site potentials, and for a strong enough quadratic on-site potential, the QBs are always stable. Furthermore, the instability threshold is highly sensitive to the seed mode, in stark contrast to systems without on-site potentials. In addition, the instability phase diagrams exhibit joint interplay between different terms in the Hamiltonian, such as the quadratic on-site and quartic inter-site interaction terms, in regulating the QB dynamics. These findings offer valuable insights into QB stability and the manipulation of localized excitations in diverse physical systems with on-site potentials.

cond-mat.stat-mech

Efficient shortcuts-to-adiabaticity for loading an ultracold Fermi gas into higher orbital bands of one-dimensional optical lattice

We propose an experimental scheme to load ultracold Fermi gases from the ground orbital band of a one-dimensional optical lattice into the first excited orbital band. Unlike the narrow momentum distribution of a Bose-Einstein Condensate, Fermi gases exhibit a broad momentum distribution. To address this, we define the average loading efficiency across all quasi-momentum states and theoretically perform the loading operation simultaneously for each Bloch state. Using a multiparameter global optimization method, we determine the loading efficiency at various lattice depths. We can enhance the loading efficiency by adjusting the phase of the lattice, which leverages the different symmetries of Bloch wavefunctions in various optical lattice orbitals. We also identified that the primary factor hindering higher loading efficiency in the Fermi gas is the multiple occupancy of the quasi-momentum states. Our simulations of various occupancies revealed a decreasing trend in mean loading efficiency as the number of occupied quasi-momentum states increases. Finally, we compare our method with other loading techniques and assess its experimental feasibility.

cond-mat.quant-gas

Advancing Out-of-Distribution Detection via Local Neuroplasticity

In the domain of machine learning, the assumption that training and test data share the same distribution is often violated in real-world scenarios, requiring effective out-of-distribution (OOD) detection. This paper presents a novel OOD detection method that leverages the unique local neuroplasticity property of Kolmogorov-Arnold Networks (KANs). Unlike traditional multilayer perceptrons, KANs exhibit local plasticity, allowing them to preserve learned information while adapting to new tasks. Our method compares the activation patterns of a trained KAN against its untrained counterpart to detect OOD samples. We validate our approach on benchmarks from image and medical domains, demonstrating superior performance and robustness compared to state-of-the-art techniques. These results underscore the potential of KANs in enhancing the reliability of machine learning systems in diverse environments.

cs.LG

Observation of quantum-classical transition behavior of LGI in a dissipative quantum gas

The Leggett-Garg inequality (LGI) is a powerful tool for distinguishing between quantum and classical properties in studies of macroscopic systems. Applying the LGI to non-Hermitian systems with dissipation presents a fascinating opportunity, as competing mechanisms can either strengthen or weaken LGI violations. On one hand, dissipation-induced nonlinear interactions amplify LGI violations compared to Hermitian systems; on the other hand, dissipation leads to decoherence, which could weaken the LGI violation. In this paper, we investigate a non-Hermitian system of ultracold Fermi gas with dissipation. Our experiments reveal that as dissipation increases, the upper bound of the third-order LGI parameter $K_3$ initially rises, reaching its maximum at the exceptional point (EP), where $K_3 = C_{21} + C_{32} - C_{31}$, encompassing three two-time correlation functions. Beyond a certain dissipation threshold, the LGI violation weakens, approaching the classical limit, indicating a quantum-to-classical transition (QCT). Furthermore, we observe that the LGI violation decreases with increasing evolution time, reinforcing the QCT in the time domain. This study provides a crucial stepping stone for using the LGI to explore the QCT in many-body open quantum systems.

cond-mat.quant-gas

Traffic and Safety Rule Compliance of Humans in Diverse Driving Situations

The increasing interest in autonomous driving systems has highlighted the need for an in-depth analysis of human driving behavior in diverse scenarios. Analyzing human data is crucial for developing autonomous systems that replicate safe driving practices and ensure seamless integration into human-dominated environments. This paper presents a comparative evaluation of human compliance with traffic and safety rules across multiple trajectory prediction datasets, including Argoverse 2, nuPlan, Lyft, and DeepUrban. By defining and leveraging existing safety and behavior-related metrics, such as time to collision, adherence to speed limits, and interactions with other traffic participants, we aim to provide a comprehensive understanding of each datasets strengths and limitations. Our analysis focuses on the distribution of data samples, identifying noise, outliers, and undesirable behaviors exhibited by human drivers in both the training and validation sets. The results underscore the need for applying robust filtering techniques to certain datasets due to high levels of noise and the presence of such undesirable behaviors.

cs.RO

CoBa: Convergence Balancer for Multitask Finetuning of Large Language Models

Multi-task learning (MTL) benefits the fine-tuning of large language models (LLMs) by providing a single model with improved performance and generalization ability across tasks, presenting a resource-efficient alternative to developing separate models for each task. Yet, existing MTL strategies for LLMs often fall short by either being computationally intensive or failing to ensure simultaneous task convergence. This paper presents CoBa, a new MTL approach designed to effectively manage task convergence balance with minimal computational overhead. Utilizing Relative Convergence Scores (RCS), Absolute Convergence Scores (ACS), and a Divergence Factor (DF), CoBa dynamically adjusts task weights during the training process, ensuring that the validation loss of all tasks progress towards convergence at an even pace while mitigating the issue of individual task divergence. The results of our experiments involving three disparate datasets underscore that this approach not only fosters equilibrium in task convergence but enhances the LLMs' performance by up to 13% relative to the second-best baselines. Code is open-sourced at https://github.com/codefuse-ai/MFTCoder.

cs.CL

$q$-Breathers in the diatomic $β$-Fermi-Pasta-Ulam- Tsingou chains

$q$-Breathers (QBs) represent a quintessential phenomenon of energy localization, manifesting as stable periodic orbits exponentially localized in normal mode space. Their existence can hinder the thermalization process in nonlinear lattices. In this study, we employ the Newton's method to identify QB solutions in the diatomic Fermi-Pasta-Ulam-Tsingou chains and perform a comprehensive analysis of their linear stability. We derive an analytical expression for the instability thresholds of low-frequency QBs, which converges to the known results of monoatomic chains as the bandgap approaches zero. The expression reveals an inverse square relationship between instability thresholds and system size, as well as a quadratic dependence on the mass difference, both of which have been corroborated through extensive numerical simulations. Our results demonstrate that the presence of a bandgap can markedly enhance QB stability, providing a novel theoretical foundation and practical framework for controlling energy transport between modes in complex lattice systems. These results not only expand the applicability of QBs but also offer significant implications for understanding the thermalization dynamics in complex lattice structures, with wide potential applications in related low-dimensional materials.

cond-mat.stat-mech

Harnessing Uncertainty-aware Bounding Boxes for Unsupervised 3D Object Detection

Unsupervised 3D object detection aims to identify objects of interest from unlabeled raw data, such as LiDAR points. Recent approaches usually adopt pseudo 3D bounding boxes (3D bboxes) from clustering algorithm to initialize the model training. However, pseudo bboxes inevitably contain noise, and such inaccuracies accumulate to the final model, compromising the performance. Therefore, in an attempt to mitigate the negative impact of inaccurate pseudo bboxes, we introduce a new uncertainty-aware framework for unsupervised 3D object detection, dubbed UA3D. In particular, our method consists of two phases: uncertainty estimation and uncertainty regularization. (1) In the uncertainty estimation phase, we incorporate an extra auxiliary detection branch alongside the original primary detector. The prediction disparity between the primary and auxiliary detectors could reflect fine-grained uncertainty at the box coordinate level. (2) Based on the assessed uncertainty, we adaptively adjust the weight of every 3D bbox coordinate via uncertainty regularization, refining the training process on pseudo bboxes. For pseudo bbox coordinate with high uncertainty, we assign a relatively low loss weight. Extensive experiments verify that the proposed method is robust against the noisy pseudo bboxes, yielding substantial improvements on nuScenes and Lyft compared to existing approaches, with increases of +6.9% AP$_{BEV}$ and +2.5% AP$_{3D}$ on nuScenes, and +4.1% AP$_{BEV}$ and +2.0% AP$_{3D}$ on Lyft.

cs.CV

Survey of Query-based Text Summarization

Query-based text summarization is an important real world problem that requires to condense the prolix text data into a summary under the guidance of the query information provided by users. The topic has been studied for a long time and there are many existing interesting research related to query-based text summarization. Yet much of the work is not systematically surveyed. This survey aims at summarizing some interesting work in query-based text summarization methods as well as related generic text summarization methods. Not all taxonomies in this paper exist the related work to the best of our knowledge and some analysis will be presented.

cs.IR

Focus On What Matters: Separated Models For Visual-Based RL Generalization

A primary challenge for visual-based Reinforcement Learning (RL) is to generalize effectively across unseen environments. Although previous studies have explored different auxiliary tasks to enhance generalization, few adopt image reconstruction due to concerns about exacerbating overfitting to task-irrelevant features during training. Perceiving the pre-eminence of image reconstruction in representation learning, we propose SMG (Separated Models for Generalization), a novel approach that exploits image reconstruction for generalization. SMG introduces two model branches to extract task-relevant and task-irrelevant representations separately from visual observations via cooperatively reconstruction. Built upon this architecture, we further emphasize the importance of task-relevant features for generalization. Specifically, SMG incorporates two additional consistency losses to guide the agent's focus toward task-relevant areas across different scenarios, thereby achieving free from overfitting. Extensive experiments in DMC demonstrate the SOTA performance of SMG in generalization, particularly excelling in video-background settings. Evaluations on robotic manipulation tasks further confirm the robustness of SMG in real-world applications.

cs.CV

SDP: Spiking Diffusion Policy for Robotic Manipulation with Learnable Channel-Wise Membrane Thresholds

This paper introduces a Spiking Diffusion Policy (SDP) learning method for robotic manipulation by integrating Spiking Neurons and Learnable Channel-wise Membrane Thresholds (LCMT) into the diffusion policy model, thereby enhancing computational efficiency and achieving high performance in evaluated tasks. Specifically, the proposed SDP model employs the U-Net architecture as the backbone for diffusion learning within the Spiking Neural Network (SNN). It strategically places residual connections between the spike convolution operations and the Leaky Integrate-and-Fire (LIF) nodes, thereby preventing disruptions to the spiking states. Additionally, we introduce a temporal encoding block and a temporal decoding block to transform static and dynamic data with timestep $T_S$ into each other, enabling the transmission of data within the SNN in spike format. Furthermore, we propose LCMT to enable the adaptive acquisition of membrane potential thresholds, thereby matching the conditions of varying membrane potentials and firing rates across channels and avoiding the cumbersome process of manually setting and tuning hyperparameters. Evaluating the SDP model on seven distinct tasks with SNN timestep $T_S=4$, we achieve results comparable to those of the ANN counterparts, along with faster convergence speeds than the baseline SNN method. This improvement is accompanied by a reduction of 94.3\% in dynamic energy consumption estimated on 45nm hardware.

cs.RO