SearcharxivSearch

arXiv subjects

Minsu Kang

Publications and source records attributed to Minsu Kang.

13 recordsLinked to original sources

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion

Advances in AI-based voice conversion have enabled a wide range of media applications, including films, audiobooks, and games. However, most research and public benchmarks still focus on natural human speech, leaving designed vocalizations, such as monster growls and robotic voices, underexplored, partly due to the lack of publicly available resources. To address this gap, we introduce the Designed Vocalizations Dataset, constructed by curating diverse raw vocal sources, including speech and animal vocalizations, and applying professional vocal effects processing to produce corresponding effect modified variants. We further provide a standardized test set with explicit seen/unseen splits over source timbre groups and preset styles to assess generalization under controlled conditions. Finally, we report baseline benchmark results to support reproducible evaluation and future research. The dataset and demo samples are available at https://ncai-official.github.io/speech/publications/designed-vocalizations-dataset/.

eess.AS

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds

Human to non-human voice conversion (H2NH-VC) transforms human speech into animal or designed vocalizations. Unlike prior studies focused on dog-sounds and 16 or 22.05kHz audio transformation, this work addresses a broader range of non-speech sounds, including natural sounds (lion-roars, birdsongs) and designed voice (synthetic growls). To accomodate generation of diverse non-speech sounds and 44.1kHz high-quality audio transformation, we introduce a preprocessing pipeline and an improved CVAE-based H2NH-VC model, both optimized for human and non-human voices. Experimental results showed that the proposed method outperformed baselines in quality, naturalness, and similarity MOS, achieving effective voice conversion across diverse non-human timbres. Demo samples are available at https://nc-ai.github.io/speech/publications/nonhuman-vc/

eess.AS

A Charged and Neutral Spin-$4$ Currents in the Grassmannian-like Coset Model

By calculating the second order pole in the operator product expansion (OPE) of the charged spin-$3$ current with the neutral spin-$3$ current in the Grassmannian-like coset model, we determine the primary charged spin-$4$ current. Similarly, by computing the second order pole in the OPE of the neutral spin-$3$ current with itself, we obtain the primary neutral spin-$4$ current. We determine the OPE of the charged spin-$2$ current with the charged spin-$3$ current for generic parameters and the large $k$ (one of the parameters) limit is also obtained for this OPE. In particular, the above primary charged spin-$4$ current appears in the first order pole of this OPE for generic parameters. We also check that the above primary charged and neutral spin-$4$ currents occur at the second order pole in the OPE of the charged spin-$3$ current with itself for fixed parameters.

hep-th

UniTTS: Residual Learning of Unified Embedding Space for Speech Style Control

We propose a novel high-fidelity expressive speech synthesis model, UniTTS, that learns and controls overlapping style attributes avoiding interference. UniTTS represents multiple style attributes in a single unified embedding space by the residuals between the phoneme embeddings before and after applying the attributes. The proposed method is especially effective in controlling multiple attributes that are difficult to separate cleanly, such as speaker ID and emotion, because it minimizes redundancy when adding variance in speaker ID and emotion, and additionally, predicts duration, pitch, and energy based on the speaker ID and emotion. In experiments, the visualization results exhibit that the proposed methods learned multiple attributes harmoniously in a manner that can be easily separated again. As well, UniTTS synthesized high-fidelity speech signals controlling multiple style attributes. The synthesized speech samples are presented at https://anonymous-authors2022.github.io/paper_works/UniTTS/demos/.

eess.AS

Fast DCTTS: Efficient Deep Convolutional Text-to-Speech

We propose an end-to-end speech synthesizer, Fast DCTTS, that synthesizes speech in real time on a single CPU thread. The proposed model is composed of a carefully-tuned lightweight network designed by applying multiple network reduction and fidelity improvement techniques. In addition, we propose a novel group highway activation that can compromise between computational efficiency and the regularization effect of the gating mechanism. As well, we introduce a new metric called Elastic mel-cepstral distortion (EMCD) to measure the fidelity of the output mel-spectrogram. In experiments, we analyze the effect of the acceleration techniques on speed and speech quality. Compared with the baseline model, the proposed model exhibits improved MOS from 2.62 to 2.74 with only 1.76% computation and 2.75% parameters. The speed on a single CPU thread was improved by 7.45 times, which is fast enough to produce mel-spectrogram in real time without GPU.

eess.AS

Lasing from complete set of topological states in two dimensional photonic crystal structure

Recently, topologically engineered photonic structures have garnered significant attention as their eigenstates may offer a new insight on photon manipulation and an unconventional route for nanophotonic devices with unprecedented functionalities and robustness. Herein, we present lasing actions at all hierarchical eigenstates that can exist in a topologically designed single two-dimensional (2D) photonic crystal (PhC) platform: 2D bulk, one-dimensional edge, and zero-dimensional corner states. In particular, multiple topological eigenstates are generated in a hierarchical manner with no bulk multipole moment. The unit cell of the topological PhC structure is a tetramer composed of four identical air holes perforated into an InGaAsP multiple-quantum-well epilayer slab. A square area of a topologically nontrivial PhC structure is surrounded by a topologically trivial counterpart, resulting in multidimensional eigenstates of one bulk, four side edges, and four corners within and at the boundaries. Spatially resolved optical excitation spontaneously results in lasing actions at all nine hierarchical topological states. Our experimental findings may provide insight into the development of sophisticated next-generation nanophotonic devices and robust integration platforms.

physics.app-ph

Quantum macroscopicity measure for arbitrary spin systems and its application to quantum phase transitions

We explore a previously unknown connection between two important problems in physics, i.e., quantum macroscopicity and the quantum phase transition. We devise a general and computable measure of quantum macroscopicity that can be applied to arbitrary spin states. We find that a macroscopic quantum superposition of an extremely large size arises during the quantum phase transition of the transverse Ising model in contrast to some seeming macroscopic quantum phenomena such as superconductivity, superfluidity and Bose-Einstein condensates. Our result may be an important step forward in understanding macroscopic quantum properties of many-body systems.

quant-ph

Is macroscopic entanglement a typical trait of many-particle quantum states?

We elucidate the relationship between Schrödinger-cat-like macroscopicity and geometric entanglement, and argue that these quantities are not interchangeable. While both properties are lost due to decoherence, we show that macroscopicity is rare in uniform and in so-called random physical ensembles of pure quantum states, despite possibly large geometric entanglement. In contrast, permutation-symmetric pure states feature rather low geometric entanglement and strong and robust macroscopicity.

quant-ph

Characterizations and Quantifications of Macroscopic Quantumness and Its Implementations using Optical Fields

We present a review and discussions on characterizations and quantifications of macroscopic quantum states as well as their implementations and applications in optical systems. We compare and criticize different measures proposed to define and quantify macroscopic quantum superpositions and extend such comparisons to several types of optical quantum states actively considered for experimental implementations within recent research topics.

quant-ph

Entangling quantum and classical states of light

Entanglement between quantum and classical objects is of special interest in the context of fundamental studies of quantum mechanics and potential applications to quantum information processing. In quantum optics, single photons are treated as light quanta while coherent states are considered the most classical among all pure states. Recently, entanglement between a single photon and a coherent state in a free-traveling field was identified to be a useful resource for optical quantum information processing. However, it was pointed out to be extremely difficult to generate such states since it requires a clean cross-Kerr nonlinear interaction. Here, we devise and experimentally demonstrate a scheme to generate such hybrid entanglement by implementing a coherent superposition of two distinct quantum operations. The generated states clearly show entanglement between the two different types of states. Our work opens a way to generate hybrid entanglement of a larger size and to develop efficient quantum information processing using such a new type of qubits.

quant-ph

Using macroscopic entanglement to close the detection loophole in Bell inequality

We consider a Bell-like inequality performed using various instances of multi-photon entangled states to demonstrate that losses occurring after the unitary transformations used in the nonlocality test can be counteracted by enhancing the "size" of such entangled states. In turn, this feature can be used to overcome detection inefficiencies affecting the test itself: a slight increase in the size of such states, pushing them towards a more "macroscopic" form of entanglement, significantly improves the state robustness against detection inefficiency, thus easing the closing of the detection loophole. Differently, losses before the unitary transformations cause decoherence effects that cannot be compensated using macroscroscopic entanglement.

quant-ph

Production of entanglement with highly-mixed states

We study production of entanglement with highly-mixed states. We find that entanglement between highly mixed states can be generated via a direct unitary interaction even when both states have purities arbitrarily close to zero. This indicates that purity of a subsystem is not required for entanglement generation. Our result is in contrast to previous studies where the importance of the subsystem purity was emphasized.

quant-ph

Reply to Comment on "Quantification of Macroscopic Quantum Superpositions within Phase Space"

Gong points out [quant-ph arXiv:1106.0062] a "direct connection" between our measure [PRL 106, 220401 (2011)] recently proposed to quantify macroscopic quantum superpositions and a previously studied quantity introduced to study classical and quantum chaos. We point out that the two measures are obviously different for some mixed states, and the previous one does not work as a sensible measure to quantify quantum superpositions.

quant-ph