SearcharxivSearch

arXiv subjects

Maksim Borisov

Publications and source records attributed to Maksim Borisov.

4 recordsLinked to original sources

NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech

Current expressive speech synthesis models are constrained by the limited availability of open-source datasets containing diverse nonverbal vocalizations (NVs). In this work, we introduce NonverbalTTS (NVTTS), a 17-hour open-access dataset annotated with 10 types of NVs (e.g., laughter, coughs) and 8 emotional categories. The dataset is derived from popular sources, VoxCeleb and Expresso, using automated detection followed by human validation. We propose a comprehensive pipeline that integrates automatic speech recognition (ASR), NV tagging, emotion classification, and a fusion algorithm to merge transcriptions from multiple annotators. Fine-tuning open-source text-to-speech (TTS) models on the NVTTS dataset achieves parity with closed-source systems such as CosyVoice2, as measured by both human evaluation and automatic metrics, including speaker similarity and NV fidelity. By releasing NVTTS and its accompanying annotation guidelines, we address a key bottleneck in expressive TTS research. The dataset is available at https://huggingface.co/datasets/deepvk/NonverbalTTS.

cs.LG

Low-resource Machine Translation for Code-switched Kazakh-Russian Language Pair

Machine translation for low resource language pairs is a challenging task. This task could become extremely difficult once a speaker uses code switching. We propose a method to build a machine translation model for code-switched Kazakh-Russian language pair with no labeled data. Our method is basing on generation of synthetic data. Additionally, we present the first codeswitching Kazakh-Russian parallel corpus and the evaluation results, which include a model achieving 16.48 BLEU almost reaching an existing commercial system and beating it by human evaluation.

cs.CL

DINO-VITS: Data-Efficient Zero-Shot TTS with Self-Supervised Speaker Verification Loss for Noise Robustness

We address zero-shot TTS systems' noise-robustness problem by proposing a dual-objective training for the speaker encoder using self-supervised DINO loss. This approach enhances the speaker encoder with the speech synthesis objective, capturing a wider range of speech characteristics beneficial for voice cloning. At the same time, the DINO objective improves speaker representation learning, ensuring robustness to noise and speaker discriminability. Experiments demonstrate significant improvements in subjective metrics under both clean and noisy conditions, outperforming traditional speaker-encoderbased TTS systems. Additionally, we explore training zeroshot TTS on noisy, unlabeled data. Our two-stage training strategy, leveraging self-supervised speech models to distinguish between noisy and clean speech, shows notable advances in similarity and naturalness, especially with noisy training datasets, compared to the ASR-transcription-based approach.

cs.SD

The influence of color on prices of abstract paintings

Determination of price of an artwork is a fundamental problem in cultural economics. In this work we investigate what impact visual characteristics of a painting have on its price. We construct a number of visual features measuring complexity of the painting, its points of interest, segmentation-based features, local color features, and features based on Itten and Kandinsky theories, and utilize mixed-effects model to study impact of these features on the painting price. We analyze the influence of the color on the example of the most complex art style - abstractionism, by created Kandinsky, for which the color is the primary basis. We use Itten's theory - the most recognized color theory in art history, from which the largest number of subtheories was born. For this day it is taken as the base for teaching artists. We utilize novel dataset of 3885 paintings collected from Christie's and Sotheby's and find that color harmony has some explanatory power, color complexity metrics are insignificant and color diversity explains price well.

econ.GN