Searcharxiv⌕ Search

arXiv subjects

Mateo Cámara

Publications and source records attributed to Mateo Cámara.

14 recordsLinked to original sources

Counting Closures in Spanish Trills: A Multi-Corpus Acoustic Study

The Spanish trill /r/ is canonically described as a short sequence of lingual closures, yet large-scale acoustic evidence across corpora is scarce, and automatic counters locating envelope peaks tend to conflate each closure with its release. We present a closure-based detector that locates closures gated by a quality filter and cross-checked against an independent autocorrelation-based period estimator. Applied to 3,560 well-formed (voiced, periodic) trill tokens from 356 speakers across six Spanish corpora, the detector yields a median of two closures and an inter-closure period near 36ms, matching the descriptive literature on all corpora. At the speaker level, phonotactic context is the only factor with a robust, medium effect: onset trills (word-initial and post-/n,l,s/) show more closures than intervocalic rr. We find no robust evidence of a sex effect once closures are counted directly. We report reference values and release a reproducible measurement pipeline for Spanish trills.

eess.AS↗

An Acoustic Landmark Database of the English Lexicon via Articulatory Synthesis

Acoustic landmark theory treats speech as organized around the acoustic consequences of articulatory gestures that shape the vocal tract and airflow. Progress is limited by the scarcity of large, unambiguously annotated landmark datasets. We invert the problem by generating speech from landmark patterns. Using the Pink Trombone physical vocal-tract synthesizer, we produce an English lexicon for two adult configurations (male, female). With direct control of gestures, we place landmark labels algorithmically at the exact times of their physical events (e.g., oral closures/releases). The corpus contains $>$200,000 synthesized words, rendered for both configurations with time-aligned annotations; intelligibility is measured with STOI. We leverage it for statistics across the lexicon from an articulatory-event view, reporting landmark frequencies and dominant cue patterns, and enabling quantitative studies plus training/benchmarking of automatic landmark detectors.

eess.AS↗

Acoustic Landmark Detector based on Conformer and HuBERT

Acoustic landmarks (abrupt acoustic changes tied to speech events) offer a linguistically grounded representation for speech analysis. We study automatic landmark detection with Conformer models, evaluating 14 configurations spanning architecture, loss, label representation, feature extractor, and data conditions on 1 839 manually annotated utterances with eight landmark types. We propose Gaussian soft labels with per-class temporal spread (sigma=10-20 ms), improving F1-at-20 ms by 7.0% absolute vs. hard labels by modeling annotation variability. Frozen HuBERT features perform best without fine-tuning (F1-at-20 ms=0.77). Stops and fricatives are reliable (F1>0.80), while vowels remain challenging (F1 approx 0.55). On our corpus, our system reaches a 13.8% Landmark Error Rate (LER). This is not directly comparable to AutoLandmark (31.3%) or SpeechMark (56.5%), evaluated on a different corpus and metric. Per-class trends show detectability increases with event abruptness, consistent with Stevens' theory.

eess.AS↗

Word Lengthening as a Function of Utterance Position: A Multi-Corpus Study

Efficient turn-taking requires interlocutors to predict turn endings within a few hundred milliseconds. Beyond syntactic and pragmatic completion, prosody (especially pre-boundary lengthening) supports projection. We test whether turn-final words are longer than mid-sentence words, whether this reflects prosodic modification rather than lexical choice, and where within the word it concentrates. We analyze four corpora spanning styles and two languages (English, Spanish): Switchboard, Columbia Games, BU Radio, and Glissando, with >500 speakers, $39{,}470$ turn-final and $206{,}268$ mid-sentence tokens across $\sim39{,}500$ turns. Turn-final words are longer (mean ${\approx}191$\,ms; $d=1.14$). The effect persists in matched-word, within-speaker comparisons ($80$\,ms; $p<0.001$) and is localized mainly to the final syllable ($d=0.89$). Turn-final lengthening thus emerges as a robust, localized cue to floor transfer.

eess.AS↗

Differentiable Articulatory Copy-Synthesis of Biphonic Singing

Sygyt is a Tuvan style of biphonic singing in which a low vocal drone is sustained while a high harmonic is selectively amplified in the 1--3\,kHz region. Copy-synthesizing this effect remains challenging for articulatory models, since it requires fine control of narrowly focused resonances that standard low-dimensional tract parameterizations cannot easily reproduce. We address this problem with a differentiable Kelly--Lochbaum waveguide augmented with a sublingual second source, cubic B-spline tract parameterization, and spatially varying learnable damping, optimized end-to-end by gradient descent from audio. On 20 segments from two independent sygyt datasets (5 singers, 10 pitches), the proposed model reduces log-spectral distance by 30--38\% relative to an articulatory baseline, with the largest gains concentrated in the overtone region. Cepstral-envelope analysis further shows more accurate recovery of the merged formant structure characteristic of sygyt production. The model also outperforms a DDSP harmonic-plus-noise baseline with direct per-harmonic spectral control, suggesting that explicit acoustic structure is a useful inductive bias for overtone-singing copy-synthesis.

eess.AS↗

Random-Effects Algorithm for Random Objects in Metric Spaces

Across many scientific disciplines, multiple observations are collected from the same experimental units, and in modern datasets these observations often arise as non-Euclidean random objects. In such settings, the incorporation of random effects is a critical modeling step for efficient estimation and personalized prediction. Although mixed-effects models are well established for scalar outcomes and, more recently, for functional data in Hilbert spaces, general random-effects frameworks for objects in metric spaces remain underdeveloped. In this paper, we propose a nonlinear Fréchet-based algorithm for random-effects modeling of arbitrary random objects defined on a metric space. Using M-estimation theory, we establish conditions under which the proposed metric-space prediction target is consistently estimated under a working random-effects formulation. We then evaluate the empirical performance of the proposed method using both synthetic data and digital health datasets that require practical tools for analyzing random objects in metric spaces, such as multivariate probability distributions and random graphs. We show that, although our method is developed beyond Hilbert spaces, it can outperform existing Hilbert space-based methods.

stat.ML↗

Open-Source System for Multilingual Translation and Cloned Speech Synthesis

We present an open-source system designed for multilingual translation and speech regeneration, addressing challenges in communication and accessibility across diverse linguistic contexts. The system integrates Whisper for speech recognition with Voice Activity Detection (VAD) to identify speaking intervals, followed by a pipeline of Large Language Models (LLMs). For multilingual applications, the first LLM segments speech into coherent, complete sentences, which a second LLM then translates. For speech regeneration, the system uses a text-to-speech (TTS) module with voice cloning capabilities to replicate the original speaker's voice, maintaining naturalness and speaker identity. The system's open-source components can operate locally or via APIs, offering cost-effective deployment across various use cases. These include real-time multilingual translation in Zoom sessions, speech regeneration for public broadcasts, and Bluetooth-enabled multilingual playback through personal devices. By preserving the speaker's voice, the system ensures a seamless and immersive experience, whether translating or regenerating speech. This open-source project is shared with the community to foster innovation and accessibility. We provide a detailed system performance analysis, including latency and word accuracy, demonstrating its potential to enable inclusive, adaptable communication solutions in real-world multilingual scenarios.

eess.AS↗

Decoding Vocal Articulations from Acoustic Latent Representations

We present a novel neural encoder system for acoustic-to-articulatory inversion. We leverage the Pink Trombone voice synthesizer that reveals articulatory parameters (e.g tongue position and vocal cord configuration). Our system is designed to identify the articulatory features responsible for producing specific acoustic characteristics contained in a neural latent representation. To generate the necessary latent embeddings, we employed two main methodologies. The first was a self-supervised variational autoencoder trained from scratch to reconstruct the input signal at the decoder stage. We conditioned its bottleneck layer with a subnetwork called the "projector," which decodes the voice synthesizer's parameters. The second methodology utilized two pretrained models: EnCodec and Wav2Vec. They eliminate the need to train the encoding process from scratch, allowing us to focus on training the projector network. This approach aimed to explore the potential of these existing models in the context of acoustic-to-articulatory inversion. By reusing the pretrained models, we significantly simplified the data processing pipeline, increasing efficiency and reducing computational overhead. The primary goal of our project was to demonstrate that these neural architectures can effectively encapsulate both acoustic and articulatory features. This prediction-based approach is much faster than traditional methods focused on acoustic feature-based parameter optimization. We validated our models by predicting six different parameters and evaluating them with objective and ViSQOL subjective-equivalent metric using both synthesizer- and human-generated sounds. The results show that the predicted parameters can generate human-like vowel sounds when input into the synthesizer. We provide the dataset, code, and detailed findings to support future research in this field.

eess.AS↗

Perceptually Equivalent Resolution in Handheld Devices for Streaming Bandwidth Saving

We present the description, results, and analysis of the experiments conducted to find the equivalent resolution associated with handheld devices. That is, the resolution from which users stop perceiving quality improvements if better resolutions are presented to them in such devices. Thus, it is the maximum resolution that it is worth considering for generating and delivering video, as long as sequences are not too intensively compressed. Therefore, the detection of the equivalent resolutions allows for notable savings in bandwidth consumption. Subjective assessments have been carried out on fifty subjects using a set of video sequences of very different nature and four handheld devices with a broad range of screen dimensions. The results prove that the equivalent resolution in current handheld devices is 720p as higher resolutions are not valued by users.

eess.IV↗

Del Visual al Auditivo: Sonorización de Escenas Guiada por Imagen

Recent advances in image, video, text and audio generative techniques, and their use by the general public, are leading to new forms of content generation. Usually, each modality was approached separately, which poses limitations. The automatic sound recording of visual sequences is one of the greatest challenges for the automatic generation of multimodal content. We present a processing flow that, starting from images extracted from videos, is able to sound them. We work with pre-trained models that employ complex encoders, contrastive learning, and multiple modalities, allowing complex representations of the sequences for their sonorization. The proposed scheme proposes different possibilities for audio mapping and text guidance. We evaluated the scheme on a dataset of frames extracted from a commercial video game and sounds extracted from the Freesound platform. Subjective tests have evidenced that the proposed scheme is able to generate and assign audios automatically and conveniently to images. Moreover, it adapts well to user preferences, and the proposed objective metrics show a high correlation with the subjective ratings.

eess.AS↗

FOLEY-VAE: Generación de efectos de audio para cine con inteligencia artificial

In this research, we present an interface based on Variational Autoencoders trained with a wide range of natural sounds for the innovative creation of Foley effects. The model can transfer new sound features to prerecorded audio or microphone-captured speech in real time. In addition, it allows interactive modification of latent variables, facilitating precise and customized artistic adjustments. Taking as a starting point our previous study on Variational Autoencoders presented at this same congress last year, we analyzed an existing implementation: RAVE [1]. This model has been specifically trained for audio effects production. Various audio effects have been successfully generated, ranging from electromagnetic, science fiction, and water sounds, among others published with this work. This innovative approach has been the basis for the artistic creation of the first Spanish short film with sound effects assisted by artificial intelligence. This milestone illustrates palpably the transformative potential of this technology in the film industry, opening the door to new possibilities for sound creation and the improvement of artistic quality in film productions.

eess.AS↗

IA Para el Mantenimiento Predictivo en Canteras: Modelado

Dependence on raw materials, especially in the mining sector, is a key part of today's economy. Aggregates are vital, being the second most used raw material after water. Digitally transforming this sector is key to optimizing operations. However, supervision and maintenance (predictive and corrective) are challenges little explored in this sector, due to the particularities of the sector, machinery and environmental conditions. All this, despite the successes achieved in other scenarios in monitoring with acoustic and contact sensors. We present an unsupervised learning scheme that trains a variational autoencoder model on a set of sound records. This is the first such dataset collected during processing plant operations, containing information from different points of the processing line. Our results demonstrate the model's ability to reconstruct and represent in latent space the recorded sounds, the differences in operating conditions and between different equipment. In the future, this should facilitate the classification of sounds, as well as the detection of anomalies and degradation patterns in the operation of the machinery.

eess.AS↗

Optimization Techniques for a Physical Model of Human Vocalisation

We present a non-supervised approach to optimize and evaluate the synthesis of non-speech audio effects from a speech production model. We use the Pink Trombone synthesizer as a case study of a simplified production model of the vocal tract to target non-speech human audio signals --yawnings. We selected and optimized the control parameters of the synthesizer to minimize the difference between real and generated audio. We validated the most common optimization techniques reported in the literature and a specifically designed neural network. We evaluated several popular quality metrics as error functions. These include both objective quality metrics and subjective-equivalent metrics. We compared the results in terms of total error and computational demand. Results show that genetic and swarm optimizers outperform least squares algorithms at the cost of executing slower and that specific combinations of optimizers and audio representations offer significantly different results. The proposed methodology could be used in benchmarking other physical models and audio types.

eess.AS↗

Vocal Tract Area Estimation by Gradient Descent

Articulatory features can provide interpretable and flexible controls for the synthesis of human vocalizations by allowing the user to directly modify parameters like vocal strain or lip position. To make this manipulation through resynthesis possible, we need to estimate the features that result in a desired vocalization directly from audio recordings. In this work, we propose a white-box optimization technique for estimating glottal source parameters and vocal tract shapes from audio recordings of human vowels. The approach is based on inverse filtering and optimizing the frequency response of a wave\-guide model of the vocal tract with gradient descent, propagating error gradients through the mapping of articulatory features to the vocal tract area function. We apply this method to the task of matching the sound of the Pink Trombone, an interactive articulatory synthesizer, to a given vocalization. We find that our method accurately recovers control functions for audio generated by the Pink Trombone itself. We then compare our technique against evolutionary optimization algorithms and a neural network trained to predict control parameters from audio. A subjective evaluation finds that our approach outperforms these black-box optimization baselines on the task of reproducing human vocalizations.

cs.SD↗