SearcharxivSearch

arXiv · 2005.06959

Consonant gemination in Italian: the affricate and fricative case

Abstract

Consonant gemination in Italian affricates and fricatives was investigated, completing the overall study of gemination of Italian consonants. Results of the analysis of other consonant categories, i.e. stops, nasals, and liquids, showed that closure duration for stops and consonant duration for nasals and liquids, form the most salient acoustic cues to gemination. Frequency and energy domain parameters were not significantly affected by gemination in a systematic way for all consonant classes. Results on fricatives and affricates confirmed the above findings, i.e., that the primary acoustic correlate of gemination is durational in nature and corresponds to a lengthened consonant duration for fricative geminates and a lengthened closure duration for affricate geminates. An inverse correlation between consonant and pre-consonant vowel durations was present for both consonant categories, and also for both singleton and geminate word sets when considered separately. This effect was reinforced for combined sets, confirming the hypothesis that a durational compensation between different phonemes may serve to preserve rhythmical structures. Classification tests of single vs. geminate consonants using the durational acoustic cues as classification parameters confirmed their validity, and highlighted peculiarities of the two consonant classes. In particular, a relatively poor classification performance was observed for affricates, which led to refining the analysis by considering dental vs. non-dental affricates in two different sets. Results support the hypothesis that dental affricates, in Italian, may not appear in intervocalic position as singletons but only in their geminate form.

Explore related subjects

Keep this discovery

BibTeXRIS

Maria Gabriella Di Benedetto, Luca De Nardis. 2020-04-19. Consonant gemination in Italian: the affricate and fricative case. https://arxiv.org/abs/2005.06959

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Diarization Error Decomposition Under Pause Annotation Ambiguity

Speaker diarization evaluation is sensitive to ambiguity in pause annotation, which can inflate diarization error rate (DER) or obscure genuine model errors. We show that morphological closing, which has been used for pause-tolerant diarization evaluation, discards segment-level distinctions. Instead, we propose an exact, overlap-aware decomposition of standard DER into a pause-attributable component, consisting of errors compatible with pause filling, and a residual core component that can serve as a proxy for intrinsic diarization errors. The decomposition leaves DER unchanged, while the pause-attributable and core components vary monotonically with the pause threshold and eventually saturate. Experiments spanning synthetic transformations, annotation mismatch, cross-domain evaluation, and tight-boundary diarization show that the decomposition reveals error sources not apparent from standard DER.

eess.AS

Less can be More: What Aspects of Speech Drive End-of-Turn Detection

In conversational AI, detecting when a speaker has finished talking is crucial for natural turn taking. While recent work incorporates semantics, the relative contribution of different modalities remains unclear. We present a controlled ablation of acoustic, prosodic, and semantic signals for streaming end of turn detection using a lightweight trimodal classifier. Under identical training conditions, the acoustic prosodic combination achieves the best balance of accuracy and latency, achieving utterance F1 of 0.93 with 7.8% false alarms at 400ms median latency. Adding text increases premature detections without improving performance. Feature space analysis confirms that prosodic features have the strongest class separability, while text representations overlap substantially. These findings suggest that turn-taking is primarily conveyed through intonation and silence patterns rather than semantic completeness, enabling faster and more reliable systems without expensive text inference.

eess.AS

Downstream-Task-Aware Unified Source Separation

Task-aware unified source separation (TUSS) enables a single model to handle diverse separation tasks by conditioning on input prompts. However, conventional TUSS does not account for downstream task requirements, such as whether the enhanced speech will be used for human listening or automatic speech recognition (ASR). In this paper, we propose a prompt extension framework for TUSS that incorporates downstream task information into the input prompts and switches the loss function according to the given prompt during training, enabling outputs with different signal characteristics at inference time. Specifically, we introduce an ASR-dedicated prompt paired with a regularized loss function that reduces speech artifacts to improve ASR robustness, while the standard prompt is paired with the conventional SNR loss function. Experiments on the LibriSpeech and JNAS corpora demonstrate that the proposed joint-training scheme enables a single model to improve ASR performance over noisy input across a wide range of SNR conditions by selecting the ASR-dedicated prompt, while maintaining general speech enhancement quality when the standard prompt is used.

eess.AS