arXiv · 2610.00852
Child-Adapted Structured Phonological Representations for Interpretable Speech Sound Analysis
Abstract
Structured phonological representations provide an interpretable alternative to generic speech embeddings, but existing models are largely trained on adult speech. We adapt PhonoQ-2.0 to child speech using CHILDES-Aligned data and compare three alignment-supervision conditions (Adult, Adult+Child, and Child-only) across two initialization strategies (Adult PhonoQ and scratch). Generalization is evaluated against manual child-speech annotations. On 1,352 consonant targets from 58 typically developing children, child-speech adaptation improves voicing recognition across all supervision conditions, from 0.922 macro-F1 for Adult PhonoQ to 0.972--0.987 after adaptation. Manner is more sensitive to alignment supervision: Adult+Child MFA reaches 0.804 and 0.796, compared to approximately 0.70 under Adult MFA supervision. Place remains comparatively strong across systems (0.871--0.902), although per-class performance varies substantially. The velar-fronting contrast is preserved across all seven model variants. Longitudinal UltraPhonix analysis further reveals speaker-specific velar and post-alveolar changes that are largely preserved across models and broadly consistent with reported clinical progress.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Abner Hernandez, Tomás Arias Vergara, Andreas Maier, Paula Andrea Pérez-Toro. 2026-10-01. Child-Adapted Structured Phonological Representations for Interpretable Speech Sound Analysis. https://arxiv.org/abs/2610.00852
Cite the original work for its findings. Save a collection to share your selection of sources.