arXiv · 2606.08748
HydraQE: OSU's Submission for the IWSLT 2026 Speech Translation Metrics Shared Task
Abstract
We present HydraQE, our contribution to the IWSLT 2026 Speech Translation Metrics shared task. HydraQE is an end-to-end, reference-free quality estimation (QE) system for speech translation built on a Qwen3-ASR backbone, which accepts source audio and a translation hypothesis as joint input. Hidden states from all backbone layers are combined via a learnable sparsemax scalar mix, then re-encoded by a lightweight bidirectional Transformer to enable full cross-modal interaction prior to pooling into a shared embedding. Three independent prediction heads are trained on complementary supervision signals: human direct assessment (DA) annotations, MetricX-24 pseudo-labels, and xCOMET pseudo-labels. To address the scarcity of human-annotated data, we train on a combination of synthetically corrupted examples and silver pseudo-labeled machine translation outputs, using a curriculum that begins on synthetic and silver data and gradually shifts toward human-annotated examples. HydraQE outperforms cascaded text-based baselines and prior direct speech QE systems, demonstrating that end-to-end speech translation QE is competitive with cascaded approaches.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Kevin Krahn, Eric Fosler-Lussier. 2026-06-07. HydraQE: OSU's Submission for the IWSLT 2026 Speech Translation Metrics Shared Task. https://arxiv.org/abs/2606.08748
Cite the original work for its findings. Save a collection to share your selection of sources.