SearcharxivSearch

arXiv subjects

James Waller

Publications and source records attributed to James Waller.

3 recordsLinked to original sources

Are Caption Metrics Broken? Latency, Deaf and Hard of Hearing User Ratings, and Bias across Technologies

Live captions on TV often contain errors and timing issues, making it hard for deaf and hard-of-hearing (DHH) viewers to follow dialog. It is essential that caption quality metrics reflect the lived DHH TV viewing experience. To this end, we describe a U.S.-based large-scale online survey with 216 validated participants, who provided 302 responses containing a cumulative 4,832 data points. Participants viewed videos drawn from a pool of 70 clips recorded from live TV, and were asked to rate the caption quality and subjective understanding of the content across four conditions: TV captions as originally recorded with up to 7-12 seconds delay, TV captions synchronized with audio, Automatic Speech Recognition (ASR)-generated captions synchronized with audio, and ASR captions with an average two-second delay. All captions were evaluated against the Word Error Rate (WER), Automated Caption Evaluation (ACE2) and Number, Edition and Recognition (NER) metrics. Results show that TV and ASR captions were rated similarly. For TV captions, all three metrics were moderately-to-highly correlated with viewer ratings, but far less so for ASR captions, making them far from technology-neutral. Additionally, caption latencies significantly impact the viewer experience, especially typical 7-12-second TV delays. We discuss the implications for the adoption of caption quality metrics.

cs.HC

Optical and magnetic response by design in GaAs quantum dots

Quantum networking technologies use spin qubits and their interface to single photons as core components of a network node. This necessitates the ability to co-design the magnetic- and optical-dipole response of a quantum system. These properties are notoriously difficult to design in many solid-state systems, where spin-orbit coupling and the crystalline environment for each qubit create inhomogeneity of electronic g-factors and optically active states. Here, we show that GaAs quantum dots (QDs) obtained via the quasi-strain-free local droplet etching epitaxy growth method provide spin and optical properties predictable from assuming the highest possible QD symmetry. Our measurements of electron and hole g-tensors and of transition dipole moment orientations for charged excitons agree with our predictions from a multiband k.p simulation constrained only by a single atomic-force-microscopy reconstruction of QD morphology. This agreement is verified across multiple wavelength-specific growth runs at different facilities within the range of 730 nm to 790 nm for the exciton emission. Remarkably, our measurements and simulations track the in-plane electron g-factors through a zero-crossing from -0.1 to 0.3 and linear optical dipole moment orientations fully determined by an external magnetic field. The robustness of our results demonstrates the capability to design - prior to growth - the properties of a spin qubit and its tunable optical interface best adapted to a target magnetic and photonic environment with direct application for high-quality spin-photon entanglement.

quant-ph

Live Captions in Virtual Reality (VR)

Few VR applications and games implement captioning of speech and audio cues, which either inhibits or prevents access of their application by deaf or hard of hearing (DHH) users, new language learners, and other caption users. Additionally, little to no guidelines exist on how to implement live captioning on VR headsets and how it may differ from traditional television captioning. To help fill the void of information behind user preferences of different VR captioning styles, we conducted a study with eight DHH participants to test three caption movement behaviors (headlocked, lag, and appear) while watching live-captioned, single-speaker presentations in VR. Participants answered a series of Likert scale and open-ended questions about their experience. Participant preferences were split, but the majority of participants reported feeling comfortable with using live captions in VR and enjoyed the experience. When participants ranked the caption behaviors, there was almost an equal divide between the three types tested. IPQ results indicated each behavior had similar immersion ratings, however participants found headlocked and lag captions more user-friendly than appear captions. We suggest that participants may vary in caption preference depending on how they use captions, and that providing opportunities for caption customization is best.

cs.HC