SearcharxivSearch

arXiv subjects

Lloyd May

Publications and source records attributed to Lloyd May.

3 recordsLinked to original sources

Seeing the Voice, Preserving the Self: A Participatory Design Approach to Deaf-Centric Text-to-Speech

We describe a participatory design approach toward developing Deaf-centric text-to-speech (TTS) technologies. While TTS is growing rapidly in the mainstream, it has received little attention to date in the deaf and hard of hearing (DHH) technology space. Critical problems have remained unaddressed for DHH users, including the ability to manipulate tone, emotions and delivery via non-auditory means. Verifying that the generated speech matches intent and is appropriate for a given situation without having to listen to it is another challenge. Respecting cultural and identity factors in the generated speech is also important. This work explores the design space with DHH participants through two focus groups, three co-design sessions, and four one-on-one early-stage design evaluation sessions. Participants included people both familiar and unfamiliar with TTS, as well as DHH content creators. We describe key findings, design ideas, results, and implications for future Deaf-centric TTS development. We also identify unmet technology requirements that pose barriers to adoption of Deaf-centric TTS technology.

cs.HC

Decoding Imagined Auditory Pitch Phenomena with an Autoencoder Based Temporal Convolutional Architecture

Stimulus decoding of functional Magnetic Resonance Imaging (fMRI) data with machine learning models has provided new insights about neural representational spaces and task-related dynamics. However, the scarcity of labelled (task-related) fMRI data is a persistent obstacle, resulting in model-underfitting and poor generalization. In this work, we mitigated data poverty by extending a recent pattern-encoding strategy from the visual memory domain to our own domain of auditory pitch tasks, which to our knowledge had not been done. Specifically, extracting preliminary information about participants' neural activation dynamics from the unlabelled fMRI data resulted in improved downstream classifier performance when decoding heard and imagined pitch. Our results demonstrate the benefits of leveraging unlabelled fMRI data against data poverty for decoding pitch based tasks, and yields novel significant evidence for both separate and overlapping pathways of heard and imagined pitch processing, deepening our understanding of auditory cognitive neuroscience.

q-bio.NC

The Role of Vocal Persona in Natural and Synthesized Speech

The inclusion of voice persona in synthesized voice can be significant in a broad range of human-computer-interaction (HCI) applications, including augmentative and assistive communication (AAC), artistic performance, and design of virtual agents. We propose a framework to imbue compelling and contextually-dependent expression within a synthesized voice by introducing the role of the vocal persona within a synthesis system. In this framework, the resultant 'tone of voice' is defined as a point existing within a continuous, contextually-dependent probability space that is traversable by the user of the voice. We also present initial findings of a thematic analysis of 10 interviews with vocal studies and performance experts to further understand the role of the vocal persona within a natural communication ecology. The themes identified are then used to inform the design of the aforementioned framework.

cs.SD