SearcharxivSearch

arXiv subjects

Anne-Lise Giraud

Publications and source records attributed to Anne-Lise Giraud.

2 recordsLinked to original sources

Transfer Learning from ImageNet for MEG-Based Decoding of Imagined Speech

Non-invasive decoding of imagined speech remains challenging due to weak, distributed signals and limited labeled data. Our paper introduces an image-based approach that transforms magnetoencephalography (MEG) signals into time-frequency representations compatible with pretrained vision models. MEG data from 21 participants performing imagined speech tasks were projected into three spatial scalogram mixtures via a learnable sensor-space convolution, producing compact image-like inputs for ImageNet-pretrained vision architectures. These models outperformed classical and non-pretrained models, achieving up to 90.4% balanced accuracy for imagery vs. silence, 81.0% vs. silent reading, and 60.6% for vowel decoding. Cross-subject evaluation confirmed that pretrained models capture shared neural representations, and temporal analyses localized discriminative information to imagery-locked intervals. These findings show that pretrained vision models applied to image-based MEG representations can effectively capture the structure of imagined speech in non-invasive neural signals.

cs.CL

Neuro-oscillatory models of cortical speech processing

In this review, we examine computational models that explore the role of neural oscillations in speech perception, spanning from early auditory processing to higher cognitive stages. We focus on models that use rhythmic brain activities, such as gamma, theta, and delta oscillations, to encode phonemes, segment speech into syllables and words, and integrate linguistic elements to infer meaning. We analyze the mechanisms underlying these models, their biological plausibility, and their potential applications in processing and understanding speech in real time, a computational feature that is achieved by the human brain but not yet implemented in speech recognition models. Real-time processing enables dynamic adaptation to incoming speech, allowing systems to handle the rapid and continuous flow of auditory information required for effective communication, interactive applications, and accurate speech recognition in a variety of real-world settings. While significant progress has been made in modeling the neural basis of speech perception, challenges remain, particularly in accounting for the complexity of semantic processing and the integration of contextual influences. Moreover, the high computational demands of biologically realistic models pose practical difficulties for their implementation and analysis. Despite these limitations, these models provide valuable insights into the neural mechanisms of speech perception. We conclude by identifying current limitations, proposing future research directions, and suggesting how these models can be further developed to achieve a more comprehensive understanding of speech processing in the human brain.

q-bio.NC