SearcharxivSearch

arXiv subjects

Katarina C. Poole

Publications and source records attributed to Katarina C. Poole.

6 recordsLinked to original sources

Beyond Localisation Accuracy: Sensorimotor Effects of HRTF Individualisation

Everyday listening requires the brain to integrate cues from the body, environment, other senses, and movement, continuously translating auditory information into action. Yet HRTF individualisation is still commonly assessed through localisation accuracy, which may not fully capture its effects on this sensorimotor process. Here, we investigate whether these effects can instead be revealed through behaviour in a more ecologically valid listening task. We used an aurally guided visual search paradigm in which listeners located a visual target using a co-located virtual sound while moving freely, comparing individualised and non-individualised HRTFs under anechoic and reverberant conditions. Performance was assessed through response times and measures of movement organisation. In anechoic conditions, individualised HRTFs produced faster responses than non-individualised HRTFs, with an average reduction of approximately 200ms and the clearest benefit for front-back source locations. This advantage was expressed primarily in movement initiation, whereas overall movement extent was only weakly affected. Under reverberant conditions, HRTF-dependent differences disappeared. These results suggest that HRTF individualisation can influence how listeners plan and initiate orienting actions even when differences in conventional localisation outcomes are limited. Assessing sensorimotor behaviour alongside localisation performance may therefore provide a more sensitive and ecologically relevant account of the perceptual benefits of HRTF individualisation.

eess.AS

Numerical and perceptual validity of synthetic Head-Related Transfer Functions at scale

Individually measuring head-related transfer functions (HRTFs) at scale remains a central challenge for personalised spatial audio, motivating growing interest in synthetic HRTFs. We evaluated the numerical, computational, and behavioural validity of synthetic HRTFs, generated through the boundary element method simulation using Mesh2HRTF, against measured and KEMAR HRTFs using the Extended SONICOM dataset. Across 200 subjects, synthetic HRTFs deviated less from measured than KEMAR in interaural time and level differences, but residual errors, together with elevated spectral distortion, concentrated at low, rear elevations. This is consistent with the omission of torso geometry from the synthesis pipeline. Two computational models revealed a corresponding pattern of predicted localisation errors, with synthetic HRTFs positioned between measured and KEMAR. In a virtual reality localisation task (N = 20), synthetic HRTFs matched measured on every polar metric, while KEMAR was significantly worse. However, behavioural error clustered around the front-back midline regardless of condition, not at the low elevations implicated numerically or by the models. A separate spatial release from masking task (N = 18) showed no effect of HRTF type. Together, these results indicate that high-resolution synthetic HRTFs preserve behavioural localisation performance, despite discrepancies between the numerical/model-predicted bias and the spatial pattern of behavioural error.

eess.AS

Evaluation of Head-Related Transfer Functions Across Five Levels of Individualisation in Virtual Reality

Head-related transfer functions (HRTFs) underpin spatial hearing in virtual and augmented reality systems. Whilst individual HRTFs capture listener-specific morphology, their practical limitations have led to widespread use of generic HRTFs and growing interest in synthetic approaches. Yet their relative perceptual impact remains rarely compared within a single study. In this study, we analysed data from 19 listeners that completed two virtual reality sound localisation experiments with complementary subsets of interleaved HRTF conditions enabling within-subject comparison of five conditions: individually measured, KEMAR, randomly selected non-individual measured, high-resolution scan-based synthetic and photogrammetry-based synthetic HRTFs. Test-retest stability of the individually measured baseline across sessions supported pooling across experiments and attributing differences to perceptual rather than session effects. Across HRTF conditions, lateral localisation metrics were largely insensitive to HRTF type, whereas polar-domain metrics and confusion rates showed strong HRTF dependence. Random HRTFs outperformed KEMAR on several polar metrics. High-resolution synthetic HRTFs matched individual measured performance, whilst photogrammetry-based synthetic HRTFs, alongside KEMAR, showed the greatest degradation. These findings clarify practical choices for non-individual baselines and highlight the importance of mesh resolution when using numerical synthesis for elevation-dependent localisation tasks.

eess.AS

Photogrammetry-Reconstructed 3D Head Meshes for Accessible Individual Head-Related Transfer Functions

Individual head-related transfer functions (HRTFs) are essential for accurate spatial audio binaural rendering but remain difficult to obtain due to measurement complexity. This study investigates whether photogrammetry-reconstructed (PR) head and ear meshes, acquired with consumer hardware, can provide a practically useful baseline for individual HRTF synthesis. Using the SONICOM HRTF dataset, 72-image photogrammetry captures per subject were processed with Apple's Object Capture API to generate PR meshes for 150 subjects. Mesh2HRTF was used to compute PR synthetic HRTFs, which were compared against measured HRTFs, high-resolution 3D scan-derived HRTFs, KEMAR, and random HRTFs through numerical evaluation, auditory models, and a behavioural sound localisation experiment (N = 27). PR synthetic HRTFs preserved ITD cues but exhibited increased ILD and spectral errors. Auditory-model predictions and behavioural data showed substantially higher quadrant error rates, reduced elevation accuracy, and greater front-back confusions than measured HRTFs, performing worse than random HRTFs on perceptual metrics. Current photogrammetry pipelines support individual HRTF synthesis but are limited by insufficient pinna morphology details and high-frequency spectral fidelity needed for accurate individual HRTFs containing monaural cues.

eess.AS

The Extended SONICOM HRTF Dataset and Spatial Audio Metrics Toolbox

Headphone-based spatial audio uses head-related transfer functions (HRTFs) to simulate real-world acoustic environments. HRTFs are unique to everyone, due to personal morphology, shaping how sound waves interact with the body before reaching the eardrums. Here we present the extended SONICOM HRTF dataset which expands on the previous version released in 2023. The total number of measured subjects has now been increased to 300, with demographic information for a subset of the participants, providing context for the dataset's population and relevance. The dataset incorporates synthesised HRTFs for 200 of the 300 subjects, generated using Mesh2HRTF, alongside pre-processed 3D scans of the head and ears, optimised for HRTF synthesis. This rich dataset facilitates rapid and iterative optimisation of HRTF synthesis algorithms, allowing the automatic generation of large data. The optimised scans enable seamless morphological modifications, providing insights into how anatomical changes impact HRTFs, and the larger sample size enhances the effectiveness of machine learning approaches. To support analysis, we also introduce the Spatial Audio Metrics (SAM) Toolbox, a Python package designed for efficient analysis and visualisation of HRTF data, offering customisable tools for advanced research. Together, the extended dataset and toolbox offer a comprehensive resource for advancing personalised spatial audio research and development.

eess.AS

Enhancing Photogrammetry Reconstruction For HRTF Synthesis Via A Graph Neural Network

Traditional Head-Related Transfer Functions (HRTFs) acquisition methods rely on specialised equipment and acoustic expertise, posing accessibility challenges. Alternatively, high-resolution 3D modelling offers a pathway to numerically synthesise HRTFs using Boundary Elements Methods and others. However, the high cost and limited availability of advanced 3D scanners restrict their applicability. Photogrammetry has been proposed as a solution for generating 3D head meshes, though its resolution limitations restrict its application for HRTF synthesis. To address these limitations, this study investigates the feasibility of using Graph Neural Networks (GNN) using neural subdivision techniques for upsampling low-resolution Photogrammetry-Reconstructed (PR) meshes into high-resolution meshes, which can then be employed to synthesise individual HRTFs. Photogrammetry data from the SONICOM dataset are processed using Apple Photogrammetry API to reconstruct low-resolution head meshes. The dataset of paired low- and high-resolution meshes is then used to train a GNN to upscale low-resolution inputs to high-resolution outputs, using a Hausdorff Distance-based loss function. The GNN's performance on unseen photogrammetry data is validated geometrically and through synthesised HRTFs generated via Mesh2HRTF. Synthesised HRTFs are evaluated against those computed from high-resolution 3D scans, to acoustically measured HRTFs, and to the KEMAR HRTF using perceptually-relevant numerical analyses as well as behavioural experiments, including localisation and Spatial Release from Masking (SRM) tasks.

eess.AS