SearcharxivSearch

arXiv subjects

Jorge Martinez

Publications and source records attributed to Jorge Martinez.

10 recordsLinked to original sources

Comparing Human and Automatic Recognition of Dutch Dysarthric Continuous Speech: A Case Study

In our goal to develop personalised dysarthric speech recognition (DSR) models, this study compared the recognition performances of human listeners and those of three state-of-the-art, off-the-shelf ASR systems (Whisper-large-V3, Google Chirp 3, and Omnilingual) on the recognition of Dutch continuous read and spontaneous speech from a single speaker with severe dysarthria. Results showed that both humans listeners and the three off-the-shelf ASR systems exhibit word error rates (WER) exceeding 70% on average, indicating that DSR is highly challenging for both humans and ASR systems. Fine-tuning on the dysarthric speech significantly reduced WER. Although overall WERs are still quite high (>23%), the personalised DSR models outperformed the human listeners, and performance is getting closer to being useful for supporting day-to-day communication of dysarthric speakers. Future research should focus on improving personalized DSR on spontaneous speech and longer utterances in the case of read speech, with a specific focus on particular phonemes.

cs.CL

A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition

We present DRES: a 1.5-hour Dutch realistic elicited (semi-spontaneous) speech dataset from 80 speakers recorded in noisy, public indoor environments. DRES was designed as a test set for the evaluation of state-of-the-art (SotA) automatic speech recognition (ASR) and speech enhancement (SE) models in a real-world scenario: a person speaking in a public indoor space with background talkers and noise. The speech was recorded with a four-channel linear microphone array. In this work we evaluate the speech quality of five well-known single-channel SE algorithms and the recognition performance of eight SotA off-the-shelf ASR models before and after applying SE on the speech of DRES. We found that five out of the eight ASR models have WERs lower than 22\% on DRES, despite the challenging conditions. In contrast to recent work, we did not find a positive effect of modern single-channel SE on ASR performance, emphasizing the importance of evaluating in realistic conditions.

eess.AS

Performance of Objective Speech Quality Metrics on Languages Beyond Validation Data: A Study of Turkish and Korean

Objective speech quality measures are widely used to assess the performance of video conferencing platforms and telecommunication systems. They predict human-rated speech quality and are crucial for assessing the systems quality of experience. Despite the widespread use, the quality measures are developed on a limited set of languages. This can be problematic since the performance on unseen languages is consequently not guaranteed or even studied. Here we raise awareness to this issue by investigating the performance of two objective speech quality measures (PESQ and ViSQOL) on Turkish and Korean. Using English as baseline, we show that Turkish samples have significantly higher ViSQOL scores and that for Turkish male speakers the correlation between PESQ and ViSQOL is highest. These results highlight the need to explore biases across metrics and to develop a labeled speech quality dataset with a variety of languages.

eess.AS

Loudspeaker Beamforming to Enhance Speech Recognition Performance of Voice Driven Applications

In this paper we propose a robust loudspeaker beamforming algorithm which is used to enhance the performance of voice driven applications in scenarios where the loudspeakers introduce the majority of the noise, e.g. when music is playing loudly. The loudspeaker beamformer modifies the loudspeaker playback signals to create a low-acoustic-energy region around the device that implements automatic speech recognition for a voice driven application (VDA). The algorithm utilises a distortion measure based on human auditory perception to limit the distortion perceived by human listeners. Simulations and real-world experiments show that the proposed loudspeaker beamformer improves the speech recognition performance in all tested scenarios. Moreover, the algorithm allows to further reduce the acoustic energy around the VDA device at the expense of reduced objective audio quality at the listener's location.

eess.AS

On the Integration of Acoustics and LiDAR: a Multi-Modal Approach to Acoustic Reflector Estimation

Having knowledge on the room acoustic properties, e.g., the location of acoustic reflectors, allows to better reproduce the sound field as intended. Current state-of-the-art methods for room boundary detection using microphone measurements typically focus on a two-dimensional setting, causing a model mismatch when employed in real-life scenarios. Detection of arbitrary reflectors in three dimensions encounters practical limitations, e.g., the need for a spherical array and the increased computational complexity. Moreover, loudspeakers may not have an omnidirectional directivity pattern, as usually assumed in the literature, making the detection of acoustic reflectors in some directions more challenging. In the proposed method, a LiDAR sensor is added to a loudspeaker to improve wall detection accuracy and robustness. This is done in two ways. First, the model mismatch introduced by horizontal reflectors can be resolved by detecting reflectors with the LiDAR sensor to enable elimination of their detrimental influence from the 2D problem in pre-processing. Second, a LiDAR-based method is proposed to compensate for the challenging directions where the directive loudspeaker emits little energy. We show via simulations that this multi-modal approach, i.e., combining microphone and LiDAR sensors, improves the robustness and accuracy of wall detection.

eess.AS

Propelling Interplanetary Spacecraft Utilizing Water-Steam

Water has been identified as a critical resource both to sustain human-life but also for use in propulsion, attitude-control, power, thermal and radiation pro-tection systems. Water may be obtained off-world through In-Situ Resource Utilization (ISRU) in the course of human or robotic space exploration that replace materials that would otherwise be shipped from Earth. Water has been highlighted by many in the space community as a credible solution for affordable/sustainable exploration. Water can be extracted from the Moon, C-class Near Earth Objects (NEOs), surface of Mars and Martian Moons Pho-bos and Deimos and from the surface of icy, rugged terrains of Ocean Worlds. However, use of water for propulsion faces some important techno-logical barriers. A technique to use water as a propellant is to electrolyze it into hydrogen and oxygen that is then pulse-detonated. High-efficiency elec-trolysis requires use of platinum-catalyst based fuel cells. Even trace ele-ments of sulfur and carbon monoxide found on planetary bodies can poison these cells making them unusable. In this work, we develop steam-based propulsion that avoids the technological barriers of electrolyzing impure water as propellant. Using a solar concentrator, heat is used to extract the water which is then condensed as a liquid and stored. Steam is then formed using the solar thermal reflectors to concentrate the light into a nanoparticle-water mix. This solar thermal heating (STH) process converts 80 to 99% of the in-coming light into heat.

astro-ph.IM

Reverberation Mapping of Luminous Quasars at High-z

We present Reverberation Mapping (RM) results for 17 high-redshift, high-luminosity quasars with good quality R-band and emission line light curves. We are able to measure statistically significant lags for Ly_alpha (11 objects), SiIV (5 objects), CIV (11 objects), and CIII] (2 objects). Using our results and previous lag determinations taken from the literature, we present an updated CIV radius--luminosity relation and provide for the first time radius--luminosity relations for Ly_alpha, SiIV and CIII]. While in all cases the slope of the correlations are statistically significant, the zero points are poorly constrained because of the lack of data at the low luminosity end. We find that the emissivity weighted distance from the central source of the Ly_alpha, SiIV and CIII] line emitting regions are all similar, which corresponds to about half that of the H_beta region. We also find that 3/17 of our sources show an unexpected behavior in some emission lines, two in the Ly_alpha light curve and one in the SiIV light curve, in that they do not seem to follow the variability of the UV continuum. Finally, we compute RM black hole masses for those quasars with highly significant lag measurements and compare them with CIV single--epoch (SE) mass determinations. We find that the RM-based black hole mass determinations seem smaller than those found using SE calibrations.

astro-ph.GA

Serendipitous discovery of RR Lyrae stars in the Leo V ultra-faint galaxy

During the analysis of RR Lyrae stars discovered in the High cadence Transient Survey (HiTS) taken with the Dark Energy Camera at the 4-m telescope at Cerro Tololo Inter-American Observatory, we found a group of three very distant, fundamental mode pulsator RR Lyrae (type ab). The location of these stars agrees with them belonging to the Leo V ultra-faint satellite galaxy, for which no variable stars have been reported to date. The heliocentric distance derived for Leo V based on these stars is 173 +/- 5 kpc. The pulsational properties (amplitudes and periods) of these stars locate them within the locus of the Oosterhoff II group, similar to most other ultra-faint galaxies with known RR Lyrae stars. This serendipitous discovery shows that distant RR Lyrae stars may be used to search for unknown faint stellar systems in the outskirts of the Milky Way.

astro-ph.GA

Unsupervised edge map scoring: a statistical complexity approach

We propose a new Statistical Complexity Measure (SCM) to qualify edge maps without Ground Truth (GT) knowledge. The measure is the product of two indices, an \emph{Equilibrium} index $\mathcal{E}$ obtained by projecting the edge map into a family of edge patterns, and an \emph{Entropy} index $\mathcal{H}$, defined as a function of the Kolmogorov Smirnov (KS) statistic. This new measure can be used for performance characterization which includes: (i)~the specific evaluation of an algorithm (intra-technique process) in order to identify its best parameters, and (ii)~the comparison of different algorithms (inter-technique process) in order to classify them according to their quality. Results made over images of the South Florida and Berkeley databases show that our approach significantly improves over Pratt's Figure of Merit (PFoM) which is the objective reference-based edge map evaluation standard, as it takes into account more features in its evaluation.

cs.CV

Accuracy of MAP segmentation with hidden Potts and Markov mesh prior models via Path Constrained Viterbi Training, Iterated Conditional Modes and Graph Cut based algorithms

In this paper, we study statistical classification accuracy of two different Markov field environments for pixelwise image segmentation, considering the labels of the image as hidden states and solving the estimation of such labels as a solution of the MAP equation. The emission distribution is assumed the same in all models, and the difference lays in the Markovian prior hypothesis made over the labeling random field. The a priori labeling knowledge will be modeled with a) a second order anisotropic Markov Mesh and b) a classical isotropic Potts model. Under such models, we will consider three different segmentation procedures, 2D Path Constrained Viterbi training for the Hidden Markov Mesh, a Graph Cut based segmentation for the first order isotropic Potts model, and ICM (Iterated Conditional Modes) for the second order isotropic Potts model. We provide a unified view of all three methods, and investigate goodness of fit for classification, studying the influence of parameter estimation, computational gain, and extent of automation in the statistical measures Overall Accuracy, Relative Improvement and Kappa coefficient, allowing robust and accurate statistical analysis on synthetic and real-life experimental data coming from the field of Dental Diagnostic Radiography. All algorithms, using the learned parameters, generate good segmentations with little interaction when the images have a clear multimodal histogram. Suboptimal learning proves to be frail in the case of non-distinctive modes, which limits the complexity of usable models, and hence the achievable error rate as well. All Matlab code written is provided in a toolbox available for download from our website, following the Reproducible Research Paradigm.

cs.LG