SearcharxivSearch

arXiv subjects

Matthieu Puigt

Publications and source records attributed to Matthieu Puigt.

5 recordsLinked to original sources

SVH-BD : Synthetic Vegetation Hyperspectral Benchmark Dataset for Emulation of Remote Sensing Images

This dataset provides a large collection of 10,915 synthetic hyperspectral image cubes paired with pixel-level vegetation trait maps, designed to support research in radiative transfer emulation, vegetation trait retrieval, and uncertainty quantification. Each hyperspectral cube contains 211 bands spanning 400--2500 nm at 10 nm resolution and a fixed spatial layout of 64 \times 64 pixels, offering continuous simulated surface reflectance spectra suitable for emulator development and machine-learning tasks requiring high spectral detail. Vegetation traits were derived by inverting Sentinel-2 Level-2A surface reflectance using a PROSAIL-based lookup-table approach, followed by forward PROSAIL simulations to generate hyperspectral reflectance under physically consistent canopy and illumination conditions. The dataset covers four ecologically diverse regions -- East Africa, Northern France, Eastern India, and Southern Spain -- and includes 5th and 95th percentile uncertainty maps as well as Sentinel-2 scene classification layers. This resource enables benchmarking of inversion methods, development of fast radiative transfer emulators, and studies of spectral--biophysical relationships under controlled yet realistic environmental variability.

cs.CV

A Latent Representation Learning Framework for Hyperspectral Image Emulation in Remote Sensing

Synthetic hyperspectral image (HSI) generation is essential for large-scale simulation, algorithm development, and mission design, yet traditional radiative transfer models remain computationally expensive and proposed emulation methods are often limited to spectrum-level outputs. In this work, we propose a latent representation-based framework for hyperspectral emulation that learns a probabilistic latent representation of hyperspectral data. The proposed approach supports both spectrum-level and spatial-spectral emulation and can be trained either in a direct one-step formulation or in a two-step strategy that couples variational autoencoder (VAE) pretraining with parameter-to-latent mapping. Experiments on PROSAIL-simulated vegetation data and Sentinel-3 OLCI imagery demonstrate that the method outperforms classical regression-based emulators in reconstruction accuracy, spectral fidelity, and robustness to real-world spatial variability. We further show that emulated HSIs preserve performance in downstream biophysical parameter retrieval, highlighting the practical relevance of emulated data for remote sensing applications.

cs.CV

Gaussian Compression Stream: Principle and Preliminary Results

Random projections became popular tools to process big data. In particular, when applied to Nonnegative Matrix Factorization (NMF), it was shown that structured random projections were far more efficient than classical strategies based on Gaussian compression. However, they remain costly and might not fully benefit from recent fast random projection techniques. In this paper, we thus investigate an alternative to structured ran-om projections-named Gaussian compression stream-which (i) is based on Gaussian compressions only, (ii) can benefit from the above fast techniques, and (iii) is shown to be well-suited to NMF.

eess.SP

Faster-than-fast NMF using random projections and Nesterov iterations

Random projections have been recently implemented in Nonnegative Matrix Factorization (NMF) to speed-up the NMF computations, with a negligible loss of performance. In this paper, we investigate the effects of such projections when the NMF technique uses the fast Nesterov gradient descent (NeNMF). We experimentally show the randomized subspace iteration to significantly speed-up NeNMF.

eess.SP

Post-Nonlinear Sparse Component Analysis Using Single-Source Zones and Functional Data Clustering

In this paper, we introduce a general extension of linear sparse component analysis (SCA) approaches to postnonlinear (PNL) mixtures. In particular, and contrary to the state-of-art methods, our approaches use a weak sparsity source assumption: we look for tiny temporal zones where only one source is active. We investigate two nonlinear single-source confidence measures, using the mutual information and a local linear tangent space approximation (LTSA). For this latter measure, we derive two extensions of linear single-source measures, respectively based on correlation (LTSA-correlation) and eigenvalues (LTSA-PCA). A second novelty of our approach consists of applying functional data clustering techniques to the scattered observations in the above single-source zones, thus allowing us to accurately estimate them.We first study a classical approach using a B-spline approximation, and then two approaches which locally approximate the nonlinear functions as lines. Finally, we extend our PNL methods to more general nonlinear mixtures. Combining single-source zones and functional data clustering allows us to tackle speech signals, which has never been performed by other PNL-SCA methods. We investigate the performance of our approaches with simulated PNL mixtures of real speech signals. Both the mutual information and the LTSA-correlation measures are better-suited to detecting single-source zones than the LTSA-PCA measure. We also find local-linear-approximation-based clustering approaches to be more flexible and more accurate than the B-spline one.

cs.IT