SearcharxivSearch

arXiv subjects

Peter Gerstoft

Publications and source records attributed to Peter Gerstoft.

At least 37 records · Page 2Linked to original sources

Robust and Sparse M-Estimation of DOA

A robust and sparse Direction of Arrival (DOA) estimator is derived for array data that follows a Complex Elliptically Symmetric (CES) distribution with zero-mean and finite second-order moments. The derivation allows to choose the loss function and four loss functions are discussed in detail: the Gauss loss which is the Maximum-Likelihood (ML) loss for the circularly symmetric complex Gaussian distribution, the ML-loss for the complex multivariate $t$-distribution (MVT) with $ν$ degrees of freedom, as well as Huber and Tyler loss functions. For Gauss loss, the method reduces to Sparse Bayesian Learning (SBL). The root mean square DOA error of the derived estimators is discussed for Gaussian, MVT, and $ε$-contaminated data. The robust SBL estimators perform well for all cases and nearly identical with classical SBL for Gaussian noise.

math.ST

Gridless DOA Estimation with Multiple Frequencies

Direction-of-arrival (DOA) estimation is widely applied in acoustic source localization. A multi-frequency model is suitable for characterizing the broadband structure in acoustic signals. In this paper, the continuous (gridless) DOA estimation problem with multiple frequencies is considered. This problem is formulated as an atomic norm minimization (ANM) problem. The ANM problem is equivalent to a semi-definite program (SDP) which can be solved by an off-the-shelf SDP solver. The dual certificate condition is provided to certify the optimality of the SDP solution so that the sources can be localized by finding the roots of a polynomial. We also construct the dual polynomial to satisfy the dual certificate condition and show that such a construction exists when the source amplitude has a uniform magnitude. In multi-frequency ANM, spatial aliasing of DOAs at higher frequencies can cause challenges. We discuss this issue extensively and propose a robust solution to combat aliasing. Numerical results support our theoretical findings and demonstrate the effectiveness of the proposed method.

eess.SP

Graph-based sequential beamforming

This paper presents a Bayesian estimation method for sequential direction finding. The proposed method estimates the number of directions of arrivals (DOAs) and their DOAs performing operations on the factor graph. The graph represents a statistical model for sequential beamforming. At each time step, belief propagation predicts the number of DOAs and their DOAs using posterior probability density functions (pdfs) from the previous time and a different Bernoulli-von Mises state transition model. Variational Bayesian inference then updates the number of DOAs and their DOAs. The method promotes sparse solutions through a Bernoulli-Gaussian amplitude model, is gridless, and provides marginal posterior pdfs from which DOA estimates and their uncertainties can be extracted. Compared to nonsequential approaches, the method can reduce DOA estimation errors in scenarios involving multiple time steps and time-varying DOAs. Simulation results demonstrate performance improvements compared to state-of-the-art methods. The proposed method is evaluated using ocean acoustic experimental data.

eess.SP

Spoofing Attack Detection in the Physical Layer with Commutative Neural Networks

In a spoofing attack, an attacker impersonates a legitimate user to access or tamper with data intended for or produced by the legitimate user. In wireless communication systems, these attacks may be detected by relying on features of the channel and transmitter radios. In this context, a popular approach is to exploit the dependence of the received signal strength (RSS) at multiple receivers or access points with respect to the spatial location of the transmitter. Existing schemes rely on long-term estimates, which makes it difficult to distinguish spoofing from movement of a legitimate user. This limitation is here addressed by means of a deep neural network that implicitly learns the distribution of pairs of short-term RSS vector estimates. The adopted network architecture imposes the invariance to permutations of the input (commutativity) that the decision problem exhibits. The merits of the proposed algorithm are corroborated on a data set that we collected.

cs.LG

Deep Learning Object Detection Approaches to Signal Identification

Traditionally source identification is solved using threshold based energy detection algorithms. These algorithms frequently sum up the activity in regions, and consider regions above a specific activity threshold to be sources. While these algorithms work for the majority of cases, they often fail to detect signals that occupy small frequency bands, fail to distinguish sources with overlapping frequency bands, and cannot detect any signals under a specified signal to noise ratio. Through the conversion of raw signal data to spectrogram, source identification can be framed as an object detection problem. By leveraging modern advancements in deep learning based object detection, we propose a system that manages to alleviate the failure cases encountered when using traditional source identification algorithms. Our contributions include framing source identification as an object detection problem, the publication of a spectrogram object detection dataset, and evaluation of the RetinaNet and YOLOv5 object detection models trained on the dataset. Our final models achieve Mean Average Precisions of up to 0.906. With such a high Mean Average Precision, these models are sufficiently robust for use in real world applications.

cs.SD

Audio scene monitoring using redundant ad-hoc microphone array networks

We present a system for localizing sound sources in a room with several ad-hoc microphone arrays. Each circular array performs direction of arrival (DOA) estimation independently using commercial software. The DOAs are fed to a fusion center, concatenated, and used to perform the localization based on two proposed methods, which require only few labeled source locations (anchor points) for training. The first proposed method is based on principal component analysis (PCA) of the observed DOA and does not require any knowledge of anchor points. The array cluster can then perform localization on a manifold defined by the PCA of concatenated DOAs over time. The second proposed method performs localization using an affine transformation between the DOA vectors and the room manifold. The PCA has fewer requirements on the training sequence, but is less robust to missing DOAs from one of the arrays. The methods are demonstrated with five IoT 8-microphone circular arrays, placed at unspecified fixed locations in an office. Both the PCA and the affine method can easily map out a rectangle based on a few anchor points with similar accuracy. The proposed methods provide a step towards monitoring activities in a smart home and require little installation effort as the array locations are not needed.

cs.SD

Semi-supervised source localization in reverberant environments with deep generative modeling

We propose a semi-supervised approach to acoustic source localization in reverberant environments based on deep generative modeling. Localization in reverberant environments remains an open challenge. Even with large data volumes, the number of labels available for supervised learning in reverberant environments is usually small. We address this issue by performing semi-supervised learning (SSL) with convolutional variational autoencoders (VAEs) on reverberant speech signals recorded with microphone arrays. The VAE is trained to generate the phase of relative transfer functions (RTFs) between microphones, in parallel with a direction of arrival (DOA) classifier based on RTF-phase. These models are trained using both labeled and unlabeled RTF-phase sequences. In learning to perform these tasks, the VAE-SSL explicitly learns to separate the physical causes of the RTF-phase (i.e., source location) from distracting signal characteristics such as noise and speech activity. Relative to existing semi-supervised localization methods in acoustics, VAE-SSL is effectively an end-to-end processing approach which relies on minimal preprocessing of RTF-phase features. As far as we are aware, our paper presents the first approach to modeling the physics of acoustic propagation using deep generative modeling. The VAE-SSL approach is compared with two signal processing-based approaches, steered response power with phase transform (SRP-PHAT) and MUltiple SIgnal Classification (MUSIC), as well as fully supervised CNNs. We find that VAE-SSL can outperform the conventional approaches and the CNN in label-limited scenarios. Further, the trained VAE-SSL system can generate new RTF-phase samples, which shows the VAE-SSL approach learns the physics of the acoustic environment. The generative modeling in VAE-SSL thus provides a means of interpreting the learned representations.

eess.SP

Deep embedded clustering of coral reef bioacoustics

Deep clustering was applied to unlabeled, automatically detected signals in a coral reef soundscape to distinguish fish pulse calls from segments of whale song. Deep embedded clustering (DEC) learned latent features and formed classification clusters using fixed-length power spectrograms of the signals. Handpicked spectral and temporal features were also extracted and clustered with Gaussian mixture models (GMM) and conventional clustering. DEC, GMM, and conventional clustering were tested on simulated datasets of fish pulse calls (fish) and whale song units (whale) with randomized bandwidth, duration, and SNR. Both GMM and DEC achieved high accuracy and identified clusters with fish, whale, and overlapping fish and whale signals. Conventional clustering methods had low accuracy in scenarios with unequal-sized clusters or overlapping signals. Fish and whale signals recorded near Hawaii in February-March 2020 were clustered with DEC, GMM, and conventional clustering. DEC features demonstrated the highest accuracy of 77.5% on a small, manually labeled dataset for classifying signals into fish and whale clusters.

cs.LG

SSLIDE: Sound Source Localization for Indoors based on Deep Learning

This paper presents SSLIDE, Sound Source Localization for Indoors using DEep learning, which applies deep neural networks (DNNs) with encoder-decoder structure to localize sound sources with random positions in a continuous space. The spatial features of sound signals received by each microphone are extracted and represented as likelihood surfaces for the sound source locations in each point. Our DNN consists of an encoder network followed by two decoders. The encoder obtains a compressed representation of the input likelihoods. One decoder resolves the multipath caused by reverberation, and the other decoder estimates the source location. Experiments based on both the simulated and experimental data show that our method can not only outperform multiple signal classification (MUSIC), steered response power with phase transform (SRP-PHAT), sparse Bayesian learning (SBL), and a competing convolutional neural network (CNN) approach in the reverberant environment but also achieve a good generalization performance.

eess.AS

Alternating projections gridless covariance-based estimation for DOA

We present a gridless sparse iterative covariance-based estimation method based on alternating projections for direction-of-arrival (DOA) estimation. The gridless DOA estimation is formulated in the reconstruction of Toeplitz-structured low rank matrix, and is solved efficiently with alternating projections. The method improves resolution by achieving sparsity, deals with single-snapshot data and coherent arrivals, and, with co-prime arrays, estimates more DOAs than the number of sensors. We evaluate the proposed method using simulation results focusing on co-prime arrays.

eess.SP

Semi-supervised source localization with deep generative modeling

We propose a semi-supervised localization approach based on deep generative modeling with variational autoencoders (VAEs). Localization in reverberant environments remains a challenge, which machine learning (ML) has shown promise in addressing. Even with large data volumes, the number of labels available for supervised learning in reverberant environments is usually small. We address this issue by performing semi-supervised learning (SSL) with convolutional VAEs. The VAE is trained to generate the phase of relative transfer functions (RTFs), in parallel with a DOA classifier, on both labeled and unlabeled RTF samples. The VAE-SSL approach is compared with SRP-PHAT and fully-supervised CNNs. We find that VAE-SSL can outperform both SRP-PHAT and CNN in label-limited scenarios.

eess.AS

Gridless DOA Estimation and Root-MUSIC for Non-Uniform Arrays

The problem of gridless direction of arrival (DOA) estimation is addressed in the non-uniform array (NUA) case. Traditionally, gridless DOA estimation and root-MUSIC are only applicable for measurements from a uniform linear array (ULA). This is because the sample covariance matrix of ULA measurements has Toeplitz structure, and both algorithms are based on the Vandermonde decomposition of a Toeplitz matrix. The Vandermonde decomposition breaks a Toeplitz matrix into its harmonic components, from which the DOAs are estimated. First, we present the `irregular' Toeplitz matrix and irregular Vandermonde decomposition (IVD), which generalizes the Vandermonde decomposition to apply to a more general set of matrices. It is shown that the IVD is related to the MUSIC and root-MUSIC algorithms. Next, gridless DOA is generalized to the NUA case using IVD. The resulting non-convex optimization problem is solved using alternating projections (AP). A numerical analysis is performed on the AP based solution which shows that the generalization to NUAs has similar performance to traditional gridless DOA.

eess.SP

Machine learning in acoustics: theory and applications

Acoustic data provide scientific and engineering insights in fields ranging from biology and communications to ocean and Earth science. We survey the recent advances and transformative potential of machine learning (ML), including deep learning, in the field of acoustics. ML is a broad family of techniques, which are often based in statistics, for automatically detecting and utilizing patterns in data. Relative to conventional acoustics and signal processing, ML is data-driven. Given sufficient training data, ML can discover complex relationships between features and desired labels or actions, or between features themselves. With large volumes of training data, ML can discover models describing complex acoustic phenomena such as human speech and reverberation. ML in acoustics is rapidly developing with compelling results and significant future promise. We first introduce ML, then highlight ML developments in four acoustics research areas: source localization in speech processing, source localization in ocean acoustics, bioacoustics, and environmental sounds in everyday scenes.

eess.SP

Deep-learning source localization using multi-frequency magnitude-only data

A deep learning approach based on big data is proposed to locate broadband acoustic sources using a single hydrophone in ocean waveguides with uncertain bottom parameters. Several 50-layer residual neural networks, trained on a huge number of sound field replicas generated by an acoustic propagation model, are used to handle the bottom uncertainty in source localization. A two-step training strategy is presented to improve the training of the deep models. First, the range is discretized in a coarse (5 km) grid. Subsequently, the source range within the selected interval and source depth are discretized on a finer (0.1 km and 2 m) grid. The deep learning methods were demonstrated for simulated magnitude-only multi-frequency data in uncertain environments. Experimental data from the China Yellow Sea also validated the approach.

physics.ao-ph

Sound source ranging using a feed-forward neural network with fitting-based early stopping

When a feed-forward neural network (FNN) is trained for source ranging in an ocean waveguide, it is difficult evaluating the range accuracy of the FNN on unlabeled test data. A fitting-based early stopping (FEAST) method is introduced to evaluate the range error of the FNN on test data where the distance of source is unknown. Based on FEAST, when the evaluated range error of the FNN reaches the minimum on test data, stopping training, which will help to improve the ranging accuracy of the FNN on the test data. The FEAST is demonstrated on simulated and experimental data.

cs.LG

Variational Bayesian Line Spectral Estimation with Multiple Measurement Vectors

In this paper, the line spectral estimation (LSE) problem with multiple measurement vectors (MMVs) is studied utilizing the Bayesian methods. Motivated by the recently proposed variational line spectral estimation (VALSE) method, we develop the multisnapshot VALSE (MVALSE) for multi snapshot scenarios, which is especially important in array signal processing. The MVALSE shares the advantages of the VALSE method, such as automatically estimating the model order, noise variance, weight variance, and providing the uncertain degrees of the frequency estimates. It is shown that the MVALSE can be viewed as applying the VALSE with single measurement vector (SMV) to each snapshot, and combining the intermediate data appropriately. Furthermore, the Seq-MVALSE is developed to perform sequential estimation. Finally, numerical results are conducted to demonstrate the effectiveness of the MVALSE method, compared to the state-of-the-art methods in the MMVs setting.

cs.IT

Gridless Line Spectral Estimation with Multiple Measurement Vector via Variational Bayesian Inference

Line spectral estimation (LSE) from multi snapshot samples is studied utilizing the variational Bayesian methods. Motivated by the recently proposed variational line spectral estimation (VALSE) method for a single snapshot, we develop the multisnapshot VALSE (MVALSE) for multi snapshot scenarios, which is important for array processing. The MVALSE shares the advantages of the VALSE method, such as automatically estimating the model order, noise variance and weight variance, closed-form updates of the posterior probability density function (PDF) of the frequencies. By using multiple snapshots, MVALSE improves the recovery performance and it encodes the prior distribution naturally. Finally, numerical results demonstrate the competitive performance of the MVALSE compared to state-of-the-art methods.

eess.SP

Travel time tomography with adaptive dictionaries

We develop a 2D travel time tomography method which regularizes the inversion by modeling groups of slowness pixels from discrete slowness maps, called patches, as sparse linear combinations of atoms from a dictionary. We propose to use dictionary learning during the inversion to adapt dictionaries to specific slowness maps. This patch regularization, called the local model, is integrated into the overall slowness map, called the global model. The local model considers small-scale variations using a sparsity constraint and the global model considers larger-scale features constrained using $\ell_2$ regularization. This strategy in a locally-sparse travel time tomography (LST) approach enables simultaneous modeling of smooth and discontinuous slowness features. This is in contrast to conventional tomography methods, which constrain models to be exclusively smooth or discontinuous. We develop a $\textit{maximum a posteriori}$ formulation for LST and exploit the sparsity of slowness patches using dictionary learning. The LST approach compares favorably with smoothness and total variation regularization methods on densely, but irregularly sampled synthetic slowness maps.

physics.geo-ph