SearcharxivSearch

arXiv subjects

Arie Yeredor

Publications and source records attributed to Arie Yeredor.

17 recordsLinked to original sources

A Unified MDL-based Binning and Tensor Factorization Framework for PDF Estimation

Reliable density estimation is fundamental for numerous applications in statistics and machine learning. In many practical scenarios, data are best modeled as mixtures of component densities that capture complex and multimodal patterns. However, conventional density estimators based on uniform histograms often fail to capture local variations, especially when the underlying distribution is highly nonuniform. Furthermore, the inherent discontinuity of histograms poses challenges for tasks requiring smooth derivatives, such as gradient-based optimization, clustering, and nonparametric discriminant analysis. In this work, we present a novel non-parametric approach for multivariate probability density function (PDF) estimation that utilizes minimum description length (MDL)-based binning with quantile cuts. Our approach builds upon tensor factorization techniques, leveraging the canonical polyadic decomposition (CPD) of a joint probability tensor. We demonstrate the effectiveness of our method on synthetic data and a challenging real dry bean classification dataset.

cs.LG

Probabilistic Position-Aided Beam Selection for mmWave MIMO Systems

Millimeter-wave (mmWave) MIMO systems rely on highly directional beamforming to overcome severe path loss and ensure robust communication links. However, selecting the optimal beam pair efficiently remains a challenge due to the large search space and the overhead of conventional methods. This paper proposes a probabilistic position-aided beam selection approach that exploits the statistical dependence between user equipment (UE) positions and optimal beam indices. We model the underlying joint probability mass function (PMF) of the positions and the beam indices as a low-rank tensor and estimate its parameters from training data using Bayesian inference. The estimated model is then used to predict the best (or a list of the top) beam pair indices for new UE positions. The proposed method is evaluated using data generated from a state-of-the-art ray tracing simulator and compared with neural network-based and fingerprinting approaches. The results show that our approach achieves a high data rate with fewer training samples and a significantly reduced beam search space. These advantages render it a promising solution for practical mmWave MIMO deployments, reducing the beam search overhead while maintaining a reliable connectivity.

eess.SP

Joint Bayesian Parameter and Model Order Estimation for Low-Rank Probability Mass Tensors

Obtaining a reliable estimate of the joint probability mass function (PMF) of a set of random variables from observed data is a significant objective in statistical signal processing and machine learning. Modelling the joint PMF as a tensor that admits a low-rank canonical polyadic decomposition (CPD) has enabled the development of efficient PMF estimation algorithms. However, these algorithms require the rank (model order) of the tensor to be specified beforehand. In real-world applications, the true rank is unknown. Therefore, an appropriate rank is usually selected from a candidate set either by observing validation errors or by computing various likelihood-based information criteria, a procedure that could be costly in terms of computational time or hardware resources, or could result in mismatched models which affect the model accuracy. This paper presents a novel Bayesian framework for estimating the low-rank components of a joint PMF tensor and simultaneously inferring its rank from the observed data. We specify a Bayesian PMF estimation model and employ appropriate prior distributions for the model parameters, allowing the rank to be inferred without cross-validation.We then derive a deterministic solution based on variational inference (VI) to approximate the posterior distributions of various model parameters. Numerical experiments involving both synthetic data and real classification and item recommendation data illustrate the advantages of our VI-based method in terms of estimation accuracy, automatic rank detection, and computational efficiency.

stat.ML

Time-Domain Based Embeddings for Spoofed Audio Representation

Anti-spoofing is the task of speech authentication. That is, identifying genuine human speech compared to spoofed speech. The main focus of this paper is to suggest new representations for genuine and spoofed speech, based on the probability mass function (PMF) estimation of the audio waveforms' amplitude. We introduce a new feature extraction method for speech audio signals: unlike traditional methods, our method is based on direct processing of time-domain audio samples. The PMF is utilized by designing a feature extractor based on different PMF distances and similarity measures. As an additional step, we used filter-bank preprocessing, which significantly affects the discriminative characteristics of the features and facilitates convenient visualization of possible clustering of spoofing attacks. Furthermore, we use diffusion maps to reveal the underlying manifold on which the data lies. The suggested embeddings allow the use of simple linear separators to achieve decent performance. In addition, we present a convenient way to visualize the data, which helps to assess the efficiency of different spoofing techniques. The experimental results show the potential of using multi-channel PMF based features for the anti-spoofing task, in addition to the benefits of using diffusion maps both as an analysis tool and as an embedding tool.

eess.AS

Monotonicity of the Trace-Inverse of Covariance Submatrices and Two-Sided Prediction

It is common to assess the "memory strength" of a stationary process looking at how fast the normalized log-determinant of its covariance submatrices (i.e., entropy rate) decreases. In this work, we propose an alternative characterization in terms of the normalized trace-inverse of the covariance submatrices. We show that this sequence is monotonically non-decreasing and is constant if and only if the process is white. Furthermore, while the entropy rate is associated with one-sided prediction errors (present from past), the new measure is associated with two-sided prediction errors (present from past and future). This measure can be used as an alternative to Burg's maximum-entropy principle for spectral estimation. We also propose a counterpart for non-stationary processes, by looking at the average trace-inverse of subsets.

eess.SP

Non-Iterative Blind Calibration of Nested Arrays with Asymptotically Optimal Weighting

Blind calibration of sensors arrays (without using calibration signals) is an important, yet challenging problem in array processing. While many methods have been proposed for "classical" array structures, such as uniform linear arrays, not as many are found in the context of the more "modern" sparse arrays. In this paper, we present a novel blind calibration method for $2$-level nested arrays. Specifically, and despite recent contradicting claims in the literature, we show that the Least-Squares (LS) approach can in fact be used for this purpose with such arrays. Moreover, the LS approach gives rise to optimally-weighted LS joint estimation of the sensors' gains and phases offsets, which leads to more accurate calibration, and in turn, to higher accuracy in subsequent estimation tasks (e.g., direction-of-arrival). Our method, which can be extended to $K$-level arrays ($K>2$), is superior to the current state of the art both in terms of accuracy and computational efficiency, as we demonstrate in simulation.

eess.SP

Enhanced Blind Calibration of Uniform Linear Arrays with One-Bit Quantization by Kullback-Leibler Divergence Covariance Fitting

One-bit quantization has recently become an attractive option for data acquisition in cutting edge applications, due to the increasing demand for low power and higher sampling rates. Subsequently, the rejuvenated one-bit array processing field is now receiving more attention, as "classical" array processing techniques are adapted / modified accordingly. However, array calibration, often an instrumental preliminary stage in array processing, has so far received little attention in its one-bit form. In this paper, we present a novel solution approach for the blind calibration problem, namely, without using known calibration signals. In order to extract information within the second-order statistics of the quantized measurements, we propose to estimate the unknown sensors' gains and phases offsets according to a Kullback-Leibler Divergence (KLD) covariance fitting criterion. We then provide a quasi-Newton solution algorithm, with a consistent initial estimate, and demonstrate the improved accuracy of our KLD-based estimates in simulations.

eess.SP

Iterative Symbol Recovery For Power Efficient DC Biased Optical OFDM Systems

Orthogonal frequency division multiplexing (OFDM) has proven itself as an effective multi-carrier digital communication technique. In recent years the interest in optical OFDM has grown significantly, due to its spectral efficiency and inherent resilience to frequency-selective channels and to narrowband interference. For these reasons it is currently considered to be one of the leading candidates for deployment in short fiber links such as the ones intended for inter-data-center communications. In this paper we present a new power-efficient symbol recovery scheme for dc-biased optical OFDM (DCO-OFDM) in an intensity-modulation direct-detection (IM/DD) system. We introduce an alternative method for clipping in order to maintain a non-negative real-valued signal and still preserve information, which is lost when using clipping, and propose an iterative detection algorithm for this method. A reduction of $50\%$ in the transmitted optical power along with an increase of signal-independent noise immunity (gaining $3$[dB] in SNR), compared to traditional DCO-OFDM with a DC bias of $2$ standard deviations of the OFDM signal, is attained by our new scheme for a symbol error rate (SER) of $10^{-3}$ in a QPSK constellation additive white gaussian noise (AWGN) flat channel model.

eess.SP

Blind Determination of the Number of Sources Using Distance Correlation

A novel blind estimate of the number of sources from noisy, linear mixtures is proposed. Based on Székely et al.'s distance correlation measure, we define the Sources' Dependency Criterion (SDC), from which our estimate arises. Unlike most previously proposed estimates, the SDC estimate exploits the full independence of the sources and noise, as well as the non-Gaussianity of the sources (as opposed to the Gaussianity of the noise), via implicit use of high-order statistics. This leads to a more robust, resilient and stable estimate w.r.t. the mixing matrix and the noise covariance structure. Empirical simulation results demonstrate these virtues, on top of superior performance in comparison with current state of the art estimates.

eess.SP

The Extended "Sequentially Drilled" Joint Congruence Transformation and its Application in Gaussian Independent Vector Analysis

Independent Vector Analysis (IVA) has emerged in recent years as an extension of Independent Component Analysis (ICA) into multiple sets of mixtures, where the source signals in each set are independent, but may depend on source signals in the other sets. In a semi-blind IVA (or ICA) framework, information regarding the probability distributions of the sources may be available, giving rise to Maximum Likelihood (ML) separation. In recent work we have shown that under the multivariate Gaussian model, with arbitrary temporal covariance matrices (stationary or non-stationary) of the source signals, ML separation requires the solution of a "Sequentially Drilled" Joint Congruence (SeDJoCo) transformation of a set of matrices, which is reminiscent of (but different from) classical joint diagonalization. In this paper we extend our results to the IVA problem, showing how the ML solution for the Gaussian model (with arbitrary covariance and cross-covariance matrices) takes the form of an extended SeDJoCo problem. We formulate the extended problem, derive a condition for the existence of a solution, and propose two iterative solution algorithms. In addition, we derive the induced Cramér-Rao Lower Bound (iCRLB) on the resulting Interference-to-Source Ratios (ISR) matrices, and demonstrate by simulation how the ML separation obtained by solving the extended SeDJoCo problem indeed attains the iCRLB (asymptotically), as opposed to other separation approaches, which cannot exploit prior knowledge regarding the sources' distributions.

eess.SP

Performance Analysis of the Gaussian Quasi-Maximum Likelihood Approach for Independent Vector Analysis

Maximum Likelihood (ML) estimation requires precise knowledge of the underlying statistical model. In Quasi ML (QML), a presumed model is used as a substitute to the (unknown) true model. In the context of Independent Vector Analysis (IVA), we consider the Gaussian QML Estimate (QMLE) of the demixing matrices set and present an (approximate) analysis of its asymptotic separation performance. In Gaussian QML the sources are presumed to be Gaussian, with covariance matrices specified by some "educated guess". The resulting quasi-likelihood equations of the demixing matrices take a special form, recently termed an extended "Sequentially Drilled" Joint Congruence (SeDJoCo) transformation, which is reminiscent of (though essentially different from) classical joint diagonalization. We show that asymptotically this QMLE, i.e., the solution of the resulting extended SeDJoCo transformation, attains perfect separation (under some mild conditions) regardless of the sources' true distributions and/or covariance matrices. In addition, based on the "small-errors" assumption, we present a first-order perturbation analysis of the extended SeDJoCo solution. Using the resulting closed-form expressions for the errors in the solution matrices, we provide closed-form expressions for the resulting Interference-to-Source Ratios (ISRs) for IVA. Moreover, we prove that asymptotically the ISRs depend only on the sources' covariances, and not on their specific distributions. As an immediate consequence of this result, we provide an asymptotically attainable lower bound on the resulting ISRs. We also present empirical results, corroborating our analytical derivations, of three simulation experiments concerning two possible model errors - inaccurate covariance matrices and sources' distribution mismodeling.

eess.SP

Asymptotically Optimal Blind Calibration of Uniform Linear Sensor Arrays for Narrowband Gaussian Signals

An asymptotically optimal blind calibration scheme of uniform linear arrays for narrowband Gaussian signals is proposed. Rather than taking the direct Maximum Likelihood (ML) approach for joint estimation of all the unknown model parameters, which leads to a multi-dimensional optimization problem with no closed-form solution, we revisit Paulraj and Kailath's (P-K's) classical approach in exploiting the special (Toeplitz) structure of the observations' covariance. However, we offer a substantial improvement over P-K's ordinary Least Squares (LS) estimates by using asymptotic approximations in order to obtain simple, non-iterative, (quasi-)linear Optimally-Weighted LS (OWLS) estimates of the sensors gains and phases offsets with asymptotically optimal weighting, based only on the empirical covariance matrix of the measurements. Moreover, we prove that our resulting estimates are also asymptotically optimal w.r.t. the raw data, and can therefore be deemed equivalent to the ML Estimates (MLE), which are otherwise obtained by joint ML estimation of all the unknown model parameters. After deriving computationally convenient expressions of the respective Cramér-Rao lower bounds, we also show that our estimates offer improved performance when applied to non-Gaussian signals (and/or noise) as quasi-MLE in a similar setting. The optimal performance of our estimates is demonstrated in simulation experiments, with a considerable improvement (reaching an order of magnitude and more) in the resulting mean squared errors w.r.t. P-K's ordinary LS estimates. We also demonstrate the improved accuracy in a multiple-sources directions-of-arrivals estimation task.

eess.SP

MultiView Diffusion Maps

In this paper, we address the challenging task of achieving multi-view dimensionality reduction. The goal is to effectively use the availability of multiple views for extracting a coherent low-dimensional representation of the data. The proposed method exploits the intrinsic relation within each view, as well as the mutual relations between views. The multi-view dimensionality reduction is achieved by defining a cross-view model in which an implied random walk process is restrained to hop between objects in the different views. The method is robust to scaling and insensitive to small structural changes in the data. We define new diffusion distances and analyze the spectra of the proposed kernel. We show that the proposed framework is useful for various machine learning applications such as clustering, classification, and manifold learning. Finally, by fusing multi-sensor seismic data we present a method for automatic identification of seismic events.

cs.LG

Kernel Scaling for Manifold Learning and Classification

Kernel methods play a critical role in many machine learning algorithms. They are useful in manifold learning, classification, clustering and other data analysis tasks. Setting the kernel's scale parameter, also referred to as the kernel's bandwidth, highly affects the performance of the task in hand. We propose to set a scale parameter that is tailored to one of two types of tasks: classification and manifold learning. For manifold learning, we seek a scale which is best at capturing the manifold's intrinsic dimension. For classification, we propose three methods for estimating the scale, which optimize the classification results in different senses. The proposed frameworks are simulated on artificial and on real datasets. The results show a high correlation between optimal classification rates and the estimated scales. Finally, we demonstrate the approach on a seismic event classification task.

cs.LG

A Maximum Likelihood-Based Minimum Mean Square Error Separation and Estimation of Stationary Gaussian Sources from Noisy Mixtures

In the context of Independent Component Analysis (ICA), noisy mixtures pose a dilemma regarding the desired objective. On one hand, a "maximally separating" solution, providing the minimal attainable Interference-to-Source-Ratio (ISR), would often suffer from significant residual noise. On the other hand, optimal Minimum Mean Square Error (MMSE) estimation would yield estimates which are the "closest possible" to the true sources, often at the cost of compromised ISR. In this work, we consider noisy mixtures of temporally-diverse stationary Gaussian sources in a semi-blind scenario, which conveniently lends itself to either one of these objectives. We begin by deriving the ML Estimates (MLEs) of the unknown (deterministic) parameters of the model: the mixing matrix and the (possibly different) noise variances in each sensor. We derive the likelihood equations for these parameters, as well as the corresponding Cramér-Rao lower bound, and propose an iterative solution for obtaining the MLEs. Based on these MLEs, the asymptotically-optimal "maximally separating" solution can be readily obtained. However, we also present the ML-based MMSE estimate of the sources, alongside a frequency-domain-based computationally efficient scheme, exploiting their stationarity. We show that this estimate is asymptotically optimal and attains the (oracle) MMSE lower bound. Furthermore, for non-Gaussian signals, we show that this estimate serves as a Quasi ML (QML)-based Linear MMSE (LMMSE) estimate, and attains the (oracle) LMMSE lower bound asymptotically. Empirical results of three simulation experiments are presented, corroborating our analytical derivations.

stat.AP

First-Order Perturbation Analysis of the SECSI Framework for the Approximate CP Decomposition of 3-D Noise-Corrupted Low-Rank Tensors

The Semi-Algebraic framework for the approximate Canonical Polyadic (CP) decomposition via SImultaneaous matrix diagonalization (SECSI) is an efficient tool for the computation of the CP decomposition. The SECSI framework reformulates the CP decomposition into a set of joint eigenvalue decomposition (JEVD) problems. Solving all JEVDs, we obtain multiple estimates of the factor matrices and the best estimate is chosen in a subsequent step by using an exhaustive search or some heuristic strategy that reduces the computational complexity. Moreover, the SECSI framework retains the option of choosing the number of JEVDs to be solved, thus providing an adjustable complexity-accuracy trade-off. In this work, we provide an analytical performance analysis of the SECSI framework for the computation of the approximate CP decomposition of a noise corrupted low-rank tensor, where we derive closed-form expressions of the relative mean square error for each of the estimated factor matrices. These expressions are obtained using a first-order perturbation analysis and are formulated in terms of the second-order moments of the noise, such that apart from a zero mean, no assumptions on the noise statistics are required. Simulation results exhibit an excellent match between the obtained closed-form expressions and the empirical results. Moreover, we propose a new Performance Analysis based Selection (PAS) scheme to choose the final factor matrix estimate. The results show that the proposed PAS scheme outperforms the existing heuristics, especially in the high SNR regime.

cs.IT

Independent Component Analysis Over Galois Fields

We consider the framework of Independent Component Analysis (ICA) for the case where the independent sources and their linear mixtures all reside in a Galois field of prime order P. Similarities and differences from the classical ICA framework (over the Real field) are explored. We show that a necessary and sufficient identifiability condition is that none of the sources should have a Uniform distribution. We also show that pairwise independence of the mixtures implies their full mutual independence (namely a non-mixing condition) in the binary (P=2) and ternary (P=3) cases, but not necessarily in higher order (P>3) cases. We propose two different iterative separation (or identification) algorithms: One is based on sequential identification of the smallest-entropy linear combinations of the mixtures, and is shown to be equivariant with respect to the mixing matrix; The other is based on sequential minimization of the pairwise mutual information measures. We provide some basic performance analysis for the binary (P=2) case, supplemented by simulation results for higher orders, demonstrating advantages and disadvantages of the proposed separation approaches.

cs.IT