Searcharxiv⌕ Search

arXiv subjects

Elsa Dupraz

Publications and source records attributed to Elsa Dupraz.

25 records · Page 2Linked to original sources

Decentralized Clustering on Compressed Data without Prior Knowledge of the Number of Clusters

In sensor networks, it is not always practical to set up a fusion center. Therefore, there is need for fully decentralized clustering algorithms. Decentralized clustering algorithms should minimize the amount of data exchanged between sensors in order to reduce sensor energy consumption. In this respect, we propose one centralized and one decentralized clustering algorithm that work on compressed data without prior knowledge of the number of clusters. In the standard K-means clustering algorithm, the number of clusters is estimated by repeating the algorithm several times, which dramatically increases the amount of exchanged data, while our algorithm can estimate this number in one run. The proposed clustering algorithms derive from a theoretical framework establishing that, under asymptotic conditions, the cluster centroids are the only fixed-point of a cost function we introduce. This cost function depends on a weight function which we choose as the p-value of a Wald hypothesis test. This p-value measures the plausibility that a given measurement vector belongs to a given cluster. Experimental results show that our two algorithms are competitive in terms of clustering performance with respect to K-means and DB-Scan, while lowering by a factor at least $2$ the amount of data exchanged between sensors.

stat.ML↗

K-means Algorithm over Compressed Binary Data

We consider a network of binary-valued sensors with a fusion center. The fusion center has to perform K-means clustering on the binary data transmitted by the sensors. In order to reduce the amount of data transmitted within the network, the sensors compress their data with a source coding scheme based on binary sparse matrices. We propose to apply the K-means algorithm directly over the compressed data without reconstructing the original sensors measurements, in order to avoid potentially complex decoding operations. We provide approximated expressions of the error probabilities of the K-means steps in the compressed domain. From these expressions, we show that applying the K-means algorithm in the compressed domain enables to recover the clusters of the original domain. Monte Carlo simulations illustrate the accuracy of the obtained approximated error probabilities, and show that the coding rate needed to perform K-means clustering in the compressed domain is lower than the rate needed to reconstruct all the measurements.

cs.IT↗

Rate-Distortion Performance of Sequential Massive Random Access to Gaussian Sources with Memory

In Sequential Massive Random Access (SMRA), a set of correlated sources is jointly encoded and stored on a server, and clients want to access to only a subset of the sources. Since the number of simultaneous clients can be huge, the server is only authorized to extract a bitstream from the stored data: no re-encoding can be performed before the transmission of a request. In this paper, we investigate the SMRA performance of lossy source coding of Gaussian sources with memory. In practical applications such as Free Viewpoint Television, this model permits to take into account not only inter but also intra correlation between sources. For this model, we provide the storage and transmission rates that are achievable for SMRA under some distortion constraint, and we consider two particular examples of Gaussian sources with memory.

cs.IT↗

Transmission and Storage Rates for Sequential Massive Random Access

This paper introduces a new source coding paradigm called Sequential Massive Random Access (SMRA). In SMRA, a set of correlated sources is encoded once for all and stored on a server, and clients want to successively access to only a subset of the sources. Since the number of simultaneous clients can be huge, the server is only allowed to extract a bitstream from the stored data: no re-encoding can be performed before the transmission of the specific client's request. In this paper, we formally define the SMRA framework and introduce both storage and transmission rates to characterize the performance of SMRA. We derive achievable transmission and storage rates for lossless source coding of i.i.d. and non i.i.d. sources, and transmission and storage rates-distortion regions for Gaussian sources. We also show two practical implementations of SMRA systems based on rate-compatible LDPC codes. Both theoretical and experimental results demonstrate that SMRA systems can reach the same transmission rates as in traditional point to point source coding schemes, while having a reasonable overhead in terms of storage rate. These results constitute a breakthrough for many recent data transmission applications in which different parts of the data are requested by the clients.

cs.IT↗

Decentralized Clustering based on Robust Estimation and Hypothesis Testing

This paper considers a network of sensors without fusion center that may be difficult to set up in applications involving sensors embedded on autonomous drones or robots. In this context, this paper considers that the sensors must perform a given clustering task in a fully decentralized setup. Standard clustering algorithms usually need to know the number of clusters and are very sensitive to initialization, which makes them difficult to use in a fully decentralized setup. In this respect, this paper proposes a decentralized model-based clustering algorithm that overcomes these issues. The proposed algorithm is based on a novel theoretical framework that relies on hypothesis testing and robust M-estimation. More particularly, the problem of deciding whether two data belong to the same cluster can be optimally solved via Wald's hypothesis test on the mean of a Gaussian random vector. The p-value of this test makes it possible to define a new type of score function, particularly suitable for devising an M-estimation of the centroids. The resulting decentralized algorithm efficiently performs clustering without prior knowledge of the number of clusters. It also turns out to be less sensitive to initialization than the already existing clustering algorithms, which makes it appropriate for use in a network of sensors without fusion center.

math.ST↗

Analysis and Design of Finite Alphabet Iterative Decoders Robust to Faulty Hardware

This paper addresses the problem of designing LDPC decoders robust to transient errors introduced by a faulty hardware. We assume that the faulty hardware introduces errors during the message passing updates and we propose a general framework for the definition of the message update faulty functions. Within this framework, we define symmetry conditions for the faulty functions, and derive two simple error models used in the analysis. With this analysis, we propose a new interpretation of the functional Density Evolution threshold previously introduced, and show its limitations in case of highly unreliable hardware. However, we show that under restricted decoder noise conditions, the functional threshold can be used to predict the convergence behavior of FAIDs under faulty hardware. In particular, we reveal the existence of robust and non-robust FAIDs and propose a framework for the design of robust decoders. We finally illustrate robust and non-robust decoders behaviors of finite length codes using Monte Carlo simulations.

cs.IT↗

Density Evolution and Functional Threshold for the Noisy Min-Sum Decoder

This paper investigates the behavior of the Min-Sum decoder running on noisy devices. The aim is to evaluate the robustness of the decoder in the presence of computation noise, e.g. due to faulty logic in the processing units, which represents a new source of errors that may occur during the decoding process. To this end, we first introduce probabilistic models for the arithmetic and logic units of the the finite-precision Min-Sum decoder, and then carry out the density evolution analysis of the noisy Min-Sum decoder. We show that in some particular cases, the noise introduced by the device can help the Min-Sum decoder to escape from fixed points attractors, and may actually result in an increased correction capacity with respect to the noiseless decoder. We also reveal the existence of a specific threshold phenomenon, referred to as functional threshold. The behavior of the noisy decoder is demonstrated in the asymptotic limit of the code-length -- by using "noisy" density evolution equations -- and it is also verified in the finite-length case by Monte-Carlo simulation.

cs.IT↗