SearcharxivSearch

arXiv subjects

Mireille Boutin

Publications and source records attributed to Mireille Boutin.

At least 19 recordsLinked to original sources

Localization from Pseudoranges: Quadrics and Duality

This paper gives a complete description of the solutions of the global positioning problem, emphasizing the under-determined case. We show that the solutions form a quadric, which may degenerate in various ways. Perhaps more surprisingly, the satellite positions also lie on a quadric, and these two quadrics exhibit a remarkable duality: They live on perpendicular affine spaces but share the same axis of symmetry. Moreover, the vertices of one quadric are the foci of the other and vice versa. The results of this paper are not only applicable to the global positioning problem, but to a wider class of problems known as pseudorange-multilateration. This includes a range of real-world localization problems where a signal is emitted at an unknown emission time, and received by sensors at known positions. In particular, the paper can be useful for solving an under-determined multilateration problem in the presence of additional constraints. We illustrate this with two examples: locating a cleaning robot on the ground and locating a raft on the ocean.

math.MG

Global Positioning on Earth

Contrary to popular belief, the global positioning problem on earth may have more than one solutions even if the user position is restricted to a sphere. With 3 satellites, we show that there can be up to 4 solutions on a sphere. With 4 or more satellites, we show that, for any pair of points on a sphere, there is a family of hyperboloids of revolution such that if the satellites are placed on one sheet of one of these hyperboloid, then the global positioning problem has both points as solutions. We give solution methods that yield the correct number of solutions on/near a sphere.

math.MG

Approximation and generalization properties of the random projection classification method

The generalization gap of a classifier is related to the complexity of the set of functions among which the classifier is chosen. We study a family of low-complexity classifiers consisting of thresholding a random one-dimensional feature. The feature is obtained by projecting the data on a random line after embedding it into a higher-dimensional space parametrized by monomials of order up to k. More specifically, the extended data is projected n-times and the best classifier among those n, based on its performance on training data, is chosen. We show that this type of classifier is extremely flexible as, given full knowledge of the class conditional densities, under mild conditions, the error of these classifiers would converge to the optimal (Bayes) error as k and n go to infinity. We also bound the generalization gap of the random classifiers. In general, these bounds are better than those for any classifier with VC dimension greater than O(ln n). In particular, the bounds imply that, unless the number of projections n is extremely large, the generalization gap of the random projection approach is significantly smaller than that of a linear classifier in the extended space. Thus, for certain classification problems (e.g., those with a large Rashomon ratio), there is a potntially large gain in generalization properties by selecting parameters at random, rather than selecting the best one amongst the class.

cs.LG

A new model for natural groupings in high-dimensional data

Clustering aims to divide a set of points into groups. The current paradigm assumes that the grouping is well-defined (unique) given the probability model from which the data is drawn. Yet, recent experiments have uncovered several high-dimensional datasets that form different binary groupings after projecting the data to randomly chosen one-dimensional subspaces. This paper describes a probability model for the data that could explain this phenomenon. It is a simple model to serve as a proof of concept for understanding the geometry of high-dimensional data. We start by building a rescaled multivariate Bernouilli model (stretched hypercube) so to create several overlapping grouping structures in the data. The size of each scaling parameter is related to the likelihood of uncovering the corresponding grouping by random 1D projection. Clusters in the original space are then created by adding noise to this cluster-free model. In high dimension, these clusters would hardly be observable given a sample set from the distribution because of the curse of dimensionality, but the binary groupings are clear. Our construction makes it clear that one needs to make a distinction between "groupings" and "clusters" in the original space. It also highlights the need to interpret any clustering found in projected data as merely one among potentially many other groupings in a dataset.

stat.ML

On the Rashomon ratio of infinite hypothesis sets

Given a classification problem and a family of classifiers, the Rashomon ratio measures the proportion of classifiers that yield less than a given loss. Previous work has explored the advantage of a large Rashomon ratio in the case of a finite family of classifiers. Here we consider the more general case of an infinite family. We show that a large Rashomon ratio guarantees that choosing the classifier with the best empirical accuracy among a random subset of the family, which is likely to improve generalizability, will not increase the empirical loss too much. We quantify the Rashomon ratio in two examples involving infinite classifier families in order to illustrate situations in which it is large. In the first example, we estimate the Rashomon ratio of the classification of normally distributed classes using an affine classifier. In the second, we obtain a lower bound for the Rashomon ratio of a classification problem with a modified Gram matrix when the classifier family consists of two-layer ReLU neural networks. In general, we show that the Rashomon ratio can be estimated using a training dataset along with random samples from the classifier family and we provide guarantees that such an estimation is close to the true value of the Rashomon ratio.

cs.LG

Path Tracking using Echoes in an Unknown Environment: the Issue of Symmetries and How to Break Them

This paper deals with the problem of reconstructing the path of a vehicle in an unknown environment consisting of planar structures using sound. Many systems in the literature do this by using a loudspeaker and microphones mounted on a vehicle. Symmetries in the environment lead to solution ambiguities for such systems. We propose to resolve this issue by placing the loudspeaker at a fixed location in the environment rather than on the vehicle. The question of whether this will remove ambiguities regardless of the environment geometry leads to a question about breaking symmetries that can be phrased in purely mathematical terms. We solve this question in the affirmative if the geometry is in dimension three or bigger, and give counterexamples in dimension two. Excluding the rare situations where the counterexamples arise, we also give an affirmative answer in dimension two. Our results lead to a simple path reconstruction algorithm for a vehicle carrying four microphones navigating within an environment in which a loudspeaker at a fixed position emits short bursts of sounds. This algorithm could be combined with other methods from the literature to construct a path tracking system for vehicles navigating within a potentially symmetric environment.

math.MG

Global Positioning: the Uniqueness Question and a New Solution Method

We provide a new algebraic solution procedure for the global positioning problem in $n$ dimensions using $m$ satellites. We also give a geometric characterization of the situations in which the problem does not have a unique solution. This characterization shows that such cases can happen in any dimension and with any number of satellites, leading to counterexamples to some open conjectures. We fill a gap in the literature by giving a proof for the long-held belief that when $m \ge n+2$, the solution is unique for almost all user positions. Even better, when $m \ge 2n+2$, almost all satellite configurations will guarantee a unique solution for all user positions. Some of our results are obtained using tools from algebraic geometry.

eess.SP

Multilateration and Signal Matching with Unknown Emission Times

Assume that a source emits a signal in $3$-dimensional space at an unknown time, which is received by at least~$5$ sensors. In almost all cases the emission time and source position can be worked out uniquely from the knowledge of the times when the sensors receive the signal. The task to do so is the multilateration problem. But when there are several emission events originating from several sources, the received signals must first be matched in order to find the emission times and source positions. In this paper, we propose to use algebraic relations between reception times to achieve this matching. A special case occurs when the signals are actually echoes from a single emission event. In this case, solving the signal matching problem allows one to reconstruct the positions of the reflecting walls. We show that, no matter where the walls are situated, our matching algorithm works correctly for almost all positions of the sensors. In the first section of this paper we consider the multilateration problem, which is equivalent to the GPS-problem, and give a simple algebraic solution that applies in all dimensions.

math.AC

Can a Ground-Based Vehicle Hear the Shape of a Room?

Assume that a ground-based vehicle moves in a room with walls or other planar surfaces. Can the vehicle reconstruct the positions of the walls from the echoes of a single sound event? We assume that the vehicle carries some microphones and that a loudspeaker is either also mounted on the vehicle or placed at a fixed location in the room. We prove that the reconstruction is almost always possible if (1) no echoes are received from floors, ceilings or sloping walls and the vehicle carries at least three non-collinear microphones, or if (2) walls of any inclination may occur, the loudspeaker is fixed in the room and there are four non-coplanar microphones. The difficulty lies in the echo-matching problem: how to determine which echoes come from the same wall. We solve this by using a Cayley-Menger determinant. Our proofs use methods from computational commutative algebra.

math.AC

The shape of a Gaussian mixture is characterized by the probability density of the distance between two samples

Let $\bf{x}$ be a random variable with density $ρ(x)$ taking values in ${\mathbb R}^d$. We are interested in finding a representation for the shape of $ρ(x)$, i.e. for the orbit $\{ ρ(g\cdot x) | g\in E(d) \}$ of $ρ$ under the Euclidean group. Let $x_1$ and $x_2$ be two random samples picked, independently, following $ρ(x)$, and let $Δ$ be the squared Euclidean distance between $x_1$ and $x_2$. We show, if $ρ(x)$ is a mixture of Gaussians whose covariance matrix is the identity, and if the means of the Gaussians are in generic position, then the density $ρ(x)$ is reconstructible, up to a rigid motion in $E(d)$, from the density of $\bfΔ$. In other words, any two such Gaussian mixtures $ρ(x)$ and $\barρ (x)$ with the same distribution of distances are guaranteed to be related by a rigid motion $g\in E(d)$ as $ρ(x)=\barρ (g\cdot x)$. We also show that a similar result holds when the distance is defined by a symmetric bilinear form.

math.PR

Clustering small datasets in high-dimension by random projection

Datasets in high-dimension do not typically form clusters in their original space; the issue is worse when the number of points in the dataset is small. We propose a low-computation method to find statistically significant clustering structures in a small dataset. The method proceeds by projecting the data on a random line and seeking binary clusterings in the resulting one-dimensional data. Non-linear separations are obtained by extending the feature space using monomials of higher degrees in the original features. The statistical validity of the clustering structures obtained is tested in the projected one-dimensional space, thus bypassing the challenge of statistical validation in high-dimension. Projecting on a random line is an extreme dimension reduction technique that has previously been used successfully as part of a hierarchical clustering method for high-dimensional data. Our experiments show that with this simplified framework, statistically significant clustering structures can be found with as few as 100-200 points, depending on the dataset. The different structures uncovered are found to persist as more points are added to the dataset.

stat.ML

A Drone Can Hear the Shape of a Room

We show that one can reconstruct the shape of a room with planar walls from the first-order echoes received by four non-planar microphones placed on a drone with generic position and orientation. Both the cases where the source is located in the room and on the drone are considered. If the microphone positions are picked at random, then with probability one, the location of any wall is correctly reconstructed as long as it is heard by four microphones. Our algorithm uses a simple echo sorting criterion to recover the wall assignments for the echoes. We prove that, if the position and orientation of the drone on which the microphones are mounted do not lie on a certain set of dimension at most 5 in the 6-dimensional space of all drone positions and orientations, then the wall assignment obtained through our echo sorting criterion must be the right one and thus the reconstruction obtained through our algorithm is correct. Our proof uses methods from computational commutative algebra.

math.AC

Three Efficient, Low-Complexity Algorithms for Automatic Color Trapping

Color separations (most often cyan, magenta, yellow, and black) are commonly used in printing to reproduce multi-color images. For mechanical reasons, these color separations are generally not perfectly aligned with respect to each other when they are rendered by their respective imaging stations. This phenomenon, called color plane misregistration, causes gap and halo artifacts in the printed image. Color trapping is an image processing technique that aims to reduce these artifacts by modifying the susceptible edge boundaries to create small, unnoticeable overlaps between the color planes. We propose three low-complexity algorithms for automatic color trapping which hide the effects of small color plane mis-registrations. Our algorithms are designed for software or embedded firmware implementation. The trapping method they follow is based on a hardware-friendly technique proposed by J. Trask (JTHBCT03) which is too computationally expensive for software or firmware implementation. The first two algorithms are based on the use of look-up tables (LUTs). The first LUT-based algorithm corrects all registration errors of one pixel in extent and reduces several cases of misregistration errors of two pixels in extent using only 727 Kbytes of storage space. This algorithm is particularly attractive for implementation in the embedded firmware of low-cost formatter-based printers. The second LUT-based algorithm corrects all types of misregistration errors of up to two pixels in extent using 3.7 Mbytes of storage space. The third algorithm is a hybrid one that combines LUTs and feature extraction to minimize the storage requirements (724 Kbytes) while still correcting all misregistration errors of up to two pixels in extent. This algorithm is suitable for both embedded firmware implementation on low-cost formatter-based printers and software implementation on host-based printers.

cs.CV

Photo-unrealistic Image Enhancement for Subject Placement in Outdoor Photography

Camera display reflections are an issue in bright light situations, as they may prevent users from correctly positioning the subject in the picture. We propose a software solution to this problem, which consists in modifying the image in the viewer, in real time. In our solution, the user is seeing a posterized image which roughly represents the contour of the objects. Five enhancement methods are compared in a user study. Our results indicate that the problem considered is a valid one, as users had problems locating landmarks nearly 37% of the time under sunny conditions, and that our proposed enhancement method using contrasting colors is a practical solution to that problem.

cs.CV

A Hand-Held Multimedia Translation and Interpretation System with Application to Diet Management

We propose a network independent, hand-held system to translate and disambiguate foreign restaurant menu items in real-time. The system is based on the use of a portable multimedia device, such as a smartphones or a PDA. An accurate and fast translation is obtained using a Machine Translation engine and a context-specific corpora to which we apply two pre-processing steps, called translation standardization and $n$-gram consolidation. The phrase-table generated is orders of magnitude lighter than the ones commonly used in market applications, thus making translations computationally less expensive, and decreasing the battery usage. Translation ambiguities are mitigated using multimedia information including images of dishes and ingredients, along with ingredient lists. We implemented a prototype of our system on an iPod Touch Second Generation for English speakers traveling in Spain. Our tests indicate that our translation method yields higher accuracy than translation engines such as Google Translate, and does so almost instantaneously. The memory requirements of the application, including the database of images, are also well within the limits of the device. By combining it with a database of nutritional information, our proposed system can be used to help individuals who follow a medical diet maintain this diet while traveling.

cs.CL

Benchmarks for Image Classification and Other High-dimensional Pattern Recognition Problems

A good classification method should yield more accurate results than simple heuristics. But there are classification problems, especially high-dimensional ones like the ones based on image/video data, for which simple heuristics can work quite accurately; the structure of the data in such problems is easy to uncover without any sophisticated or computationally expensive method. On the other hand, some problems have a structure that can only be found with sophisticated pattern recognition methods. We are interested in quantifying the difficulty of a given high-dimensional pattern recognition problem. We consider the case where the patterns come from two pre-determined classes and where the objects are represented by points in a high-dimensional vector space. However, the framework we propose is extendable to an arbitrarily large number of classes. We propose classification benchmarks based on simple random projection heuristics. Our benchmarks are 2D curves parameterized by the classification error and computational cost of these simple heuristics. Each curve divides the plane into a "positive- gain" and a "negative-gain" region. The latter contains methods that are ill-suited for the given classification problem. The former is divided into two by the curve asymptote; methods that lie in the small region under the curve but right of the asymptote merely provide a computational gain but no structural advantage over the random heuristics. We prove that the curve asymptotes are optimal (i.e. at Bayes error) in some cases, and thus no sophisticated method can provide a structural advantage over the random heuristics. Such classification problems, an example of which we present in our numerical experiments, provide poor ground for testing new pattern classification methods.

stat.ML

Pattern Dependence Detection using n-TARP Clustering

Consider an experiment involving a potentially small number of subjects. Some random variables are observed on each subject: a high-dimensional one called the "observed" random variable, and a one-dimensional one called the "outcome" random variable. We are interested in the dependencies between the observed random variable and the outcome random variable. We propose a method to quantify and validate the dependencies of the outcome random variable on the various patterns contained in the observed random variable. Different degrees of relationship are explored (linear, quadratic, cubic, ...). This work is motivated by the need to analyze educational data, which often involves high-dimensional data representing a small number of students. Thus our implementation is designed for a small number of subjects; however, it can be easily modified to handle a very large dataset. As an illustration, the proposed method is used to study the influence of certain skills on the course grade of students in a signal processing class. A valid dependency of the grade on the different skill patterns is observed in the data.

stat.ML

Distributed Reception with Spatial Multiplexing: MIMO Systems for the Internet of Things

The Internet of things (IoT) holds much commercial potential and could facilitate distributed multiple-input multiple-output (MIMO) communication in future systems. We study a distributed reception scenario in which a transmitter equipped with multiple antennas sends multiple streams via spatial multiplexing to a large number of geographically separated single antenna receive nodes. The receive nodes then quantize their received signals and forward the quantized received signals to a receive fusion center. With global channel knowledge and forwarded quantized information from the receive nodes, the fusion center attempts to decode the transmitted symbols. We assume the transmit vector consists of phase shift keying (PSK) constellation points, and each receive node quantizes its received signal with one bit for each of the real and imaginary parts of the signal to minimize the transmission overhead between the receive nodes and the fusion center. Fusing this data is a non-trivial problem because the receive nodes cannot decode the transmitted symbols before quantization. Instead, each receive node processes a single quantity, i.e., the received signal, regardless of the number of transmitted symbols. We develop an optimal maximum likelihood (ML) receiver and a low-complexity zero-forcing (ZF)-type receiver at the fusion center. Despite its suboptimality, the ZF-type receiver is simple to implement and shows comparable performance with the ML receiver in the low signal-to-noise ratio (SNR) regime but experiences an error rate floor at high SNR. It is shown that this error floor can be overcome by increasing the number of receive nodes. Hence, the ZF-type receiver would be a practical solution for distributed reception with spatial multiplexing in the era of the IoT where we can easily have a large number of receive nodes.

cs.IT