SearcharxivSearch

arXiv subjects

Mariane Pelletier

Publications and source records attributed to Mariane Pelletier.

9 recordsLinked to original sources

A Recursive Algorithm for Mining Association Rules

Mining frequent itemsets and association rules is an essential task within data mining and data analysis. In this paper, we introduce PrefRec, a recursive algorithm for finding frequent itemsets and association rules. Its main advantage is its recursiveness with respect to the items. It is particularly efficient for updating the mining process when new items are added to the database or when some are excluded. We present in a complete way the logic of the algorithm, and give some of its applications. After that, we carry out an experimental study on the effectiveness of PrefRec. We first compare the execution times with some very popular frequent itemset mining algorithms. Then, we do experiments to test the updating capabilities of our algorithm.

cs.DB

Revisiting Révész's stochastic approximation method for the estimation of a regression function

In a pioneer work, Révész (1973) introduces the stochastic approximation method to build up a recursive kernel estimator of the regression function $x\mapsto E(Y|X=x)$. However, according to Révész (1977), his estimator has two main drawbacks: on the one hand, its convergence rate is smaller than that of the nonrecursive Nadaraya-Watson's kernel regression estimator, and, on the other hand, the required assumptions on the density of the random variable $X$ are stronger than those usually needed in the framework of regression estimation. We first come back on the study of the convergence rate of Révész's estimator. An approach in the proofs completely different from that used in Révész (1977) allows us to show that Révész's recursive estimator may reach the same optimal convergence rate as Nadaraya-Watson's estimator, but the required assumptions on the density of $X$ remain stronger than the usual ones, and this is inherent to the definition of Révész's estimator. To overcome this drawback, we introduce the averaging principle of stochastic approximation algorithms to construct the averaged Révész's regression estimator, and give its asymptotic behaviour. Our assumptions on the density of $X$ are then usual in the framework of regression estimation. We prove that the averaged Révész's regression estimator may reach the same optimal convergence rate as Nadaraya-Watson's estimator. Moreover, we show that, according to the estimation by confidence intervals point of view, it is better to use the averaged Révész's estimator rather than Nadaraya-Watson's estimator.

math.ST

The stochastic approximation method for the estimation of a multivariate probability density

We apply the stochastic approximation method to construct a large class of recursive kernel estimators of a probability density, including the one introduced by Hall and Patil (1994). We study the properties of these estimators and compare them with Rosenblatt's nonrecursive estimator. It turns out that, for pointwise estimation, it is preferable to use the nonrecursive Rosenblatt's kernel estimator rather than any recursive estimator. A contrario, for estimation by confidence intervals, it is better to use a recursive estimator rather than Rosenblatt's estimator.

math.ST

Joint behaviour of semirecursive kernel estimators of the location and of the size of the mode of a probability density function

Let $θ$ and $μ$ denote the location and the size of the mode of a probability density. We study the joint convergence rates of semirecursive kernel estimators of $θ$ and $μ$. We show how the estimation of the size of the mode allows to measure the relevance of the estimation of its location. We also enlighten that, beyond their computational advantage on nonrecursive estimators, the semirecursive estimators are preferable to use for the construction on confidence regions.

math.ST

A companion for the Kiefer--Wolfowitz--Blum stochastic approximation algorithm

A stochastic algorithm for the recursive approximation of the location $θ$ of a maximum of a regression function was introduced by Kiefer and Wolfowitz [Ann. Math. Statist. 23 (1952) 462--466] in the univariate framework, and by Blum [Ann. Math. Statist. 25 (1954) 737--744] in the multivariate case. The aim of this paper is to provide a companion algorithm to the Kiefer--Wolfowitz--Blum algorithm, which allows one to simultaneously recursively approximate the size $μ$ of the maximum of the regression function. A precise study of the joint weak convergence rate of both algorithms is given; it turns out that, unlike the location of the maximum, the size of the maximum can be approximated by an algorithm which converges at the parametric rate. Moreover, averaging leads to an asymptotically efficient algorithm for the approximation of the couple $(θ,μ)$.

math.ST

Large and moderate deviations principles for kernel estimators of the multivariate regression

In this paper, we prove large deviations principle for the Nadaraya-Watson estimator and for the semi-recursive kernel estimator of the regression in the multidimensional case. Under suitable conditions, we show that the rate function is a good rate function. We thus generalize the results already obtained in the unidimensional case for the Nadaraya-Watson estimator. Moreover, we give a moderate deviations principle for these two estimators. It turns out that the rate function obtained in the moderate deviations principle for the semi-recursive estimator is larger than the one obtained for the Nadaraya-Watson estimator.

math.ST

Convergence rate and averaging of nonlinear two-time-scale stochastic approximation algorithms

The first aim of this paper is to establish the weak convergence rate of nonlinear two-time-scale stochastic approximation algorithms. Its second aim is to introduce the averaging principle in the context of two-time-scale stochastic approximation algorithms. We first define the notion of asymptotic efficiency in this framework, then introduce the averaged two-time-scale stochastic approximation algorithm, and finally establish its weak convergence rate. We show, in particular, that both components of the averaged two-time-scale stochastic approximation algorithm simultaneously converge at the optimal rate $\sqrt{n}$.

math.PR

Confidence bands for densities, logarithmic point of view

Let $f$ be a probability density and $C$ be an interval on which $f$ is bounded away from zero. By establishing the limiting distribution of the uniform error of the kernel estimates $f_n$ of $f$, Bickel and Rosenblatt (1973) provide confidence bands $B_n$ for $f$ on $C$ with asymptotic level $1-α\in]0,1[$. Each of the confidence intervals whose union gives $B_n$ has an asymptotic level equal to one; pointwise moderate deviations principles allow to prove that all these intervals share the same logarithmic asymptotic level. Now, as soon as both pointwise and uniform moderate deviations principles for $f_n$ exist, they share the same asymptotics. Taking this observation as a starting point, we present a new approach for the construction of confidence bands for $f$, based on the use of moderate deviations principles. The advantages of this approach are the following: (i) it enables to construct confidence bands, which have the same width (or even a smaller width) as the confidence bands provided by Bickel and Rosenblatt (1973), but which have a better aymptotic level; (ii) any confidence band constructed in that way shares the same logarithmic asymptotic level as all the confidence intervals, which make up this confidence band; (iii) it allows to deal with all the dimensions in the same way; (iv) it enables to sort out the problem of providing confidence bands for $f$ on compact sets on which $f$ vanishes (or on all $\bb R^d$), by introducing a truncating operation.

math.ST

Large and moderate deviations principles for recursive kernel estimators of a multivariate density and its partial derivatives

In this paper we prove large and moderate deviations principles for the recursive kernel estimator of a probability density function and its partial derivatives. Unlike the density estimator, the derivatives estimators exhibit a quadratic behavior not only for the moderate deviations scale but also for the large deviations one. We provide results both for the pointwise and the uniform deviations.

math.ST