SearcharxivSearch

arXiv subjects

Paul Rochet

Publications and source records attributed to Paul Rochet.

At least 19 recordsLinked to original sources

Cyclists route choice modeling from trip duration data in urban areas

The lack of GPS data limits the ability to reconstruct the actual routes taken by cyclists in urban areas. This article introduces an inference method based solely on trip durations and origin-destination pairs from bike-sharing system (BSS) users. Travel time distributions are modeled using log-normal mixture models, allowing us to identify the presence of distinct behaviors. The approach is applied to 3.8 million trips recorded in 2022 in the Toulouse metropolitan area, with observed durations compared against travel times estimated by OpenStreetMap (OSM). Results show that, for many station pairs, trip durations align closely with the fastest route suggested by OSM, reflecting a dominant and routine practice. In other cases, mixture models reveal more heterogeneous behaviors, including longer trips, detours, or intermediate stops. This approach highlights both the stability and diversity of cycling practices, providing a robust tool for usage analysis in data-limited contexts, and offering new insights into urban mobility dynamics without relying on spatially explicit data.

stat.AP

Importance sampling for Sobol' indices estimation

We propose a new importance sampling framework for the estimation and analysis of Sobol' indices. We focus on the estimation of the conditional second-moment quantity underlying these indices, which is the most challenging term to estimate. We show that this quantity, originally defined under a reference input distribution, can be estimated from samples drawn under auxiliary distributions by reweighting the model outputs. We derive the optimal sampling distribution that minimises the asymptotic variance of efficient estimators and demonstrate its impact on estimation. Beyond variance reduction, the framework also supports distributional sensitivity analysis through reverse importance sampling.

math.ST

Asymptotic efficiency for Sobol' and Cram{\'e}r-von Mises indices under two designs of experiments

A variety of indices aim to quantify the impact of input variables on a response, typically the output from a complex computer code or black-box model. Most commonly used, the Sobol' index typically measures the influence of some inputs from an explained variance perspective. However, some situations may require a more targeted analysis of some inputs influence. With no prior information, distribution-based measures appear to be appealing. In this purpose, so-called Cram{\'e}r-von Mises indices (and their generalization) have been proposed in the literature, defined as an excess probability integrated over the output distribution that aim to reflect influence on the whole distribution of the output rather than on the variance solely. Inference of these various indices has remained a challenging topic especially in presence of many inputs. While several Sobol' indices estimators are known to be optimal under regularity conditions, the issue of asymptotic efficiency for Cram{\'e}r-von Mises indices has been unaddressed in the literature so far. For these indices, we derive in this paper the efficiency bounds and discuss the known methods to achieve such optimal bounds. Two estimation contexts are considered: the so-called Pick-Freeze scheme and the Given-Data setting, for which the estimation is produced from a unique input-output sample.

math.ST

Efficiency of the averaged rank-based estimator for first order Sobol index inference

Among the many estimators of first order Sobol indices that have been proposed in the literature, the so-called rank-based estimator is arguably the simplest to implement. This estimator can be viewed as the empirical auto-correlation of the response variable sample obtained upon reordering the data by increasing values of the inputs. This simple idea can be extended to higher lags of autocorrelation, thus providing several competing estimators of the same parameter. We show that these estimators can be combined in a simple manner to achieve the theoretical variance efficiency bound asymptotically.

math.ST

Impact of the COVID-19 pandemic on bike-sharing uses in two french towns

Urban areas have been dramatically impacted by the sudden and fast spread of the COVID-19 pandemic. As one of the most noticeable consequences of the pandemic, people have quickly reconsidered their travel options to minimize infection risk. Many studies on the Bike Sharing System (BSS) of several towns have shown that, in this context, cycling appears as a resilient, safe and very reliable mobility option. Differences and similarities exist about how people reacted depending on the place being considered, and it is paramount to identify and understand such reactions in the aftermath of an event in order to successfully foster permanent changes. In this paper, we carry out a comparative analysis of the effects of the pandemic on BSS usage in two French towns, Toulouse and Lyon. We used Origin/Destination data for the two years 2019 (pre-pandemic) and 2020 (pandemic), and considered two complementary quantitative approaches. Our results confirm that cycling increased during the pandemic, more significantly in Lyon than in Toulouse, with rush times remaining exactly the same as during the pre-pandemic year. Among several results, we note for example that BSS usage is more evenly spread throughout the day in 2020, peripheral/city center flow is more noticeable in Toulouse than in Lyon and that student BSS usage is more specific in Lyon. We also found that trip duration during the pandemic situation was longer on working days and shorter on weekends.

stat.AP

Test comparison for Sobol Indices over nested sets of variables

Sensitivity indices are commonly used to quantify the relative influence of any specific group of input variables on the output of a computer code. One crucial question is then to decide whether a given set of variables has a significant impact on the output. Sobol indices are often used to measure this impact but their estimation can be difficult as they usually require a particular design of experiment. In this work, we take advantage of the monotonicity of Sobol indices with respect to set inclusion to test the influence of some of the input variables. The method does not rely on a direct estimation of the Sobol indices and can be performed under classical iid sampling designs.

math.ST

A coupling of the spectral measures at a vertex

Given the adjacency matrix of an undirected graph, we define a coupling of the spectral measures at the vertices, whose moments count the rooted closed paths in the graph. The resulting joint spectral measure verifies numerous interesting properties that allow to recover minors of analytical functions of the adjacency matrix from its generalized moments. We prove an extension of Obata's Central Limit Theorem in growing star-graphs to the multivariate case and discuss some combinatorial properties using Viennot's heaps of pieces point of view.

math.CO

Enumerating simple paths from connected induced subgraphs

We present an exact formula for the ordinary generating series of the simple paths between any two vertices of a graph. Our formula involves the adjacency matrix of the connected induced subgraphs and remains valid on weighted and directed graphs. As a particular case, we obtain a relation linking the Hamiltonian paths and cycles of a graph to its dominating connected sets.

math.CO

Reconstructing undirected graphs from eigenspaces

In this paper, we aim at recovering an undirected weighted graph of $N$ vertices from the knowledge of a perturbed version of the eigenspaces of its adjacency matrix $W$. For instance, this situation arises for stationary signals on graphs or for Markov chains observed at random times. Our approach is based on minimizing a cost function given by the Frobenius norm of the commutator $\mathsf{A} \mathsf{B}-\mathsf{B} \mathsf{A}$ between symmetric matrices $\mathsf{A}$ and $\mathsf{B}$. In the Erdős-Rényi model with no self-loops, we show that identifiability (i.e., the ability to reconstruct $W$ from the knowledge of its eigenspaces) follows a sharp phase transition on the expected number of edges with threshold function $N\log N/2$. Given an estimation of the eigenspaces based on a $n$-sample, we provide support selection procedures from theoretical and practical point of views. In particular, when deleting an edge from the active support, our study unveils that our test statistic is the order of $\mathcal O(1/n)$ when we overestimate the true support and lower bounded by a positive constant when the estimated support is smaller than the true support. This feature leads to a powerful practical support estimation procedure. Simulated and real life numerical experiments assert our new methodology.

math.ST

An Hopf algebra for counting simple cycles

Simple cycles, also known as self-avoiding polygons, are cycles on graphs which are not allowed to visit any vertex more than once. We present an exact formula for enumerating the simple cycles of any length on any directed graph involving a sum over its induced subgraphs. This result stems from an Hopf algebra, which we construct explicitly, and which provides further means of counting simple cycles. Finally, we obtain a more general theorem asserting that any Lie idempotent can be used to enumerate simple cycles.

math.AC

The Mean/Max Statistic in Extreme Value Analysis

Most extreme events in real life can be faithfully modeled as random realizations from a Generalized Pareto distribution, which depends on two parameters: the scale and the shape. In many actual situations, one is mostly concerned with the shape parameter, also called tail index, as it contains the main information on the likelihood of extreme events. In this paper, we show that the mean/max statistic, that is the empirical mean divided by the maximal value of the sample, constitutes an ideal normalization to study the tail index independently of the scale. This statistic appears naturally when trying to distinguish between uniform and exponential distributions, the two transitional phases of the Generalized Pareto model. We propose a simple methodology based on the mean/max statistic to detect, classify and infer on the tail of the distribution of a sample. Applications to seismic events and detection of saturation in experimental measurements are presented.

math.ST

A new class of graphs that satisfies the Chen-Chvátal Conjecture

A well-known combinatorial theorem says that a set of n non-collinear points in the plane determines at least n distinct lines. Chen and Chvátal conjectured that this theorem extends to metric spaces, with an appropriated definition of line. In this work we prove a slightly stronger version of Chen and Chvátal conjecture for a family of graphs containing chordal graphs and distance-hereditary graphs.

math.CO

Relations between connected and self-avoiding walks in a digraph

Walks in a directed graph can be given a partially ordered structure that extends to possibly unconnected objects, called hikes. Studying the incidence algebra on this poset reveals unsuspected relations between walks and self-avoiding hikes. These relations are derived by considering truncated versions of the characteristic polynomial of the weighted adjacency matrix, resulting in a collection of matrices whose entries enumerate the self-avoiding hikes of length $\ell$ from one vertex to another.

math.CO

A general procedure to combine estimators

A general method to combine several estimators of the same quantity is investigated. In the spirit of model and forecast averaging, the final estimator is computed as a weighted average of the initial ones, where the weights are constrained to sum to one. In this framework, the optimal weights, minimizing the quadratic loss, are entirely determined by the mean square error matrix of the vector of initial estimators. The averaging estimator is built using an estimation of this matrix, which can be computed from the same dataset. A non-asymptotic error bound on the averaging estimator is derived, leading to asymptotic optimality under mild conditions on the estimated mean square error matrix. This method is illustrated on standard statistical problems in parametric and semi-parametric models where the averaging estimator outperforms the initial estimators in most cases.

stat.ME

Hypothesis testing for markovian models with random time observations

The aim of this paper is to propose a methodology for testing general hypothesis in a Markovian setting with random sampling. A discrete Markov chain X is observed at random time intervals $τ$ k, assumed to be iid with unknown distribution $μ$. Two test procedures are investigated. The first one is devoted to testing if the transition matrix P of the Markov chain X satisfies specific affine constraints, covering a wide range of situations such as symmetry or sparsity. The second procedure is a goodness-of-fit test on the distribution $μ$, which reveals to be consistent under mild assumptions even though the time gaps are not observed. The theoretical results are supported by a Monte Carlo simulation study to show the performance and robustness of the proposed methodologies on specific numerical examples.

math.ST

Estimating the transition matrix of a Markov chain observed at random times

In this paper we develop a statistical estimation technique to recover the transition kernel $P$ of a Markov chain $X=(X_m)_{m \in \mathbb N}$ in presence of censored data. We consider the situation where only a sub-sequence of $X$ is available and the time gaps between the observations are iid random variables. Under the assumption that neither the time gaps nor their distribution are known, we provide an estimation method which applies when some transitions in the initial Markov chain $X$ are known to be unfeasible. A consistent estimator of $P$ is derived in closed form as a solution of a minimization problem. The asymptotic performance of the estimator is then discussed in theory and through numerical simulations.

math.ST

A Cramér-Rao inequality for non differentiable models

We compute a variance lower bound for unbiased estimators in specified statistical models. The construction of the bound is related to the original Cramér-Rao bound, although it does not require the differentiability of the model. Moreover, we show our efficiency bound to be always greater than the Cramér-Rao bound in smooth models, thus providing a sharper result.

math.ST

Bayesian interpretation of Generalized empirical likelihood by maximum entropy

We study a parametric estimation problem related to moment condition models. As an alternative to the generalized empirical likelihood (GEL) and the generalized method of moments (GMM), a Bayesian approach to the problem can be adopted, extending the MEM procedure to parametric moment conditions. We show in particular that a large number of GEL estimators can be interpreted as a maximum entropy solution. Moreover, we provide a more general field of applications by proving the method to be robust to approximate moment conditions.

math.ST