SearcharxivSearch

arXiv subjects

Luis Mendo

Publications and source records attributed to Luis Mendo.

13 recordsLinked to original sources

Efficient estimation of relative risk, odds ratio and their logarithms for rare events

Sequential estimators are proposed for the relative risk, odds ratio, log relative risk or log odds ratio of a dichotomous attribute in two populations. The estimators take the same number of observations from each population, and guarantee that the relative mean-square error for the relative risk or odds ratio, or the mean-square error for their logarithmic versions, is less than a given target. The efficiency of the estimators, defined in terms of the Cram\'er-Rao bound, is high when the considered attribute is rare or moderately rare.

stat.ME

Characterizing 5G User Throughput via Uncertainty Modeling and Crowdsourced Measurements

Characterizing application-layer user throughput in next-generation networks is increasingly challenging as the higher capacity of the 5G Radio Access Network (RAN) shifts connectivity bottlenecks towards deeper parts of the network. Traditional methods, such as drive tests and operator equipment counters, are costly, limited, or fail to capture end-to-end (E2E) Quality of Service (QoS) and its variability. In this work, we leverage large-scale crowdsourced measurements-including E2E, radio, contextual and network deployment features collected by the user equipment (UE)-to propose an uncertainty-aware and explainable approach for downlink user throughput estimation. We first validate prior 4G methods, improving R^2 by 8.7%, and then extend them to 5G NSA and 5G SA, providing the first benchmarks for 5G crowdsourced datasets. To address the variability of throughput, we apply NGBoost, a model that outputs both point estimates and calibrated confidence intervals, representing its first use in the field of computer communications. Finally, we use the proposed model to analyze the evolution from 4G to 5G SA, and show that throughput bottlenecks move from the RAN to transport and service layers, as seen by E2E metrics gaining importance over radio-related features.

cs.NI

Direct-to-Cell: A First Look into Starlink's Direct Satellite-to-Device Radio Access Network through Crowdsourced Measurements

Low Earth Orbit (LEO) satellite mega-constellations have emerged as a viable access solution for broadband connectivity in underserved areas. In 2024, Starlink, in partnership with T-Mobile, began beta testing an SMS-only Supplemental Coverage from Space (SCS) service. This marks the first large-scale deployment of Direct Satellite-to-Device (DS2D) communications, allowing unmodified smartphones to connect directly to spaceborne base stations. This paper presents the first measurement study of deployed DS2D technologies. Using crowdsourced mobile network data from the U.S. between October 2024 and July 2025, we provide evidence-based insights into the capabilities, limitations, and future evolution of DS2D technologies for extending mobile connectivity. We find a strong correlation between the number of satellites deployed, the number of unique cell identifiers measured, and the volume of measurements, concentrated in accessible areas with poor terrestrial network coverage, such as national parks and sparsely populated counties. Stable physical-layer measurements were observed throughout the period, with a 24-dB lower median RSRP and a 3-dB higher RSRQ compared to terrestrial networks, reflecting the SMS-only usage of the DS2D network during this period. Based on the SINR measurements collected, we estimate the expected performance of the announced DS2D mobile data service to be around 3 Mbps per beam in outdoor conditions. We also discuss strategies to expand this capacity up to 18 Mbps in the future, depending on key regulatory and business decisions, including allowable out-of-band emissions, permitted number of satellites, and availability of spectrum and orbital resources.

cs.NI

Estimation of relative risk, odds ratio and their logarithms with guaranteed accuracy and controlled sample size ratio

Given two populations from which independent binary observations are taken with parameters $p_1$ and $p_2$ respectively, estimators are proposed for the relative risk $p_1/p_2$, the odds ratio $p_1(1-p_2)/(p_2(1-p_1))$ and their logarithms. The sampling strategy used by the estimators is based on two-stage sequential sampling applied to each population, where the sample sizes of the second stage depend on the results observed in the first stage. The estimators guarantee that the relative mean-square error, or the mean-square error for the logarithmic versions, is less than a target value for any $p_1, p_2 \in (0,1)$, and the ratio of average sample sizes from the two populations is close to a prescribed value. The estimators can also be used with group sampling, whereby samples are taken in batches of fixed size from the two populations simultaneously, each batch containing samples from the two populations. The efficiency of the estimators with respect to the Cram\'er-Rao bound is good, and in particular it is close to $1$ for small values of the target error.

stat.ME

Estimating odds and log odds with guaranteed accuracy

Two sequential estimators are proposed for the odds p/(1-p) and log odds log(p/(1-p)) respectively, using independent Bernoulli random variables with parameter p as inputs. The estimators are unbiased, and guarantee that the variance of the estimation error divided by the true value of the odds, or the variance of the estimation error of the log odds, are less than a target value for any p in (0,1). The estimators are close to optimal in the sense of Wolfowitz's bound.

math.ST

On the number of tiles visited by a line segment on a rectangular grid

Consider a line segment placed on a two-dimensional grid of rectangular tiles. This paper addresses the relationship between the length of the segment and the number of tiles it visits (i.e. has intersection with). The square grid is also considered explicitly, as some of the specific problems studied are more tractable in that particular case. The segment position and orientation can be modelled as either deterministic or random. In the deterministic setting, the maximum possible number of visited tiles is characterized for a given length, and conversely, the infimum segment length needed to visit a desired number of tiles is analyzed. In the random setting, the average number of visited tiles and the probability of visiting the maximum number of tiles on a square grid are studied as a function of segment length. These questions are related to Buffon's needle problem and its extension by Laplace.

math.MG

Simulating a coin with irrational bias using rational arithmetic

An algorithm is presented that, taking a sequence of independent Bernoulli random variables with parameter $1/2$ as inputs and using only rational arithmetic, simulates a Bernoulli random variable with possibly irrational parameter $\tau$. It requires a series representation of $\tau$ with positive, rational terms, and a rational bound on its truncation error that converges to $0$. The number of required inputs has an exponentially bounded tail, and its mean is at most $3$. The number of arithmetic operations has a tail that can be bounded in terms of the sequence of truncation error bounds. The algorithm is applied to two specific values of $\tau$, including Euler's constant, for which obtaining a simple simulation algorithm was an open problem.

math.PR

Estimation of a Probability with Guaranteed Normalized Mean Absolute Error

The estimation of a probability p from repeated Bernoulli trials is considered in this paper. A sequential approach is followed, using a simple stopping rule. A closed-form expression and an upper bound are obtained for the mean absolute error of the unbiased estimator of p. The results given permit the estimation of an arbitrary probability with a prescribed level of normalized mean absolute error.

math.ST

An asymptotically optimal Bernoulli factory for certain functions that can be expressed as power series

Given a sequence of independent Bernoulli variables with unknown parameter $p$, and a function $f$ expressed as a power series with non-negative coefficients that sum to at most $1$, an algorithm is presented that produces a Bernoulli variable with parameter $f(p)$. In particular, the algorithm can simulate $f(p)=p^a$, $a\in(0,1)$. For functions with a derivative growing at least as $f(p)/p$ for $p\rightarrow 0$, the average number of inputs required by the algorithm is asymptotically optimal among all simulations that are fast in the sense of Nacu and Peres. A non-randomized version of the algorithm is also given. Some extensions are discussed.

math.ST

Estimation of a probability in inverse binomial sampling under normalized linear-linear and inverse-linear loss

Sequential estimation of the success probability $p$ in inverse binomial sampling is considered in this paper. For any estimator $\hat p$, its quality is measured by the risk associated with normalized loss functions of linear-linear or inverse-linear form. These functions are possibly asymmetric, with arbitrary slope parameters $a$ and $b$ for $\hat p p$ respectively. Interest in these functions is motivated by their significance and potential uses, which are briefly discussed. Estimators are given for which the risk has an asymptotic value as $p$ tends to $0$, and which guarantee that, for any $p$ in $(0,1)$, the risk is lower than its asymptotic value. This allows selecting the required number of successes, $r$, to meet a prescribed quality irrespective of the unknown $p$. In addition, the proposed estimators are shown to be approximately minimax when $a/b$ does not deviate too much from $1$, and asymptotically minimax as $r$ tends to infinity when $a=b$.

math.ST

Asymptotically optimum estimation of a probability in inverse binomial sampling under general loss functions

The optimum quality that can be asymptotically achieved in the estimation of a probability p using inverse binomial sampling is addressed. A general definition of quality is used in terms of the risk associated with a loss function that satisfies certain assumptions. It is shown that the limit superior of the risk for p asymptotically small has a minimum over all (possibly randomized) estimators. This minimum is achieved by certain non-randomized estimators. The model includes commonly used quality criteria as particular cases. Applications to the non-asymptotic regime are discussed considering specific loss functions, for which minimax estimators are derived.

math.ST

Estimation of a probability with optimum guaranteed confidence in inverse binomial sampling

Sequential estimation of a probability $p$ by means of inverse binomial sampling is considered. For $μ_1,μ_2>1$ given, the accuracy of an estimator $\hat{p}$ is measured by the confidence level $P[p/μ_2\leq\hat{p}\leq pμ_1]$. The confidence levels $c_0$ that can be guaranteed for $p$ unknown, that is, such that $P[p/μ_2\leq \hat{p}\leq pμ_1]\geq c_0$ for all $p\in(0,1)$, are investigated. It is shown that within the general class of randomized or non-randomized estimators based on inverse binomial sampling, there is a maximum $c_0$ that can be guaranteed for arbitrary $p$. A non-randomized estimator is given that achieves this maximum guaranteed confidence under mild conditions on $μ_1$, $μ_2$.

math.ST

Improved Sequential Stopping Rule for Monte Carlo Simulation

This paper presents an improved result on the negative-binomial Monte Carlo technique analyzed in a previous paper for the estimation of an unknown probability p. Specifically, the confidence level associated to a relative interval [p/μ_2, pμ_1], with μ_1, μ_2 > 1, is proved to exceed its asymptotic value for a broader range of intervals than that given in the referred paper, and for any value of p. This extends the applicability of the estimator, relaxing the conditions that guarantee a given confidence level.

stat.CO