Searcharxiv⌕ Search

arXiv subjects

C. Houdré

Publications and source records attributed to C. Houdré.

13 recordsLinked to original sources

Sparse long blocks and the variance of the longest common subsequences in random words

Consider two independent random strings having same length and taking values uniformly in a common finite alphabet. We study the order of the variance of the length of the longest common subsequences (LCS) of these strings when long blocks, or other types of atypical substrings, are sparsely added into one of them. Under weak conditions on the derivative of the mean LCS-curve, the order of the variance of the LCS is shown to be linear in the length of the strings. We also argue that our proofs carry over to many models used by computational biologists to simulate DNA-sequences. This is the first result where the open question of the order of the fluctuation of the LCS of random strings is solved for a realistic model. Until now, this type of result had only been established for low entropy cases.

math.PR↗

Closeness to the Diagonal for Longest Common Subsequences in Random Words

The nature of the alignment with gaps corresponding to a longest common subsequence (LCS) of two independent iid random sequences drawn from a finite alphabet is investigated. It is shown that such an optimal alignment typically matches pieces of similar short-length. This is of importance in understanding the structure of optimal alignments of two sequences. Moreover, it is also shown that any property, common to two subsequences, typically holds in most parts of the optimal alignment whenever this same property holds, with high probability, for strings of similar short-length. Our results should, in particular, prove useful for simulations since they imply that the re-scaled two dimensional representation of a LCS gets uniformly close to the diagonal as the length of the sequences grows without bound.

math.PR↗

Sparse Long Blocks and the Micro-Structure of the Longest Common Subsequences

Consider two random strings having the same length and generated by an iid sequence taking its values uniformly in a fixed finite alphabet. Artificially place a long constant block into one of the strings, where a constant block is a contiguous substring consisting only of one type of symbol. The long block replaces a segment of equal size and its length is smaller than the length of the strings, but larger than its square-root. We show that for sufficiently long strings the optimal alignment corresponding to a Longest Common Subsequence (LCS) treats the inserted block very differently depending on the size of the alphabet. For two-letter alphabets, the long constant block gets mainly aligned with the same symbol from the other string, while for three or more letters the opposite is true and the block gets mainly aligned with gaps. We further provide simulation results on the proportion of gaps in blocks of various lengths. In our simulations, the blocks are "regular blocks" in an iid sequence, and are not artificially inserted. Nonetheless, we observe for these natural blocks a phenomenon similar to the one shown in case of artificially-inserted blocks: with two letters, the long blocks get aligned with a smaller proportion of gaps; for three or more letters, the opposite is true. It thus appears that the microscopic nature of two-letter optimal alignments and three-letter optimal alignments are entirely different from each other.

math.PR↗

Small-time expansions of the distributions, densities, and option prices of stochastic volatility models with Lévy jumps

We consider a stochastic volatility model with Lévy jumps for a log-return process $Z=(Z_{t})_{t\geq 0}$ of the form $Z=U+X$, where $U=(U_{t})_{t\geq 0}$ is a classical stochastic volatility process and $X=(X_{t})_{t\geq 0}$ is an independent Lévy process with absolutely continuous Lévy measure $ν$. Small-time expansions, of arbitrary polynomial order, in time-$t$, are obtained for the tails $\bbp(Z_{t}\geq z)$, $z>0$, and for the call-option prices $\bbe(e^{z+Z_{t}}-1)_{+}$, $z\neq 0$, assuming smoothness conditions on the {\PaleGrey density of $ν$} away from the origin and a small-time large deviation principle on $U$. Our approach allows for a unified treatment of general payoff functions of the form $ϕ(x){\bf 1}_{x\geq{}z}$ for smooth functions $ϕ$ and $z>0$. As a consequence of our tail expansions, the polynomial expansions in $t$ of the transition densities $f_{t}$ are also {\Green obtained} under mild conditions.

q-fin.PR↗

Transportation Distance and the Central Limit Theorem

For probability measures on a complete separable metric space, we present sufficient conditions for the existence of a solution to the Kantorovich transportation problem. We also obtain sufficient conditions (which sometimes also become necessary) for the convergence, in transportation, of probability measures when the cost function is continuous, non-decreasing and depends on the distance. As an application, the CLT in the transportation distance is proved for independent and some dependent stationary sequences.

math.PR↗

Median, Concentration and Fluctuation for Lévy Processes

We estimate a median of $f(X_t)$ where $f$ is a Lipschitz function, $X$ is a Lévy process and $t$ an arbitrary time. This leads to concentration inequalities for $f(X_t)$. In turn, corresponding fluctuation estimates are obtained under assumptions typically satisfied if the process has a regular behavior in small time and a, possibly different, regular behavior in large time.

math.PR↗

Wavelet thresholding for nonnecessarily Gaussian noise: functionality

For signals belonging to balls in smoothness classes and noise with enough moments, the asymptotic behavior of the minimax quadratic risk among soft-threshold estimates is investigated. In turn, these results, combined with a median filtering method, lead to asymptotics for denoising heavy tails via wavelet thresholding. Some further comparisons of wavelet thresholding and of kernel estimators are also briefly discussed.

math.ST↗

On Fractional Tempered Stable Motion

Fractional tempered stable motion (fTSm)} is defined and studied. FTSm has the same covariance structure as fractional Brownian motion, while having tails heavier than Gaussian but lighter than stable. Moreover, in short time it is close to fractional stable Lévy motion, while it is approximately fractional Brownian motion in long time. A series representation of fTSm is derived and used for simulation and to study some of its sample path properties.

math.PR↗

On Layered Stable Processes

Layered stable (multivariate) distributions and processes are defined and studied. A layered stable process combines stable trends of two different indices, one of them possibly Gaussian. More precisely, in short time, it is close to a stable process while, in long time, it approximates another stable (possibly Gaussian) process. We also investigate the absolute continuity of a layered stable process with respect to its short time limiting stable process. A series representation of layered stable processes is derived, giving insights into both the structure of the sample paths and of the short and long time behaviors. This series is further used for sample paths simulation.

math.PR↗

Dimension free and infinite variance tail estimates on Poisson space

Concentration inequalities are obtained on Poisson space, for random functionals with finite or infinite variance. In particular, dimension free tail estimates and exponential integrability results are given for the Euclidean norm of vectors of independent functionals. In the finite variance case these results are applied to infinitely divisible random variables such as quadratic Wiener functionals, including Lévy's stochastic area and the square norm of Brownian paths. In the infinite variance case, various tail estimates such as stable ones are also presented.

math.PR↗

On finite range stable type concentration

The purpose of these notes is to further complete our understanding of the stable concentration phenomenon, by obtaining the finite range behavior of $P(F-E[F]\geq x)$, with $F=f(X)$ where $f$ is a Lipschitz function and $X$ is a stable random vector or with $F$ a stochastic functional on the Poisson space equipped with a stable Lévy measure.

math.PR↗