Searcharxiv⌕ Search

arXiv subjects

Jens Ledet Jensen

Publications and source records attributed to Jens Ledet Jensen.

5 recordsLinked to original sources

Using Topology to Estimate Structural Similarities of Proteins

An effective model for protein structures is important for the study of protein geometry, which, to a large extent, determine the functions of proteins. There are a number of approaches for modelling; one might focus on the conformation of the backbone or H-bonds, and the model may be based on the geometry or the topology of the structure in focus. We focus on the topology of H-bonds in proteins, and explore the link between the topology and the geometry of protein structures. More specifically, we take inspiration from CASP Evaluation of Model Accuracy and investigate the extent to which structural similarities, via GDT_TS, can be estimated from the topology of H-bonds. We report on two experiments; one where we attempt to mimic the computation of GDT_TS based solely on the topology of H-bonds, and the other where we perform linear regression where the independent variables are various scores computed from the topology of H-bonds. We achieved an average $Δ\text{GDT}$ of 6.45 with 54.5% of predictions inside 2 $Δ\mathrm{GDT}$ for the first method, and an average $Δ\mathrm{GDT}$ of 4.41 with 72.7% of predictions inside 2 $Δ\mathrm{GDT}$ for the second method.

q-bio.BM↗

Approximating the Laplace transform of the sum of dependent lognormals

Let $(X_1, \dots, X_n)$ be multivariate normal, with mean vector $\boldsymbolμ$ and covariance matrix $\boldsymbolΣ$, and $S_n=\mathrm{e}^{X_1}+\cdots+\mathrm{e}^{X_n}$. The Laplace transform ${\cal L}(θ)=\mathbb{E}\mathrm{e}^{-θS_n} \propto \int \exp\{-h_θ(\boldsymbol{x})\} \,\mathrm{d} \boldsymbol{x}$ is represented as $\tilde{\cal L}(θ)I(θ)$, where $\tilde{\cal L}(θ)$ is given in closed-form and $I(θ)$ is the error factor ($\approx 1$). We obtain $\tilde{\cal L}(θ)$ by replacing $h_θ(\boldsymbol{x})$ with a second order Taylor expansion around its minimiser $\boldsymbol{x}^*$. An algorithm for calculating the asymptotic expansion of $\boldsymbol{x}^*$ is presented, and it is shown that $I(θ)\to 1$ as $θ\to\infty$. A variety of numerical methods for evaluating $I(θ)$ are discussed, including Monte Carlo with importance sampling and quasi-Monte Carlo. Numerical examples (including Laplace transform inversion for the density of $S_n$) are also given.

math.PR↗

On oracle efficiency of the ROAD classification rule

For high-dimensional classification Fishers rule performs poorly due to noise from estimation of the covariance matrix. Fan, Feng and Tong (2012) introduced the ROAD classifier that puts an $L_1$-constraint on the classification vector. In their Theorem 1 Fan, Feng and Tong (2012) show that the ROAD classifier asymptotically has the same misclassification rate as the corresponding oracle based classifier. Unfortunately, the proof contains an error. Here we restate the theorem and provide a new proof.

math.ST↗

Exponential Family Techniques for the Lognormal Left Tail

Let $X$ be lognormal$(μ,σ^2)$ with density $f(x)$, let $θ>0$ and define ${L}(θ)=E e^{-θX}$. We study properties of the exponentially tilted density (Esscher transform) $f_θ(x) =e^{-θx}f(x)/{L}(θ)$, in particular its moments, its asymptotic form as $θ\to\infty$ and asymptotics for the Cramér function; the asymptotic formulas involve the Lambert W function. This is used to provide two different numerical methods for evaluating the left tail probability of lognormal sum $S_n=X_1+\cdots+X_n$: a saddlepoint approximation and an exponential twisting importance sampling estimator. For the latter we demonstrate the asymptotic consistency by proving logarithmic efficiency in terms of the mean square error. Numerical examples for the c.d.f.\ $F_n(x)$ and the p.d.f.\ $f_n(x)$ of $S_n$ are given in a range of values of $σ^2,n,x$ motivated from portfolio Value-at-Risk calculations.

math.PR↗