Searcharxiv⌕ Search

arXiv subjects

Mark Kon

Publications and source records attributed to Mark Kon.

23 records · Page 2Linked to original sources

Absorption Probabilities of Quantum Walks

Quantum walks are known to have nontrivial interaction with absorbing boundaries. In particular, Ambainis et.\ al.\ \cite{ambainis01} showed that in the $(\Z ,C_1,H)$ quantum walk (one-dimensional Hadamard walk) an absorbing boundary partially reflects information. These authors also conjectured that the left absorption probabilities $P_n^{(1)}(1,0)$ related to the finite absorbing Hadamard walks $(\Z ,C_1,H,\{ 0,n\} )$ satisfy a linear fractional recurrence in $n$ (here $P_n(1,0)$ is the probability that a Hadamard walk particle initialized in $|1\rangle |R\rangle$ is eventually absorbed at $|0\rangle$ and not at $|n\rangle$). This result, as well as a third order linear recurrence in initial position $m$ of $P_n^{(m)}(1,0)$, was later proved by Bach and Borisov \cite{bach09} using techniques from complex analysis. In this paper we extend these results to general two state quantum walks and three-state Grover walks, while providing a partial calculation for absorption in $d$-dimensional Grover walks by a $d-1$-dimensional wall. In the one-dimensional cases, we prove partial reflection of information, a linear fractional recurrence in lattice size, and a linear recurrence in initial position.

quant-ph↗

Transcription Factor-DNA Binding Via Machine Learning Ensembles

We present ensemble methods in a machine learning (ML) framework combining predictions from five known motif/binding site exploration algorithms. For a given TF the ensemble starts with position weight matrices (PWM's) for the motif, collected from the component algorithms. Using dimension reduction, we identify significant PWM-based subspaces for analysis. Within each subspace a machine classifier is built for identifying the TF's gene (promoter) targets (Problem 1). These PWM-based subspaces form an ML-based sequence analysis tool. Problem 2 (finding binding motifs) is solved by agglomerating k-mer (string) feature PWM-based subspaces that stand out in identifying gene targets. We approach Problem 3 (binding sites) with a novel machine learning approach that uses promoter string features and ML importance scores in a classification algorithm locating binding sites across the genome. For target gene identification this method improves performance (measured by the F1 score) by about 10 percentage points over the (a) motif scanning method and (b) the coexpression-based association method. Top motif outperformed 5 component algorithms as well as two other common algorithms (BEST and DEME). For identifying individual binding sites on a benchmark cross species database (Tompa et al., 2005) we match the best performer without much human intervention. It also improved the performance on mammalian TFs. The ensemble can integrate orthogonal information from different weak learners (potentially using entirely different types of features) into a machine learner that can perform consistently better for more TFs. The TF gene target identification component (problem 1 above) is useful in constructing a transcriptional regulatory network from known TF-target associations. The ensemble is easily extendable to include more tools as well as future PWM-based information.

q-bio.GN↗

Real-time simulation of dissipation-driven quantum systems

We set up a real-time path integral to study the evolution of quantum systems driven in real-time completely by the coupling of the system to the environment. For specifically chosen interactions, this can be interpreted as measurements being performed on the system. For a spin-1/2 system, in particular, when the measurement results are averaged over, the resulting sign problem completely disappears, and the system can be simulated with an efficient cluster algorithm.

quant-ph↗

The Marr Conjecture and Uniqueness of Wavelet Transforms

The inverse question of identifying a function from the nodes (zeroes) of its wavelet transform arises in a number of fields. These include whether the nodes of a heat or hypoelliptic equation solution determine its initial conditions, and in mathematical vision theory the Marr conjecture, on whether an image is mathematically determined by its edge information. We prove a general version of this conjecture by reducing it to the moment problem, using a basis dual to the Taylor monomial basis $x^α$ on $\mathbb {R}^n$.

math.FA↗

Feature vector regularization in machine learning

Problems in machine learning (ML) can involve noisy input data, and ML classification methods have reached limiting accuracies when based on standard ML data sets consisting of feature vectors and their classes. Greater accuracy will require incorporation of prior structural information on data into learning. We study methods to regularize feature vectors (unsupervised regularization methods), analogous to supervised regularization for estimating functions in ML. We study regularization (denoising) of ML feature vectors using Tikhonov and other regularization methods for functions on ${\bf R}^n$. A feature vector ${\bf x}=(x_1,\ldots,x_n)=\{x_q\}_{q=1}^n$ is viewed as a function of its index $q$, and smoothed using prior information on its structure. This can involve a penalty functional on feature vectors analogous to those in statistical learning, or use of proximity (e.g. graph) structure on the set of indices. Such feature vector regularization inherits a property from function denoising on ${\bf R}^n$, in that accuracy is non-monotonic in the denoising (regularization) parameter $α$. Under some assumptions about the noise level and the data structure, we show that the best reconstruction accuracy also occurs at a finite positive $α$ in index spaces with graph structures. We adapt two standard function denoising methods used on ${\bf R}^n$, local averaging and kernel regression. In general the index space can be any discrete set with a notion of proximity, e.g. a metric space, a subset of ${\bf R}^n$, or a graph/network, with feature vectors as functions with some notion of continuity. We show this improves feature vector recovery, and thus the subsequent classification or regression done on them. We give an example in gene expression analysis for cancer classification with the genome as an index space and network structure based protein-protein interactions.

stat.ML↗