SearcharxivSearch

arXiv subjects

Yong Kong

Publications and source records attributed to Yong Kong.

At least 19 recordsLinked to original sources

Regular structures of an intractable enumeration problem: a diagonal recurrence relation of monomer-polymer coverings on two-dimensional rectangular lattices

In the monomer-polymer model, a linear rigid polymer covers $k$ adjacent lattice sites, with no lattice site occupied by more than one polymer. The polymers are called $k$-mers, and those unoccupied lattice sites are called monomers. The well-known monomer-dimer model is a special case of the monomer-polymer model with $k=2$. The enumeration of polymer coverings on two-dimensional rectangular lattices is considered as "intractable". We prove that the number of coverings of $s$ polymer satisfies a simple recurrence relation $\sum_{i=0}^{2s} (-1)^i \binom{2s}{i} a_{n-i, m-i} = 2^s {(2s)!} / {s!}$ on a $n \times m$ rectangular lattice with open boundary conditions in both directions.

math.CO

Recurrence solution of monomer-polymer models on two-dimensional rectangular lattices

The problem of counting polymer coverings on the rectangular lattices is investigated. In this model, a linear rigid polymer covers $k$ adjacent lattice sites such that no two polymers occupy a common site. Those unoccupied lattice sites are considered as monomers. We prove that for a given number of polymers ($k$-mers), the number of arrangements for the polymers on two-dimensional rectangular lattices satisfies simple recurrence relations. These recurrence relations are quite general and apply for arbitrary polymer length ($k$) and the width of the lattices ($n$). The well-studied monomer-dimer problem is a special case of the monomer-polymer model when $k=2$. It is known the enumeration of monomer-dimer configurations in planar lattices is #P-complete. The recurrence relations shown here have the potential for hints for the solution of long-standing problems in this class of computational complexity.

cond-mat.stat-mech

Multiple consecutive runs of multi-state trials: distributions of $(k_1, k_2, \dots, k_\ell)$ patterns

The pattern $(k_1, k_2, \dots, k_\ell)$ is defined to have at least $k_1$ consecutive $1$'s followed by at least $k_2$ consecutive $2$'s, $\dots$, followed by at least $k_\ell$ consecutive $\ell$'s. By iteratively applying the method that was developed previously to decouple the combinatorial complexity involved in studying complicated patterns in random sequences, the distribution of pattern $(k_1, k_2, \dots, k_\ell)$ is derived for arbitrary $\ell$. Numerical examples are provided to illustrate the results.

math.CO

The m-th Longest Runs of Multivariate Random Sequences

The distributions of the $m$-th longest runs of multivariate random sequences are considered. For random sequences made up of $k$ kinds of letters, the lengths of the runs are sorted in two ways to give two definitions of run length ordering. In one definition, the lengths of the runs are sorted separately for each letter type. In the second definition, the lengths of all the runs are sorted together. Exact formulas are developed for the distributions of the m-th longest runs for both definitions. The derivations are based on a two-step method that is applicable to various other runs-related distributions, such as joint distributions of several letter types and multiple run lengths of a single letter type.

math.CO

Joint distribution of rises, falls, and number of runs in random sequences

By using the matrix formulation of the two-step approach to the distributions of runs, a recursive relation and an explicit expression are derived for the generating function of the joint distribution of rises and falls for multivariate random sequences in terms of generating functions of individual letters, from which the generating functions of the joint distribution of rises, falls, and number of runs are obtained. An explicit formula for the joint distribution of rises and falls with arbitrary specification is also obtained.

math.CO

Distributions of successions of arbitrary multisets

By using the matrix formulation of the two-step approach to distributions of patterns in random sequences, recurrence and explicit formulas for the generating functions of successions in random permutations of arbitrary multisets are derived. Explicit formulas for the mean and variance are also obtained.

math.CO

Adventitious angles problem: the lonely fractional derived angle

In the "classical" adventitious angle problem, for a given set of three angles $a$, $b$, and $c$ measured in integral degrees in an isosceles triangle, a fourth angle $θ$ (the derived angle), also measured in integral degrees, is sought. We generalize the problem to find $θ$ in fractional degrees. We show that the triplet $(a, b, c) = (45^\circ, 45^\circ, 15^\circ)$ is the only combination that leads to $θ= 7\frac{1}{2}^\circ$ as the fractional derived angle.

math.CA

Fast- and thermal-neutron detection with common NaI(Tl) detectors

Radionuclide Identification Devices (RIDs) or Backpack Radiation Detection Systems (BRDs) are often equipped with NaI(Tl) detectors. We demonstrate that such instruments could be provided with reasonable thermal- and fast-neutron sensitivity by means of an improved and sophisticated processing of the digitized detector signals: Fast neutrons produce nuclear recoils in the scintillation crystal. Corresponding signals are detectible and can be distinguished from that of electronic interactions by pulse-shape discrimination (PSD) techniques as used in experiments searching for weakly interacting massive particles (WIMPs). Thermal neutrons are often captured in iodine nuclei of the scintillator. The gamma-ray cascades following such captures comprise a sum energy of almost 7 MeV, and some of them involve isomeric states leading to delayed gamma emissions. Both features can be used to distinguish corresponding detector signals from responses to ambient gamma radiation. The experimental proof was adduced by offline analyses of pulse records taken with a commercial RID. An implementation of such techniques in commercial RIDs is feasible.

physics.ins-det

Length distribution of sequencing by synthesis: fixed flow cycle model

Sequencing by synthesis is the underlying technology for many next-generation DNA sequencing platforms. We developed a new model, the fixed flow cycle model, to derive the distributions of sequence length for a given number of flow cycles under the general conditions where the nucleotide incorporation is probabilistic and may be incomplete, as in some single-molecule sequencing technologies. Unlike the previous model, the new model yields the probability distribution for the sequence length. Explicit closed form formulas are derived for the mean and variance of the distribution.

q-bio.GN

Distributions of positive signals in pyrosequencing

Pyrosequencing is one of the important next-generation sequencing technologies. We derive the distribution of the number of positive signals in pyrograms of this sequencing technology as a function of flow cycle numbers and nucleotide probabilities of the target sequences. As for the distribution of sequence length, we also derive the distribution of positive signals for the fixed flow cycle model. Explicit formulas are derived for the mean and variance of the distributions. A simple result for the mean of the distribution is that the mean number of positive signals in a pyrogram is approximately twice the number of flow cycles, regardless of nucleotide probabilities. The statistical distributions will be useful for instrument and software development for pyrosequencing and other related platforms.

q-bio.GN

Statistical distributions of pyrosequencing

Pyrosequencing is emerging as one of the important next-generation sequencing technologies. We derive the statistical distributions of this technique in terms of nucleotide probabilities of the target sequences. We give exact distributions both for fixed number of flow cycles and for fixed sequence length. Explicit formulas are derived for the mean and variance of these distributions. In both cases, the distributions can be approximated accurately by normal distributions with the same mean and variance. The statistical distributions will be useful for instrument and software development for pyrosequencing platforms.

q-bio.GN

Statistical distributions of sequencing by synthesis with probabilistic nucleotide incorporation

Sequencing by synthesis is used in many next-generation DNA sequencing technologies. Some of the technologies, especially those exploring the principle of single-molecule sequencing, allow incomplete nucleotide incorporation in each cycle. We derive statistical distributions for sequencing by synthesis by taking into account the possibility that nucleotide incorporation may not be complete in each flow cycle. The statistical distributions are expressed in terms of nucleotide probabilities of the target sequences and the nucleotide incorporation probabilities for each nucleotide. We give exact distributions both for fixed number of flow cycles and for fixed sequence length. Explicit formulas are derived for the mean and variance of these distributions. The results are generalizations of our previous work for pyrosequencing. Incomplete nucleotide incorporation leads to significant change in the mean and variance of the distributions, but still they can be approximated by normal distributions with the same mean and variance. The results are also generalized to handle sequence context dependent incorporation. The statistical distributions will be useful for instrument and software development for sequencing by synthesis platforms.

q-bio.GN

Packing dimers on $(2p + 1) \times (2q + 1) $ lattices

We use computational method to investigate the number of ways to pack dimers on \emph{odd-by-odd} lattices. In this case, there is always a single vacancy in the lattices. We show that the dimer configuration numbers on $(2k+1) \times (2k+1)$ \emph{odd} square lattices have some remarkable number-theoretical properties in parallel to those of close-packed dimers on $2k \times 2k$ \emph{even} square lattices, for which exact solution exists. Furthermore, we demonstrate that there is an unambiguous logarithm term in the finite size correction of free energy of odd-by-odd lattice strips with any width $n \ge 1$. This logarithm term determines the distinct behavior of the free energy of odd square lattices. These findings reveal a deep and previously unexplored connection between statistical physics models and number theory, and indicate the possibility that the monomer-dimer problem might be solvable.

cond-mat.stat-mech

Btrim: A fast, lightweight adapter and quality trimming program for next-generation sequencing technologies

Btrim is a fast and lightweight software to trim adapters and low quality regions in reads from ultra high-throughput next-generation sequencing machines. It also can reliably identify barcodes and assign the reads to the original samples. Based on a modified Myers's bit-vector dynamic programming algorithm, Btrim can handle indels in adapters and barcodes. It removes low quality regions and trims off adapters at both or either end of the reads. A typical trimming of 30M reads with two sets of adapter pairs can be done in about a minute with a small memory footprint. Btrim is a versatile stand-alone tool that can be used as the first step in virtually all next-generation sequence analysis pipelines. The program is available at \url{http://graphics.med.yale.edu/trim/}.

q-bio.GN

Calculating complexity of large randomized libraries

Randomized libraries are increasingly popular in protein engineering and other biomedical research fields. Statistics of the libraries are useful to guide and evaluate randomized library construction. Previous works only give the mean of the number of unique sequences in the library, and they can only handle equal molar ratio of the four nucleotides at a small number of mutation sites. We derive formulas to calculate the mean and variance of the number of unique sequences in libraries generated by cassette mutagenesis with mixtures of arbitrary nucleotide ratios. Computer program was developed which utilizes arbitrary numerical precision software package to calculate the statistics of large libraries. The statistics of library with mutations in more than $20$ amino acids can be calculated easily. Results show that the nucleotide ratios have significant effects on these statistics. The more skewed the ratio, the larger the library size is needed to obtain the same expected number of unique sequences. The program is freely available at \url{http://graphics.med.yale.edu/cgi-bin/lib_comp.pl}.

q-bio.QM

Exact asymptotics of monomer-dimer model on rectangular semi-infinite lattices

By using the asymptotic theory of Pemantle and Wilson, exact asymptotic expansions of the free energy of the monomer-dimer model on rectangular $n \times \infty$ lattices in terms of dimer density are obtained for small values of $n$, at both high and low dimer density limits. In the high dimer density limit, the theoretical results confirm the dependence of the free energy on the parity of $n$, a result obtained previously by computational methods. In the low dimer density limit, the free energy on a cylinder $n \times \infty$ lattice strip has exactly the same first $n$ terms in the series expansion as that of infinite $\infty \times \infty$ lattice.

cond-mat.stat-mech

Monomer-dimer model in two-dimensional rectangular lattices with fixed dimer density

The classical monomer-dimer model in two-dimensional lattices has been shown to belong to the \emph{``#P-complete''} class, which indicates the problem is computationally ``intractable''. We use exact computational method to investigate the number of ways to arrange dimers on $m \times n$ two-dimensional rectangular lattice strips with fixed dimer density $ρ$. For any dimer density $0 < ρ< 1$, we find a logarithmic correction term in the finite-size correction of the free energy per lattice site. The coefficient of the logarithmic correction term is exactly -1/2. This logarithmic correction term is explained by the newly developed asymptotic theory of Pemantle and Wilson. The sequence of the free energy of lattice strips with cylinder boundary condition converges so fast that very accurate free energy $f_2(ρ)$ for large lattices can be obtained. For example, for a half-filled lattice, $f_2(1/2) = 0.633195588930$, while $f_2(1/4) = 0.4413453753046$ and $f_2(3/4) = 0.64039026$. For $ρ< 0.65$, $f_2(ρ)$ is accurate at least to 10 decimal digits. The function $f_2(ρ)$ reaches the maximum value $f_2(ρ^*) = 0.662798972834$ at $ρ^* = 0.6381231$, with 11 correct digits. This is also the \md constant for two-dimensional rectangular lattices. The asymptotic expressions of free energy near close packing are investigated for finite and infinite lattice widths. For lattices with finite width, dependence on the parity of the lattice width is found. For infinite lattices, the data support the functional form obtained previously through series expansions.

cond-mat.stat-mech

Logarithmic corrections in the free energy of monomer-dimer model on plane lattices with free boundaries

Using exact computations we study the classical hard-core monomer-dimer models on m x n plane lattice strips with free boundaries. For an arbitrary number v of monomers (or vacancies), we found a logarithmic correction term in the finite-size correction of the free energy. The coefficient of the logarithmic correction term depends on the number of monomers present (v) and the parity of the width n of the lattice strip: the coefficient equals to v when n is odd, and v/2 when n is even. The results are generalizations of the previous results for a single monomer in an otherwise fully packed lattice of dimers.

cond-mat.stat-mech