Searcharxiv⌕ Search

arXiv subjects

Joel S. Bader

Publications and source records attributed to Joel S. Bader.

3 recordsLinked to original sources

Fast and reliable inference algorithm for hierarchical stochastic block models

Network clustering reveals the organization of a network or corresponding complex system with elements represented as vertices and interactions as edges in a (directed, weighted) graph. Although the notion of clustering can be somewhat loose, network clusters or groups are generally considered as nodes with enriched interactions and edges sharing common patterns. Statistical inference often treats groups as latent variables, with observed networks generated from latent group structure, termed a stochastic block model. Regardless of the definitions, statistical inference can be either translated to modularity maximization, which is provably an NP-complete problem. Here we present scalable and reliable algorithms that recover hierarchical stochastic block models fast and accurately. Our algorithm scales almost linearly in number of edges, and inferred models were more accurate that other scalable methods.

stat.ML↗

DNA transport by a micromachined Brownian ratchet device

We have micromachined a silicon-chip device that transports DNA with a Brownian ratchet that rectifies the Brownian motion of microscopic particles. Transport properties for a DNA 50mer agree with theoretical predictions, and the DNA diffusion constant agrees with previous experiments. This type of micromachine could provide a generic pump or separation component for DNA or other charged species as part of a microscale lab-on-a-chip. A device with reduced feature size could produce a size-based separation of DNA molecules, with applications including the detection of single nucleotide polymorphisms.

physics.bio-ph↗

Evolutionary implications of a power-law distribution of protein family sizes

Current-day genomes bear the mark of the evolutionary processes. One of the strongest indications is the sequence homology among families of proteins that perform similar biological functions in different species. The number of proteins in a family can grow over time as genetic information is duplicated through evolution. We explore how evolution directs the size distribution of these families. Theoretical predictions for family sizes are obtained from two models, one in which individual genes duplicate and a second in which the entire genome duplicates. Predictions from these models are compared with the family size distributions for several organisms whose complete genome sequence is known. We find that protein family size distributions in nature follow a power-law distribution. Comparing these results to the model systems, we conclude that genome duplication is the dominant mechanism leading to increased genetic material in the species considered.

physics.bio-ph↗