SearcharxivSearch

arXiv subjects

Partha P. Mitra

Publications and source records attributed to Partha P. Mitra.

At least 19 recordsLinked to original sources

Skeletonization of neuronal processes using Discrete Morse techniques from computational topology

To understand biological intelligence we need to map neuronal networks in vertebrate brains. Mapping mesoscale neural circuitry is done using injections of tracers that label groups of neurons whose axons project to different brain regions. Since many neurons are labeled, it is difficult to follow individual axons. Previous approaches have instead quantified the regional projections using the total label intensity within a region. However, such a quantification is not biologically meaningful. We propose a new approach better connected to the underlying neurons by skeletonizing labeled axon fragments and then estimating a volumetric length density. Our approach uses a combination of deep nets and the Discrete Morse (DM) technique from computational topology. This technique takes into account nonlocal connectivity information and therefore provides noise-robustness. We demonstrate the utility and scalability of the approach on whole-brain tracer injected data. We also define and illustrate an information theoretic measure that quantifies the additional information obtained, compared to the skeletonized tracer injection fragments, when individual axon morphologies are available. Our approach is the first application of the DM technique to computational neuroanatomy. It can help bridge between single-axon skeletons and tracer injections, two important data types in mapping neural networks in vertebrates.

q-bio.NC

Detection and skeletonization of single neurons and tracer injections using topological methods

Neuroscientific data analysis has traditionally relied on linear algebra and stochastic process theory. However, the tree-like shapes of neurons cannot be described easily as points in a vector space (the subtraction of two neuronal shapes is not a meaningful operation), and methods from computational topology are better suited to their analysis. Here we introduce methods from Discrete Morse (DM) Theory to extract the tree-skeletons of individual neurons from volumetric brain image data, and to summarize collections of neurons labelled by tracer injections. Since individual neurons are topologically trees, it is sensible to summarize the collection of neurons using a consensus tree-shape that provides a richer information summary than the traditional regional 'connectivity matrix' approach. The conceptually elegant DM approach lacks hand-tuned parameters and captures global properties of the data as opposed to previous approaches which are inherently local. For individual skeletonization of sparsely labelled neurons we obtain substantial performance gains over state-of-the-art non-topological methods (over 10% improvements in precision and faster proofreading). The consensus-tree summary of tracer injections incorporates the regional connectivity matrix information, but in addition captures the collective collateral branching patterns of the set of neurons connected to the injection site, and provides a bridge between single-neuron morphology and tracer-injection data.

cs.CV

SSFN -- Self Size-estimating Feed-forward Network with Low Complexity, Limited Need for Human Intervention, and Consistent Behaviour across Trials

We design a self size-estimating feed-forward network (SSFN) using a joint optimization approach for estimation of number of layers, number of nodes and learning of weight matrices. The learning algorithm has a low computational complexity, preferably within few minutes using a laptop. In addition the algorithm has a limited need for human intervention to tune parameters. SSFN grows from a small-size network to a large-size network, guaranteeing a monotonically non-increasing cost with addition of nodes and layers. The learning approach uses judicious a combination of `lossless flow property' of some activation functions, convex optimization and instance of random matrix. Consistent performance -- low variation across Monte-Carlo trials -- is found for inference performance (classification accuracy) and estimation of network size.

cs.LG

Critical Behavior and Universality Classes for an Algorithmic Phase Transition in Sparse Reconstruction

Recovery of an $N$-dimensional, $K$-sparse solution $\mathbf{x}$ from an $M$-dimensional vector of measurements $\mathbf{y}$ for multivariate linear regression can be accomplished by minimizing a suitably penalized least-mean-square cost $||\mathbf{y}-\mathbf{H} \mathbf{x}||_2^2+λV(\mathbf{x})$. Here $\mathbf{H}$ is a known matrix and $V(\mathbf{x})$ is an algorithm-dependent sparsity-inducing penalty. For `random' $\mathbf{H}$, in the limit $λ\rightarrow 0$ and $M,N,K\rightarrow \infty$, keeping $ρ=K/N$ and $α=M/N$ fixed, exact recovery is possible for $α$ past a critical value $α_c = α(ρ)$. Assuming $\mathbf{x}$ has iid entries, the critical curve exhibits some universality, in that its shape does not depend on the distribution of $\mathbf{x}$. However, the algorithmic phase transition occurring at $α=α_c$ and associated universality classes remain ill-understood from a statistical physics perspective, i.e. in terms of scaling exponents near the critical curve. In this article, we analyze the mean-field equations for two algorithms, Basis Pursuit ($V(\mathbf{x})=||\mathbf{x}||_{1} $) and Elastic Net ($V(\mathbf{x})= ||\mathbf{x}||_{1} + \tfrac{g}{2} ||\mathbf{x}||_{2}^2$) and show that they belong to different universality classes in the sense of scaling exponents, with Mean Squared Error (MSE) of the recovered vector scaling as $λ^\frac{4}{3}$ and $λ$ respectively, for small $λ$ on the critical line. In the presence of additive noise, we find that, when $α>α_c$, MSE is minimized at a non-zero value for $λ$, whereas at $α=α_c$, MSE always increases with $λ$.

cs.IT

Multimodal Cross-registration and Quantification of Metric Distortions in Whole Brain Histology of Marmoset using Diffeomorphic Mappings

Whole brain neuroanatomy using tera-voxel light-microscopic data sets is of much current interest. A fundamental problem in this field is the mapping of individual brain data sets to a reference space. Previous work has not rigorously quantified the distortions in brain geometry from in-vivo to ex-vivo brains due to the tissue processing, which will be important when computing properties such as local cell and process densities at the voxel level in creating reference brain maps. Further, existing approaches focus on registering uni-modal volumetric data; however, given the increasing interest in the marmoset model for neuroscience research, it is necessary to cross-register multi-modal data sets including MRIs and multiple histological series that can help address individual variations in brain architecture. Here we present a computational approach for same-subject multimodal MRI guided reconstruction of a histological series, jointly with diffeomorphic mapping to a reference atlas. We quantify the scale change during the different stages of histological processing of the brains using the Jacobian determinant of the diffeomorphic transformations involved. There are two major steps in the histology process with associated scale distortions (a) brain perfusion (b) histological sectioning and reassembly. By mapping the final image stacks to the ex-vivo post fixation MRI, we show that tape-transfer histology can be reassembled accurately into 3D volumes with a local scale change of 2.0 $\pm$ 0.4% per axis dimension. In contrast, the perfusion step, as assessed by mapping the in-vivo MRIs to the ex-vivo post fixation MRIs, shows a larger local scale change of 6.9 $\pm$ 2.1% per axis dimension. This is the first systematic quantification of the local metric distortions associated with whole-brain histological processing, and we expect that the results will generalize to other species.

q-bio.NC

Locally Convex Sparse Learning over Networks

We consider a distributed learning setup where a sparse signal is estimated over a network. Our main interest is to save communication resource for information exchange over the network and reduce processing time. Each node of the network uses a convex optimization based algorithm that provides a locally optimum solution for that node. The nodes exchange their signal estimates over the network in order to refine their local estimates. At a node, the optimization algorithm is based on an $\ell_1$-norm minimization with appropriate modifications to promote sparsity as well as to include influence of estimates from neighboring nodes. Our expectation is that local estimates in each node improve fast and converge, resulting in a limited demand for communication of estimates between nodes and reducing the processing time. We provide restricted-isometry-property (RIP)-based theoretical analysis on estimation quality. In the scenario of clean observation, it is shown that the local estimates converge to the exact sparse signal under certain technical conditions. Simulation results show that the proposed algorithms show competitive performance compared to a globally optimum distributed LASSO algorithm in the sense of convergence speed and estimation error.

stat.ML

On variational solutions for whole brain serial-section histology using the computational anatomy random orbit model

This paper presents a variational framework for dense diffeomorphic atlas-mapping onto high-throughput histology stacks at the 20 um meso-scale. The observed sections are modelled as Gaussian random fields conditioned on a sequence of unknown section by section rigid motions and unknown diffeomorphic transformation of a three-dimensional atlas. To regularize over the high-dimensionality of our parameter space (which is a product space of the rigid motion dimensions and the diffeomorphism dimensions), the histology stacks are modelled as arising from a first order Sobolev space smoothness prior. We show that the joint maximum a-posteriori, penalized-likelihood estimator of our high dimensional parameter space emerges as a joint optimization interleaving rigid motion estimation for histology restacking and large deformation diffeomorphic metric mapping to atlas coordinates. We show that joint optimization in this parameter space solves the classical curvature non-identifiability of the histology stacking problem. The algorithms are demonstrated on a collection of whole-brain histological image stacks from the Mouse Brain Architecture Project.

eess.IV

Progressive Learning for Systematic Design of Large Neural Networks

We develop an algorithm for systematic design of a large artificial neural network using a progression property. We find that some non-linear functions, such as the rectifier linear unit and its derivatives, hold the property. The systematic design addresses the choice of network size and regularization of parameters. The number of nodes and layers in network increases in progression with the objective of consistently reducing an appropriate cost. Each layer is optimized at a time, where appropriate parameters are learned using convex optimization. Regularization parameters for convex optimization do not need a significant manual effort for tuning. We also use random instances for some weight matrices, and that helps to reduce the number of parameters we learn. The developed network is expected to show good generalization power due to appropriate regularization and use of random weights in the layers. This expectation is verified by extensive experiments for classification and regression problems, using standard databases.

cs.NE

Estimate Exchange over Network is Good for Distributed Hard Thresholding Pursuit

We investigate an existing distributed algorithm for learning sparse signals or data over networks. The algorithm is iterative and exchanges intermediate estimates of a sparse signal over a network. This learning strategy using exchange of intermediate estimates over the network requires a limited communication overhead for information transmission. Our objective in this article is to show that the strategy is good for learning in spite of limited communication. In pursuit of this objective, we first provide a restricted isometry property (RIP)-based theoretical analysis on convergence of the iterative algorithm. Then, using simulations, we show that the algorithm provides competitive performance in learning sparse signals vis-a-vis an existing alternate distributed algorithm. The alternate distributed algorithm exchanges more information including observations and system parameters.

stat.ML

Brain Gene Expression Analysis: a MATLAB toolbox for the analysis of brain-wide gene-expression data

The Allen Brain Atlas project (ABA) generated a genome-scale collection of gene-expression profiles using in-situ hybridization. These profiles were co-registered to the three-dimensional Allen Reference Atlas (ARA) of the adult mouse brain. A set of more than 4,000 such volumetric data are available for the full brain, at a resolution of 200 microns. These data are presented in a voxel-by-gene matrix. The ARA comes with several systems of annotation, hierarchical (40 cortical regions, 209 sub-cortical regions in the whole brain), or non-hierarchical (12 regions in the left hemisphere, with refinement into 94 regions, and cortical layers). The high-dimensional nature of this dataset and the possible connection between anatomy and gene expression pose challenges to data analysis. We developed the Brain Gene Expression Analysis Toolbox, whose functionalities include: determination of marker genes for brain regions, statistical analysis of brain-wide co-expression patterns, and the computation of brain-wide correlation maps with cell-type specific microarray data.

q-bio.NC

Phase transitions in distributed control systems with multiplicative noise

Contemporary technological challenges often involve many degrees of freedom in a distributed or networked setting. Three aspects are notable: the variables are usually associated with the nodes of a graph with limited communication resources, hindering centralized control; the communication is subjected to noise; and the number of variables can be very large. These three aspects make tools and techniques from statistical physics particularly suitable for the performance analysis of such networked systems in the limit of many variables (analogous to the thermodynamic limit in statistical physics). Perhaps not surprisingly, phase-transition like phenomena appear in these systems, where a sharp change in performance can be observed with a smooth parameter variation, with the change becoming discontinuous or singular in the limit of infinite system size. In this paper we analyze the so called network consensus problem, prototypical of the above considerations, that has been previously analyzed mostly in the context of additive noise. We show that qualitatively new phase-transition like phenomena appear for this problem in the presence of multiplicative noise. Depending on dimensions and on the presence or absence of a conservation law, the system performance shows a discontinuous change at a threshold value of the multiplicative noise strength. In the absence of the conservation law, and for graph spectral dimension less than two, the multiplicative noise threshold (the stability margin of the control problem) is zero. This is reminiscent of the absence of robust controllers for certain classes of centralized control problems. Although our study involves a toy model we believe that the qualitative features are generic, with implication for the robust stability of distributed control systems, as well as the effect of roundoff errors and communication noise on distributed algorithms.

cond-mat.stat-mech

The cavity method for analysis of large-scale penalized regression

Penalized regression methods aim to retrieve reliable predictors among a large set of putative ones from a limited amount of measurements. In particular, penalized regression with singular penalty functions is important for sparse reconstruction algorithms. For large-scale problems, these algorithms exhibit sharp phase transition boundaries where sparse retrieval breaks down. Large optimization problems associated with sparse reconstruction have been analyzed in the literature by setting up corresponding statistical mechanical models at a finite temperature. Using replica method for mean field approximation, and subsequently taking a zero temperature limit, this approach reproduces the algorithmic phase transition boundaries. Unfortunately, the replica trick and the non-trivial zero temperature limit obscure the underlying reasons for the failure of a sparse reconstruction algorithm, and of penalized regression methods, in general. In this paper, we employ the ``cavity method'' to give an alternative derivation of the mean field equations, working directly in the zero-temperature limit. This derivation provides insight into the origin of the different terms in the self-consistency conditions. The cavity method naturally involves a quantity, the average local susceptibility, whose behavior distinguishes different phases in this system. This susceptibility can be generalized for analysis of a broader class of sparse reconstruction algorithms.

cs.IT

Cell-type-specific transcriptomes and the Allen Atlas (II): discussion of the linear model of brain-wide densities of cell types

The voxelized Allen Atlas of the adult mouse brain (at a resolution of 200 microns) has been used in [arXiv:1303.0013] to estimate the region-specificity of 64 cell types whose transcriptional profile in the mouse brain has been measured in microarray experiments. In particular, the model yields estimates for the brain-wide density of each of these cell types. We conduct numerical experiments to estimate the errors in the estimated density profiles. First of all, we check that a simulated thalamic profile based on 200 well-chosen genes can transfer signal from cerebellar Purkinje cells to the thalamus. This inspires us to sub-sample the atlas of genes by repeatedly drawing random sets of 200 genes and refitting the model. This results in a random distribution of density profiles, that can be compared to the predictions of the model. This results in a ranking of cell types by the overlap between the original and sub-sampled density profiles. Cell types with high rank include medium spiny neurons, several samples of cortical pyramidal neurons, hippocampal pyramidal neurons, granule cells and cholinergic neurons from the brain stem. In some cases with lower rank, the average sub-sample can have better contrast properties than the original model (this is the case for amygdalar neurons and dopaminergic neurons from the ventral midbrain). Finally, we add some noise to the cell-type-specific transcriptomes by mixing them using a scalar parameter weighing a random matrix. After refitting the model, we observe than a mixing parameter of $5\%$ leads to modifications of density profiles that span the same interval as the ones resulting from sub-sampling.

q-bio.NC

Cell-type-specific microarray data and the Allen atlas: quantitative analysis of brain-wide patterns of correlation and density

The Allen Atlas of the adult mouse brain is used to estimate the region-specificity of 64 cell types whose transcriptional profile in the mouse brain has been measured in microarray experiments. We systematically analyze the preliminary results presented in [arXiv:1111.6217], using the techniques implemented in the Brain Gene Expression Analysis toolbox. In particular, for each cell-type-specific sample in the study, we compute a brain-wide correlation profile to the Allen Atlas, and estimate a brain-wide density profile by solving a quadratic optimization problem at each voxel in the mouse brain. We characterize the neuroanatomical properties of the correlation and density profiles by ranking the regions of the left hemisphere delineated in the Allen Reference Atlas. We compare these rankings to prior biological knowledge of the brain region from which the cell-type-specific sample was extracted.

q-bio.NC

Computational neuroanatomy and co-expression of genes in the adult mouse brain, analysis tools for the Allen Brain Atlas

We review quantitative methods and software developed to analyze genome-scale, brain-wide spatially-mapped gene-expression data. We expose new methods based on the underlying high-dimensional geometry of voxel space and gene space, and on simulations of the distribution of co-expression networks of a given size. We apply them to the Allen Atlas of the adult mouse brain, and to the co-expression network of a set of genes related to nicotine addiction retrieved from the NicSNP database. The computational methods are implemented in {\ttfamily{BrainGeneExpressionAnalysis}}, a Matlab toolbox available for download.

q-bio.QM

What does the Allen Gene Expression Atlas tell us about mouse brain evolution?

We use the Allen Gene Expression Atlas (AGEA) and the OMA ortholog dataset to investigate the evolution of mouse-brain neuroanatomy from the standpoint of the molecular evolution of brain-specific genes. For each such gene, using the phylogenetic tree for all fully sequenced species and the presence of orthologs of the gene in these species, we construct and assign a discrete measure of evolutionary age. The gene expression profile of all gene of similar age, relative to the average gene expression profile, distinguish regions of the brain that are over-represented in the corresponding evolutionary timescale. We argue that the conclusions one can draw on evolution of twelve major brain regions from such a molecular level analysis supplements existing knowledge of mouse brain evolution and introduces new quantitative tools, especially for comparative studies, when AGEA-like data sets for other species become available. Using the functional role of the genes representational of a certain evolutionary timescale and brain region we compare and contrast, wherever possible, our observations with existing knowledge in evolutionary neuroanatomy.

q-bio.QM

Computational neuroanatomy and gene expression: optimal sets of marker genes for brain regions

The three-dimensional data-driven Allen Gene Expression Atlas of the adult mouse brain consists of numerized in-situ hybridization data for thousands of genes, co-registered to the Allen Reference Atlas. We propose quantitative criteria to rank genes as markers of a brain region, based on the localization of the gene expression and on its functional fitting to the shape of the region. These criteria lead to natural generalizations to sets of genes. We find sets of genes weighted with coefficients of both signs with almost perfect localization in all major regions of the left hemisphere of the brain, except the pallidum. Generalization of the fitting criterion with positivity constraint provides a lesser improvement of the markers, but requires sparser sets of genes.

q-bio.QM

A cell-type based model explaining co-expression patterns of genes in the brain

Much of the genome is expressed in the vertebrate brain, with individual genes exhibiting different spatially-varying patterns of expression. These variations are not independent, with pairs of genes exhibiting complex patterns of co-expression, such that two genes may be similarly expressed in one region, but differentially expressed in other regions. These correlations have been previously studied quantitatively, particularly for the gene expression atlas of the mouse brain, but the biological meaning of the co-expression patterns remains obscure. We propose a simple model of the co-expression patterns in terms of spatial distributions of underlying cell types. We establish the plausibility of the model in terms of a test set of cell types for which both the gene expression profiles and the spatial distributions are known.

q-bio.QM