SearcharxivSearch

arXiv subjects

Y-h. Taguchi

Publications and source records attributed to Y-h. Taguchi.

15 recordsLinked to original sources

Unsupervised feature selection using Bayesian Tucker decomposition

In this paper, we proposed Bayesian Tucker decomposition (BTuD) in which residual is supposed to obey Gaussian distribution analogous to linear regression. Although we have proposed an algorithm to perform the proposed BTuD, the conventional higher-order orthogonal iteration can generate Tucker decomposition consistent with the present implementation. Using the proposed BTuD, we can perform unsupervised feature selection successfully applied to various synthetic datasets, global coupled maps with randomized coupling strength, and gene expression profiles. Thus we can conclude that our newly proposed unsupervised feature selection method is promising. In addition to this, BTuD based unsupervised FE is expected to coincide with TD based unsupervised FE that were previously proposed and successfully applied to a wide range of problems.

stat.ML

Dynamics-Based Intrinsic Signal Model for High-Dimensional, Small-Sample Data

Signal extraction is difficult when the number of variables $N$ is much larger than the number of observations $M$. We address this problem under the working hypothesis that an empirical dataset consists of states sampled from underlying multivariate dynamics. Instead of treating the $M$ observations as points in an $N$-dimensional variable space, we treat the $N$ variables as points in an $M$-dimensional sample-coordinate space, interpreted as an effective time-delay coordinate space of the latent dynamics. This representation allows a large $N$ to provide many points for estimating the variable distribution even when $M$ is small. Transposed singular value decomposition (SVD) and the unsupervised feature-selection method of Taguchi are used to extract variable-side deviations from an estimated Gaussian background as signal candidates. As $M$ is reduced, the effective separation between sampled states increases; contributions from finite-correlation components are expected to decay, whereas sufficiently long-correlation components can persist toward the small-sample limit. We define these persistent components as intrinsic signals and estimate them by extrapolation toward $M=0$. We first tested the method on high-dimensional, small-sample data explicitly generated by a randomised coupling strength globally coupled map (RCS-GCM), for which the long- and short-correlation components were known. The extracted intrinsic signals corresponded to the known long-correlation variables. We then applied the method to The Cancer Genome Atlas (TCGA) pan-kidney gene-expression data, which are not ordinarily treated as data explicitly generated by a dynamical system. Using an SVD component associated with the pathologic-M category, variable-side signal components were extracted from the 20,531-dimensional data even under small-sample conditions.

physics.data-an

Comparison of amino acid occurrence and composition for predicting protein folds

Background:Prediction of protein three-dimensional structures from amino acid sequences is a long-standing goal in computational/molecular biology. The successful discrimination of protein folds would help to improve the accuracy of protein 3D structure prediction. Results: In this work, we propose a method based on linear discriminant analysis (LDA) for recognizing proteins belonging to 30 different folds using the occurrence of amino acid residues in a set of 1612 proteins. The present method could discriminate the globular proteins from 30 major folding types with the sensitivity of 37%, which is comparable to or better than other methods in the literature. A web server has been developed for predicting the folding type of the protein from amino acid sequence and it is available at http://granular.com/PROLDA/. Conclusions:Linear discriminant analysis based on amino acid occurrence could successfully recognize protein folds. The present method has several advantages such as, (i) it directly predicts the folding type of a protein without performing pair-wise comparisons, (ii) it can discriminate folds among large number of proteins and (iii) it is very fast to obtain the results. This is a simple method, which can be easily incorporated in any other structure prediction algorithms.

q-bio.BM

Temperature measurement in the convective and segregated vibrated bed of powder : A numerical study

In numerically simulated vibrated beds of powder, we measure temperature under convection by the generalized Einstein's relation. The spatial temperature distribution turns out to be quite uniform except for the boundary layers. In addition to this, temperature remains uniform even if segregation occurs. This suggests the possibility that there exists some "thermal equilibrium state" even in a vibrated bed of powder. This finding may lead to a unified view of the dynamic steady state of granular matter.

nlin.PS

Can Neural Networks Recognize Parts?

We have demonstrated neural networks can recognize parts by visual images. Input signals are gray scale photographs of objects consisting of some parts and output signals are their shapes. By training neural networks by a few set of images, without any supervision they become to be able to recognize the boundary between parts.

q-bio.NC

Temporal patterns of gene expression via nonmetric multidimensional scaling analysis

Motivation: Microarray experiments result in large scale data sets that require extensive mining and refining to extract useful information. We have been developing an efficient novel algorithm for nonmetric multidimensional scaling (nMDS) analysis for very large data sets as a maximally unsupervised data mining device. We wish to demonstrate its usefulness in the context of bioinformatics. In our motivation is also an aim to demonstrate that intrinsically nonlinear methods are generally advantageous in data mining. Results: The Pearson correlation distance measure is used to indicate the dissimilarity of the gene activities in transcriptional response of cell cycle-synchronized human fibroblasts to serum [Iyer et al., Science vol. 283, p83 (1999)]. These dissimilarity data have been analyzed with our nMDS algorithm to produce an almost circular arrangement of the genes. The temporal expression patterns of the genes rotate along this circular arrangement. If an appropriate preparation procedure may be applied to the original data set, linear methods such as the principal component analysis (PCA) could achieve reasonable results, but without data preprocessing linear methods such as PCA cannot achieve a useful picture. Furthermore, even with an appropriate data preprocessing, the outcomes of linear procedures are not as clearcut as those by nMDS without preprocessing.

nlin.PS

A Toy Model of Flying Snake's Glide

We have developed a toy model of flying snake's glide [J.J. Socha, Nature vol. 418 (2002) 603.] by modifying a model for a falling paper. We have found that asymmetric oscillation is a key about why snake can glide. Further investigation for snake's glide will provide us details about how it can glide without a wing.

nlin.PS

Inelastic clump collision model for non-Gaussian velocity distribution in molecular clouds

Non-Gaussian velocity distribution in star forming region is reproduced by inelastic clump collision model. We numerically calculated the evolution of inelastic hard spheres in sheared flow, which corresponds to cloud clumps in differential galactic rotation. This system fluctuates largely around equilibrium state, creating clusters with inelastic collisions and destroying them with shear motion. The fluctuation makes spheres have non-Gaussian velocity distribution with nearly exponential tail. How far from Gaussian distribution depends upon coefficient of restitution, which can produce the variety of degree of deviation from Gaussian among regions.

astro-ph

Power law velocity fluctuations due to inelastic collisions in numerically simulated vibrated bed of powder}

Distribution functions of relative velocities among particles in a vibrated bed of powder are studied both numerically and theoretically. In the solid phase where granular particles remain near their local stable states, the probability distribution is Gaussian. On the other hand, in the fluidized phase, where the particles can exchange their positions, the distribution clearly deviates from Gaussian. This is interpreted with two analogies; aggregation processes and soft-to-hard turbulence transition in thermal convection. The non-Gaussian distribution is well-approximated by the t-distribution which is derived theoretically by considering the effect of clustering by inelastic collisions in the former analogy.

chao-dyn

A set of hard spheres with tangential inelastic collision as a model of granular matter: $1/f^α$ fluctuation, non-Gaussian distribution, and convective motion

A set of hard spheres with tangential inelastic collision is found to reproduce observations of real and numerical granular matter. After time is scaled so as to cancel energy dissipation due to inelastic collisions out, inelastically colliding hard spheres in two dimensional space come to have $1/f^α$ fluctuation of total energy, non-Gaussian distribution of displacement vectors, and convective motion of spheres, which hard spheres with elastic collision, a conventional model of granular matter, cannot reproduce.

adap-org

Numerical Study of Granular Turbulence and the appearance of $k^{-5/3}$ energy spectrum without flow

The vibrated bed of powder, a vessel that mono disperse glass beads fill and a loud speaker shakes, is investigated numerically with the distinct element method, a kind of molecular dynamics. When the bed is heavily shaken, the displacement vectors of powder have the power spectrum with the dependence upon the wave number $k$ as $k^{-5/3}$. The origin of this spectrum is suggested to be the balance between the injected and dissipative energy, analogous to the proposal by Kolmogorov to explain the $k^{-5/3}$ energy spectrum observed in the fluid turbulence. Furthermore, the same spectrum still appears even without flows of powder. Thus Kolmogorov's argument is more universal than believed before.

chao-dyn

Non-Gaussian distribution in Random advection dynamics

Simulations of vortex tube dynamics reveal that the non-Gaussian nature of turbulent fluctuation originates in the effect of random advection. A similar non-Gaussian distribution is found numerically in a simplified statistical model of random advection. An analytical solution is obtained in the mean-field case.

cond-mat

Homogeneous Isotropic Fluid Turbulence Simulated with the Lattice Vortex Tube Model

Fully developed turbulence is analised with the lattice model employing vortex tube representation which is introduced recently by the authors. Several characteric features observed in experiments and direct numeric integrations are reproduced. Not only Kolmogorov's inertial range is observed, but also several local probability distribution functions are obtained as well. Those of the local velocities are close to the Gaussian and exponential-like distributions appear in local vorticity, relative velocities and local velocity consisting of only higher wave number components. Coherent structure of vortex tubes is seen, too. Moreover required cpu-time and memory-size are very little comparing with the conventional pseudo-spectral method. keywords : fluid turbulence, vortex tube, lattice model, numerical technic.

cond-mat