SearcharxivSearch

arXiv subjects

Doreswamy

Publications and source records attributed to Doreswamy.

13 recordsLinked to original sources

Predictive Comparative QSAR analysis of Sulfathiazole Analogues as Mycobacterium Tuberculosis H37RV Inhabitors

Antitubercular activity of Sulfathiazole Derivitives series were subjected to Quantitative Structure Activity Relationship (QSAR) Analysis with an attempt to derive and understand a correlation between the Biologically Activity as dependent variable and various descriptors as independent variables. QSAR models generated using 28 compounds. Several statistical regression expressions were obtained using Partial Least Squares (PLS) Regression, Multiple Linear Regression (MLR) and Principal Component Regression (PCR) methods. The among these methods, Partial Least Square Regression (PLS) method has shown very promising result as compare to other two methods. A QSAR model was generated by a training set of 18 molecules with correlation coefficient r (r square) of 0.9191, significant cross validated correlation coefficient (q square) of 0.8300, F test of 53.5783, r square for external test set pred_r square -3.6132, coefficient of correlation of predicted data set pred_r_se square 1.4859 and degree of freedom 14 by Partial Least Squares Regression Method.

cs.CE

Non linear Prediction of Antitubercular Activity Of Oxazolines and Oxazoles derivatives Making Use of Compact TS-Fuzzy models Through Clustering with orthogonal least sqaure technique and Fuzzy identification system

The prediction of uncertain and predictive nonlinear systems is an important and challenging problem. Fuzzy logic models are often a good choice to describe such systems however in many cases these become complex soon. commonlly, too less effort is put into descriptor selection and in the creation of suitable local rules. Moreover, in common no model reduction is applied, while this may analyze the model by removing redundant data. This paper suggests a combined method that deal with these issues in order to create compact Takagi Sugeno (TS) models that can be effectively used to represent complex predictive systems. A new fuzzy clustering method is come up with for the identification of compact TS-fuzzy models. The best relevant consequent variables of the TS model are choosen by an orthogonal least squares technique based on the obtained clusters.For the selection of the relevant antecedent (scheduling) variables a new method has been developed based on Fisher's interclass separability basis. This complete approach is demonstrated by means of the Oxazolines and Oxazoles derivatives as antituberculosis agent for nonlinear regression benchmark. The results are compared with results obtained by neuro-fuzzy i.e. ANFIS algorithm and advanced fuzzyy clustering techniques i.e FMID toolbox .

cs.CE

Important Molecular Descriptors Selection Using Self Tuned Reweighted Sampling Method for Prediction of Antituberculosis Activity

In this paper, a new descriptor selection method for selecting an optimal combination of important descriptors of sulfonamide derivatives data, named self tuned reweighted sampling (STRS), is developed. descriptors are defined as the descriptors with large absolute coefficients in a multivariate linear regression model such as partial least squares(PLS). In this study, the absolute values of regression coefficients of PLS model are used as an index for evaluating the importance of each descriptor Then, based on the importance level of each descriptor, STRS sequentially selects N subsets of descriptors from N Monte Carlo (MC) sampling runs in an iterative and competitive manner. In each sampling run, a fixed ratio (e.g. 80%) of samples is first randomly selected to establish a regresson model. Next, based on the regression coefficients, a two-step procedure including rapidly decreasing function (RDF) based enforced descriptor selection and self tuned sampling (STS) based competitive descriptor selection is adopted to select the important descriptorss. After running the loops, a number of subsets of descriptors are obtained and root mean squared error of cross validation (RMSECV) of PLS models established with subsets of descriptors is computed. The subset of descriptors with the lowest RMSECV is considered as the optimal descriptor subset. The performance of the proposed algorithm is evaluated by sulfanomide derivative dataset. The results reveal an good characteristic of STRS that it can usually locate an optimal combination of some important descriptors which are interpretable to the biologically of interest. Additionally, our study shows that better prediction is obtained by STRS when compared to full descriptor set PLS modeling, Monte Carlo uninformative variable elimination (MC-UVE).

cs.LG

Predictive Comparative QSAR Analysis Of As 5-Nitofuran-2-YL Derivatives Myco bacterium tuberculosis H37RV Inhibitors Bacterium Tuberculosis H37RV Inhibitors

Antitubercular activity of 5-nitrofuran-2-yl Derivatives series were subjected to Quantitative Structure Activity Relationship (QSAR) Analysis with an effort to derive and understand a correlation between the biological activity as response variable and different molecular descriptors as independent variables. QSAR models are built using 40 molecular descriptor dataset. Different statistical regression expressions were got using Partial Least Squares (PLS),Multiple Linear Regression (MLR) and Principal Component Regression (PCR) techniques. The among these technique, Partial Least Square Regression (PLS) technique has shown very promising result as compared to MLR technique A QSAR model was build by a training set of 30 molecules with correlation coefficient ($r^2$) of 0.8484, significant cross validated correlation coefficient ($q^2$) is 0.0939, F test is 48.5187, ($r^2$) for external test set (pred$_r^2$) is -0.5604, coefficient of correlation of predicted data set (pred$_r^2se$) is 0.7252 and degree of freedom is 26 by Partial Least Squares Regression technique.

cs.CE

A Robust Missing Value Imputation Method MifImpute For Incomplete Molecular Descriptor Data And Comparative Analysis With Other Missing Value Imputation Methods

Missing data imputation is an important research topic in data mining. Large-scale Molecular descriptor data may contains missing values (MVs). However, some methods for downstream analyses, including some prediction tools, require a complete descriptor data matrix. We propose and evaluate an iterative imputation method MiFoImpute based on a random forest. By averaging over many unpruned regression trees, random forest intrinsically constitutes a multiple imputation scheme. Using the NRMSE and NMAE estimates of random forest, we are able to estimate the imputation error. Evaluation is performed on two molecular descriptor datasets generated from a diverse selection of pharmaceutical fields with artificially introduced missing values ranging from 10% to 30%. The experimental result demonstrates that missing values has a great impact on the effectiveness of imputation techniques and our method MiFoImpute is more robust to missing value than the other ten imputation methods used as benchmark. Additionally, MiFoImpute exhibits attractive computational efficiency and can cope with high-dimensional data.

cs.CE

Performance Analysis Of Regularized Linear Regression Models For Oxazolines And Oxazoles Derivitive Descriptor Dataset

Regularized regression techniques for linear regression have been created the last few ten years to reduce the flaws of ordinary least squares regression with regard to prediction accuracy. In this paper, new methods for using regularized regression in model choice are introduced, and we distinguish the conditions in which regularized regression develops our ability to discriminate models. We applied all the five methods that use penalty-based (regularization) shrinkage to handle Oxazolines and Oxazoles derivatives descriptor dataset with far more predictors than observations. The lasso, ridge, elasticnet, lars and relaxed lasso further possess the desirable property that they simultaneously select relevant predictive descriptors and optimally estimate their effects. Here, we comparatively evaluate the performance of five regularized linear regression methods The assessment of the performance of each model by means of benchmark experiments is an established exercise. Cross-validation and resampling methods are generally used to arrive point evaluates the efficiencies which are compared to recognize methods with acceptable features. Predictive accuracy was evaluated using the root mean squared error (RMSE) and Square of usual correlation between predictors and observed mean inhibitory concentration of antitubercular activity (R square). We found that all five regularized regression models were able to produce feasible models and efficient capturing the linearity in the data. The elastic net and lars had similar accuracies as well as lasso and relaxed lasso had similar accuracies but outperformed ridge regression in terms of the RMSE and R square metrics.

cs.LG

Identification Of Outliers In Oxazolines AND Oxazoles High Dimension Molecular Descriptor Dataset Using Principal Component Outlier Detection Algorithm And Comparative Numerical Study Of Other Robust Estimators

From the past decade outlier detection has been in use. Detection of outliers is an emerging topic and is having robust applications in medical sciences and pharmaceutical sciences. Outlier detection is used to detect anomalous behaviour of data. Typical problems in Bioinformatics can be addressed by outlier detection. A computationally fast method for detecting outliers is shown, that is particularly effective in high dimensions. PrCmpOut algorithm make use of simple properties of principal components to detect outliers in the transformed space, leading to significant computational advantages for high dimensional data. This procedure requires considerably less computational time than existing methods for outlier detection. The properties of this estimator (Outlier error rate (FN), Non-Outlier error rate(FP) and computational costs) are analyzed and compared with those of other robust estimators described in the literature through simulation studies. Numerical evidence based Oxazolines and Oxazoles molecular descriptor dataset shows that the proposed method performs well in a variety of situations of practical interest. It is thus a valuable companion to the existing outlier detection methods.

cs.CE

Performance Analysis Of Neural Network Models For Oxazolines And Oxazoles Derivatives Descriptor Dataset

Neural networks have been used successfully to a broad range of areas such as business, data mining, drug discovery and biology. In medicine, neural networks have been applied widely in medical diagnosis, detection and evaluation of new drugs and treatment cost estimation. In addition, neural networks have begin practice in data mining strategies for the aim of prediction, knowledge discovery. This paper will present the application of neural networks for the prediction and analysis of antitubercular activity of Oxazolines and Oxazoles derivatives. This study presents techniques based on the development of Single hidden layer neural network (SHLFFNN), Gradient Descent Back propagation neural network (GDBPNN), Gradient Descent Back propagation with momentum neural network (GDBPMNN), Back propagation with Weight decay neural network (BPWDNN) and Quantile regression neural network (QRNN) of artificial neural network (ANN) models Here, we comparatively evaluate the performance of five neural network techniques. The evaluation of the efficiency of each model by ways of benchmark experiments is an accepted application. Cross-validation and resampling techniques are commonly used to derive point estimates of the performances which are compared to identify methods with good properties. Predictive accuracy was evaluated using the root mean squared error (RMSE), Coefficient determination(???), mean absolute error(MAE), mean percentage error(MPE) and relative square error(RSE). We found that all five neural network models were able to produce feasible models. QRNN model is outperforms with all statistical tests amongst other four models.

cs.CE

Study Of E-Smooth Support Vector Regression And Comparison With E- Support Vector Regression And Potential Support Vector Machines For Prediction For The Antitubercular Activity Of Oxazolines And Oxazoles Derivatives

A new smoothing method for solving ? -support vector regression (?-SVR), tolerating a small error in fitting a given data sets nonlinearly is proposed in this study. Which is a smooth unconstrained optimization reformulation of the traditional linear programming associated with a ?-insensitive support vector regression. We term this redeveloped problem as ?-smooth support vector regression (?-SSVR). The performance and predictive ability of ?-SSVR are investigated and compared with other methods such as LIBSVM (?-SVR) and P-SVM methods. In the present study, two Oxazolines and Oxazoles molecular descriptor data sets were evaluated. We demonstrate the merits of our algorithm in a series of experiments. Primary experimental results illustrate that our proposed approach improves the regression performance and the learning efficiency. In both studied cases, the predictive ability of the ?- SSVR model is comparable or superior to those obtained by LIBSVM and P-SVM. The results indicate that ?-SSVR can be used as an alternative powerful modeling method for regression studies. The experimental results show that the presented algorithm ?-SSVR, plays better precisely and effectively than LIBSVMand P-SVM in predicting antitubercular activity.

cs.CE

Knowledge Discovery System For Fiber Reinforced Polymer Matrix Composite Laminate

In this paper Knowledge Discovery System (KDS) is proposed and implemented for the extraction of knowledge-mean stiffness of a polymer composite material in which when fibers are placed at different orientations. Cosine amplitude method is implemented for retrieving compatible polymer matrix and reinforcement fiber which is coming under predicted fiber class, from the polymer and reinforcement database respectively, based on the design requirements. Fuzzy classification rules to classify fibers into short, medium and long fiber classes are derived based on the fiber length and the computed or derive critical length of fiber. Longitudinal and Transverse module of Polymer Matrix Composite consisting of seven layers with different fiber volume fractions and different fibers orientations at 0,15,30,45,60,75 and 90 degrees are analyzed through Rule-of Mixture material design model. The analysis results are represented in different graphical steps and have been measured with statistical parameters. This data mining application implemented here has focused the mechanical problems of material design and analysis. Therefore, this system is an expert decision support system for optimizing the materials performance for designing light-weight and strong, and cost effective polymer composite materials.

cs.AI

Similarity Measuring Approuch for Engineering Materials Selection

Advanced engineering materials design involves the exploration of massive multidimensional feature spaces, the correlation of materials properties and the processing parameters derived from disparate sources. The search for alternative materials or processing property strategies, whether through analytical, experimental or simulation approaches, has been a slow and arduous task, punctuated by infrequent and often expected discoveries. A few systematic efforts have been made to analyze the trends in data as a basis for classifications and predictions. This is particularly due to the lack of large amounts of organized data and more importantly the challenging of shifting through them in a timely and efficient manner. The application of recent advances in Data Mining on materials informatics is the state of art of computational and experimental approaches for materials discovery. In this paper similarity based engineering materials selection model is proposed and implemented to select engineering materials based on the composite materials constraints. The result reviewed from this model is sustainable for effective decision making in advanced engineering materials design applications.

cs.AI

A Novel Design Specification Distance(DSD) Based K-Mean Clustering Performace Evluation on Engineering Materials Database

Organizing data into semantically more meaningful is one of the fundamental modes of understanding and learning. Cluster analysis is a formal study of methods for understanding and algorithm for learning. K-mean clustering algorithm is one of the most fundamental and simple clustering algorithms. When there is no prior knowledge about the distribution of data sets, K-mean is the first choice for clustering with an initial number of clusters. In this paper a novel distance metric called Design Specification (DS) distance measure function is integrated with K-mean clustering algorithm to improve cluster accuracy. The K-means algorithm with proposed distance measure maximizes the cluster accuracy to 99.98% at P = 1.525, which is determined through the iterative procedure. The performance of Design Specification (DS) distance measure function with K - mean algorithm is compared with the performances of other standard distance functions such as Euclidian, squared Euclidean, City Block, and Chebshew similarity measures deployed with K-mean algorithm.The proposed method is evaluated on the engineering materials database. The experiments on cluster analysis and the outlier profiling show that these is an excellent improvement in the performance of the proposed method.

cs.LG

Hybrid Data Mining Technique for Knowledge Discovery from Engineering Materials' Data sets

Studying materials informatics from a data mining perspective can be beneficial for manufacturing and other industrial engineering applications. Predictive data mining technique and machine learning algorithm are combined to design a knowledge discovery system for the selection of engineering materials that meet the design specifications. Predictive method-Naive Bayesian classifier and Machine learning Algorithm - Pearson correlation coefficient method were implemented respectively for materials classification and selection. The knowledge extracted from the engineering materials data sets is proposed for effective decision making in advanced engineering materials design applications.

cs.DB