SearcharxivSearch

arXiv subjects

Shankar Narasimhan

Publications and source records attributed to Shankar Narasimhan.

13 recordsLinked to original sources

Multivariate linear regression without prior assumptions

Recovering the linear relationships that govern a system from noisy measurements is a basic task across the physical and engineering sciences. Because every measured variable may carry an unknown amount of noise, classical regression must commit in advance to a set of structural assumptions: ordinary least squares requires a declared input-output partition with input variables being noise-free, total least squares assumes equal noise variance across all variables, and generalized total least squares additionally requires the noisy-variable partition and variances to be known beforehand. Kalman~\cite{Kalman:1982} showed that any procedure returning a unique linear model from inexact data must rest on such unverifiable a priori assumptions -- ``prejudices'' -- that cannot be checked against the data itself, and that removing them leaves the identification problem fundamentally indeterminate. Whether these prejudices can instead be resolved directly from the data has remain unresolved. Here we show that an iterative generalized-eigenvalue algorithm, QZ-IPCA, recovers the noisy-variable partition, noise variances, number of linear relations, and regression coefficients of a multivariate linear system simultaneously, using only the raw data. Across all possible exhaustive noise configurations of a five-variable benchmark network, QZ-IPCA correctly identifies model structure and recovers coefficients with error below 6.4\%. It outperforms ordinary least squares even when given the best partition, and succeeds in rank identification precisely where standard total least squares falls once noise variances differ across variables. These results show that the assumptions conventionally required for multivariate regression are not necessary, recasting model identification as a problem solvable from data geometry alone.

eess.SY

Subspace Based Identification of Errors-in-Variables Linear Descriptor Systems

The identification of linear descriptor systems (DAEs) from noise-corrupted data makes two critical assumptions: requirement of an \textit{a priori} classification of variables into inputs and outputs, and a pre-specified structural assumption with respect to the index of the system. This paper proposes a data-driven methodology for identifying index-0 and index-1 DAEs within an errors-in-variables framework. We extend a subspace-based iterative PCA (SMI-IPCA) approach to the behavioral setting, treating all measured variables as a unified augmented vector to avoid classification bias. This method enables systematic estimation of the noise variances, the number of algebraic and differential output variables, while simultaneously identifying the algebraic constraints and kernel representation of the dynamic system corresponding to its minimal realization order without prior structural knowledge. Simulation studies on index-0 and index-1 systems demonstrate the effectiveness of the proposed approach and its practical applicability.

eess.SY

A recursive subspace based method for errors-in-variables model identification of time-varying systems

The Subspace-based Model Identification algorithm using a modified Iterative Principal Component Analysis (SMI-IPCA) is a theoretically rigorous method for identifying a linear state-space model of a multi-input multi-output (MIMO) process, in an errors-in-variables (EIV) setting. The method can simultaneously estimate unknown heteroskedastic noise variances corrupting the input and output measurements, along with the state space model. This work proposes a recursive formulation of SMI-IPCA (RSMI-IPCA) enabling online identification and adaptive model updates as and when new data arrive. By maintaining a fixed length lag window rather than storing the complete historical data, RSMI-IPCA estimates measurement noise variances, process order, while simultaneously identifying the state-space matrices, making it suitable to monitor time-varying systems, whether the induced changes are slow or abrupt. The algorithm gradually adapts to slow sensor degradation (time-varying noise variances), changes in process operating conditions (time-varying model parameters), and structural modifications (varying model order). Simulation studies are presented to demonstrate the efficacy and practical applicability of the proposed algorithm.

eess.SY

DISPCA : A hybrid iterative-sequential approach for the identification of errors-in-variables model of linear DAE systems

The dynamic behavior of numerous engineering processes is effectively characterized through differential-algebraic equations (DAEs), commonly referred to as descriptor systems. While substantial progress has been achieved in identifying dynamic models governed by ordinary differential equations (ODEs), limited research has addressed the identification of descriptor systems from measured data. This work presents a systematic methodology for identifying the DAE model of a linear descriptor system in discrete difference equation form under errors-in-variables (EIV) setting, where both input and output measurements are corrupted by random noise. The proposed methodology generalizes the identification framework to handle scenarios where the system contains multiple algebraic and different ordered differential relations. The key innovation involves a partial stacking procedure of lagged data matrix with a sequentially increasing lag window that identifies all the differential relations individually. This is preceded by an iterative estimation of the measurement error covariance matrix that is diagonal and heteroskedastic, under large sample conditions. The algorithm simultaneously estimates the number of differential and algebraic relations, observability indices and delay parameters of the differential equations, and all the model coefficients directly from measured data without requiring prior specification from the user. The framework addresses the increased complexity arising from multiple dynamic coupled interactions while maintaining computational tractability through systematic decomposition of the identification problem. Effectiveness of the proposed methodology is demonstrated through several simulation studies.

eess.SY

Recursive Identification of EIV-ARX Models for Time Varying SISO Processes

This paper proposes a recursive algorithm, rARX-DIPCA, for identifying errors-in-variables autoregressive models with exogenous input (EIV-ARX), for tracking time-varying SISO processes. Building on a recently developed recursive iterative PCA method, the proposed algorithm recursively updates model parameters and noise variances as new measurements arrive, without storing historical data beyond a specified lag window. The method enables real-time adaptation to sensor degradation, and changes in model coefficients. The algorithm simultaneously identifies process order, time delay, and noise variances while maintaining computational efficiency through online covariance updates. Simulation studies on benchmark systems demonstrate effective tracking performance and practical applicability.

eess.SY

Identification of Errors-in-Variables ARX Models Using Modified Dynamic Iterative PCA

Identification of autoregressive models with exogenous input (ARX) is a classical problem in system identification. This article considers the errors-in-variables (EIV) ARX model identification problem, where input measurements are also corrupted with noise. The recently proposed DIPCA technique solves the EIV identification problem but is only applicable to white measurement errors. We propose a novel identification algorithm based on a modified Dynamic Iterative Principal Components Analysis (DIPCA) approach for identifying the EIV-ARX model for single-input, single-output (SISO) systems where the output measurements are corrupted with coloured noise consistent with the ARX model. Most of the existing methods assume important parameters like input-output orders, delay, or noise-variances to be known. This work's novelty lies in the joint estimation of error variances, process order, delay, and model parameters. The central idea used to obtain all these parameters in a theoretically rigorous manner is based on transforming the lagged measurements using the appropriate error covariance matrix, which is obtained using estimated error variances and model parameters. Simulation studies on two systems are presented to demonstrate the efficacy of the proposed algorithm.

eess.SY

Identification of MISO systems in Minimal Realization Form

The paper is concerned with identifying transfer functions of individual input channels in minimal realization form of a Multi-Input Single Output (MISO) from the input-output data corrupted by the error in all the variables. Such a framework is commonly referred to as error-in-variables (EIV). A common approach in the existing methods for identification of MISO systems is to estimate a non-minimal order transfer function under a subset of simplistic assumptions like homoskedastic error variances, known order, and delay. In this work, we deal with the challenging problem of identifying order, delay in each input of minimal realization form separately while estimating the transfer functions. We also estimate the heteroskedastic noise variances in each of the multiple inputs and output variables. An automated approach for the identification of MISO systems of minimal realization form in the EIV framework is proposed. Numerical case studies are presented to illustrate the efficacy of the proposed algorithm in identifying the transfer function along with the order, delay, and noise variances.

eess.SY

ARX Model Identification using Generalized Spectral Decomposition

This article is concerned with the identification of autoregressive with exogenous inputs (ARX) models. Most of the existing approaches like prediction error minimization and state-space framework are widely accepted and utilized for the estimation of ARX models but are known to deliver unbiased and consistent parameter estimates for a correctly supplied guess of input-output orders and delay. In this paper, we propose a novel automated framework which recovers orders, delay, output noise distribution along with parameter estimates. The primary tool utilized in the proposed framework is generalized spectral decomposition. The proposed algorithm systematically estimates all the parameters in two steps. The first step utilizes estimates of the order by examining the generalized eigenvalues, and the second step estimates the parameter from the generalized eigenvectors. Simulation studies are presented to demonstrate the efficacy of the proposed method and are observed to deliver consistent estimates even at low signal to noise ratio (SNR).

eess.SY

Learning Conserved Networks from Flows

A challenging problem in complex networks is the network reconstruction problem from data. This work deals with a class of networks denoted as conserved networks, in which a flow associated with every edge and the flows are conserved at all non-source and non-sink nodes. We propose a novel polynomial time algorithm to reconstruct conserved networks from flow data by exploiting graph theoretic properties of conserved networks combined with learning techniques. We prove that exact network reconstruction is possible for arborescence networks. We also extend the methodology for reconstructing networks from noisy data and explore the reconstruction performance on arborescence networks with different structural characteristics.

cs.LG

A Graph Partitioning Algorithm for Leak Detection in Water Distribution Networks

Leak detection in urban water distribution networks (WDNs) is challenging given their scale, complexity, and limited instrumentation. We present an algorithm for leak detection in WDNs, which involves making additional flow measurements on-demand, and repeated use of water balance. Graph partitioning is used to determine the location of flow measurements, with the objective to minimize the measurement cost. We follow a multi-stage divide and conquer approach. In every stage, a section of the WDN identified to contain the leak is partitioned into two or more sub-networks, and water balance is used to trace the leak to one of these sub-networks. This process is recursively continued until the desired resolution is achieved. We investigate different methods for solving the arising graph partitioning problem like integer linear programming (ILP) and spectral bisection. The proposed methods are tested on large scale benchmark networks, and our results indicate that on average, less than 3% of the pipes need to be measured for finding the leak in large networks.

cs.DS

Optimal Power Distribution Control for a Network of Fuel Cell Stacks

In power networks where multiple fuel cell stacks are employed to deliver the required power, optimal sharing of the power demand between different stacks is an important problem. This is because the total current collectively produced by all the stacks is directly proportional to the fuel utilization, through stoichiometry. As a result, one would like to produce the required power while minimizing the total current produced. In this paper, an optimization formulation is proposed for this power distribution control problem. An algorithm that identifies the globally optimal solution for this problem is developed. Through an analysis of the KKT conditions, the solution to the optimization problem is decomposed into on-line and on-line computations. The on-line computations reduce to simple equation solving. For an application with a specific v-i function derived from data, we show that analytical solutions exist for on-line computations. We also discuss the wider applicability of the proposed approach for similar problems in other domains.

eess.SY

Network Topology Identification using PCA and its Graph Theoretic Interpretations

We solve the problem of identifying (reconstructing) network topology from steady state network measurements. Concretely, given only a data matrix $\mathbf{X}$ where the $X_{ij}$ entry corresponds to flow in edge $i$ in configuration (steady-state) $j$, we wish to find a network structure for which flow conservation is obeyed at all the nodes. This models many network problems involving conserved quantities like water, power, and metabolic networks. We show that identification is equivalent to learning a model $\mathbf{A_n}$ which captures the approximate linear relationships between the different variables comprising $\mathbf{X}$ (i.e. of the form $\mathbf{A_n X \approx 0}$) such that $\mathbf{A_n}$ is full rank (highest possible) and consistent with a network node-edge incidence structure. The problem is solved through a sequence of steps like estimating approximate linear relationships using Principal Component Analysis, obtaining f-cut-sets from these approximate relationships, and graph realization from f-cut-sets (or equivalently f-circuits). Each step and the overall process is polynomial time. The method is illustrated by identifying topology of a water distribution network. We also study the extent of identifiability from steady-state data.

cs.LG

Deconstructing Principal Component Analysis Using a Data Reconciliation Perspective

Data reconciliation (DR) and Principal Component Analysis (PCA) are two popular data analysis techniques in process industries. Data reconciliation is used to obtain accurate and consistent estimates of variables and parameters from erroneous measurements. PCA is primarily used as a method for reducing the dimensionality of high dimensional data and as a preprocessing technique for denoising measurements. These techniques have been developed and deployed independently of each other. The primary purpose of this article is to elucidate the close relationship between these two seemingly disparate techniques. This leads to a unified framework for applying PCA and DR. Further, we show how the two techniques can be deployed together in a collaborative and consistent manner to process data. The framework has been extended to deal with partially measured systems and to incorporate partial knowledge available about the process model.

cs.LG