SearcharxivSearch

arXiv subjects

Sriram Nagaraj

Publications and source records attributed to Sriram Nagaraj.

14 recordsLinked to original sources

Telemetry and Concealment in Self-Adapting Generative AI: Logging Architecture, Adversarial Model Hiding, and the Limits of Detection

Model risk management (MRM) guidance assumes a static model lifecycle, in which models are developed, independently validated, and implemented without further autonomous modification. Continually self-adapting generative AI systems --- models that update their own weights during production deployment --- fundamentally violate this assumption and render point-in-time validation inadequate. This paper addresses the resulting governance problem in two parts. Part I develops a rigorous telemetry architecture for such models, operating simultaneously in discrete and continuous time. We establish a Minimal Sufficient Statistic for audit purposes, construct a tamper-evident Merkle chain for discrete weight sequences, derive the appropriate continuous-time generalization via the Ito formula, and propose event-driven logging via KL divergence stopping times that is both computationally tractable and meaningful for validation. Part II asks what happens when the model provider is adversarial. A firm deploying such a model has strong incentives to conceal learning updates that would trigger mandatory validation review. We formalize this as the Model Hiding Problem and provide a systematic taxonomy of six distinct attack strategies against the Part I architecture, spanning discrete and continuous time, with a formal countermeasure for each. Together the two parts establish a dual-regime architecture in which continuous telemetry is necessary but not sufficient, narrowing---but never replacing---periodic invasive audit. The framework is model-architecture-agnostic and is designed to satisfy the three pillars of traditional MRM.

cs.CR

Joint Lyapunov Certificates for K-Agent Generative AI Governance: Stochastic Stability, Emergent Ensemble Risk, and Zero-Knowledge Governance Attestation

We develop a rigorous mathematical framework for the governance of systems of K self-adapting generative AI models under the principles of Model Risk Management (MRM). When multiple models share a meta-learning coupling through an interaction matrix, the per-agent Lyapunov analysis that underpins standard MRM is provably insufficient: individual agents can each satisfy their declared stability bounds while the joint system is in a regime of emergent ensemble-level drift. We formalize this gap through the Joint Lyapunov Proof (JLP)---a cryptographic and stochastic protocol that attests, without revealing proprietary weights, that the aggregate dynamics satisfy MRM Ongoing Monitoring standard at every validation epoch. Our main contributions are the following. We give a complete characterization of the infinitesimal generator of the joint quadratic Lyapunov function. We derive the exact critical coupling threshold above which the system loses mean-square stability. We prove a Noise-Floor Theorem and identify the correct target for zero-knowledge attestation. A per-epoch Succinct Non-Interactive Argument of Knowledge (SNARK) on the live weights is derived. All theoretical claims are validated against five numerical studies using a multi-agent softmax system.

math.NA

Recursive Governance: A Graph-Theoretic Framework for Risk Propagation and Drift Detection in Agentic AI Systems

As financial institutions transition from traditional predictive models to autonomous agentic systems, the static model inventory requirements of traditional model risk management (MRM) face structural obsolescence. This paper proposes a dynamic Inventory-as-Code (IaC) governance loop that treats the model inventory as a living architectural component rather than a periodic documentation artifact. We make four principal contributions. First, we introduce a calibrated Degree of Autonomy (DoA) materiality score with an explicit, taxonomized tool-complexity weighting scheme that addresses the dominance problem of naive additive risk formulations. Second, we construct the agent inventory as a Directed Acyclic Graph (DAG) and define a formal Composite Risk Propagation algorithm under which upstream validation failures induce risk penalties on all reachable descendants. Third, we develop a Trajectory Monitoring protocol based on distributional cosine drift across ensembled Chain-of-Thought (CoT) embeddings, with an explicit procedure for constructing and certifying the Golden Path baseline, a matched-bootstrap calibration that we show is necessary to avoid a severe false-positive artifact in the naive alternative, and a two-stage response that separates legitimate reasoning variation from detrimental drift without demanding that a single exceedance event trigger an irreversible action. Fourth, we address practical complications largely absent from the prior literature: LLM base-model version changes, latent feedback loops in nominally acyclic agent graphs, and a two-pass execution structure that stages validation ahead of risk propagation.

math.NA

Nonlocal Neural Tangent Kernels via Parameter-Space Interactions

The Neural Tangent Kernel (NTK) framework has provided deep insights into the training dynamics of neural networks under gradient flow. However, it relies on the assumption that the network is differentiable with respect to its parameters, an assumption that breaks down when considering non-smooth target functions or parameterized models exhibiting non-differentiable behavior. In this work, we propose a Nonlocal Neural Tangent Kernel (NNTK) that replaces the local gradient with a nonlocal interaction-based approximation in parameter space. Nonlocal gradients are known to exist for a wider class of functions than the standard gradient. This allows NTK theory to be extended to nonsmooth functions, stochastic estimators, and broader families of models. We explore both fixed-kernel and attention-based formulations of this nonlocal operator. We illustrate the new formulation with numerical studies.

cs.LG

BrowNNe: Brownian Nonlocal Neurons & Activation Functions

It is generally thought that the use of stochastic activation functions in deep learning architectures yield models with superior generalization abilities. However, a sufficiently rigorous statement and theoretical proof of this heuristic is lacking in the literature. In this paper, we provide several novel contributions to the literature in this regard. Defining a new notion of nonlocal directional derivative, we analyze its theoretical properties (existence and convergence). Second, using a probabilistic reformulation, we show that nonlocal derivatives are epsilon-sub gradients, and derive sample complexity results for convergence of stochastic gradient descent-like methods using nonlocal derivatives. Finally, using our analysis of the nonlocal gradient of Holder continuous functions, we observe that sample paths of Brownian motion admit nonlocal directional derivatives, and the nonlocal derivatives of Brownian motion are seen to be Gaussian processes with computable mean and standard deviation. Using the theory of nonlocal directional derivatives, we solve a highly nondifferentiable and nonconvex model problem of parameter estimation on image articulation manifolds. Using Brownian motion infused ReLU activation functions with the nonlocal gradient in place of the usual gradient during backpropagation, we also perform experiments on multiple well-studied deep learning architectures. Our experiments indicate the superior generalization capabilities of Brownian neural activation functions in low-training data regimes, where the use of stochastic neurons beats the deterministic ReLU counterpart.

cs.LG

Physics Informed Machine Learning (PIML) methods for estimating the remaining useful lifetime (RUL) of aircraft engines

This paper is aimed at using the newly developing field of physics informed machine learning (PIML) to develop models for predicting the remaining useful lifetime (RUL) aircraft engines. We consider the well-known benchmark NASA Commercial Modular Aero-Propulsion System Simulation (C-MAPSS) data as the main data for this paper, which consists of sensor outputs in a variety of different operating modes. C-MAPSS is a well-studied dataset with much existing work in the literature that address RUL prediction with classical and deep learning methods. In the absence of published empirical physical laws governing the C-MAPSS data, our approach first uses stochastic methods to estimate the governing physics models from the noisy time series data. In our approach, we model the various sensor readings as being governed by stochastic differential equations, and we estimate the corresponding transition density mean and variance functions of the underlying processes. We then augment LSTM (long-short term memory) models with the learned mean and variance functions during training and inferencing. Our PIML based approach is different from previous methods, and we use the data to first learn the physics. Our results indicate that PIML discovery and solutions methods are well suited for this problem and outperform previous data-only deep learning methods for this data set and task. Moreover, the framework developed herein is flexible, and can be adapted to other situations (other sensor modalities or combined multi-physics environments), including cases where the underlying physics is only partially observed or known.

cs.LG

Marrying Compressed Sensing and Deep Signal Separation

Blind signal separation (BSS) is an important and challenging signal processing task. Given an observed signal which is a superposition of a collection of unknown (hidden/latent) signals, BSS aims at recovering the separate, underlying signals from only the observed mixed signal. As an underdetermined problem, BSS is notoriously difficult to solve in general, and modern deep learning has provided engineers with an effective set of tools to solve this problem. For example, autoencoders learn a low-dimensional hidden encoding of the input data which can then be used to perform signal separation. In real-time systems, a common bottleneck is the transmission of data (communications) to a central command in order to await decisions. Bandwidth limits dictate the frequency and resolution of the data being transmitted. To overcome this, compressed sensing (CS) technology allows for the direct acquisition of compressed data with a near optimal reconstruction guarantee. This paper addresses the question: can compressive acquisition be combined with deep learning for BSS to provide a complete acquire-separate-predict pipeline? In other words, the aim is to perform BSS on a compressively acquired signal directly without ever having to decompress the signal. We consider image data (MNIST and E-MNIST) and show how our compressive autoencoder approach solves the problem of compressive BSS. We also provide some theoretical insights into the problem.

math.NA

Optimization and Learning With Nonlocal Calculus

Nonlocal models have recently had a major impact in nonlinear continuum mechanics and are used to describe physical systems/processes which cannot be accurately described by classical, calculus based "local" approaches. In part, this is due to their multiscale nature that enables aggregation of micro-level behavior to obtain a macro-level description of singular/irregular phenomena such as peridynamics, crack propagation, anomalous diffusion and transport phenomena. At the core of these models are nonlocal differential operators, including nonlocal analogs of the gradient/Hessian. This paper initiates the use of such nonlocal operators in the context of optimization and learning. We define and analyze the convergence properties of nonlocal analogs of (stochastic) gradient descent and Newton's method on Euclidean spaces. Our results indicate that as the nonlocal interactions become less noticeable, the optima corresponding to nonlocal optimization converge to the "usual" optima. At the same time, we argue that nonlocal learning is possible in situations where standard calculus fails. As a stylized numerical example of this, we consider the problem of non-differentiable parameter estimation on a non-smooth translation manifold and show that our nonlocal gradient descent recovers the unknown translation parameter from a non-differentiable objective function.

math.OC

Zheghalkin-Boolean Calculus

Boolean calculus has been studied extensively in the past in the context of switching circuits, error-correcting codes etc. This work generalizes several approaches to defining a differential calculus for Boolean functions. A unified theory of Boolean calculus, complete with k-forms and integration, is presented through the use of Zhegalkin algebras (i.e., algebraic normal forms), culminating in a Stokes-like theorem for Boolean functions.

math.RA

A spacetime DPG method for the Schrodinger equation

A spacetime Discontinuous Petrov Galerkin (DPG) method for the linear time-dependent Schrodinger equation is proposed. The spacetime approach is particularly attractive for capturing irregular solutions. Motivated by the fact that some irregular Schrodinger solutions cannot be solutions of certain first order reformulations, the proposed spacetime method uses the second order Schrodinger operator. Two variational formulations are proved to be well-posed: a strong formulation (with no relaxation of the original equation) and a weak formulation (also called the ultraweak formulation, that transfers all derivatives onto test functions). The convergence of the DPG method based on the ultraweak formulation is investigated using an interpolation operator. A standalone appendix analyzes the ultraweak formulation for general differential operators. Reports of numerical experiments motivated by pulse propagation in dispersive optical fibers are also included.

math.NA

Orientation Embedded High Order Shape Functions for the Exact Sequence Elements of All Shapes

A unified construction of high order shape functions is given for all four classical energy spaces ($H^1$, $H(\mathrm{curl})$, $H(\mathrm{div})$ and $L^2$) and for elements of "all" shapes (segment, quadrilateral, triangle, hexahedron, tetrahedron, triangular prism and pyramid). The discrete spaces spanned by the shape functions satisfy the commuting exact sequence property for each element. The shape functions are conforming, hierarchical and compatible with other neighboring elements across shared boundaries so they may be used in hybrid meshes. Expressions for the shape functions are given in coordinate free format in terms of the relevant affine coordinates of each element shape. The polynomial order is allowed to differ for each separate topological entity (vertex, edge, face or interior) in the mesh, so the shape functions can be used to implement local $p$ adaptive finite element methods. Each topological entity may have its own orientation, and the shape functions can have that orientation embedded by a simple permutation of arguments.

math.NA

A Theory for Optical flow-based Transport on Image Manifolds

An image articulation manifold (IAM) is the collection of images formed when an object is articulated in front of a camera. IAMs arise in a variety of image processing and computer vision applications, where they provide a natural low-dimensional embedding of the collection of high-dimensional images. To date IAMs have been studied as embedded submanifolds of Euclidean spaces. Unfortunately, their promise has not been realized in practice, because real world imagery typically contains sharp edges that render an IAM non-differentiable and hence non-isometric to the low-dimensional parameter space under the Euclidean metric. As a result, the standard tools from differential geometry, in particular using linear tangent spaces to transport along the IAM, have limited utility. In this paper, we explore a nonlinear transport operator for IAMs based on the optical flow between images and develop new analytical tools reminiscent of those from differential geometry using the idea of optical flow manifolds (OFMs). We define a new metric for IAMs that satisfies certain local isometry conditions, and we show how to use this metric to develop a new tools such as flow fields on IAMs, parallel flow fields, parallel transport, as well as a intuitive notion of curvature. The space of optical flow fields along a path of constant curvature has a natural multi-scale structure via a monoid structure on the space of all flow fields along a path. We also develop lower bounds on approximation errors while approximating non-parallel flow fields by parallel flow fields.

cs.CV

The Orbit Group of a Quandle

We define the notion of the orbit group of a quandle via its connectivity and compute the orbit groups for some basic quandles. We also show that the orbit group counts the number of orbits of certain quandles.

math.GT