SearcharxivSearch

arXiv subjects

Matthew Francis Dixon

Publications and source records attributed to Matthew Francis Dixon.

6 recordsLinked to original sources

Identification and Honest Recovery from Semantic Observation Kernels: Operator Error, Coarsening, and Stability

Probabilistic text generators, such as large language models, assign probabilities to phrases, but consequential decisions require posterior uncertainty over meaningful states. These are not interchangeable: language probabilities depend on the prompt, may be incomplete and need not reliably identify state uncertainty. Without a statistical bridge, fluent responses and numerical confidence are insufficient for inference or governance. We formulate recovery of the target posterior as a semiparametric inverse problem and develop honest recovery guarantees that account jointly for calibration error, measurement noise, incomplete probabilities and weak identification. Simulations demonstrate the predicted coverage and stability behaviour, while two frozen language-model studies demonstrate held-out recovery. The resulting method determines when a semantic measurement can be trusted for inference and when use, review, recalibration or abstention is warranted, providing a statistical foundation for runtime AI governance.

stat.ME

Schur-Riesz Variational Enrichment: A Generalized Refinement Framework for Finite Elements

Width theory identifies economical spaces for compact PDE solution families, but does not provide a stable adaptive selection rule. We introduce Schur-Riesz refinement, which compares ordinary h/p refinement with operator-informed functions on a common variational scale. Projection removes content already represented by the incumbent, Riesz bounds test coefficient stability, and exact Schur gain ranks the surviving directions. New non-regression, bulk-contraction, and mixed near-oracle results provide finite-run error, dimension, and work certificates for coercive Galerkin and noncoercive minimum-residual formulations. We applied Schur-Riesz refinement across coercive and noncoercive PDEs. At matched dimension, automatic modes reduce held-out error by $50.2\%$ for 2D heterogeneous Helmholtz and $30.9\%$ for 2D Darcy. At matched Darcy error, the selected space uses $38.2\%$ fewer coordinates, reducing deployment memory by $38.9\%$ and online time by $30.1\%$. The additional offline construction is reusable and can be amortized across multiple right-hand sides for a fixed operator. On locked 3D heterogeneous Helmholtz tests, the method reduces mean error by $25.1$-$47.3\%$ against matched spectral-polynomial spaces and meets target errors with 24-96 coordinates, versus 1,331-6,859 for native MFEM hp-refinement. In an adaptive audit, 48 coordinates attain mean error $.1092$, compared with $.0967$ using 878 polynomial coordinates, while making online solves $6.8\times$ faster. Thus, the experiments demonstrate that Schur-Riesz refinement can construct smaller stable spaces with lower deployment memory and repeated-solve cost, while preserving a certified incumbent when richer functions do not help.

math.NA

Successive Schur-Riesz Analysis for Approximation

Many approximation methods enlarge a trial space by adjoining function blocks generated by different operators. Exact redundancy and strong cross-level interaction can make coefficients nonunique and render pairwise or diagonal-dominance tests needlessly pessimistic. For \(V_m=\sum_{\ell\leq m}S_\ell(E_\ell)\) in a Hilbert space $\mathcal H$, we quotient coefficients representing the same function and control successive orthogonal innovations to obtain Riesz bounds independent of \(m\). The setting includes factored operators \(S_\ell=T_\ell\circ\cdots\circ T_1:E_\ell\to\mathcal H\), with compatible intermediate spaces, but the theorem allows arbitrary bounded \(S_\ell\). The same constants control approximation, truncation, perturbation, and levelwise error. A block Schur complement identifies the intrinsic new dimension and gives the exact reduction in squared best-approximation error, leading to a constructive enrichment procedure. Nonstationary and lifted examples give positive intrinsic bounds where diagonal-dominance estimates are negative or labelled Gram matrices are singular; adaptive and recycled-subspace calculations illustrate the distinct roles of representation stability and application-specific utility.

math.NA

Adaptive AI Delegation under Uncertainty: A Bayesian Governance Policy for Sequential Decision Authority

Organizations increasingly use large language models and agentic AI systems to generate probabilistic assessments and candidate actions in high-consequence settings. This creates a managerial problem distinct from prediction: how should organizations allocate decision authority to AI-generated recommendations as evidence quality, uncertainty, and organizational objectives evolve over time? Existing AI governance frameworks emphasize transparency, documentation, oversight, and regulatory compliance, but provide limited quantitative guidance for dynamically allocating decision authority under uncertainty. To address this challenge, we formulate adaptive AI delegation as a Governance-Aware Partially Observable Markov Decision Process (POMDP) in which Bayesian inference estimates the informational state and sequential optimization determines delegated AI authority. The paper also develops a quantitative validation and benchmarking framework for governance policies. Synthetic stress tests, reported LLM-confidence robustness, forecast-accuracy validation, governance-appetite sensitivity, and fragile-AI early-warning experiments evaluate whether the proposed policy exhibits graceful degradation, robustness to confidence-only perturbations, adaptive delegation under improving evidence quality, and interpretable calibration of institutional conservatism. The Governance-Aware POMDP is further benchmarked against five representative governance strategies operating under identical Bayesian beliefs, information, and governance objectives. The results show that while specialized heuristics perform well in stationary settings, sequential Bayesian governance provides the strongest general-purpose governance policy across heterogeneous AI-quality regimes by adaptively allocating organizational decision authority under uncertainty.

q-fin.RM

Model Validation of Agentic AI Systems: A POMDP-Based Framework for Belief-State, Forecast, and Policy Validation

Agentic artificial intelligence systems introduce a new class of model risk. Unlike traditional predictive models, autonomous agents continuously acquire information, form beliefs regarding latent states of the environment, generate forecasts, select actions, and adapt their behavior over time. Existing validation methodologies focus primarily on predictive accuracy and therefore provide limited insight into the quality of the underlying decision process. This paper proposes a model validation framework for agentic AI based on Partially Observable Markov Decision Processes (POMDPs). The framework decomposes autonomous decision making into information, beliefs, forecasts, actions, and utility, allowing each component to be validated independently. Large language models (LLMs) are formalized as approximate Bayesian filtering operators, and a model-risk taxonomy is developed encompassing state-space, filtering, forecast, policy, utility-specification, and parameter risks. The model risk validation methodology is demonstrated through a portfolio-management case study in which an agent infers latent market regimes from market and macroeconomic information, generates belief-conditioned forecasts, and constructs portfolios using a Black--Litterman framework. Empirical validation combines performance analysis, belief calibration diagnostics, coverage tests, ablation studies, and parameter-sensitivity analysis. The results indicate that latent-state inference contributes independently to decision quality and that the principal conclusions remain robust across a broad range of parameter values. The principal contribution of the paper is a practical framework for extending established model risk management concepts to autonomous AI systems and providing a rigorous foundation for their validation, governance, and monitoring.

q-fin.RM

Belief at Risk: Quantifying Agentic AI Model Risk with LLM-Inferred Bayesian State Filters

Agentic AI systems create model risk because uncertain beliefs are coupled to autonomous actions. This paper develops a mathematical framework for quantifying agentic AI risk by representing the system as a partially observed Markov decision process with latent states, Bayesian belief updates, control-dependent losses, and tail-risk functionals. The main methodological contribution is to treat a large language model as an uncertain semantic observation model: the LLM maps high-dimensional evidence into a probability vector over latent regimes, while a Bayesian filter imposes temporal coherence and produces auditable posterior beliefs. The resulting framework separates uncertainty quantification from risk measurement. Uncertainty is represented by posterior entropy, belief drift, and calibration error; risk is represented by the distribution of losses induced by decisions taken under those beliefs. The paper connects this construction to model risk management, coherent risk measures, Bayesian filtering, POMDP theory, robust control, and quantitative portfolio risk. An empirical case study using adjusted daily equity returns from Massive.com illustrates how LLM-inferred belief states can be combined with Bayesian filtering to produce regime probabilities, uncertainty diagnostics, calibration statistics, and VaR/CVaR-style risk measures. The framework is intended as a rigorous foundation for validating agentic AI in financial and other regulated decision environments.

q-fin.RM