SearcharxivSearch

arXiv subjects

Hongru Zhao

Publications and source records attributed to Hongru Zhao.

17 recordsLinked to original sources

Shifted Anticoncentration for Real Gram Hafnians and Symmetric Gaussian Hafnians

We prove a uniform shifted anticoncentration theorem for the hafnian of a real Gaussian Gram matrix under an explicit condition on the row dimension. After normalization by its root mean square, the law has a bounded continuous density, maximal at zero, and every interval has probability bounded by an explicit coefficient times its radius. Under suitable growth conditions on the row dimension, this coefficient grows at most polynomially in the hafnian order, meaning half the dimension of the Gram matrix. We also compute the exact second moment. The proof exploits the perfect matching structure, combining conditional Gaussian representations, row suspension, and bilinear interpolation to control an inverse moment of the conditional variance. At fixed hafnian order, a rescaled limit as the row dimension grows yields corresponding bounds for the hafnian of a real symmetric Gaussian matrix with independent entries above the diagonal.

math.PR

Uniform Hiding of Haar Block Transpose Gram Matrices

Gaussian boson sampling with equally squeezed active inputs assigns collision free probabilities through a complex symmetric transpose Gram matrix formed from a rectangular block of a Haar interferometer. We prove a finite total variation comparison with the corresponding complex Gaussian transpose Gram law. The error bound is explicit, quadratic in the number of selected output modes, inversely proportional to the interferometer size, and uniform in the number of squeezed active inputs. The proof combines a centered circular orthogonal ensemble score analysis in the dense regime with a rectangular relative entropy bound. This result supplies a random matrix replacement component of Gaussian boson sampling hardness arguments.

quant-ph

Convex Reparameterization and Self-Concordant Algorithms for Multivariate Regression with Covariance Estimation

Building on a reparameterization for multivariate linear regression that yields a jointly convex penalized likelihood in the reparameterized regression coefficient matrix and the precision matrix, we show that the resulting scaled Gaussian loss is standard self-concordant. This places the joint estimation problem within composite self-concordant optimization and leads to two algorithms: a proximal gradient method and a damped proximal Newton method. In simulations, we evaluate algorithmic robustness, iterations to convergence, and elapsed time. In a protein expression application, compared with the classical-parameterization formulation, the proposed convex formulation attains similar mean squared prediction error and can be substantially faster when the fitted precision matrix is dense.

stat.CO

Weak Typicality of von Neumann Entanglement Entropy in Gaussian Boson Sampling

We study the von Neumann entanglement entropy generated by a Haar distributed passive interferometer acting on $n$ equally squeezed input modes with fixed nonzero squeezing strength $s$. Previous work established proportional weak typicality for integer R'enyi orders $\alpha\geq 2$ and stated a sublinear von Neumann result, while the proportional von Neumann case remained open. For a subsystem of $k_n$ modes satisfying $k_n/n\to r\in(0,1)$, we prove that, for every $\varepsilon>0$ and all sufficiently large $n$, $\mathbb{P}\left(\left|\frac{S_{1,n}}{\mathbb{E}S_{1,n}}-1\right|\geq\varepsilon\right)\leq2\exp\left[-\frac{c_{s,r}\varepsilon^2n^2}{\log^2(en)}\right].$ The proof represents the entropy as a singular value statistic of a principal block of $UU^{\mathsf T}$, where $U$ denotes the unitary interferometer. It regularizes the logarithmic singularity at the endpoint corresponding to a pure Gaussian mode and applies concentration on the unitary group. The result establishes proportional von Neumann weak typicality and further implies almost sure convergence of $S_{1,n}/\mathbb{E}S_{1,n}$ to $1$, a typical volume law, and the variance bound $\mathrm{Var}(S_{1,n})=O_s(\log^2 n)$. An accompanying Lean 4 development verifies the proof chain.

quant-ph

Exact Moments of Gaussian Gram Hafnians Reveal an $n^2/\log n$ Threshold for Weak Anticoncentration

Anticoncentration is central to hardness arguments for approximate sampling. In the independent Gaussian surrogate for collision free Gaussian boson sampling, the moment ratio studied here also determines the averaged ideal linear cross entropy reference value. Let $H_{k,n}=\mathrm{haf}(X^{\mathsf T}X)$, where $X\in\mathbb{C}^{k\times 2n}$ has independent standard circular complex Gaussian entries. We evaluate $\mathbb{E}|H_{k,n}|^2$ and $\mathbb{E}|H_{k,n}|^4$ exactly by reducing four hafnian copies to a rank two Gaussian integral. For $R_{k,n}=(\mathbb{E}|H_{k,n}|^2)^2/\mathbb{E}|H_{k,n}|^4$, we obtain $R_{k,n}=4^{-n}\binom{2n}{n}/F_{k,n}$, where $F_{k,n}={}3F_2(-n,-n,1/2;1,k/2;1)$ is a terminating generalized hypergeometric polynomial. If $k/n^2\to c>0$, then $F{k,n}\to e^{1/c}I_0(1/c)$, where $I_0$ is the modified Bessel function of the first kind of order zero, and consequently $R_{k,n}\sqrt{\pi n}\to[e^{1/c}I_0(1/c)]^{-1}$. Thus $k\asymp n^2$ is a smooth Bessel crossover, whereas the scaling order boundary for inverse polynomial weak anticoncentration is $k\asymp n^2/\log n$. These conclusions concern the Gaussian surrogate moment criterion; finite dimensional Haar moment transfer and high probability small ball anticoncentration remain separate problems.

quant-ph

Sharp Berry-Esseen Bounds for the Log Determinant of a Gaussian Sample Correlation Matrix

Let $\widehat R$ be the Pearson sample correlation matrix formed from $n$ independent Gaussian observations in $p$ dimensions, and write $m=n-1\ge p$. Under the null correlation $R=I_p$, the classical independent beta product, exact cumulants, and full Fourier inversion yield, along every sequence $p\to\infty$ with $m\ge p$, a uniform first Edgeworth expansion for $\log\det\widehat R$, centered by its exact mean and scaled by its exact standard deviation. The expansion identifies the exact finite dimensional skewness correction and gives the sharp Kolmogorov equivalent $A_{m,p}/\{6\sqrt{2\pi}V_{m,p}^{3/2}\}$, where $V_{m,p}$ is the exact variance and $A_{m,p}$ is the absolute third cumulant. This equivalent unifies the square, fixed gap, growing gap, proportional, and dilute regimes; in the square regime the error has order $(\log p)^{-3/2}$ with an exact constant. For every positive definite population correlation matrix $R$, we prove a uniform finite sample Berry-Esseen bound that explicitly tracks population dependence. All theoretical results have exact or proved equivalent Lean 4 formulations whose declarations and dependencies are kernel checked.

math.PR

On the Log Determinant of Sample Correlation Matrices under Gaussianity

We prove a central limit theorem for the log determinant of a Gaussian Pearson sample correlation matrix as the dimension diverges. Only two conditions are imposed: the population correlation matrix is positive definite, and the sample degrees of freedom are at least the dimension. Both are necessary for the ordinary log determinant to be finite. To the best of our knowledge, no previous central limit theorem covers this full nonsingular domain. It covers every aspect ratio from dilute growth to the square hard edge. No uniform lower or upper bound is imposed on the eigenvalues of the population correlation matrices: the smallest may approach zero and the largest may diverge. The proof develops a coordinatewise Wiener chaos reduction for the random diagonal normalization and combines it with an exact Wishart transform comparison. Geometrically, the statistic is twice the log volume of a random parallelotope spanned by standardized Gaussian coordinate vectors.

math.ST

Convergence and Stability Analysis of Self-Consuming Generative Models with Heterogeneous Human Curation

Self-consuming generative models have received significant attention over the last few years. In this paper, we study a self-consuming generative model with heterogeneous preferences that is a generalization of the model in Ferbach et al. (2024). The model is retrained round by round using real data and its previous-round synthetic outputs. The asymptotic behavior of the retraining dynamics is investigated across four regimes using different techniques including the nonlinear Perron--Frobenius theory. Our analyses improve upon that of Ferbach et al. (2024) and provide convergence results in settings where the well-known Banach contraction mapping arguments do not apply. Stability and non-stability results regarding the retraining dynamics are also given.

stat.ML

InSight-R: A Framework for Risk-informed Human Failure Event Identification and Interface-Induced Risk Assessment Driven by AutoGraph

Human reliability remains a critical concern in safety-critical domains such as nuclear power, where operational failures are often linked to human error. While conventional human reliability analysis (HRA) methods have been widely adopted, they rely heavily on expert judgment for identifying human failure events (HFEs) and assigning performance influencing factors (PIFs). This reliance introduces challenges related to reproducibility, subjectivity, and limited integration of interface-level data. In particular, current approaches lack the capacity to rigorously assess how human-machine interface design contributes to operator performance variability and error susceptibility. To address these limitations, this study proposes a framework for risk-informed human failure event identification and interface-induced risk assessment driven by AutoGraph (InSight-R). By linking empirical behavioral data to the interface-embedded knowledge graph (IE-KG) constructed by the automated graph-based execution framework (AutoGraph), the InSight-R framework enables automated HFE identification based on both error-prone and time-deviated operational paths. Furthermore, we discuss the relationship between designer-user conflicts and human error. The results demonstrate that InSight-R not only enhances the objectivity and interpretability of HFE identification but also provides a scalable pathway toward dynamic, real-time human reliability assessment in digitalized control environments. This framework offers actionable insights for interface design optimization and contributes to the advancement of mechanism-driven HRA methodologies.

cs.HC

AutoGraph: A Knowledge-Graph Framework for Modeling Interface Interaction and Automating Procedure Execution in Digital Nuclear Control Rooms

Digitalization in nuclear power plant (NPP) control rooms is reshaping how operators interact with procedures and interface elements. However, existing computer-based procedures (CBPs) often lack semantic integration with human-system interfaces (HSIs), limiting their capacity to support intelligent automation and increasing the risk of human error, particularly under dynamic or complex operating conditions. In this study, we present AutoGraph, a knowledge-graph-based framework designed to formalize and automate procedure execution in digitalized NPP environments.AutoGraph integrates (1) a proposed HTRPM tracking module to capture operator interactions and interface element locations; (2) an Interface Element Knowledge Graph (IE-KG) encoding spatial, semantic, and structural properties of HSIs; (3) automatic mapping from textual procedures to executable interface paths; and (4) an execution engine that maps textual procedures to executable interface paths. This enables the identification of cognitively demanding multi-action steps and supports fully automated execution with minimal operator input. We validate the framework through representative control room scenarios, demonstrating significant reductions in task completion time and the potential to support real-time human reliability assessment. Further integration into dynamic HRA frameworks (e.g., COGMIF) and real-time decision support systems (e.g., DRIF) illustrates AutoGraph extensibility in enhancing procedural safety and cognitive performance in complex socio-technical systems.

cs.HC

A Cognitive-Mechanistic Human Reliability Analysis Framework: A Nuclear Power Plant Case Study

Traditional human reliability analysis (HRA) methods, such as IDHEAS-ECA, rely on expert judgment and empirical rules that often overlook the cognitive underpinnings of human error. Moreover, conducting human-in-the-loop experiments for advanced nuclear power plants is increasingly impractical due to novel interfaces and limited operational data. This study proposes a cognitive-mechanistic framework (COGMIF) that enhances the IDHEAS-ECA methodology by integrating an ACT-R-based human digital twin (HDT) with TimeGAN-augmented simulation. The ACT-R model simulates operator cognition, including memory retrieval, goal-directed procedural reasoning, and perceptual-motor execution, under high-fidelity scenarios derived from a high-temperature gas-cooled reactor (HTGR) simulator. To overcome the resource constraints of large-scale cognitive modeling, TimeGAN is trained on ACT-R-generated time-series data to produce high-fidelity synthetic operator behavior datasets. These simulations are then used to drive IDHEAS-ECA assessments, enabling scalable, mechanism-informed estimation of human error probabilities (HEPs). Comparative analyses with SPAR-H and sensitivity assessments demonstrate the robustness and practical advantages of the proposed COGMIF. Finally, procedural features are mapped onto a Bayesian network to quantify the influence of contributing factors, revealing key drivers of operational risk. This work offers a credible and computationally efficient pathway to integrate cognitive theory into industrial HRA practices.

cs.AI

On Validating Angular Power Spectral Models for the Stochastic Gravitational-Wave Background Without Distributional Assumptions

It is demonstrated that estimators of the angular power spectrum commonly used for the stochastic gravitational-wave background (SGWB) lack a closed-form analytical expression for the likelihood function and, typically, cannot be accurately approximated by a Gaussian likelihood. Nevertheless, a robust statistical analysis can be performed to enable the estimation and testing of angular power spectral models for the SGWB without specifying distributional assumptions. Here, the technical aspects of the method are discussed in detail. Moreover, a new, consistent estimator for the covariance of the angular power spectrum is derived. The proposed approach is applied to data from the third observing run (O3) of Advanced LIGO and Advanced Virgo.

astro-ph.IM

Testing models for angular power spectra: A distribution-free approach

A novel goodness-of-fit strategy is introduced for testing models of angular power spectra with unknown parameters. Using this strategy, it is possible to assess the validity of such models without specifying the distribution of the angular power spectrum estimators. This holds under general conditions, ensuring the method's applicability in diverse applications. Moreover, the proposed solution overcomes the need for case-by-case simulations when testing different models, leading to notable computational advantages.

physics.data-an

KRAIL: A Knowledge-Driven Framework for Base Human Reliability Analysis Integrating IDHEAS and Large Language Models

Human reliability analysis (HRA) is crucial for evaluating and improving the safety of complex systems. Recent efforts have focused on estimating human error probability (HEP), but existing methods often rely heavily on expert knowledge,which can be subjective and time-consuming. Inspired by the success of large language models (LLMs) in natural language processing, this paper introduces a novel two-stage framework for knowledge-driven reliability analysis, integrating IDHEAS and LLMs (KRAIL). This innovative framework enables the semi-automated computation of base HEP values. Additionally, knowledge graphs are utilized as a form of retrieval-augmented generation (RAG) for enhancing the framework' s capability to retrieve and process relevant data efficiently. Experiments are systematically conducted and evaluated on authoritative datasets of human reliability. The experimental results of the proposed methodology demonstrate its superior performance on base HEP estimation under partial information for reliability assessment.

cs.CL

Subspace decompositions for association structure learning in multivariate categorical response regression

Modeling the complex relationships between multiple categorical response variables as a function of predictors is a fundamental task in the analysis of categorical data. However, existing methods can be difficult to interpret and may lack flexibility. To address these challenges, we introduce a penalized likelihood method for multivariate categorical response regression that relies on a novel subspace decomposition to parameterize interpretable association structures. Our approach models the relationships between categorical responses by identifying mutual, joint, and conditionally independent associations, which yields a linear problem within a tensor product space. We establish theoretical guarantees for our estimator, including error bounds in high-dimensional settings, and demonstrate the method's interpretability and prediction accuracy through comprehensive simulation studies.

stat.ME

Globally-Optimal Greedy Experiment Selection for Active Sequential Estimation

Motivated by modern applications such as computerized adaptive testing, sequential rank aggregation, and heterogeneous data source selection, we study the problem of active sequential estimation, which involves adaptively selecting experiments for sequentially collected data. The goal is to design experiment selection rules for more accurate model estimation. Greedy information-based experiment selection methods, optimizing the information gain for one-step ahead, have been employed in practice thanks to their computational convenience, flexibility to context or task changes, and broad applicability. However, statistical analysis is restricted to one-dimensional cases due to the problem's combinatorial nature and the seemingly limited capacity of greedy algorithms, leaving the multidimensional problem open. In this study, we close the gap for multidimensional problems. In particular, we propose adopting a class of greedy experiment selection methods and provide statistical analysis for the maximum likelihood estimator following these selection rules. This class encompasses both existing methods and introduces new methods with improved numerical efficiency. We prove that these methods produce consistent and asymptotically normal estimators. Additionally, within a decision theory framework, we establish that the proposed methods achieve asymptotic optimality when the risk measure aligns with the selection rule. We also conduct extensive numerical studies on both simulated and real data to illustrate the efficacy of the proposed methods. From a technical perspective, we devise new analytical tools to address theoretical challenges. These analytical tools are of independent theoretical interest and may be reused in related problems involving stochastic approximation and sequential designs.

math.ST

Limiting Empirical Spectral Distribution for Products of Rectangular Matrices

In this paper, we consider $m$ independent random rectangular matrices whose entries are independent and identically distributed standard complex Gaussian random variables and assume the product of the $m$ rectangular matrices is an $n$ by $n$ square matrix. We study the limiting empirical spectral distributions of the product where the dimension of the product matrix goes to infinity, and $m$ may change with the dimension of the product matrix and diverge. We give a complete description for the limiting distribution of the empirical spectral distributions for the product matrix and illustrate some examples.

math.PR