SearcharxivSearch

arXiv subjects

Ting Yan

Publications and source records attributed to Ting Yan.

At least 19 recordsLinked to original sources

Do User-Authored Permission Policies Improve Protection Against AI Agent Overreach?

AI agents are poised to become a primary interface to digital products, acting across email, files, payments, and personal data. People without professional software backgrounds need understandable, reusable ways to control actions across services. We examine a mechanism in which a language model maps actions to plain-language consequence categories with user-authored "allow", "ask", or "never" rules. We ask what is gained and lost when decisions are made in advance as reusable rules rather than separately for each action. We analyzed 113 participants without professional software backgrounds across three conditions: per-action human-in-the-loop approval (HITL), automated per-action model review (AUTO), or user-authored consequence policy (POLICY). Participants judged 2 examples in each of 4 consequence categories; POLICY participants then set one rule per category. All supervised an 18-action simulated day, including 7 overreach actions. POLICY blocked less overreach than HITL (-20.1 percentage points, 95% CI [-32.1, -8.1]) and AUTO (-14.5 points, 95% CI [-25.8, -3.2]). POLICY lowered runtime prompts from 18.0 to 10.9, but total intervention time was not reliably lower when rule setup was included. Exploratory analysis showed that participants chose "ask" for 114 of 140 POLICY rules, returning most overreach actions to runtime. Of the 148 overreach actions executed in POLICY, 133 followed human approval and 15 ran automatically under "allow" rules. Across all 7 overreach actions, POLICY had the highest approval rate. Counterintuitively, user-authored rules did not by themselves provide stronger protection: many actions outside users' original requests went through after users approved them. These results reveal a gap between preference and commitment: repeatedly choosing "ask" preserves case-by-case choice but prevents a standing policy from settling decisions in advance.

cs.HC

When "Do Not" Is Not Deny: Security Rules in CLAUDE.md vs Built-In Controls

In CLAUDE.md, "do not" is a natural-language instruction that the model interprets. Claude Code's deny is a built-in control that blocks an action before the agent can take it. Both can express the same security goal, but they control the agent in different ways. We measure this gap in 481 public CLAUDE.md files. An LLM matched the extracted candidate rules against Claude Code's documented controls, and two security practitioners independently checked a sample without seeing the model's answers or each other's labels. Depending on how closely a control had to match the written rule, only about 4-16% of the retrieved security rules had a matching built-in control. Under the strictest standard the estimate was 4.4% (95% CI: 2.6-6.7%), and the two annotators agreed closely on which rules had a match. A manual review of complete files found that our extraction method captured 66.3% of eligible security rules; the reported rates therefore apply to the rules it captured. This is a usable security problem: CLAUDE.md is a write-only channel. A developer writes a security rule but gets no feedback on whether a control will enforce it. The same plain-text form hides two kinds of rule: those a permission rule, mode, or sandbox can enforce, and those left to the model to interpret.

cs.HC

Subgraph counting estimation for the $\beta$-model in sparse networks

The $\beta$-model is popular for characterizing the commonly observed degree heterogeneity phenomenon in real-world networks. In this study, we develop a cycle counting approach to estimate $n$ node-specific parameters in the $\beta$-model for moderate or extremely sparse networks. Our proposed estimators, called \emph{Cycle Counting Ratio (CCR) Estimator}, are based on the log-ratios of two network cycle counting statistics with explicit expressions and therefore easy to compute. We focus on conditions to guarantee statistical properties of the single estimator for each node. Under the very weak conditions that $\max_t \theta_t \to 0$ and $\theta_t \|\theta\|_1 \to \infty$, we show that the CCR estimator is consistent and achieves the minimax rate in terms of the mean squared error, which is the squared signal-to-noise ratio for $\hat{\beta}_t$ up to a constant factor. Here, $\hat{\beta}_t$ is the CCR estimator of the node-specific parameter $\beta_t$, $\theta_t = \exp(\beta_t)$ and $\theta=(\theta_1, \ldots, \theta_n)$. Even if the whole network density is close to the Erd\H{o}s-R\'{e}nyi lower bound $\log n/n$, the CCR estimator for the single parameter $\beta_t$ is still consistent as long as $\theta_t \|\theta\|_1 \to \infty$. To the best of our knowledge, this is the first time to derive the minimax rate and consistency result under such weak conditions. Under a slight stronger condition, we further establish its uniform consistency and asymptotic normality, whose asymptotic variance is $\theta_t \|\theta\|_1$. Numerical studies and an application to a sparse network data set demonstrate our theoretical findings.

stat.ME

Semiparametric analysis for paired comparisons with covariates

Statistical inference in parametric models (e.g., the Bradley--Terry model and its variants) for paired-comparison data has been explored in the high-dimensional regime, in which the number of items involving in paired comparisons diverges. However, parametric models are highly susceptible to model misspecification. To relax the assumption of known distributions and provide flexibility, we propose a semiparametric framework for modeling the merits of items and covariate effects (e.g., home-field advantage) by introducing latent random variables with unspecified distributions. As the number of parameters increases with the number of items, semiparametric inference is highly nontrivial. To address this issue, we employ a kernel-based least squares approach to estimate all unknown parameters. When each pair of items has a fixed number of comparisons and the number of items tends to infinity, we prove the consistency of all resulting estimators and derive their asymptotic normal distributions. To the best of our knowledge, this is the first study to conduct a semiparametric analysis of paired comparisons with an increasing dimension. We conduct simulations to evaluate the finite-sample performance of the proposed method and illustrate its practical utility by analyzing an NBA dataset.

stat.ME

Triple-dyad ratio estimation for the $p_1$ model

Although the $p_1$ model was proposed 40 years ago, little progress has been made to address asymptotic theories in this model, that is, neither consistency of the maximum likelihood estimator (MLE) nor other parameter estimation with statistical guarantees is understood. This problem has been acknowledged as a long-standing open problem. To address it, we propose a novel parametric estimation method based on the ratios of the sum of a sequence of triple-dyad indicators to another one, where a triple-dyad indicator means the product of three dyad indicators. Our proposed estimators, called \emph{triple-dyad ratio estimator}, have explicit expressions and can be scaled to very large networks with millions of nodes. We establish the consistency and asymptotic normality of the triple-dyad ratio estimator when the number of nodes reaches infinity. Based on the asymptotic results, we develop a test statistic for evaluating whether is a reciprocity effect in directed networks. The estimators for the density and reciprocity parameters contain bias terms, where analytical bias correction formulas are proposed to make valid inference. Numerical studies demonstrate the findings of our theories and show that the estimator is comparable to the MLE in large networks.

stat.ME

Optimal estimators and tests for reciprocal effects

The $p_1$ model plays a fundamental role in modeling directed networks, where the reciprocal effect parameter $\rho$ is of special interest in practice. However, due to nonlinear factors in this model, how to estimate $\rho$ efficiently is a long-standing open problem. We tackle the problem by the cycle count approach. The challenge is, due to the nonlinear factors in the model, for any given type of generalized cycles, the expected count is a complicated function of many parameters in the model, so it is unclear how to use cycle counts to estimate $\rho$. However, somewhat surprisingly, we discover that, among many types of generalized cycles with the same length, we can carefully pick a pair of them such that in the ratio between the expected cycle counts of the two types, the non-linear factors cancel out nicely with each other, and as a result, the ratio equals to $\mathrm{exp}(\rho)$ exactly. Therefore, though the expected count of cycles of any type is not tractable, the ratio between the expected cycle counts of a (carefully chosen) pair of generalized cycles may have an utterly simple form. We study to what extent such pairs exist, and use our discovery to derive both an estimate for $\rho$ and a testing procedure for testing $\rho = \rho_0$. In a setting where we allow a wide range of reciprocal effects and a wide variety of network sparsity and degree heterogeneity, we show that our estimator achieves the optimal rate and our test achieves the optimal phase transition. Technically, first, motivated by what we observe on real networks, we do not want to impose strong conditions on reciprocal effects, network sparsity, and degree heterogeneity. Second, our proposed statistic is a type of $U$-statistic, the analysis of which involves complex combinatorics and is error-prone. For these reasons, our analysis is long and delicate.

math.ST

Inference in the $p_0$ model for directed networks under local differential privacy

We explore the edge-flipping mechanism, a type of input perturbation, to release the directed graph under edge-local differential privacy. By using the noisy bi-degree sequence from the output graph, we construct the moment equations to estimate the unknown parameters in the $p_0$ model, which is an exponential family distribution with the bi-degree sequence as the natural sufficient statistic. We show that the resulting private estimator is asymptotically consistent and normally distributed under some conditions. In addition, we compare the performance of input and output perturbation mechanisms for releasing bi-degree sequences in terms of parameter estimation accuracy and privacy protection. Numerical studies demonstrate our theoretical findings and compare the performance of the private estimates obtained by different types of perturbation methods. We apply the proposed method to analyze the UC Irvine message network.

math.ST

Moment estimation in paired comparison models with a growing number of subjects

When the number of subjects, $n$, is large, paired comparisons are often sparse. Here, we study statistical inference in a class of paired comparison models parameterized by a set of merit parameters, under an Erd\"{o}s--R\'{e}nyi comparison graph, where the sparsity is measured by a probability $p_n$ tending to zero. We use the moment estimation base on the scores of subjects to infer the merit parameters. We establish a unified theoretical framework in which the uniform consistency and asymptotic normality of the moment estimator hold as the number of subjects goes to infinity. A key idea for the proof of the consistency is that we obtain the convergence rate of the Newton iterative sequence for solving the estimator. We use the Thurstone model to illustrate the unified theoretical results. Further extensions to a fixed sparse comparison graph are also provided. Numerical studies and a real data analysis illustrate our theoretical findings.

math.ST

Inference in a generalized Bradley-Terry model for paired comparisons with covariates and a growing number of subjects

Motivated by the home-field advantage in sports, we propose a generalized Bradley--Terry model that incorporates covariate information for paired comparisons. It has an $n$-dimensional merit parameter $\bs{\beta}$ and a fixed-dimensional regression coefficient $\bs{\gamma}$ for covariates. When the number of subjects $n$ approaches infinity and the number of comparisons between any two subjects is fixed, we show the uniform consistency of the maximum likelihood estimator (MLE) $(\widehat{\bs{\beta}}, \widehat{\bs{\gamma}})$ of $(\bs{\beta}, \bs{\gamma})$ Furthermore, we derive the asymptotic normal distribution of the MLE by characterizing its asymptotic representation. The asymptotic distribution of $\widehat{\bs{\gamma}}$ is biased, while that of $\widehat{\bs{\beta}}$ is not. This phenomenon can be attributed to the different convergence rates of $\widehat{\bs{\gamma}}$ and $\widehat{\bs{\beta}}$. To the best of our knowledge, this is the first study to explore the asymptotic theory in paired comparison models with covariates in a high-dimensional setting. The consistency result is further extended to an Erd\H{o}s--R\'{e}nyi comparison graph with a diverging number of covariates. Numerical studies and a real data analysis demonstrate our theoretical findings.

stat.ME

Temporal network analysis via a degree-corrected Cox model

Temporal dynamics, characterised by time-varying degree heterogeneity and homophily effects, are often exhibited in many real-world networks. As observed in an MIT Social Evolution study, the in-degree and out-degree of the nodes show considerable heterogeneity that varies with time. Concurrently, homophily effects, which explain why nodes with similar characteristics are more likely to connect with each other, are also time-dependent. To facilitate the exploration and understanding of these dynamics, we propose a novel degree-corrected Cox model for directed networks, where the way for degree-heterogeneity or homophily effects to change with time is left completely unspecified. Because each node has individual-specific in- and out-degree parameters that vary over time, the number of unknown parameters grows with the number of nodes, leading to a high-dimensional estimation problem. Therefore, it is highly nontrivial to make inference. We develop a local estimating equations approach to estimate the unknown parameters and establish the consistency and asymptotic normality of the proposed estimators in the high-dimensional regime. We further propose test statistics to check whether temporal variation or degree heterogeneity is present in the network and develop a graphically diagnostic method to evaluate goodness-of-fit for dynamic network models. Simulation studies and two real data analyses are provided to assess the finite sample performance of the proposed method and illustrate its practical utility.

stat.ME

3D surface profiling via photonic integrated geometric sensor

Measurements of microscale surface patterns are essential for process and quality control in industries across semiconductors, micro-machining, and biomedicines. However, the development of miniaturized and intelligent profiling systems remains a longstanding challenge, primarily due to the complexity and bulkiness of existing benchtop systems required to scan large-area samples. A real-time, in-situ, and fast detection alternative is therefore highly desirable for predicting surface topography on the fly. In this paper, we present an ultracompact geometric profiler based on photonic integrated circuits, which directly encodes the optical reflectance of the sample and decodes it with a neural network. This platform is free of complex interferometric configurations and avoids time-consuming nonlinear fitting algorithms. We show that a silicon programmable circuit can generate pseudo-random kernels to project input data into higher dimensions, enabling efficient feature extraction via a lightweight one-dimensional convolutional neural network. Our device is capable of high-fidelity, fast-scanning-rate thickness identification for both smoothly varying samples and intricate 3D printed emblem structures, paving the way for a new class of compact geometric sensors.

physics.optics

Testing degree heterogeneity in directed networks

In this study, we focus on the likelihood ratio tests in the $p_0$ model for testing degree heterogeneity in directed networks, which is an exponential family distribution on directed graphs with the bi-degree sequence as the naturally sufficient statistic. For testing the homogeneous null hypotheses $H_0: \alpha_1 = \cdots = \alpha_r$, we establish Wilks-type results in both increasing-dimensional and fixed-dimensional settings. For increasing dimensions, the normalized log-likelihood ratio statistic $[2\{\ell(\widehat{\mathbf{\theta}})-\ell(\widehat{\mathbf{\theta}}^0)\}-r]/(2r)^{1/2}$ converges in distribution to a standard normal distribution. For fixed dimensions, $2\{\ell(\widehat{\mathbf{\theta}})-\ell(\widehat{\mathbf{\theta}}^0)\}$ converges in distribution to a chi-square distribution with $r-1$ degrees of freedom as $n\rightarrow \infty$, independent of the nuisance parameters. Additionally, we present a Wilks-type theorem for the specified null $H_0: \alpha_i=\alpha_i^0$, $i=1,\ldots, r$ in high-dimensional settings, where the normalized log-likelihood ratio statistic also converges in distribution to a standard normal distribution. These results extend the work of \cite{yan2025likelihood} to directed graphs in a highly non-trivial way, where we need to analyze much more expansion terms in the fourth-order asymptotic expansions of the likelihood function and develop new approximate inverse matrices under the null restricted parameter spaces for approximating the inverse of the Fisher information matrices in the $p_0$ model. Simulation studies and real data analyses are presented to verify our theoretical results.

math.ST

Optical Convolutional Spectrometer

Optical spectrometers are fundamental across numerous disciplines in science and technology. However, miniaturized versions, while essential for in situ measurements, are often restricted to coarse identification of signature peaks and inadequate for metrological purposes. Here, we introduce a new class of spectrometer, leveraging the convolution theorem as its mathematical foundation. Our convolutional spectrometer offers unmatched performance for miniaturized systems and distinct structural and computational simplicity, featuring a centimeter-scale footprint for the fully packaged unit, low cost (~$10) and a 2400 cm-1 (approximately 500 nm) bandwidth. We achieve excellent precision in resolving complex spectra with sub-second sampling and processing time, demonstrating a wide range of applications from industrial and agricultural analysis to healthcare monitoring. Specifically, our spectrometer system classifies diverse solid samples, including plastics, pharmaceuticals, coffee, flour and tea, with 100% success rate, and quantifies concentrations of aqueous and organic solutions with detection accuracy surpassing commercial benchtop spectrometers. We also realize the non-invasive sensing of human biomarkers, such as skin moisture (mean absolute error; MAE = 2.49%), blood alcohol (1.70 mg/dL), blood lactate (0.81 mmol/L), and blood glucose (0.36 mmol/L), highlighting the potential of this new class of spectrometers for low-cost, high-precision, portable/wearable spectral metrology.

physics.optics

Maximum likelihood estimation in the sparse Rasch model

The Rasch model has been widely used to analyse item response data in psychometrics and educational assessments. When the number of individuals and items are large, it may be impractical to provide all possible responses. It is desirable to study sparse item response experiments. Here, we propose to use the Erd\H{o}s\textendash R\'enyi random sampling design, where an individual responds to an item with low probability $p$. We prove the uniform consistency of the maximum likelihood estimator %by developing a leave-one-out method for the Rasch model when both the number of individuals, $r$, and the number of items, $t$, approach infinity. Sampling probability $p$ can be as small as $\max\{\log r/r, \log t/t\}$ up to a constant factor, which is a fundamental requirement to guarantee the connection of the sampling graph by the theory of the Erd\H{o}s\textendash R\'enyi graph. The key technique behind this significant advancement is a powerful leave-one-out method for the Rasch model. We further establish the asymptotical normality of the MLE by using a simple matrix to approximate the inverse of the Fisher information matrix. The theoretical results are corroborated by simulation studies and an analysis of a large item-response dataset.

math.ST

Chip-scale sensor for spectroscopic metrology

Miniaturized spectrometers hold great promise for in situ, in vitro, and even in vivo sensing applications. However, their size reduction imposes vital performance constraints in meeting the rigorous demands of spectroscopy, including fine resolution, high accuracy, and ultra-wide observation window. The prevailing view in the community holds that miniaturized spectrometers are most suitable for the coarse identification of signature peaks. In this paper, we present an integrated reconstructive spectrometer that enables near-infrared (NIR) spectroscopic metrology, and demonstrate a fully packaged sensor with auxiliary electronics. Such a sensor operates over a 520 nm bandwidth together with a resolution of less than 8 pm, which translates into a record-breaking bandwidth-to-resolution ratio of over 65,000. The classification of different types of solid substances and the concentration measurement of aqueous and organic solutions are performed, all achieving approximately 100% accuracy. Notably, the detection limit of our sensor matches that of the commercial benchtop counterparts, which is as low as 0.1% (i.e. 100 mg/dL) for identifying the concentration of glucose solution.

physics.optics

Inference in semiparametric formation models for directed networks

We propose a semiparametric model for dyadic link formations in directed networks. The model contains a set of degree parameters that measure different effects of popularity or outgoingness across nodes, a regression parameter vector that reflects the homophily effect resulting from the nodal attributes or pairwise covariates associated with edges, and a set of latent random noises with unknown distributions. Our interest lies in inferring the unknown degree parameters and homophily parameters. The dimension of the degree parameters increases with the number of nodes. Under the high-dimensional regime, we develop a kernel-based least squares approach to estimate the unknown parameters. The major advantage of our estimator is that it does not encounter the incidental parameter problem for the homophily parameters. We prove consistency of all the resulting estimators of the degree parameters and homophily parameters. We establish high-dimensional central limit theorems for the proposed estimators and provide several applications of our general theory, including testing the existence of degree heterogeneity, testing sparse signals and recovering the support. Simulation studies and a real data application are conducted to illustrate the finite sample performance of the proposed methods.

stat.ME

Asymmetrical estimator for training encapsulated deep photonic neural networks

Photonic neural networks (PNNs) are fast in-propagation and high bandwidth paradigms that aim to popularize reproducible NN acceleration with higher efficiency and lower cost. However, the training of PNN is known to be challenging, where the device-to-device and system-to-system variations create imperfect knowledge of the PNN. Despite backpropagation (BP)-based training algorithms being the industry standard for their robustness, generality, and fast gradient convergence for digital training, existing PNN-BP methods rely heavily on accurate intermediate state extraction or extensive computational resources for deep PNNs (DPNNs). The truncated photonic signal propagation and the computation overhead bottleneck DPNN's operation efficiency and increase system construction cost. Here, we introduce the asymmetrical training (AsyT) method, tailored for encapsulated DPNNs, where the signal is preserved in the analogue photonic domain for the entire structure. AsyT offers a lightweight solution for DPNNs with minimum readouts, fast and energy-efficient operation, and minimum system footprint. AsyT's ease of operation, error tolerance, and generality aim to promote PNN acceleration in a widened operational scenario despite the fabrication variations and imperfect controls. We demonstrated AsyT for encapsulated DPNN with integrated photonic chips, repeatably enhancing the performance from in-silico BP for different network structures and datasets.

cs.LG

Differentially private analysis of networks with covariates via a generalized $\beta$-model

How to achieve the tradeoff between privacy and utility is one of fundamental problems in private data analysis.In this paper, we give a rigourous differential privacy analysis of networks in the appearance of covariates via a generalized $\beta$-model, which has an $n$-dimensional degree parameter $\beta$ and a $p$-dimensional homophily parameter $\gamma$.Under $(k_n, \epsilon_n)$-edge differential privacy, we use the popular Laplace mechanism to release the network statistics.The method of moments is used to estimate the unknown model parameters. We establish the conditions guaranteeing consistency of the differentially private estimators $\widehat{\beta}$ and $\widehat{\gamma}$ as the number of nodes $n$ goes to infinity, which reveal an interesting tradeoff between a privacy parameter and model parameters. The consistency is shown by applying a two-stage Newton's method to obtain the upper bound of the error between $(\widehat{\beta},\widehat{\gamma})$ and its true value $(\beta, \gamma)$ in terms of the $\ell_\infty$ distance, which has a convergence rate of rough order $1/n^{1/2}$ for $\widehat{\beta}$ and $1/n$ for $\widehat{\gamma}$, respectively. Further, we derive the asymptotic normalities of $\widehat{\beta}$ and $\widehat{\gamma}$, whose asymptotic variances are the same as those of the non-private estimators under some conditions. Our paper sheds light on how to explore asymptotic theory under differential privacy in a principled manner; these principled methods should be applicable to a class of network models with covariates beyond the generalized $\beta$-model. Numerical studies and a real data analysis demonstrate our theoretical findings.

stat.ME