SearcharxivSearch

arXiv subjects

Ronald Sielinski

Publications and source records attributed to Ronald Sielinski.

3 recordsLinked to original sources

From Stochastic to Stable: Rank Stability and Structural Sufficiency in AI Visibility Measurement

AI visibility measurement is comparative: practitioners want to know which domains generative search engines cite most often and whether observed differences are large enough to support decisions. Yet the industry lacks a principled way to determine whether enough data has been collected. Collection budgets vary widely across studies and platforms, and conclusions are often drawn from rankings whose stability and precision are unknown. We introduce a sequential convergence framework based on two complementary criteria: rank stability evaluates whether the rank-correlation trajectory has reached a structural plateau, while structural sufficiency evaluates whether the spread of citation shares among established domains -- those whose confidence intervals exclude zero -- exceeds the uncertainty of those estimates. Together, these criteria distinguish rankings that have merely stabilized from those sufficiently resolved to support inference. Both are derived from regularities in the observed citation distribution, including its rank structure, uncertainty profile, and the boundary between observed and established domains. The framework retains a small number of structural constants but requires no externally specified query count, correlation target, or confidence-interval width target; stopping is driven by observed measurement uncertainty and remains robust across a range of sufficiency thresholds. Applied across 30 platform-topic combinations spanning Gemini, SearchGPT, and Perplexity, the framework adapts to platform- and topic-specific citation distributions. Results show that no fixed collection budget can be justified across contexts and that convergence can instead be evaluated from the structure of the observed distribution. The framework provides a practical basis for determining when AI visibility measurements are ready to support comparative analysis.

stat.AP

Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement

AI-powered answer engines are inherently non-deterministic: identical queries submitted at different times can produce different responses and cite different sources. Despite this stochastic behavior, current approaches to measuring domain visibility in generative search typically rely on single-run point estimates of citation share and prevalence, implicitly treating them as fixed values. This paper argues that citation visibility metrics should be treated as sample estimators of an underlying response distribution rather than fixed values. We conduct an empirical study of citation variability across three generative search platforms--Perplexity Search, OpenAI SearchGPT, and Google Gemini--using repeated sampling across three consumer product topics. Two sampling regimes are employed: daily collections over nine days and high-frequency sampling at ten-minute intervals. We show that citation distributions follow a power-law form and exhibit substantial variability across repeated samples. Bootstrap confidence intervals reveal that many apparent differences between domains fall within the noise floor of the measurement process. Distribution-wide rank stability analysis further demonstrates that citation rankings are unstable across samples, not only among top-ranked domains but throughout the frequently cited domain set. These findings demonstrate that single-run visibility metrics provide a misleadingly precise picture of domain performance in generative search. We argue that citation visibility must be reported with uncertainty estimates and provide practical guidance for sample sizes required to achieve interpretable confidence intervals.

stat.AP

The BAD Paradox: A Critical Assessment of the Belin/Ambrósio Deviation Model

The Belin/Ambrósio Deviation (BAD) model is a widely used diagnostic tool for detecting keratoconus and corneal ectasia. The input to the model is a set of z-score normalized $D$ indices that represent physical characteristics of the cornea. Paradoxically, the output of the model, Total Deviation Value ($D_{\text{final}}$), is reported in standard deviations from the mean, but $D_{\text{final}}$ does not behave like a z-score normalized value. Although thresholds like $D_{\text{final}} \ge 1.6$ for "suspicious" and $D_{\text{final}} \ge 3.0$ for "abnormal" are commonly cited, there is little explanation on how to interpret values outside of those thresholds or to understand how they relate to physical characteristics of the cornea. This study explores the reasons for $D_{\text{final}}$'s apparent inconsistency through a meta-analysis of published data and a more detailed statistical analysis of over 1,600 Pentacam exams. The results reveal that systematic bias in the BAD regression model, multicollinearity among predictors, and inconsistencies in normative datasets contribute to the non-zero mean of $D_{\text{final}}$, complicating its clinical interpretation. These findings highlight critical limitations in the model's design and underscore the need for recalibration to enhance its transparency and diagnostic reliability.

stat.AP