SearcharxivSearch

arXiv subjects

Regina Ruane

Publications and source records attributed to Regina Ruane.

At least 19 recordsLinked to original sources

Identifiability of Latent Space Network Models on Anisotropic Thurston Geometries

A latent space network model places the nodes in a metric space and lets the probability of a tie decrease with distance. In a space of constant curvature, pairwise distances determine the positions up to an isometry. In the products and in the three remaining three-dimensional model geometries they do not. We study the three geometries that are neither of constant curvature nor products: the Heisenberg group, the solvable group and the universal cover of the unit tangent bundle of the hyperbolic plane. We ask what one network identifies about positions in them and when their geometry is detectable. Two anchors remove the isometry ambiguity. Small configurations are not determined by their distances, and generic local identification holds beyond a finite threshold, certified in the Heisenberg group. We derive the posterior on the quotient by the isometry group. For small configurations, the divergence to the nearest product or constant-curvature competitor is the stress component of the curvature difference and vanishes at high order in the scale; undirected ties therefore detect the geometry only in large networks with large enclosed areas. Directed ties expose it at first order: in the two twisted geometries, asymmetric preferences circulate around triangles in proportion to enclosed area, which no additive ranking produces. On dense competitive-game counter networks, the coupled model beats rankings on every geometry, degree-corrected rankings and free antisymmetric terms. A Euclidean model with the same coupled term matches it, so the gain is the coupling of similarity and circulation through shared coordinates. An additive-and-multiplicative-effects model predicts better still.

physics.soc-ph

Replicable Conformal Prediction

Two analysts who calibrate the same predictive model on independent samples will deploy different prediction sets every time, because the calibration threshold inherits the randomness of the data. Wherever deployments must be audited, cached, or approved across sites, this instability is costly: no one can verify that two calibrations produced the same object. We ask two questions: when can independent calibrations yield the identical classifier, and what must that agreement cost? Perfect agreement is impossible, since a procedure that almost always returns one fixed answer cannot remain valid for every distribution, and exact agreement through shared randomness forces the procedure to ignore its data. Sharing a single random seed and rounding the calibrated threshold up to a coarse shared grid resolves the tension: the deployed classifier becomes identical across analysts with any desired probability, coverage guarantees survive, and the price is a quantified increase in set size and calibration data. Matching lower bounds show that no threshold method can pay less, and the method's one tuning constant vanishes asymptotically. Without any shared seed, a fixed grid still confines all analysts to two adjacent classifiers, and no method does better. Replicability also blocks gaming: selecting the most favorable of many recalibrations barely moves a replicable classifier, while the same selection silently undercovers standard conformal prediction. Experiments on real ImageNet outputs, a four-hospital site split, and four language-model families match the theory, including the measured sample-cost frontier.

stat.ML

Separating Time-Varying Network Composition from Predictive Dependence under Noisy Network Measurement

A common question about networked time series is whether outcomes changed because shocks transmit more strongly or because the pattern of connections changed. Standard practice inserts a recorded network into an outcome regression and reads movements of the fitted coefficient as changes in transmission strength. When the network is latent, time varying, and measured with error, this reading fails: changes in strength and changes in composition can produce the same outcome distribution at a single date, and the population coefficient moves under composition changes alone. The question becomes answerable when outcomes are analyzed jointly with repeated noisy measurements of the network, such as paired reports of bilateral trade flows. For the joint model we establish necessary and sufficient conditions for local identification, estimators of the strength and composition paths, a simultaneous confidence band for the strength path, confidence sets that remain exact under weak identification, breakdown bounds under common reporting bias, and an exactly sized test that detects changes on the observed path and attributes them to strength or to composition. Simulations assess each procedure at its stated boundary. On a mirror-reported trade panel of eighteen economies over 1995 to 2020, the diagnostics flag exactly the crisis years and the composition coordinate attached to European Union membership declines by roughly two thirds. The estimand is predictive dependence, not a causal effect.

stat.ME

Which Question Is Your Attention Metric Answering? Attention Rows as Compositional Data

Each row of a transformer's attention matrix is a probability distribution over tokens, and in trained models most of that probability lands on a single \emph{sink} token, usually the first. Standard tools for comparing attention rows (cosine similarity, Jensen--Shannon divergence, Shannon entropy) therefore hinge on a choice papers rarely report: keep the sink, or drop it and renormalize. This choice can reverse conclusions. On ten pretrained models from five families, 17--47% of verdicts about which of two heads is more similar flip with the convention, and the most prominent structure in a standard BERT head-clustering pipeline is an artifact of it. The reason is that one-number summaries mix two questions: how much attention the sink takes, and how the rest is divided among the content tokens. Treating rows as compositional data separates them exactly: the Aitchison distance splits orthogonally into a sink term and a content term, entropy splits by an exact identity, and the content distance is characterized by invariances the transformer itself possesses. The separation matters in practice: most measured entropy collapse during training is the sink growing, not attention sharpening (30% of the drop at 70M parameters, 95% at 1B, 79% at 1.4B), and pruning heads with the wrong channel can inflate perplexity more than a hundredfold. We map where each convention is safe, test a frozen out-of-sample predictor (one confirmation, one abstention, one failure), and release code regenerating every number.

cs.CL

Bayesian Predictive Synthesis for Dynamic Networks: Forecasting and Identifying Structural Mechanisms

Networks are shaped by competing structural mechanisms, such as communities, geometry, or hubs. In a dynamic network the most predictive mechanism can change, and a model tied to one mechanism, or to fixed weights, cannot adapt as the dominant structure shifts. We develop dynamic Bayesian predictive synthesis for networks, in which a mechanism is an agent forecasting the next snapshot's edges and a synthesis layer combines them with time-varying weights. At each step the method returns a calibrated edge forecast and inference on the mechanism weights, with intervals valid given the fitted agents, so it also reports which mechanism is most informative. Inference of this kind requires a sparse-safe parametrization and an identification theory, under which a single graph identifies and estimates the weights. A sharp threshold separates distinguishable from indistinguishable mechanisms, a change in the active mechanism is tracked at an optimal per-switch cost, and for a single snapshot the method reduces to calibrated link prediction. On real networks, simulations, and benchmarks, the synthesis gives accurate, calibrated forecasts and recovers the leading mechanism when

cs.SI

Minimax Synthesis of Network Mechanisms

A single observed network reflects several mechanisms at once: communities, hubs, and clustering coexist in one graph, each a different model. We treat the network as a combination of candidate mechanisms and study, from a single graph, how strongly each mechanism contributes and how they combine. We address two questions. The first is how to measure each mechanism's contribution when the mechanisms must themselves be estimated from the graph: fitting the mechanisms and their strengths from the same data biases the strengths toward zero, and a correction removes this bias and yields valid confidence intervals. The second is whether the rule of combination is itself recoverable: when a graph is generated by two mechanisms acting together, the graph alone determines whether they combine additively or interact, exactly when the graph is dense enough, a sharp threshold below which no test can decide. The estimate calibrates the candidate mechanisms against the observed edges. We establish matching minimax rate, against a known-design benchmark and the estimated-design problem itself, confirm the methods in simulation, and apply them to real networks, where the signed coefficients recover known structure and, in one case, a confidence interval excludes any positive contribution from a candidate mechanism.

math.ST

How Eviction Court Governs: A Statistical Analysis of Bargaining, Templates, and Debt in Philadelphia

We analyze downstream courtroom governance in Philadelphia eviction cases using 755,004 Municipal Court landlord--tenant records filed from 1969 through 2022. Post-filing case processing is organized by repeated courtroom relationships, judge and tenant-attorney regimes, reusable agreement templates, and repeated team-property units. Among both-represented, both-attorney-named cases, 58.2% involve a plaintiff-side and tenant-side attorney pair that had appeared against one another in the prior year, and greater prior pair exposure predicts lower default, higher judgment-by-agreement, and higher served-writ rates. Judge-linked cases display statistically distinct baseline outcome, continuance, fee, and award regimes; tenant-attorney identity explains meaningful variance in both case outcomes and agreement terms. Settlement text is highly standardized: reusable templates explain strictness, waiver, lockout-trigger, payment-plan, deadline, and time-is-essence language far more strongly than raw attorney identity. Monetary burden concentrates in repeated plaintiff-attorney-property units. Assignment-cell support and balance audits indicate that judge-linked evidence reflects institutional heterogeneity rather than a clean judge lottery, and judge--triad interactions are not estimable in this docket. Eviction court emerges as a repeated institutional field that organizes bargaining, text, debt, and enforcement after cases enter the courtroom pipeline.

stat.AP

High-Volume Plaintiff-Side Counsel and Single-Appearance Eviction Cases in Philadelphia

Among 755,004 Philadelphia landlord--tenant records filed during 1969-2022, 396,163 residential cases involve tenants who appear exactly once in the observed docket. In unadjusted comparisons, single-appearance cases handled by high-volume plaintiff-side counsel are more likely to advance to the writ-of-possession and served-writ stages, but no more likely to end in default. Comparisons within the same plaintiff, and within the same plaintiff at the same property, show no broad premium on adverse case outcomes such as default, judgment, or fees. The clearer pattern is organizational: after a plaintiff adopts or switches into high-volume counsel, monthly filings rise by about 2-5% and the number of distinct buildings reached rises by a similar margin; near the prior-year top-10 attorney threshold, cases display local differences in default and enforcement; and continuances under specialist counsel are more closely linked to default. Non-flat pre-treatment trends and imprecise reverse-direction estimates from attorney exits restrict the strength of any causal claim. High-volume plaintiff-side counsel therefore functions as a mechanism of filing scale and procedural sequence, not as a uniform escalator of case outcomes or as a cause of any individual tenant becoming single-appearance.

stat.AP

Support-Safe Variational Hybrid Filtering for Contact-Mode and Sparse-Law Recovery

Contact-rich robot dynamics are hybrid: a single observation can match several latent states and contact regimes (free, impact, stick--slip). A standard amortized filter that places no probability on a feasible contact transition will permanently lose the branch the robot actually follows. We introduce VHYDRO, a variational hybrid dynamics learner that prevents this branch loss. At each step, VHYDRO mixes the learned proposal with a feasible transition law before sampling and importance weighting, ensuring that every transition retained by the model-feasible carrier remains covered. VHYDRO jointly infers a continuous latent state and a discrete contact mode, and fits a sparse port-Hamiltonian law to each recovered regime. On top of this, three guarantees connect: support coverage stabilizes filtering, the stabilized filter concentrates the discrete contact posterior on coherent regimes, and mode-pure segments admit sparse port-Hamiltonian recovery. The recovery error separates cleanly into filtering, derivative, mode-impurity, and physics-residual parts. Three empirical findings track the same mechanism. Under heavy occlusion the support-safe filter stays usable while a non-defensive proposal collapses. On ManiSkill demonstrations and on four Sawyer/BridgeData task families the discrete state forms temporally coherent contact-regime segments that the discrete state yields a stronger joint profile across ARI, change-point F1, and segment purity than post-hoc and mode-free baselines. On hybrid systems with known equations the mode-conditioned sparse fit recovers the active physical terms; purely predictive baselines do not.

cs.RO

Legal Infrastructure Organizes Eviction: Evidence from Philadelphia

We analyze the filing-side legal infrastructure of eviction using 755,004 Philadelphia Municipal Court landlord-tenant records filed between 1969 and 2022, of which 747,125 are residential. Eviction in Philadelphia is organized upstream by a concentrated plaintiff-side bar, durable plaintiff-attorney dependence, repeated use of the same properties, and recurring tenant-name exposure. Between 1983 and 2022, the ten most active plaintiff attorneys handled 82.2% of represented plaintiff-side cases per year on average, compared with 14.8% for the ten most active plaintiffs. Large plaintiffs depend heavily on a single attorney: among plaintiffs filing at least 101 cases, 78.3% of each plaintiff's filings are handled by that plaintiff's most-used attorney, on average. Repetition is likewise central to the docket. Across the residential filing universe, 48.8% of cases occur at addresses with a prior filing in the preceding year, and 23.6% at addresses with six or more prior filings; these repeats are usually filed by the same plaintiff and follow a more default-heavy, less agreement-heavy pathway. We further examine a narrower mechanism: strict switches into specialist plaintiff-side counsel, defined as a plaintiff changing attorney to one in the prior-year top ten. Filing counts rise around the switch with non-flat pre-trends, indicating organizational reconfiguration rather than a clean exogenous shock. Within-plaintiff and within-plaintiff-property comparisons yield more stable estimates: judgment by agreement, fee share, waiver language, and corrected lockout-trigger language decline, while deadline language rises. We interpret eviction as a layered upstream process in which concentrated counsel, repeated places, and recurring tenants produce filings before any courtroom bargaining or adjudication occurs.

stat.AP

Effects of Training Data Quality on Classifier Performance

We describe extensive numerical experiments assessing and quantifying how classifier performance depends on the quality of the training data, a frequently neglected component of the analysis of classifiers. More specifically, in the scientific context of metagenomic assembly of short DNA reads into "contigs," we examine the effects of degrading the quality of the training data by multiple mechanisms, and for four classifiers -- Bayes classifiers, neural nets, partition models and random forests. We investigate both individual behavior and congruence among the classifiers. We find breakdown-like behavior that holds for all four classifiers, as degradation increases and they move from being mostly correct to only coincidentally correct, because they are wrong in the same way. In the process, a picture of spatial heterogeneity emerges: as the training data move farther from analysis data, classifier decisions degenerate, the boundary becomes less dense, and congruence increases.

cs.LG

Decision-Theoretic Robustness for Network Models

Bayesian network models (Erdos Renyi, stochastic block models, random dot product graphs, graphons) are widely used in neuroscience, epidemiology, and the social sciences, yet real networks are sparse, heterogeneous, and exhibit higher-order dependence. How stable are network-based decisions, model selection, and policy recommendations to small model misspecification? We study local decision-theoretic robustness by allowing the posterior to vary within a small Kullback-Leibler neighborhood and choosing actions that minimize worst-case posterior expected loss. Exploiting low-dimensional functionals available under exchangeability, we (i) adapt decision-theoretic robustness to exchangeable graphs via graphon limits and derive sharp small-radius expansions of robust posterior risk; under squared loss the leading inflation is controlled by the posterior variance of the loss, and for robustness indices that diverge at percolation/fragmentation thresholds we obtain a universal critical exponent describing the explosion of decision uncertainty near criticality. (ii) Develop a nonparametric minimax theory for robust model selection between sparse Erdos-Renyi and block models, showing-via robustness error exponents-that no Bayesian or frequentist method can uniformly improve upon the decision-theoretic limits over configuration models and sparse graphon classes for percolation-type functionals. (iii) Propose a practical algorithm based on entropic tilting of posterior or variational samples, and demonstrate it on functional brain connectivity and Karnataka village social networks.

math.ST

Collapsed Structured Block Models for Community Detection in Complex Networks

Community detection seeks to recover mesoscopic structure from network data that may be binary, count-valued, signed, directed, weighted, or multilayer. The stochastic block model (SBM) explains such structure by positing a latent partition of nodes and block-specific edge distributions. In Bayesian SBMs, standard MCMC alternates between updating the partition and sampling block parameters, which can hinder mixing and complicate principled comparison across different partitions and numbers of communities. We develop a collapsed Bayesian SBM framework in which block-specific nuisance parameters are analytically integrated out under conjugate priors, so the marginal likelihood p(Y|z) depends only on the partition z and blockwise sufficient statistics. This yields fast local Gibbs/Metropolis updates based on ratios of closed-form integrated likelihoods and provides evidence-based complexity control that discourages gratuitous over-partitioning. We derive exact collapsed marginals for the most common SBM edge types-Beta-Bernoulli (binary), Gamma-Poisson (counts), and Normal-Inverse-Gamma (Gaussian weights)-and we extend collapsing to gap-constrained SBMs via truncated conjugate priors that enforce explicit upper bounds on between-community connectivity. We further show that the same collapsed strategy supports directed SBMs that model reciprocity through dyad states, signed SBMs via categorical block models, and multiplex SBMs where multiple layers contribute additive evidence for a shared partition. Across synthetic benchmarks and real networks (including email communication, hospital contact counts, and citation graphs), collapsed inference produces accurate partitions and interpretable posterior block summaries of within- and between-community interaction strengths while remaining computationally simple and modular.

math.ST

Decomposing Degree Assortativity in Sparse Spatial Networks

Spatial networks are typically assortative: well-connected nodes link to other well-connected nodes, and the usual reading is sorting, popular nodes seeking each other out. In space there is a rival explanation: nearby nodes draw on the same pool of potential neighbors, so their degrees move together even when popularity and location are unrelated. This paper asks when the two can be told apart from network data, and gives a three-part answer. First, in a sparse random-connection model in which latent popularity and position are coupled by a copula, we prove that the limiting assortativity splits exactly into an intensity-covariance channel and a shared-neighbor channel; because the first mixes sorting with density effects, we define the sorting contribution as the increment produced by the dependence when mean degree and transitivity are held fixed, exactly zero at independence. Second, we prove that three observable summaries (mean degree, transitivity, assortativity) identify the model globally, including on the boundary of no dependence, by a computer-assisted proof in exact rational arithmetic that also verifies the conditions for valid inference; simulations confirm the asymptotics. Third, on data the machinery acts as a gate: no sorting estimate is reported unless the model first fits. In eleven metropolitan location-based networks the model class is rejected and direct estimates find essentially no dependence; the observed assortativity is evidence of neither sorting nor shared opportunity. In a national co-authorship network the verdict reverses: productive authors concentrate where researcher density is high, yet graph-only attribution is again refused, the misfit pointing to the team structure of multi-author papers.

math.ST

State-Space Modeling of Time-Varying Spillovers on Networks

Crime counts in city neighbourhoods, disease counts in counties, and sales at firms joined by trade are naturally represented as counts on the nodes of a network. In each case a high count at one node can raise the counts at the nodes linked to it next period. The strength of that spillover changes over time, and standard network autoregressions hold it fixed. We therefore use a network state-space model, in which the spillover is a coefficient that drifts and a filter estimates its value in each period. What the data reveal about that coefficient depends on the network. It is learned by contrasting nodes whose neighbours have high counts with nodes whose neighbours have low ones. If every node is linked to every other, all nodes share the same neighbours, the contrasts vanish, and the spillover is not identified. Robustness is often checked by refitting with the links spread evenly, and a stable coefficient is read as reassurance. That refit is the same model with the spillover rescaled, so it cannot disagree. Forecasts carry a second warning: beyond two steps ahead, simulation averages a quantity with no finite mean, and the output gives no sign of it. We give a measure of what a network and a data set reveal about the spillover, the accuracy the filter can reach, and an exact test of whether the network matters. Burglaries in Chicago, COVID-19 cases in Texas counties, and measles cases in the Weser--Ems districts illustrate all three. On both disease datasets the model outperforms every competing forecast in the comparison.

stat.ME

Graphon-Level Bayesian Predictive Synthesis for Random Network

Bayesian predictive synthesis provides a coherent Bayesian framework for combining multiple predictive distributions, or agents, into a single updated prediction, extending Bayesian model averaging to allow general pooling of full predictive densities. This paper develops a static, graphon level version of Bayesian predictive synthesis for random networks. At the graphon level we show that Bayesian predictive synthesis corresponds to the integrated squared error projection of the true graphon onto the linear span of the agent graphons. We derive nonasymptotic oracle inequalities and prove that least-squares graphon-BPS, based on a finite number of edge observations, achieves the minimax L^2 rate over this agent span. Moreover, we show that any estimator that selects a single agent graphon is uniformly inconsistent on a nontrivial subset of the convex hull of the agents, whereas graphon-level Bayesian predictive synthesis remains minimax-rate optimal-formalizing a combination beats components phenomenon. Structural properties of the underlying random graphs are controlled through explicit Lipschitz bounds that transfer graphon error into error for edge density, degree distributions, subgraph densities, clustering coefficients, and giant component phase transitions. Finally, we develop a heavy tail theory for Bayesian predictive synthesis, showing how mixtures and entropic tilts preserve regularly varying degree distributions and how exponential random graph model agents remain within their family under log linear tilting with Kullback-Leibler optimal moment calibration.

math.ST

Wavelet Latent Position Exponential Random Graphs

Many network datasets exhibit connectivity with variance by resolution and large-scale organization that coexists with localized departures. When vertices have observed ordering or embedding, such as geography in spatial and village networks, or anatomical coordinates in connectomes, learning where and at what resolution connectivity departs from a baseline is crucial. Standard models typically emphasize a single representation, i.e. stochastic block models prioritize coarse partitions, latent space models prioritize global geometry, small-world generators capture local clustering with random shortcuts, and graphon formulations are fully general and do not solely supply a canonical multiresolution parameterization for interpretation and regularization. We introduce wavelet latent position exponential random graphs (WL-ERGs), an exchangeable logistic-graphon framework in which the log-odds connectivity kernel is represented in compactly supported orthonormal wavelet coordinates and mapped to edge probabilities through a logistic link. Wavelet coefficients are indexed by resolution and location, which allows multiscale structure to become sparse and directly interpretable. Although edges remain independent given latent coordinates, any finite truncation yields a conditional exponential family whose sufficient statistics are multiscale wavelet interaction counts and conditional laws admit a maximum-entropy characterization. These characteristics enable likelihood-based regularization and testing directly in coefficient space. The theory is naturally scale-resolved and includes universality for broad classes of logistic graphons, near-minimax estimation under multiscale sparsity, scale-indexed recovery and detection thresholds, and a band-limited regime in which canonical coefficient-space tilts are non-degenerate and satisfy a finite-dimensional large deviation principle.

math.ST

Radial Compensation: The Inverse Base-Distribution Problem for Chart-Based Generative Models on Riemannian Manifolds

Latent-variable models on spheres and hyperbolic spaces usually draw a Gaussian in the tangent space at a base point and push it onto the manifold. On these spaces the distance from the base point is the coordinate that carries meaning: depth in a hierarchy, the angle of a rotation, the deviation of a protein frame from a reference. We show that the standard construction silently replaces whatever distance distribution the modeler intended with a fixed one, a scaled chi law whose shape no setting of the scale can change. We then solve the reverse problem. Given the intended distance distribution, we derive in closed form the tangent density that realizes it, prove it is the only isotropic choice with chart-independent likelihoods for a broad class of charts, and prove a lower bound with explicit constants on what ignoring the problem costs a variational autoencoder. Experiments backed by an exact per-run normalization audit confirm that the compensated prior is invariant to the chart and stable across scales, every wrapped baseline we train collapses to the boundary of its chart, curvature becomes recoverable where the wrapped prior fails and protein-orientation likelihood improves from 2.58 to 0.87 nats at identical accuracy.

cs.LG