SearcharxivSearch

arXiv subjects

Sophie H. Yu

Publications and source records attributed to Sophie H. Yu.

10 recordsLinked to original sources

Flexibility allocation in random bipartite matching markets: exact matching rates and dominance regimes

This paper studies how a fixed flexibility budget should be allocated across the two sides of a balanced bipartite matching market. We model compatibilities via a sparse bipartite stochastic block model in which flexible agents are more likely to connect with agents on the opposite side, and derive an exact variational formula for the asymptotic matching rate under any flexibility allocation. The derivation extends the local weak convergence framework of [BLS11] from single-type to multi-type unimodular Galton-Watson trees, reducing the matching rate to an explicit low-dimensional optimization problem. Using this formula, we analytically investigate when the one-sided allocation, which concentrates all flexibility on one side, dominates the two-sided allocation and vice versa, sharpening and extending the comparisons of [FMZ26] which relied on approximate algorithmic bounds rather than an exact characterization of the matching rate.

math.PR

Detection of local geometry in random graphs: information-theoretic and computational limits

We study the problem of detecting local geometry in random graphs. We introduce a model $\mathcal{G}(n, p, d, k)$, where a hidden community of average size $k$ has edges drawn as a random geometric graph on $\mathbb{S}^{d-1}$, while all remaining edges follow the Erdős--Rényi model $\mathcal{G}(n, p)$. The random geometric graph is generated by thresholding inner products of latent vectors on $\mathbb{S}^{d-1}$, with each edge having marginal probability equal to $p$. This implies that $\mathcal{G}(n, p, d, k)$ and $\mathcal{G}(n, p)$ are indistinguishable at the level of the marginals, and the signal lies entirely in the edge dependencies induced by the local geometry. We investigate both the information-theoretic and computational limits of detection. On the information-theoretic side, our upper bounds follow from three tests based on signed triangle counts: a global test, a scan test, and a constrained scan test; our lower bounds follow from two complementary methods: truncated second moment via Wishart--GOE comparison, and tensorization of KL divergence. These results together settle the detection threshold at $d = \widetildeΘ(k^2 \vee k^6/n^3)$ for fixed $p$, and extend the state-of-the-art bounds from the full model (i.e., $k = n$) for vanishing $p$. On the computational side, we identify a computational--statistical gap and provide evidence via the low-degree polynomial framework, as well as the suboptimality of signed cycle counts of length $\ell \geq 4$.

math.ST

Availability is all you need: achieving optimal regret with minimal information for dynamic matching

We study a centralized discrete-time dynamic two-way matching model with finitely many agent types. Agents arrive stochastically over time and join their type-dedicated queues waiting to be matched. We focus on availability-based policies that make matching decisions based solely on agent availability across types (i.e., whether queues are empty or not), rather than relying on complete queue-length information (e.g., the longest-queue policy). We aim to achieve constant regret at all times with optimal scaling in terms of the general position gap, $ε$, which measures the distance of the fluid relaxation from degeneracy. We classify availability-based policies into global and local policies based on the scope of information they utilize. First, for general networks (possibly cyclic), we propose a global availability-based policy, probabilistic matching, and prove that it achieves the optimal all-time regret scaling of $O(ε^{-1})$, matching the known lower bound established by [KAG24]. Second, for acyclic networks, we focus on the class of local availability-based policies, specifically static priority policies that prioritize matches based on a fixed order. Within this class, we derive the first explicit regret bound for the previously proposed tree priority policy, showing all-time regret scaling of $O(ε^{-(d+1)/2})$, where $d$ is the network depth. Next, we introduce a new truncated tree priority policy and prove that it is the first static priority policy to achieve the optimal all-time regret scaling of $O(ε^{-1})$. These policies are appealing for matching systems such as queueing and load balancing; they reduce operational costs by using minimal information while effectively balancing the trade-off between immediate and future rewards.

cs.DS

A uniformity principle for spatial matching

Platforms matching spatially distributed supply to demand face a fundamental design choice: given a fixed total budget of service range, how should it be allocated across supply nodes ex ante, i.e. before supply and demand locations are realized, to maximize fulfilled demand? We model this problem using bipartite random geometric graphs where $n$ supply and $m$ demand nodes are uniformly distributed on $[0,1]^k$ ($k \ge 1$), and edges form when demand falls within a supply node's service region, the volume of which is determined by its service range. Since each supply node serves at most one demand, platform performance is determined by the expected size of a maximum matching. We establish a uniformity principle: whenever one service range allocation is more uniform than the other, the more uniform allocation yields a larger expected matching. This principle emerges from diminishing marginal returns to range expanding service range, and limited interference between supply nodes due to bounded ranges naturally fragmenting the graph. For $k=1$, we further characterize the expected matching size through a Markov chain embedding and derive closed-form expressions for special cases. Our results provide theoretical guidance for service-range allocation and incentive design in ride-hailing, on-demand labor markets, and drone delivery platforms, highlighting the benefits of reducing disparities in supply-side flexibility.

math.PR

Online Metric Matching: Beyond the Worst Case

We study the online metric matching problem. There are $m$ servers and $n$ requests located in a metric space, where all servers are available upfront and requests arrive one at a time. Upon the arrival of a new request, it needs to be immediately and irrevocably matched to an available server, resulting in a cost of their distance. The objective is to minimize the total matching cost. When servers are adversarial and requests are independently drawn from a known distribution, we reduce the problem to a more tractable setting where servers and requests are all independently drawn from the same distribution. Applying our reduction, for $[0, 1]^d$ with various choices of distributions, we achieve improved competitive ratios and nearly optimal regret in both balanced and unbalanced markets. In particular, we give $O(1)$-competitive algorithms for $d \geq 3$ in both balanced and unbalanced markets with smooth distributions. Our algorithms improve on the $O((\log \log \log n)^2)$ competitive ratio of Gupta et al. (ICALP'19) for balanced markets in various regimes, and provide the first positive results for unbalanced markets. Moreover, when servers and requests are all adversarial, and a prediction of request locations is provided, we present a general framework for transforming an arbitrary algorithm that does not use predictions into an algorithm that leverages predictions. The transformation applies the given algorithm in a black-box manner, and the performance of the resulting algorithm degrades smoothly as the prediction accuracy deteriorates while preserving the worst-case guarantee.

cs.DS

From signaling to interviews in random matching markets

In many two-sided labor markets, interviews are conducted before matches are formed. The growing number of interviews in medical residency markets has increased demand for signaling mechanisms, where applicants send a limited number of signals to communicate interest. We study the role of signaling mechanisms to reduce interviews in centralized random matching markets where initial preferences are refined through interviews. Agents can only match with those they interview. For the market to clear, we focus on perfect interim stability: no pair of agents-even if they never interviewed each other-prefers each other to their assigned partners under their interim preferences. A matching is almost interim stable if it is perfect interim stable after removing a vanishingly small fraction of agents. We analyze signaling mechanisms in random matching markets with $n$ agents where agents on the short side, long side, or both sides signal their top $d$ preferred partners. The interview graph connects pairs where at least one party signaled the other. We reveal a fundamental trade-off between almost and perfect interim stability. For almost interim stability, $d=ω(1)$ signals suffice: short-side signaling is always effective, whereas long-side signaling is effective only when the market is weakly imbalanced, i.e., when any size difference between the two sides becomes negligible as the market grows. For perfect interim stability, at least $d=Ω(\log^2 n)$ signals are necessary, and short-side signaling becomes crucial in any imbalanced market. We establish that truthful signaling is a Bayes-Nash equilibrium and extend our analysis to markets with hierarchical structure. As a technical contribution, we develop a message-passing algorithm that efficiently determines interim stability by leveraging local neighborhood structures.

cs.GT

Random graph matching at Otter's threshold via counting chandeliers

We propose an efficient algorithm for graph matching based on similarity scores constructed from counting a certain family of weighted trees rooted at each vertex. For two Erdős-Rényi graphs $\mathcal{G}(n,q)$ whose edges are correlated through a latent vertex correspondence, we show that this algorithm correctly matches all but a vanishing fraction of the vertices with high probability, provided that $nq\to\infty$ and the edge correlation coefficient $ρ$ satisfies $ρ^2>α\approx 0.338$, where $α$ is Otter's tree-counting constant. Moreover, this almost exact matching can be made exact under an extra condition that is information-theoretically necessary. This is the first polynomial-time graph matching algorithm that succeeds at an explicit constant correlation and applies to both sparse and dense graphs. In comparison, previous methods either require $ρ=1-o(1)$ or are restricted to sparse graphs. The crux of the algorithm is a carefully curated family of rooted trees called chandeliers, which allows effective extraction of the graph correlation from the counts of the same tree while suppressing the undesirable correlation between those of different trees.

cs.DS

Testing network correlation efficiently via counting trees

We propose a new procedure for testing whether two networks are edge-correlated through some latent vertex correspondence. The test statistic is based on counting the co-occurrences of signed trees for a family of non-isomorphic trees. When the two networks are Erdős-Rényi random graphs $\mathcal{G}(n,q)$ that are either independent or correlated with correlation coefficient $ρ$, our test runs in $n^{2+o(1)}$ time and succeeds with high probability as $n\to\infty$, provided that $n\min\{q,1-q\} \ge n^{-o(1)}$ and $ρ^2>α\approx 0.338$, where $α$ is Otter's constant so that the number of unlabeled trees with $K$ edges grows as $(1/α)^K$. This significantly improves the prior work in terms of statistical accuracy, running time, and graph sparsity.

math.ST

Settling the Sharp Reconstruction Thresholds of Random Graph Matching

This paper studies the problem of recovering the hidden vertex correspondence between two edge-correlated random graphs. We focus on the Gaussian model where the two graphs are complete graphs with correlated Gaussian weights and the Erdős-Rényi model where the two graphs are subsampled from a common parent Erdős-Rényi graph $\mathcal{G}(n,p)$. For dense graphs with $p=n^{-o(1)}$, we prove that there exists a sharp threshold, above which one can correctly match all but a vanishing fraction of vertices and below which correctly matching any positive fraction is impossible, a phenomenon known as the "all-or-nothing" phase transition. Even more strikingly, in the Gaussian setting, above the threshold all vertices can be exactly matched with high probability. In contrast, for sparse Erdős-Rényi graphs with $p=n^{-Θ(1)}$, we show that the all-or-nothing phenomenon no longer holds and we determine the thresholds up to a constant factor. Along the way, we also derive the sharp threshold for exact recovery, sharpening the existing results in Erdős-Rényi graphs. The proof of the negative results builds upon a tight characterization of the mutual information based on the truncated second-moment computation and an "area theorem" that relates the mutual information to the integral of the reconstruction error. The positive results follows from a tight analysis of the maximum likelihood estimator that takes into account the cycle structure of the induced permutation on the edges.

math.ST

Testing correlation of unlabeled random graphs

We study the problem of detecting the edge correlation between two random graphs with $n$ unlabeled nodes. This is formalized as a hypothesis testing problem, where under the null hypothesis, the two graphs are independently generated; under the alternative, the two graphs are edge-correlated under some latent node correspondence, but have the same marginal distributions as the null. For both Gaussian-weighted complete graphs and dense Erdős-Rényi graphs (with edge probability $n^{-o(1)}$), we determine the sharp threshold at which the optimal testing error probability exhibits a phase transition from zero to one as $n\to \infty$. For sparse Erdős-Rényi graphs with edge probability $n^{-Ω(1)}$, we determine the threshold within a constant factor. The proof of the impossibility results is an application of the conditional second-moment method, where we bound the truncated second moment of the likelihood ratio by carefully conditioning on the typical behavior of the intersection graph (consisting of edges in both observed graphs) and taking into account the cycle structure of the induced random permutation on the edges. Notably, in the sparse regime, this is accomplished by leveraging the pseudoforest structure of subcritical Erdős-Rényi graphs and a careful enumeration of subpseudoforests that can be assembled from short orbits of the edge permutation.

math.ST