SearcharxivSearch

arXiv subjects

Geng Zhao

Publications and source records attributed to Geng Zhao.

At least 19 recordsLinked to original sources

Efficiency Adjustments Break the Logarithmic Rank Barrier

We study the expected average rank achieved by the Efficiency-Adjusted Deferred Acceptance (EADA) mechanism in i.i.d.\ matching markets. While student-proposing Deferred Acceptance gives students an expected average rank of logarithmic order, we prove that EADA's expected average rank is at most $4\log\log n+O(1)$. Therefore, EADA improves the asymptotic order of students' assignments. At the cost of a weaker bound, $O((\log\log n)^2)$, we extend this conclusion to a much larger class of mechanisms. Namely, every Pareto-efficient mechanism that weakly Pareto-dominates DA breaks DA's logarithmic barrier. These are the first asymptotic guarantees for the expected average rank of EADA and of the broader class of Pareto-efficient improvements of DA. The conclusions extend to many-to-one markets with bounded quotas and random markets with correlated preferences.

cs.GT

The Distribution of Envy in Matching Markets

We study the distribution of envy in random matching markets under the Deferred Acceptance (DA) algorithm. Using tools from applied probability, we compute the expected number of proposing agents whom nobody envies and those who envy nobody. We obtain an exact finite-market expression for the former, based on a connection with the coupon collector problem, and asymptotic bounds for the latter. To put these quantities into perspective, we compare them to their counterparts under Random Serial Dictatorship (RSD): while RSD assigns a constant fraction of agents to their top choice, both DA and RSD leave exactly $H_n$ proposing agents unenvied in expectation. Our results show that these clearly unimprovable proposing agents constitute a vanishing fraction of the market.

econ.TH

Tunable Single- and Multiphoton Bundles in Cavity-Coupled Atomic Arrays

We propose an experimentally accessible scheme for realizing tunable nonclassical light in cavity-coupled reconfigurable atomic arrays. By coherently controlling the collective interference phase, the system switches from single-photon blockade to high-purity multiphoton bundle emission, unveiling a hierarchical structure of photon correlations dictated by atom-number parity and cavity detuning. The scaling of photon population identifies the transition between superradiant and subradiant regimes, while parity- and phase-dependent spin correlations elucidate the microscopic interference processes enabling coherent multiphoton generation. This work establishes a unified framework connecting cooperative atomic interactions to controllable nonclassical photon statistics and introduces a distinct interference-enabled mechanism that provides a practical route toward high-fidelity multiphoton sources in scalable cavity QEDs.

quant-ph

Travel Bans vs. Other Disease Mitigation Measures: A Mathematical Analysis

As the world grows increasingly connected, infectious disease transmission and outbreaks have become a pressing global concern for public health officials and policymakers. While policy interventions to contain and prevent the spread of disease have been proposed and implemented, there has been little rigorous quantitative analysis of the effectiveness of such interventions. In this paper, we study the susceptible-infected-recovered (SIR) infection process on a dynamic network model that models two communities with travel between them with the infection starting in one of them. In particular, we consider two Erd\H{o}s--R\'enyi graphs where edges are dynamically changing based on node travel between the graphs. We characterize the time evolution of the outbreaks in both communities and pin down the time for when the infection first reaches the second community. Finally, we analyze two types of interventions--travel bans and intra-community interventions in the second community--and prove that travel bans are not effective, while the second type are effective even without travel bans, provided they sufficiently reduce the effective reproduction number. We complement our analytic results by numerical simulations on large networks with realistic degree distributions and disease recovery times, showing that these results are robust, and hold for settings that model actual contact networks and disease spread more closely.

math.PR

What Pareto-Efficiency Adjustments Cannot Fix

The Deferred Acceptance (DA) algorithm is stable and strategy-proof, but can produce outcomes that are Pareto-inefficient for students, and thus several alternative mechanisms have been proposed to correct this inefficiency. However, we show that these mechanisms cannot correct DA's rank-inefficiency and inequality, because these shortcomings can arise even in cases where DA is Pareto-efficient. We also examine students' segregation in settings with advantaged and marginalized students. We prove that the demographic composition of every school is perfectly preserved under any Pareto-efficient mechanism that dominates DA, and consequently fully segregated schools under DA maintain their extreme homogeneity.

econ.TH

JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift

Safety and security remain critical concerns in AI deployment. Despite safety training through reinforcement learning with human feedback (RLHF) [ 32], language models remain vulnerable to jailbreak attacks that bypass safety guardrails. Universal jailbreaks - prefixes that can circumvent alignment for any payload - are particularly concerning. We show empirically that jailbreak detection systems face distribution shift, with detectors trained at one point in time performing poorly against newer exploits. To study this problem, we release JailbreaksOverTime, a comprehensive dataset of timestamped real user interactions containing both benign requests and jailbreak attempts collected over 10 months. We propose a two-pronged method for defenders to detect new jailbreaks and continuously update their detectors. First, we show how to use continuous learning to detect jailbreaks and adapt rapidly to new emerging jailbreaks. While detectors trained at a single point in time eventually fail due to drift, we find that universal jailbreaks evolve slowly enough for self-training to be effective. Retraining our detection model weekly using its own labels - with no new human labels - reduces the false negative rate from 4% to 0.3% at a false positive rate of 0.1%. Second, we introduce an unsupervised active monitoring approach to identify novel jailbreaks. Rather than classifying inputs directly, we recognize jailbreaks by their behavior, specifically, their ability to trigger models to respond to known-harmful prompts. This approach has a higher false negative rate (4.1%) than supervised methods, but it successfully identified some out-of-distribution attacks that were missed by the continuous learning approach.

cs.CR

The Large and Likely Inefficiency of Stable Matching Mechanisms

We prove that any stable matching mechanism suffers from systematic inefficiency of striking magnitude: in large random markets, any stable allocation is Pareto-inefficient with high probability, and almost all students can simultaneously improve their placements without harming anyone else. We establish this result by showing that the envy digraph generated by the student-proposing Deferred Acceptance mechanism contains a unique giant strongly connected component, implying that nearly all students are improvable via trading cycles. Finally, we show that every maximal cycle packing covers almost all students, revealing a surprising asymptotic equivalence among all efficient mechanisms that Pareto-dominate DA.

econ.TH

Toxicity Detection for Free

Current LLMs are generally aligned to follow safety requirements and tend to refuse toxic prompts. However, LLMs can fail to refuse toxic prompts or be overcautious and refuse benign examples. In addition, state-of-the-art toxicity detectors have low TPRs at low FPR, incurring high costs in real-world applications where toxic examples are rare. In this paper, we introduce Moderation Using LLM Introspection (MULI), which detects toxic prompts using the information extracted directly from LLMs themselves. We found we can distinguish between benign and toxic prompts from the distribution of the first response token's logits. Using this idea, we build a robust detector of toxic prompts using a sparse logistic regression model on the first response token logits. Our scheme outperforms SOTA detectors under multiple metrics.

cs.CL

The Power of Two in Token Systems

In economies without monetary transfers, token systems serve as an alternative to sustain cooperation, alleviate free riding, and increase efficiency. This paper studies whether a token-based economy can be effective in marketplaces with thin exogenous supply. We consider a marketplace in which at each time period one agent requests a service, one agent provides the service, and one token (artificial currency) is used to pay for service provision. The number of tokens each agent has represents the difference between the amount of service provisions and service requests by the agent. We are interested in the behavior of this economy when very few agents are available to provide the requested service. Since balancing the number of tokens across agents is key to sustain cooperation, the agent with the minimum amount of tokens is selected to provide service among the available agents. When exactly one random agent is available to provide service, we show that the token distribution is unstable. However, already when just two random agents are available to provide service, the token distribution is stable, in the sense that agents' token balance is unlikely to deviate much from their initial endowment, and agents return to their initial endowment in finite expected time. Our results mirror the power of two choices paradigm in load balancing problems. Supported by numerical simulations using kidney exchange data, our findings suggest that token systems may generate efficient outcomes in kidney exchange marketplaces by sustaining cooperation between hospitals.

cs.GT

Locality via Global Ties: Stability of the 2-Core Against Misspecification

For many random graph models, the analysis of a related birth process suggests local sampling algorithms for the size of, e.g., the giant connected component, the $k$-core, the size and probability of an epidemic outbreak, etc. In this paper, we study the question of when these algorithms are robust against misspecification of the graph model, for the special case of the 2-core. We show that, for locally converging graphs with bounded average degrees, under a weak notion of expansion, a local sampling algorithm provides robust estimates for the size of both the 2-core and its largest component. Our weak notion of expansion generalizes the classical definition of expansion, while holding for many well-studied random graph models. Our method involves a two-step sprinkling argument. In the first step, we use sprinkling to establish the existence of a non-empty $2$-core inside the giant, while in the second, we use this non-empty $2$-core as seed for a second sprinkling argument to establish that the giant contains a linear sized $2$-core. The second step is based on a novel coloring scheme for the vertices in the tree-part. Our algorithmic results follow from the structural properties for the $2$-core established in the course of our sprinkling arguments. The run-time of our local algorithm is constant independent of the graph size, with the value of the constant depending on the desired asymptotic accuracy $\epsilon$. But given the existential nature of local limits, our arguments do not give any bound on the functional dependence of this constant on $\epsilon$, nor do they give a bound on how large the graph has to be for the asymptotic additive error bound $\epsilon$ to hold.

cs.DS

Welfare Distribution in Two-sided Random Matching Markets

We study the welfare structure in two-sided large random matching markets. In the model, each agent has a latent personal score for every agent on the other side of the market and her preferences follow a logit model based on these scores. Under a contiguity condition, we provide a tight description of stable outcomes. First, we identify an intrinsic fitness for each agent that represents her relative competitiveness in the market, independent of the realized stable outcome. The intrinsic fitness values correspond to scaling coefficients needed to make a latent mutual matrix bi-stochastic, where the latent scores can be interpreted as a-priori probabilities of a pair being matched. Second, in every stable (or even approximately stable) matching, the welfare or the ranks of the agents on each side of the market, when scaled by their intrinsic fitness, have an approximately exponential empirical distribution. Moreover, the average welfare of agents on one side of the market is sufficient to determine the average on the other side. Overall, each agent's welfare is determined by a global parameter, her intrinsic fitness, and an extrinsic factor with exponential distribution across the population.

econ.TH

Online Learning in Stackelberg Games with an Omniscient Follower

We study the problem of online learning in a two-player decentralized cooperative Stackelberg game. In each round, the leader first takes an action, followed by the follower who takes their action after observing the leader's move. The goal of the leader is to learn to minimize the cumulative regret based on the history of interactions. Differing from the traditional formulation of repeated Stackelberg games, we assume the follower is omniscient, with full knowledge of the true reward, and that they always best-respond to the leader's actions. We analyze the sample complexity of regret minimization in this repeated Stackelberg game. We show that depending on the reward structure, the existence of the omniscient follower may change the sample complexity drastically, from constant to exponential, even for linear cooperative Stackelberg games. This poses unique challenges for the learning process of the leader and the subsequent regret analysis.

cs.LG

Memristive Computing for Efficient Inference on Resource Constrained Devices

The advent of deep learning has resulted in a number of applications which have transformed the landscape of the research area in which it has been applied. However, with an increase in popularity, the complexity of classical deep neural networks has increased over the years. As a result, this has leads to considerable problems during deployment on devices with space and time constraints. In this work, we perform a review of the present advancements in non-volatile memory and how the use of resistive RAM memory, particularly memristors, can help to progress the state of research in deep learning. In other words, we wish to present an ideology that advances in the field of memristive technology can greatly influence and impact deep learning inference on edge devices.

cs.ET

Interference, Bias, and Variance in Two-Sided Marketplace Experimentation: Guidance for Platforms

Two-sided marketplace platforms often run experiments to test the effect of an intervention before launching it platform-wide. A typical approach is to randomize individuals into the treatment group, which receives the intervention, and the control group, which does not. The platform then compares the performance in the two groups to estimate the effect if the intervention were launched to everyone. We focus on two common experiment types, where the platform randomizes individuals either on the supply side or on the demand side. The resulting estimates of the treatment effect in these experiments are typically biased: because individuals in the market compete with each other, individuals in the treatment group affect those in the control group and vice versa, creating interference. We develop a simple tractable market model to study bias and variance in these experiments with interference. We focus on two choices available to the platform: (1) Which side of the platform should it randomize on (supply or demand)? (2) What proportion of individuals should be allocated to treatment? We find that both choices affect the bias and variance of the resulting estimators but in different ways. The bias-optimal choice of experiment type depends on the relative amounts of supply and demand in the market, and we discuss how a platform can use market data to select the experiment type. Importantly, we find in many circumstances, choosing the bias-optimal experiment type has little effect on variance. On the other hand, the choice of treatment proportion can induce a bias-variance tradeoff, where the bias-minimizing proportion increases variance. We discuss how a platform can navigate this tradeoff and best choose the treatment proportion, using a combination of modeling as well as contextual knowledge about the market, the risk of the intervention, and reasonable effect sizes of the intervention.

stat.ME

Tiered Random Matching Markets: Rank is Proportional to Popularity

We study the stable marriage problem in two-sided markets with randomly generated preferences. We consider agents on each side divided into a constant number of "soft tiers", which intuitively indicate the quality of the agent. Specifically, every agent within a tier has the same public score, and agents on each side have preferences independently generated proportionally to the public scores of the other side. We compute the expected average rank which agents in each tier have for their partners in the men-optimal stable matching, and prove concentration results for the average rank in asymptotically large markets. Furthermore, we show that despite having a significant effect on ranks, public scores do not strongly influence the probability of an agent matching to a given tier of the other side. This generalizes results of [Pittel 1989] which correspond to uniform preferences. The results quantitatively demonstrate the effect of competition due to the heterogeneous attractiveness of agents in the market, and we give the first explicit calculations of rank beyond uniform markets.

cs.GT

Multi-level Chaotic Maps for 3D Textured Model Encryption

With rapid progress of Virtual Reality and Augmented Reality technologies, 3D contents are the next widespread media in many applications. Thus, the protection of 3D models is primarily important. Encryption of 3D models is essential to maintain confidentiality. Previous work on encryption of 3D surface model often consider the point clouds, the meshes and the textures individually. In this work, a multi-level chaotic maps models for 3D textured encryption was presented by observing the different contributions for recognizing cipher 3D models between vertices (point cloud), polygons and textures. For vertices which make main contribution for recognizing, we use high level 3D Lu chaotic map to encrypt them. For polygons and textures which make relatively smaller contributions for recognizing, we use 2D Arnold's cat map and 1D Logistic map to encrypt them, respectively. The experimental results show that our method can get similar performance with the other method use the same high level chaotic map for point cloud, polygons and textures, while we use less time. Besides, our method can resist more method of attacks such as statistic attack, brute-force attack, correlation attack.

cs.CV

Aesthetic Attributes Assessment of Images

Image aesthetic quality assessment has been a relatively hot topic during the last decade. Most recently, comments type assessment (aesthetic captions) has been proposed to describe the general aesthetic impression of an image using text. In this paper, we propose Aesthetic Attributes Assessment of Images, which means the aesthetic attributes captioning. This is a new formula of image aesthetic assessment, which predicts aesthetic attributes captions together with the aesthetic score of each attribute. We introduce a new dataset named \emph{DPC-Captions} which contains comments of up to 5 aesthetic attributes of one image through knowledge transfer from a full-annotated small-scale dataset. Then, we propose Aesthetic Multi-Attribute Network (AMAN), which is trained on a mixture of fully-annotated small-scale PCCD dataset and weakly-annotated large-scale DPC-Captions dataset. Our AMAN makes full use of transfer learning and attention model in a single framework. The experimental results on our DPC-Captions and PCCD dataset reveal that our method can predict captions of 5 aesthetic attributes together with numerical score assessment of each attribute. We use the evaluation criteria used in image captions to prove that our specially designed AMAN model outperforms traditional CNN-LSTM model and modern SCA-CNN model of image captions.

cs.CV

ILGNet: Inception Modules with Connected Local and Global Features for Efficient Image Aesthetic Quality Classification using Domain Adaptation

In this paper, we address a challenging problem of aesthetic image classification, which is to label an input image as high or low aesthetic quality. We take both the local and global features of images into consideration. A novel deep convolutional neural network named ILGNet is proposed, which combines both the Inception modules and an connected layer of both Local and Global features. The ILGnet is based on GoogLeNet. Thus, it is easy to use a pre-trained GoogLeNet for large-scale image classification problem and fine tune our connected layers on an large scale database of aesthetic related images: AVA, i.e. \emph{domain adaptation}. The experiments reveal that our model achieves the state of the arts in AVA database. Both the training and testing speeds of our model are higher than those of the original GoogLeNet.

cs.CV