SearcharxivSearch

arXiv subjects

Frank Yang

Publications and source records attributed to Frank Yang.

At least 19 recordsLinked to original sources

Searchable Menus

Multidimensional screening is (in)famously intractable. In this paper, we study optimal screening mechanisms subject to a tractability constraint from the agent's perspective. Specifically, we require that the menu of options offered by the designer can be ordered so that, regardless of her preference type, the agent can find a utility-maximizing option via greedy search: any locally optimal choice must also be globally optimal. In one-dimensional screening with the single-crossing property, this requirement has no bite. In multidimensional environments, however, searchability restricts the set of implementable outcomes. In the multiproduct monopoly problem, the optimal searchable menu is a sparse upgrade menu: higher tiers offer higher allocation probabilities for every good, and the number of tiers is at most the number of goods. In a multidimensional screening problem with money and ordeals, the optimal searchable menu offers the agent a single way to obtain the good. In income taxation with rich multidimensional heterogeneity, a tax schedule is searchable if and only if it is progressive.

econ.TH

Many-body quantum optics in a cascaded chiral network

Chiral quantum emitters interact with light only in one propagation direction, allowing them to be linked into cascaded systems in which photons mediate ordered, long-range interactions. Such systems are predicted to host novel regimes of many-body physics of light and matter. Exploring these regimes requires arrays of identical quantum emitters with directional, low-loss coupling to guided photons, a combination that has thus far remained experimentally out of reach. Here we realize a cascaded network of superconducting qubits using an architecture that overcomes these bottlenecks. We implement a four-qubit chain spanning two modules, with separations ranging from millimeters to half a meter, and exploit the shared waveguide as a dissipative resource to stabilize reconfigurable entanglement, reaching a genuinely multipartite regime unavailable in reciprocal baths. By scattering weak pulses off the chain, we observe photons sorted in time by photon number, a signature of the strong photon-photon interactions mediated by the emitters. Together, these results provide experimental access to many-body light-matter regimes that are beyond the reach of reciprocal systems.

quant-ph

Stochastic Optimization and Coupling

We study optimization problems in which a linear functional is maximized over probability measures that are dominated by a given measure according to an integral stochastic order in an arbitrary dimension. We show that the following four properties are equivalent for any such order: (i) the test function cone is closed under pointwise minimum, (ii) the value function is affine, (iii) the solution correspondence has a convex graph with decomposable extreme points, and (iv) every ordered pair of measures admits an order-preserving coupling. As corollaries, we derive the extreme and exposed point properties involving integral stochastic orders such as multidimensional mean-preserving spreads and stochastic dominance. Applying these results, we generalize Blackwell's theorem by completely characterizing the comparisons of experiments that admit two equivalent descriptions -- through instrumental values and through information technologies. We also show that these results immediately yield new insights into information design, mechanism design, and decision theory.

econ.TH

SODA: Semantic-Oriented Distributional Alignment for Generative Recommendation

Generative recommendation has emerged as a scalable alternative to traditional retrieve-and-rank pipelines by operating in a compact token space. However, existing methods mainly rely on discrete code-level supervision, which leads to information loss and limits the joint optimization between the tokenizer and the generative recommender. In this work, we propose a distribution-level supervision paradigm that leverages probability distributions over multi-layer codebooks as soft and information-rich representations. Building on this idea, we introduce Semantic-Oriented Distributional Alignment (SODA), a plug-and-play contrastive supervision framework based on Bayesian Personalized Ranking, which aligns semantically rich distributions via negative KL divergence while enabling end-to-end differentiable training. Extensive experiments on multiple real-world datasets demonstrate that SODA consistently improves the performance of various generative recommender backbones, validating its effectiveness and generality. Codes will be available upon acceptance.

cs.IR

Screening Frontiers

A principal screens an agent with an arbitrary set of allocations $X$. The agent's values for the allocations are comonotonic. A subset of allocations $X^* \subseteq X$ is a surplus-elasticity frontier if (i) any other allocation has a demand curve that is pointwise lower and less elastic than that of some allocation in $X^*$ and (ii) the allocations in $X^*$ can be ordered in terms of their demand curves such that a higher demand curve is more inelastic. We show that any surplus-elasticity frontier is an optimal menu. Moreover, if the incremental demand curves along the frontier are also ordered by their elasticities, then the frontier remains optimal even if randomization is allowed. The frontier is agnostic to type distributions and redistributive welfare weights---the same menu remains optimal for a broad class of objectives, yielding a decomposition between menu design and optimal pricing. Under the same incremental-elasticity condition, a generalized notion of the frontier is not only sufficient but also necessary for this robust optimality. As applications, we derive new results on optimal bundling, taxation, sequential screening, selling information, and regulating a data-rich monopolist.

econ.TH

UniGRec: Unified Generative Recommendation with Soft Identifiers for End-to-End Optimization

Generative recommendation has recently emerged as a transformative paradigm that directly generates target items, surpassing traditional cascaded approaches. It typically involves two components: a tokenizer that learns item identifiers and a recommender trained on them. Existing methods often decouple tokenization from recommendation or rely on asynchronous alternating optimization, limiting full end-to-end alignment. To address this, we unify the tokenizer and recommender under the ultimate recommendation objective via differentiable soft item identifiers, enabling joint end-to-end training. However, this introduces three challenges: training-inference discrepancy due to soft-to-hard mismatch, item identifier collapse from codeword usage imbalance, and collaborative signal deficiency due to an overemphasis on fine-grained token-level semantics. To tackle these challenges, we propose UniGRec, a unified generative recommendation framework that addresses them from three perspectives. UniGRec employs Annealed Inference Alignment during tokenization to smoothly bridge soft training and hard inference, a Codeword Uniformity Regularization to prevent identifier collapse and encourage codebook diversity, and a Dual Collaborative Distillation mechanism that distills collaborative priors from a lightweight teacher model to jointly guide both the tokenizer and the recommender. Extensive experiments on real-world datasets demonstrate that UniGRec consistently outperforms state-of-the-art baseline methods. Our codes are available at https://github.com/Jialei-03/UniGRec.

cs.IR

Dynamic stimulated emission for deterministic addition and subtraction of propagating photons

Photon subtraction and addition are essential non-Gaussian processes in quantum optics, where conventional methods using linear optics and number-resolving detection often suffer from low success probability. Here, we introduce the concept of \textit{dynamic stimulated emission}, whereby a quantum emitter undergoes stimulated emission with a time-dependent coupling. We show that, for both two- and three-level emitters, this process can be used to deterministically add or subtract a photon to a single propagating optical mode. We provide semi-analytic solutions to this problem for Fock states, enabling deterministic and unconditional single-photon subtraction and addition with fidelity ${\cal F}>0.996$. Our semi-analytic solutions are provided for both dynamically coupled two-level systems and for three-level systems whose dynamical coupling is controlled by a coherent laser drive. Moving beyond individual Fock states, we further showcase the ability to subtract and add single photons to photon-number superposition states. We show that Schr\"{o}dinger cat states can be prepared from squeezed vacuum input via cascaded subtraction or cascaded addition. Finally, we show that our photon-addition process can be used to add a photon to any squeezed and displaced state with high success probability and fidelity ${\cal F}>0.99$, thereby potentially converting quantum emitters from single-photon sources to sources of single-photon-added Gaussian states without the need for inline squeezing. Our protocols provide a path towards integrating quantum emitters to construct efficient sources of single-mode non-Gaussian light beyond single photons.

quant-ph

Bi-Level Optimization for Generative Recommendation: Bridging Tokenization and Generation

Generative recommendation is emerging as a transformative paradigm by directly generating recommended items, rather than relying on matching. Building such a system typically involves two key components: (1) optimizing the tokenizer to derive suitable item identifiers, and (2) training the recommender based on those identifiers. Existing approaches often treat these components separately--either sequentially or in alternation--overlooking their interdependence. This separation can lead to misalignment: the tokenizer is trained without direct guidance from the recommendation objective, potentially yielding suboptimal identifiers that degrade recommendation performance. To address this, we propose BLOGER, a Bi-Level Optimization for GEnerative Recommendation framework, which explicitly models the interdependence between the tokenizer and the recommender in a unified optimization process. The lower level trains the recommender using tokenized sequences, while the upper level optimizes the tokenizer based on both the tokenization loss and recommendation loss. We adopt a meta-learning approach to solve this bi-level optimization efficiently, and introduce gradient surgery to mitigate gradient conflicts in the upper-level updates, thereby ensuring that item identifiers are both informative and recommendation-aligned. Extensive experiments on multiple real-world datasets demonstrate that BLOGER consistently outperforms state-of-the-art generative recommendation methods while maintaining practical efficiency with no significant additional computational overhead, effectively bridging the gap between item tokenization and autoregressive generation. We release our code at https://github.com/Ten-Mao/BLOGER.

cs.IR

SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents

We present SENTINEL, a framework for formally evaluating the physical safety of foundation model (FM)-based embodied agents. SENTINEL is the first to provide multi-level safety evaluation across semantic interpretation, plan generation, and physical execution within a unified formal framework. Unlike prior methods that rely on heuristic rules or subjective FM judgments, SENTINEL grounds practical safety requirements in formal temporal logic (TL) semantics that can precisely specify state invariants, temporal dependencies, and timing constraints. It employs a multi-level verification pipeline where (i) at the semantic level, intuitive natural language safety requirements are formalized into TL formulas and the agent's understanding of these requirements is probed for alignment with the TL formulas; (ii) at the plan level, high-level action plans and subgoals generated by the agent are verified against the TL formulas to detect unsafe plans before execution; and (iii) at the trajectory level, multiple execution trajectories are merged into a computation tree and efficiently verified against physically-detailed TL specifications for a final safety check. We apply SENTINEL in VirtualHome and AI2-THOR, and formally evaluate multiple FM-based embodied agents against diverse safety requirements. Our experiments show that by grounding physical safety in temporal logic and applying verification methods across multiple levels, SENTINEL provides a rigorous foundation for systematically evaluating the safety of FM-based embodied agents in simulation-based physical environments, and can effectively expose potential safety violations in interpreting, planning, and executing the tasks.

cs.AI

Dynamic Threats to Credible Auctions

A seller wants to sell a good to a set of bidders using a credible mechanism. We show that when the seller has private information about her cost, it is impossible for a static mechanism to achieve the optimal revenue. In particular, even the optimal first-price auction is not credible. We show that the English auction can credibly implement the optimal mechanism, unlike the optimal Dutch auction. For symmetric mechanisms in which only winners pay, we also characterize all the static auctions that are credible: They are first-price auctions that depend only on the seller's cost ex post via a secret reserve, and may profitably pool bidders via a bid restriction. Our impossibility result highlights the role of public institutions and helps explain the use of dynamic mechanisms in informal auctions.

econ.TH

Bundling against Learning

A monopolist sells multiple goods to an uninformed buyer. The buyer chooses to learn any one-dimensional linear signal of their values for the goods, anticipating the seller's mechanism. The seller designs an optimal mechanism, anticipating the buyer's learning choice. In a generalized Gaussian environment, we show that every equilibrium has vertical learning where the buyer's posterior means are comonotonic, and every equilibrium is outcome-equivalent to nested bundling where the seller offers a menu of nested bundles. In equilibrium, the buyer learns more about a higher-tier good, resulting in a higher posterior variance on the log scale.

econ.TH

Enabling Self-Improving Agents to Learn at Test Time With Human-In-The-Loop Guidance

Large language model (LLM) agents often struggle in environments where rules and required domain knowledge frequently change, such as regulatory compliance and user risk screening. Current approaches, like offline fine-tuning and standard prompting, are insufficient because they cannot effectively adapt to new knowledge during actual operation. To address this limitation, we propose the Adaptive Reflective Interactive Agent (ARIA), an LLM agent framework designed specifically to continuously learn updated domain knowledge at test time. ARIA assesses its own uncertainty through structured self-dialogue, proactively identifying knowledge gaps and requesting targeted explanations or corrections from human experts. It then systematically updates an internal, timestamped knowledge repository with provided human guidance, detecting and resolving conflicting or outdated knowledge through comparisons and clarification queries. We evaluate ARIA on the realistic customer due diligence name screening task on TikTok Pay, alongside publicly available dynamic knowledge tasks. Results demonstrate significant improvements in adaptability and accuracy compared to baselines using standard offline fine-tuning and existing self-improving agents. ARIA is deployed within TikTok Pay serving over 150 million monthly active users, confirming its practicality and effectiveness for operational use in rapidly evolving environments.

cs.LG

Belief-Based Offline Reinforcement Learning for Delay-Robust Policy Optimization

Offline-to-online deployment of reinforcement-learning (RL) agents must bridge two gaps: (1) the sim-to-real gap, where real systems add latency and other imperfections not present in simulation, and (2) the interaction gap, where policies trained purely offline face out-of-distribution states during online execution because gathering new interaction data is costly or risky. Agents therefore have to generalize from static, delay-free datasets to dynamic, delay-prone environments. Standard offline RL learns from delay-free logs yet must act under delays that break the Markov assumption and hurt performance. We introduce DT-CORL (Delay-Transformer belief policy Constrained Offline RL), an offline-RL framework built to cope with delayed dynamics at deployment. DT-CORL (i) produces delay-robust actions with a transformer-based belief predictor even though it never sees delayed observations during training, and (ii) is markedly more sample-efficient than na\"ive history-augmentation baselines. Experiments on D4RL benchmarks with several delay settings show that DT-CORL consistently outperforms both history-augmentation and vanilla belief-based methods, narrowing the sim-to-real latency gap while preserving data efficiency.

cs.LG

Multidimensional Monotonicity and Economic Applications

We characterize the extreme points of multidimensional monotone functions from $[0,1]^n$ to $[0,1]$, as well as the extreme points of the set of one-dimensional marginals of these functions. These characterizations lead to new results for various mechanism design and information design problems, including public good provision with interdependent values; interim efficient bilateral trade mechanisms; mechanism (anti) equivalence; asymmetric reduced form auctions; and optimal private private information structure.

econ.TH

Delayed Feedback Modeling with Influence Functions

In online advertising under the cost-per-conversion (CPA) model, accurate conversion rate (CVR) prediction is crucial. A major challenge is delayed feedback, where conversions may occur long after user interactions, leading to incomplete recent data and biased model training. Existing solutions partially mitigate this issue but often rely on auxiliary models, making them computationally inefficient and less adaptive to user interest shifts. We propose IF-DFM, an \underline{I}nfluence \underline{F}unction-empowered for \underline{D}elayed \underline{F}eedback \underline{M}odeling which estimates the impact of newly arrived and delayed conversions on model parameters, enabling efficient updates without full retraining. By reformulating the inverse Hessian-vector product as an optimization problem, IF-DFM achieves a favorable trade-off between scalability and effectiveness. Experiments on benchmark datasets show that IF-DFM outperforms prior methods in both accuracy and adaptability.

cs.LG

Inverse Delayed Reinforcement Learning

Inverse Reinforcement Learning (IRL) has demonstrated effectiveness in a variety of imitation tasks. In this paper, we introduce an IRL framework designed to extract rewarding features from expert trajectories affected by delayed disturbances. Instead of relying on direct observations, our approach employs an efficient off-policy adversarial training framework to derive expert features and recover optimal policies from augmented delayed observations. Empirical evaluations in the MuJoCo environment under diverse delay settings validate the effectiveness of our method. Furthermore, we provide a theoretical analysis showing that recovering expert policies from augmented delayed observations outperforms using direct delayed observations.

cs.LG

Case Study: Runtime Safety Verification of Neural Network Controlled System

Neural networks are increasingly used in safety-critical applications such as robotics and autonomous vehicles. However, the deployment of neural-network-controlled systems (NNCSs) raises significant safety concerns. Many recent advances overlook critical aspects of verifying control and ensuring safety in real-time scenarios. This paper presents a case study on using POLAR-Express, a state-of-the-art NNCS reachability analysis tool, for runtime safety verification in a Turtlebot navigation system using LiDAR. The Turtlebot, equipped with a neural network controller for steering, operates in a complex environment with obstacles. We developed a safe online controller switching strategy that switches between the original NNCS controller and an obstacle avoidance controller based on the verification results. Our experiments, conducted in a ROS2 Flatland simulation environment, explore the capabilities and limitations of using POLAR-Express for runtime verification and demonstrate the effectiveness of our switching strategy.

cs.RO

Stabilizing remote entanglement via waveguide dissipation

Distributing entanglement between remote sites is integral to quantum networks. Here, we demonstrate the autonomous stabilization of remote entanglement between a pair of non-interacting superconducting qubits connected by an open waveguide on a chip. In this setting, the interplay between a classical continuous drive - supplied through the waveguide - and dissipation into the waveguide stabilizes the qubit pair in a dark state, which, asymptotically, takes the form of a Bell state. We use field-quadrature measurements of the photons emitted to the waveguide to perform quantum state tomography on the stabilized states, where we find a concurrence of $0.504^{+0.007}_{-0.029}$ in the optimal setting with a stabilization time constant of 56 $\pm$ 4 ns. We examine the imperfections within our system and discuss avenues for enhancing fidelities and achieving scalability in future work. The decoherence-protected, steady-state remote entanglement offered via dissipative stabilization may find applications in distributed quantum computing, sensing, and communication.

quant-ph