SearcharxivSearch

arXiv subjects

Wei-Yu Lin

Publications and source records attributed to Wei-Yu Lin.

2 recordsLinked to original sources

Compass: SLO-aware Query Planner for Compound AI Serving at Scale

The rise of compound AI serving that integrates multiple operators in a pipeline enables end-user applications such as generative AI-powered meeting companions, autonomous driving, and immersive gaming. These workloads span diverse deployment spaces, from cloud-only queries to edge-assisted ones across infrastructure tiers, often including both within an application. Achieving high service goodput -- i.e., meeting service level objectives (SLOs) for pipeline latency, accuracy, and costs -- requires joint planning of operators' placement, configuration, and resource allocation. However, diverse SLOs, varying runtime environments (e.g., heterogeneous device speeds), and a large volume of queries competing for shared infrastructure explode the planning space, making real-time serving and cost-efficient deployment intractable with existing advances. This paper presents Compass, the first SLO-aware query planner that optimizes large-scale compound AI workloads across diverse deployment spaces. Compass decomposes the many-query, multi-SLO planning problem into tractable subproblems while preserving global decision quality, exploiting plan similarities within and across queries to slash the search steps. It further improves per-step efficiency with a plan profiler that performs selective profiling to achieve high-fidelity performance estimates at a fraction of the profiling cost. At runtime, Compass performs query-plan bipartite matching to maximize SLO goodput under resource contentions. Real-world evaluations show that Compass improves service goodput by 2.4--5.1x, reduces deployment costs by 3.8--4.5x, and accelerates planning by 4.2--10.5x, achieving service responsiveness within seconds and near-optimal decision quality.

cs.DB

Quantum steering as a witness of quantum scrambling

Quantum information scrambling describes the delocalization of local information to global information in the form of entanglement throughout all possible degrees of freedom. A natural measure of scrambling is the tripartite mutual information (TMI), which quantifies the amount of delocalized information for a given quantum channel with its state representation, i.e., the Choi state. In this work, we show that quantum information scrambling can also be witnessed by temporal quantum steering for qubit systems. We can do so because there is a fundamental equivalence between the Choi state and the pseudo-density matrix formalism used in temporal quantum correlations. In particular, we propose a quantity as a scrambling witness, based on a measure of temporal steering called temporal steerable weight. We justify the scrambling witness for unitary qubit channels by proving that the quantity vanishes whenever the channel is non-scrambling.

quant-ph