SearcharxivSearch

arXiv subjects

Weitao Duan

Publications and source records attributed to Weitao Duan.

6 recordsLinked to original sources

Improving Ego-Cluster for Network Effect Measurement

The network effect, wherein one user's activity impacts another user, is common in social network platforms. Many new features in social networks are specifically designed to create a network effect, enhancing user engagement. For instance, content creators tend to produce more when their articles and posts receive positive feedback from followers. This paper discusses a new cluster-level experimentation methodology for measuring creator-side metrics in the context of A/B experiments. The methodology is designed to address cases where the experiment randomization unit and the metric measurement unit differ. It is a crucial part of LinkedIn's overall strategy to foster a robust creator community and ecosystem. The method is developed based on widely-cited research at LinkedIn but significantly improves the efficiency and flexibility of the clustering algorithm. This improvement results in a stronger capability for measuring creator-side metrics and an increased velocity for creator-related experiments.

cs.SI

Online Experimentation with Surrogate Metrics: Guidelines and a Case Study

A/B tests have been widely adopted across industries as the golden rule that guides decision making. However, the long-term true north metrics we ultimately want to drive through A/B test may take a long time to mature. In these situations, a surrogate metric which predicts the long-term metric is often used instead to conclude whether the treatment is effective. However, because the surrogate rarely predicts the true north perfectly, a regular A/B test based on surrogate metrics tends to have high false positive rate and the treatment variant deemed favorable from the test may not be the winning one. In this paper, we discuss how to adjust the A/B testing comparison to ensure experiment results are trustworthy. We also provide practical guidelines on the choice of good surrogate metrics. To provide a concrete example of how to leverage surrogate metrics for fast decision making, we present a case study on developing and evaluating the predicted confirmed hire surrogate metric in LinkedIn job marketplace.

stat.AP

Scalable Online Survey Framework: from Sampling to Analysis

With the advancement in technology, raw event data generated by the digital world have grown tremendously. However, such data tend to be insufficient and noisy when it comes to measuring user intention or satisfaction. One effective way to measure user experience directly is through surveys. In particular, with the popularity of online surveys, extensive work has been put in to study this field. Surveys at LinkedIn play a major role in influencing product and marketing decisions and supporting our sales efforts. We run an increasing number of surveys that help us understand shifts in awareness and perceptions with regards to our own products and also to peer companies. As the need to survey grows, both sampling and analysis of surveys have become more challenging. Instead of simply multiplying the number of surveys each user takes, we need a scalable approach to collect enough and representative samples for each survey analysis while maintaining good user experience. In this paper, we start with discussions on how we handle multiple email surveys under such constraints. We then shift our discussions to challenges of in-product surveys and how we address them at LinkedIn through a survey study conducted across two mobile apps. Finally, we share how in-product surveys can be utilized as monitoring tools and connect surveys with A/B testing.

stat.AP

Testing for arbitrary interference on experimentation platforms

Experimentation platforms are essential to modern large technology companies, as they are used to carry out many randomized experiments daily. The classic assumption of no interference among users, under which the outcome of one user does not depend on the treatment assigned to other users, is rarely tenable on such platforms. Here, we introduce an experimental design strategy for testing whether this assumption holds. Our approach is in the spirit of the Durbin-Wu-Hausman test for endogeneity in econometrics, where multiple estimators return the same estimate if and only if the null hypothesis holds. The design that we introduce makes no assumptions on the interference model between units, nor on the network among the units, and has a sharp bound on the variance and an implied analytical bound on the type I error rate. We discuss how to apply the proposed design strategy to large experimentation platforms, and we illustrate it in the context of an experiment on the LinkedIn platform.

stat.ME

SQR: Balancing Speed, Quality and Risk in Online Experiments

Controlled experimentation, also called A/B testing, is widely adopted to accelerate product innovations in the online world. However, how fast we innovate can be limited by how we run experiments. Most experiments go through a "ramp up" process where we gradually increase the traffic to the new treatment to 100%. We have seen huge inefficiency and risk in how experiments are ramped, and it is getting in the way of innovation. This can go both ways: we ramp too slowly and much time and resource is wasted; or we ramp too fast and suboptimal decisions are made. In this paper, we build up a ramping framework that can effectively balance among Speed, Quality and Risk (SQR). We start out by identifying the top common mistakes experimenters make, and then introduce the four SQR principles corresponding to the four ramp phases of an experiment. To truly scale SQR to all experiments, we develop a statistical algorithm that is embedded into the process of running every experiment to automatically recommend ramp decisions. Finally, to complete the whole picture, we briefly cover the auto-ramp engineering infrastructure that can collect inputs and execute on the recommendations timely and reliably.

stat.AP

Test of Equivalence Principle at $10^{-8}$ Level by a Dual-species Double-diffraction Raman Atom Interferometer

We report an improved test of the weak equivalence principle by using a simultaneous $^{85}$Rb-$^{87}$Rb dual-species atom interferometer. We propose and implement a four-wave double-diffraction Raman transition scheme for the interferometer, and demonstrate its ability in suppressing common-mode phase noise of Raman lasers after their frequencies and intensity ratios are optimized. The statistical uncertainty of the experimental data for Eötvös parameter $η$ is $0.8\times10^{-8}$ at 3200 s. With various systematic errors corrected the final value is $η=(2.8\pm3.0)\times10^{-8}$. The major uncertainty is attributed to the Coriolis effect.

physics.atom-ph