SearcharxivSearch

arXiv subjects

Ya Xu

Publications and source records attributed to Ya Xu.

17 recordsLinked to original sources

Terrestrial Matter Effects on Reactor Antineutrino Oscillations: Constant vs. Fluctuated Density Profiles

The JUNO Collaboration has recently released its first reactor antineutrino oscillation result, achieving unprecedented precision in the measurement of $\Delta m^2_{21}$ and $\sin^2\theta_{12}$. We emphasize that the accurate determination and modeling of the terrestrial matter density profile are fundamental for extracting the oscillation parameters and probing the neutrino mass ordering. This paper presents a realistic piecewise-constant model for the shallow crustal density profile along the baselines from Taishan and Yangjiang to the experimental hall, based on geological and petrophysical information. The uncertainty in the density profiles arises from variations in the density and length of each segment, both of which are conservatively estimated to be 10%. A careful comparison of constant and fluctuated density profiles is provided and the implications for the precision measurement of oscillation parameters are discussed. Finally, we also discuss the prospect of shallow crust tomography in future reactor neutrino experiments.

hep-ph

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our most capable model yet, achieving SoTA performance on frontier coding and reasoning benchmarks. In addition to its incredible coding and reasoning skills, Gemini 2.5 Pro is a thinking model that excels at multimodal understanding and it is now able to process up to 3 hours of video content. Its unique combination of long context, multimodal and reasoning capabilities can be combined to unlock new agentic workflows. Gemini 2.5 Flash provides excellent reasoning abilities at a fraction of the compute and latency requirements and Gemini 2.0 Flash and Flash-Lite provide high performance at low latency and cost. Taken together, the Gemini 2.X model generation spans the full Pareto frontier of model capability vs cost, allowing users to explore the boundaries of what is possible with complex agentic problem solving.

cs.CL

360Brew: A Decoder-only Foundation Model for Personalized Ranking and Recommendation

Ranking and recommendation systems are the foundation for numerous online experiences, ranging from search results to personalized content delivery. These systems have evolved into complex, multilayered architectures that leverage vast datasets and often incorporate thousands of predictive models. The maintenance and enhancement of these models is a labor intensive process that requires extensive feature engineering. This approach not only exacerbates technical debt but also hampers innovation in extending these systems to emerging problem domains. In this report, we present our research to address these challenges by utilizing a large foundation model with a textual interface for ranking and recommendation tasks. We illustrate several key advantages of our approach: (1) a single model can manage multiple predictive tasks involved in ranking and recommendation, (2) decoder models with textual interface due to their comprehension of reasoning capabilities, can generalize to new recommendation surfaces and out-of-domain problems, and (3) by employing natural language interfaces for task definitions and verbalizing member behaviors and their social connections, we eliminate the need for feature engineering and the maintenance of complex directed acyclic graphs of model dependencies. We introduce our research pre-production model, 360Brew V1.0, a 150B parameter, decoder-only model that has been trained and fine-tuned on LinkedIn's data and tasks. This model is capable of solving over 30 predictive tasks across various segments of the LinkedIn platform, achieving performance levels comparable to or exceeding those of current production systems based on offline metrics, without task-specific fine-tuning. Notably, each of these tasks is conventionally addressed by dedicated models that have been developed and maintained over multiple years by teams of a similar or larger size than our own.

cs.IR

Control LLM: Controlled Evolution for Intelligence Retention in LLM

Large Language Models (LLMs) demand significant computational resources, making it essential to enhance their capabilities without retraining from scratch. A key challenge in this domain is \textit{catastrophic forgetting} (CF), which hampers performance during Continuous Pre-training (CPT) and Continuous Supervised Fine-Tuning (CSFT). We propose \textbf{Control LLM}, a novel approach that leverages parallel pre-trained and expanded transformer blocks, aligning their hidden-states through interpolation strategies This method effectively preserves performance on existing tasks while seamlessly integrating new knowledge. Extensive experiments demonstrate the effectiveness of Control LLM in both CPT and CSFT. On Llama3.1-8B-Instruct, it achieves significant improvements in mathematical reasoning ($+14.4\%$ on Math-Hard) and coding performance ($+10\%$ on MBPP-PLUS). On Llama3.1-8B, it enhances multilingual capabilities ($+10.6\%$ on C-Eval, $+6.8\%$ on CMMLU, and $+30.2\%$ on CMMLU-0shot-CoT). It surpasses existing methods and achieves SOTA among open-source models tuned from the same base model, using substantially less data and compute. Crucially, these gains are realized while preserving strong original capabilities, with minimal degradation ($<4.3\% \text{on MMLU}$) compared to $>35\%$ in open-source Math and Coding models. This approach has been successfully deployed in LinkedIn's GenAI-powered job seeker and Ads unit products. To support further research, we release the training and evaluation code (https://github.com/linkedin/ControlLLM) along with models trained on public datasets (https://huggingface.co/ControlLLM) to the community.

cs.LG

Persistent but weak magnetic field at Moon's midlife revealed by Chang'e-5 basalt

The evolution of the lunar magnetic field can reveal the Moon's interior structure, thermal history, and surface environment. The mid-to-late stage evolution of the lunar magnetic field is poorly constrained, and thus the existence of a long-lived lunar dynamo remains controversial. The Chang'e-5 mission returned the heretofore youngest mare basalts from Oceanus Procellarum uniquely positioned at mid-latitude. We recovered weak paleointensities of 2-4 uT from the Chang'e-5 basalt clasts at 2 billion years ago, attestting to the longevity of a lunar dynamo until at least the Moon's midlife. This paleomagnetic result implies the existence of thermal convection in the lunar deep interior at the lunar mid-stage which may have supplied mantle heat flux for the young volcanism.

physics.geo-ph

Expected geoneutrino signal at JUNO using local integrated 3-D refined crustal model

Geoneutrinos serve as a potent tool for comprehending the radiogenic power and composition of Earth. Although geoneutrinos have been observed in prior experiments, the forthcoming generation of experiments,such as JUNO, will be necessary for fully harnessing their potential. Precise prediction of the crustal contribution is vital for interpreting particlephysics measurements in the context of geo-scientific inquiries. Nonetheless, existing models such as JULOC and GIGJ have limitations in accurately forecasting the crustal contribution. This paper introduces JULOCI, the novel 3-D integrated crustal model of JUNO, which employs seismic, gravity, rock sample, and heat flow data to precisely estimate the geoneutrino signal of the lithosphere. The model indicates elevated concentrations of uranium and thorium in southern China, resulting in unexpectedly strong geoneutrino signals.The accuracy of JULOC-I, coupled with a decade of experimental data, affords JUNO the opportunity to test multiple mantle models. Once operational, JUNO can validate the model predictions and enhance the precision of mantle measurements. All in all, the improved accuracy ofJULOC-I represents a substantial stride towards comprehending the geochemical distribution of the South China crust, offering a valuable tool for investigating the composition and evolution of the Earth through geoneutrinos.

physics.geo-ph

JULOC: A Local 3-D Refined Crust Model for the Geoneutrino Measurement at JUNO

Geothermal energy is the key to drive the plate tectonics and interior thermodynamics of the Earth. The surface heat flux, as measured in boreholes, provide limited insights into the relative contributions of primordial versus radiogenic sources of the heat budget of the mantle. Geoneutrinos, electron antineutrinos that produced from the radioactive decay of the heat producing elements, are unique probes that bring direct information about the amount and distribution of heat producing elements in the crust and mantle. Cosmochemical, geochemical, and geodynamic compositional models of the Bulk Silicate Earth (BSE) individually predicts different mantle neutrino fluxes, and therefore can be distinguished by the direct measurement of geoneutrinos. The 20 kton detector of the Jiangmen Underground Neutrino Observatory (JUNO), currently under construction in the Guangdong Province (China), is expected to provide an exciting opportunity to obtain a high statistics measurement, which will produce sufficient data to address several key questions of geological importance. To test different compositional models of the mantle, an accurate estimation of the crust geoneutrino flux based on a three-dimensional (3-D) crust model in advance is important. This paper presents a 3-D crust model over a surface area of 10-degrees-times-10-degrees grid surrounding the JUNO detector and a depth down to the Moho discontinuity, based on the geological, geophysical and geochemistry properties. The 3-D model provides a distinction of the volumes of the different geological layers together with the corresponding Th and U abundances. We also present our predicted local contribution to the total geoneutrino flux and the corresponding radiogenic heat.

physics.geo-ph

Scalable Online Survey Framework: from Sampling to Analysis

With the advancement in technology, raw event data generated by the digital world have grown tremendously. However, such data tend to be insufficient and noisy when it comes to measuring user intention or satisfaction. One effective way to measure user experience directly is through surveys. In particular, with the popularity of online surveys, extensive work has been put in to study this field. Surveys at LinkedIn play a major role in influencing product and marketing decisions and supporting our sales efforts. We run an increasing number of surveys that help us understand shifts in awareness and perceptions with regards to our own products and also to peer companies. As the need to survey grows, both sampling and analysis of surveys have become more challenging. Instead of simply multiplying the number of surveys each user takes, we need a scalable approach to collect enough and representative samples for each survey analysis while maintaining good user experience. In this paper, we start with discussions on how we handle multiple email surveys under such constraints. We then shift our discussions to challenges of in-product surveys and how we address them at LinkedIn through a survey study conducted across two mobile apps. Finally, we share how in-product surveys can be utilized as monitoring tools and connect surveys with A/B testing.

stat.AP

Using Ego-Clusters to Measure Network Effects at LinkedIn

A network effect is said to take place when a new feature not only impacts the people who receive it, but also other users of the platform, like their connections or the people who follow them. This very common phenomenon violates the fundamental assumption underpinning nearly all enterprise experimentation systems, the stable unit treatment value assumption (SUTVA). When this assumption is broken, a typical experimentation platform, which relies on Bernoulli randomization for assignment and two-sample t-test for assessment of significance, will not only fail to account for the network effect, but potentially give highly biased results. This paper outlines a simple and scalable solution to measuring network effects, using ego-network randomization, where a cluster is comprised of an "ego" (a focal individual), and her "alters" (the individuals she is immediately connected to). Our approach aims at maintaining representativity of clusters, avoiding strong modeling assumption, and significantly increasing power compared to traditional cluster-based randomization. In particular, it does not require product-specific experiment design, or high levels of investment from engineering teams, and does not require any changes to experimentation and analysis platforms, as it only requires assigning treatment an individual level. Each user either has the feature or does not, and no complex manipulation of interactions between users is needed. It focuses on measuring the one-out network effect (i.e the effect of my immediate connection's treatment on me), and gives reasonable estimates at a very low setup cost, allowing us to run such experiments dozens of times a year.

cs.SI

Large-Scale Online Experimentation with Quantile Metrics

Online experimentation (or A/B testing) has been widely adopted in industry as the gold standard for measuring product impacts. Despite the wide adoption, few literatures discuss A/B testing with quantile metrics. Quantile metrics, such as 90th percentile page load time, are crucial to A/B testing as many key performance metrics including site speed and service latency are defined as quantiles. However, with LinkedIn's data size, quantile metric A/B testing is extremely challenging because there is no statistically valid and scalable variance estimator for the quantile of dependent samples: the bootstrap estimator is statistically valid, but takes days to compute; the standard asymptotic variance estimate is scalable but results in order-of-magnitude underestimation. In this paper, we present a statistically valid and scalable methodology for A/B testing with quantiles that is fully generalizable to other A/B testing platforms. It achieves over 500 times speed up compared to bootstrap and has only $2\%$ chance to differ from bootstrap estimates. Beyond methodology, we also share the implementation of a data pipeline using this methodology and insights on pipeline optimization.

stat.AP

A Method for Measuring Network Effects of One-to-One Communication Features in Online A/B Tests

A/B testing is an important decision making tool in product development because can provide an accurate estimate of the average treatment effect of a new features, which allows developers to understand how the business impact of new changes to products or algorithms. However, an important assumption of A/B testing, Stable Unit Treatment Value Assumption (SUTVA), is not always a valid assumption to make, especially for products that facilitate interactions between individuals. In contexts like one-to-one messaging we should expect network interference; if an experimental manipulation is effective, behavior of the treatment group is likely to influence members in the control group by sending them messages, violating this assumption. In this paper, we propose a novel method that can be used to account for network effects when A/B testing changes to one-to-one interactions. Our method is an edge-based analysis that can be applied to standard Bernoulli randomized experiments to retrieve an average treatment effect that is not influenced by network interference. We develop a theoretical model, and methods for computing point estimates and variances of effects of interest via network-consistent permutation testing. We then apply our technique to real data from experiments conducted on the messaging product at LinkedIn. We find empirical support for our model, and evidence that the standard method of analysis for A/B tests underestimates the impact of new features in one-to-one messaging contexts.

stat.AP

Causal inference from observational data: Estimating the effect of contributions on visitation frequency atLinkedIn

Randomized experiments (A/B testings) have become the standard way for web-facing companies to guide innovation, evaluate new products, and prioritize ideas. There are times, however, when running an experiment is too complicated (e.g., we have not built the infrastructure), costly (e.g., the intervention will have a substantial negative impact on revenue), and time-consuming (e.g., the effect may take months to materialize). Even if we can run an experiment, knowing the magnitude of the impact will significantly accelerate the product development life cycle by helping us prioritize tests and determine the appropriate traffic allocation for different treatment groups. In this setting, we should leverage observational data to quickly and cost-efficiently obtain a reliable estimate of the causal effect. Although causal inference from observational data has a long history, its adoption by data scientist in technology companies has been slow. In this paper, we rectify this by providing a brief introduction to the vast field of causal inference with a specific focus on the tools and techniques that data scientist can directly leverage. We illustrate how to apply some of these methodologies to measure the effect of contributions (e.g., post, comment, like or send private messages) on engagement metrics. Evaluating the impact of contributions on engagement through an A/B test requires encouragement design and the development of non-standard experimentation infrastructure, which can consume a tremendous amount of time and financial resources. We present multiple efficient strategies that exploit historical data to accurately estimate the contemporaneous (or instantaneous) causal effect of a user's contribution on her own and her neighbors' (i.e., the users she is connected to) subsequent visitation frequency. We apply these tools to LinkedIn data for several million members.

stat.AP

Testing for arbitrary interference on experimentation platforms

Experimentation platforms are essential to modern large technology companies, as they are used to carry out many randomized experiments daily. The classic assumption of no interference among users, under which the outcome of one user does not depend on the treatment assigned to other users, is rarely tenable on such platforms. Here, we introduce an experimental design strategy for testing whether this assumption holds. Our approach is in the spirit of the Durbin-Wu-Hausman test for endogeneity in econometrics, where multiple estimators return the same estimate if and only if the null hypothesis holds. The design that we introduce makes no assumptions on the interference model between units, nor on the network among the units, and has a sharp bound on the variance and an implied analytical bound on the type I error rate. We discuss how to apply the proposed design strategy to large experimentation platforms, and we illustrate it in the context of an experiment on the LinkedIn platform.

stat.ME

Automatic Detection and Diagnosis of Biased Online Experiments

We have seen a massive growth of online experiments at LinkedIn, and in industry at large. It is now more important than ever to create an intelligent A/B platform that can truly democratize A/B testing by allowing everyone to make quality decisions, regardless of their skillset. With the tremendous knowledge base created around experimentation, we are able to mine through historical data, and discover the most common causes for biased experiments. In this paper, we share four of such common causes, and how we build into our A/B testing platform the automatic detection and diagnosis of such root causes. These root causes range from design-imposed bias, self-selection bias, novelty effect and trigger-day effect. We will discuss in detail what each bias is and the scalable algorithm we developed to detect the bias. Surfacing up the existence and root cause of bias automatically for every experiment is an important milestone towards intelligent A/B testing.

stat.AP

SQR: Balancing Speed, Quality and Risk in Online Experiments

Controlled experimentation, also called A/B testing, is widely adopted to accelerate product innovations in the online world. However, how fast we innovate can be limited by how we run experiments. Most experiments go through a "ramp up" process where we gradually increase the traffic to the new treatment to 100%. We have seen huge inefficiency and risk in how experiments are ramped, and it is getting in the way of innovation. This can go both ways: we ramp too slowly and much time and resource is wasted; or we ramp too fast and suboptimal decisions are made. In this paper, we build up a ramping framework that can effectively balance among Speed, Quality and Risk (SQR). We start out by identifying the top common mistakes experimenters make, and then introduce the four SQR principles corresponding to the four ramp phases of an experiment. To truly scale SQR to all experiments, we develop a statistical algorithm that is embedded into the process of running every experiment to automatically recommend ramp decisions. Finally, to complete the whole picture, we briefly cover the auto-ramp engineering infrastructure that can collect inputs and execute on the recommendations timely and reliably.

stat.AP

Empirical stationary correlations for semi-supervised learning on graphs

In semi-supervised learning on graphs, response variables observed at one node are used to estimate missing values at other nodes. The methods exploit correlations between nearby nodes in the graph. In this paper we prove that many such proposals are equivalent to kriging predictors based on a fixed covariance matrix driven by the link structure of the graph. We then propose a data-driven estimator of the correlation structure that exploits patterns among the observed response values. By incorporating even a small fraction of observed covariation into the predictions, we are able to obtain much improved prediction on two graph data sets.

stat.AP

CUR from a Sparse Optimization Viewpoint

The CUR decomposition provides an approximation of a matrix $X$ that has low reconstruction error and that is sparse in the sense that the resulting approximation lies in the span of only a few columns of $X$. In this regard, it appears to be similar to many sparse PCA methods. However, CUR takes a randomized algorithmic approach, whereas most sparse PCA methods are framed as convex optimization problems. In this paper, we try to understand CUR from a sparse optimization viewpoint. We show that CUR is implicitly optimizing a sparse regression objective and, furthermore, cannot be directly cast as a sparse PCA method. We also observe that the sparsity attained by CUR possesses an interesting structure, which leads us to formulate a sparse PCA method that achieves a CUR-like sparsity.

cs.DS