SearcharxivSearch

arXiv subjects

Song Yao

Publications and source records attributed to Song Yao.

At least 19 recordsLinked to original sources

Stranded credentials: how a skill-signaling market absorbed generative AI

Generative AI can now perform many tasks that credentialing institutions count on to assess skill. During the AI era, do credentials retain their signaling value for subsequent performance? Mostly, yes. We audit the 2010-2026 archive of Kaggle, the largest data science competition platform, which ran two evaluation formats concurrently: upload-competitions, which directly score entrants' predictions computed on published data, and code-competitions, which score predictions by executing entrants' code on hidden data. Across 444,698 participations, competition medals predict subsequent leaderboard performance almost entirely in the first year after being earned, in both formats. Fresh medals retained most of their signaling value through the AI transition; credential stocks are only as informative as their replenishment. Although upload-competition medal stocks lost 82% of their informativeness, institutional stranding explains half to three quarters of the loss: upload-competitions had exited for reasons predating AI, and their frozen medal stock aged out under the pre-existing decay pattern. Old upload-competition medals look more valuable only in isolation, by proxying for the rest of the holder's record (e.g., experience). The measured changes are institutional rather than personal: an AI-like working style predicts performance similarly in both formats. The platform's official credential tiers, based on lifetime medal counts, discard 13-16% of the medals' information; an index weighting recent medals more heavily, built on pre-AI-era data alone, outperforms the official tiers in predicting AI-era performance. In conclusion, credentials are informative, perishable, institution-bound, and interdependent; sustaining their value under AI is a high-stakes, socio-economic problem of institutional design.

econ.GN

Unintended Consequences of Recommender System Interventions: Evidence from a Field Experiment

Platform content interventions in recommendation systems are typically evaluated as static "nudges", ignoring that the systems adaptively learn from the resulting user behavior. We investigate this dynamic through a large-scale field experiment on a short-video platform. The experiment involves a "sleep reminder" campaign designed to reduce late-night usage. Paradoxically, the intervention increased late-night engagement by 14.75% and overall platform usage by 2.18%, and the effects persisted for weeks even after the experiment. We explain this through a forced-exploration mechanism, showing that by revealing high latent demand for the promoted content, the intervention triggers a recommendation policy update that routine user behavior would not produce. The data generated by the intervention induced the algorithm to update its post-campaign policy, reinforcing the very engagement loops the campaign aimed to mitigate. Our findings demonstrate that user-facing interventions can effectively retrain the underlying algorithm, triggering durable, system-wide shifts in content distribution that challenge standard evaluation metrics in platform governance and social responsibility initiatives.

econ.GN

When the Scaffold Stays On: AI, Practice Style, and Screening in Elite Skill Formation

Generative AI raises short-term productivity by completing tasks learners would otherwise practice on their own. Whether this exchange erodes frontier skill depends on the mode of use: substitute-users let AI stand in for practice and fail to develop skills, while complement-users use AI to learn faster. The modes look alike in AI-aided output, so organizations screening on that output cannot tell them apart. We ask whether the AI-prohibited evaluation gates organizations already operate can separate the modes. In elite competitive programming, ICPC and IOI contests prohibit AI under in-person proctoring, with qualification-round entry, whereas Codeforces (CF) practice and contests are unproctored and open to all. From CF practice histories we build an AI-prompt signature consistent with AI usage, more first-attempt acceptances, fewer attempts and debugging retries. CF practice has shifted toward this signature across entry cohorts spanning two AI rollouts. On CF, a stronger signature predicts smaller rating gains for users with no ICPC-IOI affiliation, but not for those who qualified. Inside the AI-prohibited ICPC environment, AI-era entrants show no skill erosion, and shifts toward AI-style practice predict higher non-AI-aided scores. One screening mechanism fits both: where the modes mix, a stronger signature flags substitute-users; among those who qualified, a strengthening signature marks adoption of the complement mode. The message is constructive: AI-style practice is compatible with frontier skill; the erosion risk links to the substitute mode; and separating the modes is a design question for the exams organizations regularly administer, from medical and legal boards to professional certification.

econ.EM

Tianyu: search for the second solar system and explore the dynamic universe

Giant planets like Jupiter and Saturn, play important roles in the formation and habitability of Earth-like planets. The detection of solar system analogs that have multiple cold giant planets is essential for our understanding of planet habitability and planet formation. Although transit surveys such as Kepler and TESS have discovered thousands of exoplanets, these missions are not sensitive to long period planets due to their limited observation baseline. The Tianyu project, comprising two 1-meter telescopes (Tianyu-I and II), is designed to detect transiting cold giant planets in order to find solar system analogs. Featuring a large field of view and equipped with a high-speed CMOS camera, Tianyu-I will perform a high-precision photometric survey of about 100 million stars, measuring light curves at hour-long cadence. The candidates found by Tianyu-I will be confirmed by Tianyu-II and other surveys and follow-up facilities through multi-band photometry, spectroscopy, and high resolution imaging. Tianyu telescopes will be situated at an elevation about 4000 meters in Lenghu, China. With a photometric precision of 1% for stars with V < 18 mag, Tianyu is expected to find more than 300 transiting exoplanets, including about 12 cold giant planets, over five years. A five-year survey of Tianyu would discover 1-2 solar system analogs. Moreover, Tianyu is also designed for non-exoplanetary exploration, incorporating multiple survey modes covering timescales from sub-seconds to months, with a particular emphasis on events occurring within the sub-second to hour range. It excels in observing areas such as infant supernovae, rare variable stars and binaries, tidal disruption events, Be stars, cometary activities, and interstellar objects. These discoveries not only enhance our comprehension of the universe but also offer compelling opportunities for public engagement in scientific exploration.

astro-ph.IM

Stochastic Control/Stopping Problem with Expectation Constraints

We study a stochastic control/stopping problem with a series of inequality-type and equality-type expectation constraints in a general non-Markovian framework. We demonstrate that the stochastic control/stopping problem with expectation constraints (CSEC) is independent of a specific probability setting and is equivalent to the constrained stochastic control/stopping problem in weak formulation (an optimization over joint laws of Brownian motion, state dynamics, diffusion controls and stopping rules on an enlarged canonical space). Using a martingale-problem formulation of controlled SDEs in spirit of \cite{Stroock_Varadhan}, we characterize the probability classes in weak formulation by countably many actions of canonical processes, and thus obtain the upper semi-analyticity of the CSEC value function. Then we employ a measurable selection argument to establish a dynamic programming principle (DPP) in weak formulation for the CSEC value function, in which the conditional expected costs act as additional states for constraint levels at the intermediate horizon. This article extends the results of \cite{Elk_Tan_2013b} to the expectation-constraint case. We extend our previous work \cite{OSEC_stopping} to the more complicated setting where the diffusion is controlled. Compared to that paper the topological properties of diffusion-control spaces and the corresponding measurability are more technically involved which complicate the arguments especially for the measurable selection for the super-solution side of DPP in the weak formulation.

math.OC

Optimal Stopping with Expectation Constraints

We analyze an optimal stopping problem with a series of inequality-type and equality-type expectation constraints in a general non-Markovian framework. We show that the optimal stopping problem with expectation constraints (OSEC) in an arbitrary probability setting is equivalent to the constrained problem in weak formulation (optimization over joint laws of stopping rules with Brownian motion and state dynamics on an enlarged canonical space) and thus the OSEC value. Using a martingale-problem formulation, we make an equivalent characterization of the probability classes in weak formulation, which implies that the OSEC value function s upper semi-analytic. Then we exploit a measurable selection argument to establish a dynamic programming principle in weak formulation for the OSEC value function, in which the conditional expected costs act as additional states for constraint levels at the intermediate horizon.

math.OC

Number of New Top 2% Researchers from China and USA Over Time

In this paper we compare the numbers of new top 2% researchers from China and USA annually since 1980. We find that the log ratio of the numbers decreases almost linearly over time. As early as 2009, the total number of new top 2% researchers across all subfields from China exceeds that of USA. In particular, such trend is more striking in many subfields, e.g., Engineering, Chemistry, and Enabling & Strategic Technologies.

stat.AP

Robust optimization Design of a New Combined Median Barrier Based on Taguchi method and Grey Relational Analysis

Accidents that vehicles cross median and enter opposite lane happen frequently, and the existing median barrier has weak anti-collision strength. A new combined median barrier (NCMB) consisted of W-beam guardrail and concrete structure was proposed to decrease deformation and enhance anti-collision strength in this paper. However, there were some uncertainties in the initial design of the NCMB. If the uncertainties were not considered in the design process, the optimization objectives were especially sensitive to the small fluctuation of the variables, and it might result in design failure. For this purpose, the acceleration and deflection were taken as objectives; post thickness, W-beam thickness and post spacing were chosen as design variables; the velocity, mass of vehicle and the yield stress of barrier components were taken as noise factors, a multi-objective robust optimization is carried out for the NCMB based on Taguchi and grey relational analysis (GRA). The results indicate that the acceleration and deflection after optimization are reduced by 47.3% and 76.7% respectively; Signal-to-noise ratio (SNR) of objectives after optimization are increased, it greatly enhances the robustness of the NCMB. The results demonstrate that the effectiveness of the methodology that based on Taguchi method and grey relational analysis.

math.NA

Dynamic Programming Principles for Optimal Stopping with Expectation Constraint

We analyze an optimal stopping problem with a constraint on the expected cost. When the reward function and cost function are Lipschitz continuous in state variable, we show that the value of such an optimal stopping problem is a continuous function in current state and in budget level. Then we derive a dynamic programming principle (DPP) for the value function in which the conditional expected cost acts as an additional state process. As the optimal stopping problem with expectation constraint can be transformed to a stochastic optimization problem with supermartingale controls, we explore a second DPP of the value function and thus resolve an open question recently raised in [S. Ankirchner, M. Klein, and T. Kruse, A verification theorem for optimal stopping problems with expectation constraints, Appl. Math. Optim., 2017, pp. 1-33]. Based on these two DPPs, we characterize the value function as a viscosity solution to the related fully non-linear parabolic Hamilton-Jacobi-Bellman equation.

math.OC

ESE: Efficient Speech Recognition Engine with Sparse LSTM on FPGA

Long Short-Term Memory (LSTM) is widely used in speech recognition. In order to achieve higher prediction accuracy, machine learning scientists have built larger and larger models. Such large model is both computation intensive and memory intensive. Deploying such bulky model results in high power consumption and leads to high total cost of ownership (TCO) of a data center. In order to speedup the prediction and make it energy efficient, we first propose a load-balance-aware pruning method that can compress the LSTM model size by 20x (10x from pruning and 2x from quantization) with negligible loss of the prediction accuracy. The pruned model is friendly for parallel processing. Next, we propose scheduler that encodes and partitions the compressed model to each PE for parallelism, and schedule the complicated LSTM data flow. Finally, we design the hardware architecture, named Efficient Speech Recognition Engine (ESE) that works directly on the compressed model. Implemented on Xilinx XCKU060 FPGA running at 200MHz, ESE has a performance of 282 GOPS working directly on the compressed LSTM network, corresponding to 2.52 TOPS on the uncompressed one, and processes a full LSTM for speech recognition with a power dissipation of 41 Watts. Evaluated on the LSTM for speech recognition benchmark, ESE is 43x and 3x faster than Core i7 5930k CPU and Pascal Titan X GPU implementations. It achieves 40x and 11.5x higher energy efficiency compared with the CPU and GPU respectively.

cs.CL

$L^p$ Solutions of Backward Stochastic Differential Equations with Jumps

Given $p \in (1, 2)$, we study $L^p$-solutions of a multi-dimensional backward stochastic differential equation with jumps (BSDEJ) whose generator may not be Lipschitz continuous in $(y,z)-$variables. We show that such a BSDEJ with a p-integrable terminal data admits a unique $L^p$ solution by approximating the monotonic generator by a sequence of Lipschitz generators via convolution with mollifiers and using a stability result.

math.PR

On the Robust Dynkin Game

We study a robust Dynkin game over a set of mutually singular probabilities. We first prove that for the conservative player of the game, her lower and upper value processes coincide (i.e. She has a value process $V $ in the game). Such a result helps people connect the robust Dynkin game with second-order doubly reflected backward stochastic differential equations. Also, we show that the value process $V$ is a submartingale under an appropriately defined nonlinear expectations up to the first time $τ_*$ when $V$ meets the lower payoff process $L$. If the probability set is weakly compact, one can even find an optimal triplet. The mutual singularity of probabilities in causes major technical difficulties. To deal with them, we use some new methods including two approximations with respect to the set of stopping times. The mutual singularity of probabilities causes major technical difficulties. To deal with them, we use some new methods including two approximations with respect to the set of stopping times

math.PR

Optimal Stopping with Random Maturity under Nonlinear Expectations

We analyze an optimal stopping problem with random maturity under a nonlinear expectation with respect to a weakly compact set of mutually singular probabilities $\mathcal{P}$. The maturity is specified as the hitting time to level $0$ of some continuous index process at which the payoff process is even allowed to have a positive jump. When $\mathcal{P}$ is a collection of semimartingale measures, the optimal stopping problem can be viewed as a {\it discretionary} stopping problem for a player who can influence both drift and volatility of the dynamic of underlying stochastic flow.

math.PR

Doubly Reflected BSDEs with Integrable Parameters and Related Dynkin Games

We study a doubly reflected backward stochastic differential equation (BSDE) with integrable parameters and the related Dynkin game. When the lower obstacle $L$ and the upper obstacle $U$ of the equation are completely separated, we construct a unique solution of the doubly reflected BSDE by pasting local solutions and show that the $Y-$component of the unique solution represents the value process of the corresponding Dynkin game under $g-$evaluation, a nonlinear expectation induced by BSDEs with the same generator $g$ as the doubly reflected BSDE concerned. In particular, the first time when process $Y $ meets $L$ and the first time when process $Y $ meets $U$ form a saddle point of the Dynkin game.

math.PR

Almost sure existence of Navier-Stokes Equations with randomized data in the whole space

This paper considers the supercritical Navier-Stokes equations posed in the whole space $\R^d$, with suitably randomized initial data, in the weak solution setting. The global weak solutions are constructed for a large set of initial data in $H^{-s}(\R^d)$ for some $s>0$ via a probabilistic argument, and this in turn implies the almost sure existence.

math.AP

Moon night sky brightness simulation for Xinglong station

With a sky brightness monitor in Xinglong station of National Astronomical Observatories of China (NAOC), we collected data from 22 dark clear nights and 90 lunar nights. We first measured the sky brightness variation with time in dark nights, found a clear correlation between the sky brightness and human activity. Then with a modified sky brightness model of moon night and data from moon night, we derived the typical value for several important parameters in the model. With these results, we calculated the sky brightness distribution under a given moon condition for Xinglong station. Furthermore, we simulated the moon night sky brightness distribution in a 5 degree field of view telescope (such as LAMOST). These simulations will be helpful to determine the magnitude limit, exposure time as well as the survey design for LAMOST at lunar night.

astro-ph.IM

A Weak Dynamic Programming Principle for Zero-Sum Stochastic Differential Games with Unbounded Controls

We analyze a zero-sum stochastic differential game between two competing players who can choose unbounded controls. The payoffs of the game are defined through backward stochastic differential equations. We prove that each player's priority value satisfies a weak dynamic programming principle and thus solves the associated fully non-linear partial differential equation in the viscosity sense.

math.PR

On the Robust Optimal Stopping Problem

We study a robust optimal stopping problem with respect to a set $\cP$ of mutually singular probabilities. This can be interpreted as a zero-sum controller-stopper game in which the stopper is trying to maximize its pay-off while an adverse player wants to minimize this payoff by choosing an evaluation criteria from $\cP$. We show that the \emph{upper Snell envelope $\ol{Z}$} of the reward process $Y$ is a supermartingale with respect to an appropriately defined nonlinear expectation $\ul{\sE}$, and $\ol{Z}$ is further an $\ul{\sE}-$martingale up to the first time $\t^*$ when $\ol{Z}$ meets $Y$. Consequently, $\t^*$ is the optimal stopping time for the robust optimal stopping problem and the corresponding zero-sum game has a value. Although the result seems similar to the one obtained in the classical optimal stopping theory, the mutual singularity of probabilities and the game aspect of the problem give rise to major technical hurdles, which we circumvent using some new methods.

math.PR