SearcharxivSearch

arXiv subjects

Leo Yang Yang

Publications and source records attributed to Leo Yang Yang.

4 recordsLinked to original sources

The Pulse Beneath the Job Title: Monthly Readings of Requirements and Tasks from 750 Million Chinese Job Ads

How do we define an occupation? By its job title? An accountant at a small trading company keeps the books; at a listed firm the same title demands a certified-accountant licence, and the week goes to the reports that regulators and the board read. Same title, different bar, different work. What defines an occupation is who it lets in and what it asks them to do. In a rapidly changing labor market, tracking those requirements and tasks is how to take the market's pulse. Yet no instrument reads both at the speed they change. Official occupational directories like O*NET report one national average per occupation, updated every few years. Job postings are timely but unstructured. Research built on them works from job titles plus proprietary skill keywords, which blur what is asked of a candidate into what a candidate is asked to do. The blur matters, because rising requirements and changing tasks are different events with different causes. We separate them. From 752.6 million job ads posted on China's five leading recruitment platforms between 2022 and 2026, we extract the phrases employers write, unify those that name the same thing, and validate the mapping from text back to entry. By doing so we construct two catalogs, 20,721 requirements a candidate must meet and 44,479 tasks the hire will do. With the entries standardized, we annotate them further. Each task, for example, carries a score for how far a language model could absorb it. Matched back onto every ad, the catalogs read the market month by month. Two examples show what the layer beneath the job title buys. First, the occupational registry records one accountant where the ads record a staircase, the junior certificate at the bottom of the wage range and the intermediate one at the top. Second, counting occupations says the work most exposed to language models is disappearing, and counting tasks says far less of it is.

econ.GN

Scaling Reproducibility: An AI-Assisted Workflow for Large-Scale Replication and Reanalysis

Computational reproducibility is central to scientific credibility, yet verifying published results at scale remains costly. We develop an AI-assisted workflow for automated full-paper replication -- retrieving materials, reconstructing environments, executing code, and matching outputs to point estimates reported in regression tables. We define a universe of all empirical and quantitative papers from the three top political science journals (2010--2025) and measure stated data availability using automated extraction. For a stratified sample of 384 studies, we apply the workflow to conduct full-paper replication, totaling 3,523 empirical models. We find that journal verification requirements, combined with data archiving mandates, drive reproducibility: the share of fully or largely reproducible papers rises from 20.8% before DA-RT adoption to 82.5% after, and conditional on accessible replication packages, 92.1% of papers are fully or largely reproducible (234/254). As a secondary application, we apply standardized IV diagnostics to 84 studies (597 IV specifications among 1,910 replicated models), illustrating how automated execution enables systematic reanalysis across heterogeneous empirical settings.

econ.EM

Interpretable Discriminative Text Representations via Agreement and Label Disentanglement

Interpretable text representations should expose coordinates that are not only predictive, but also meaningful enough for independent auditors to apply. Existing discriminative representations often use anonymous embedding directions, while concept-bottleneck and LLM-assisted methods attach natural-language names to features without ensuring that those definitions are reproducible or distinct from the target label. We propose an operational criterion for interpretable discriminative text representations: each coordinate should satisfy conceptual clarity, measured by chance-adjusted agreement between independent annotators applying the feature definition, and label disentanglement, meaning the feature should not merely paraphrase the prediction target. We instantiate this criterion in LLM-assisted Feature Discovery (LFD), an iterative method that proposes lexical and semantic features from contrastive outcome-opposed text pairs, screens candidates using cross-LLM Cohen's $κ$, and selects features by residual held-out predictive gain. A stylized analysis connects the $κ$ screen to a per-feature annotation-noise bound, formalizing agreement as a reliability check. Across ten text-classification tasks spanning seven corpora, LFD matches the predictive performance of a strong text bottleneck baseline while producing substantially clearer and less label-entangled features. Human audits with 232 raters show that LFD features achieve higher human--human and human--LLM agreement than baseline concepts, and raters consistently judge them as less label-leaking. These results suggest that agreement-tested, label-disentangled coordinates provide a practical auditability standard for interpretable text classification.

cs.CL

User Location Disclosure Fails to Deter Overseas Criticism but Amplifies Regional Divisions on Chinese Social Media

We examine the behavioral impact of a user location disclosure policy on Sina Weibo, China's largest microblogging platform, using a unique high-frequency dataset of uncensored engagement, including tens of thousands of comments and replies, on prominent government and media accounts. The policy, publicly justified as a measure to curb misinformation and counter foreign influence, was abruptly rolled out on April 28, 2022. Using an interrupted time series design, we find no decline in participation by overseas users. Instead, it significantly reduced domestic engagement with local issues outside users' home provinces, particularly among critical comments. Evidence indicates this decline was not driven by generalized fear or concerns about credibility, but by a surge in regionally discriminatory replies that raised the social cost of cross-provincial engagement. Our findings suggest that identity disclosure tools can have unintended consequences by activating existing social divisions in ways that reinforce state control without direct censorship.

econ.GN