SearcharxivSearch

arXiv subjects

Chenxin Liu

Publications and source records attributed to Chenxin Liu.

6 recordsLinked to original sources

HEFT: Heavy-Payload Full-size Humanoid Teleoperation with Privileged Motion Guidance and Windowed Payload Curriculum

General motion tracking and teleoperation offer a promising path to scalable humanoid skill acquisition, yet most existing frameworks are validated on compact platforms or without real payload interaction, leaving full-size humanoids with real payloads largely unexplored. Scaling to full-size humanoids introduces two compounding challenges: their larger inertia and tighter balance margins make tracking highly sensitive to noise, drift, and retargeting errors from commodity VR trackers, while their payload potential remains largely underutilized. We present HEFT, a heavy-payload full-size humanoid teleoperation framework that addresses both challenges. HEFT learns from deployable noisy VR references with physically plausible reconstructed references through Privileged Motion Guidance (PMG), and uses a Windowed Payload Curriculum (WPC) with expert-guided payload caps to acquire robust heavy-payload tracking. We deploy HEFT on L7, a 175cm, 65kg humanoid. The robot tracks motions including turns, forward/backward locomotion, and squats under payloads up to 24kg.

cs.RO

MADE: Beyond Scoring via a Multilingual Agentic Diagnosing Engine for Fine-Grained Evaluation Insights

Multilingual and multicultural benchmarks now cover dozens of languages and model families, but the resulting score landscapes remain metric-rich and insight-poor, necessitating fine-grained multilingual post-evaluation diagnosis. However, single LLMs and open-ended agents are easily swamped by the long, noisy diagnostic input, and no reusable taxonomy exists for it. To address this, we propose MADE, a Multilingual Agentic Diagnosing Engine that decomposes post-evaluation analysis into planning, aggregate analysis, instance-level case inspection, multilingual and cultural reflection, and grounded report synthesis. MADE is paired with an expert-led 54-query and 15-language diagnostic set, evaluated on top of a large-scale multilingual evaluation substrate (33 model families, 11 benchmarks, 26 languages, 34 cultures, 8.66M evaluation records). Experiments show that MADE outperforms the strongest shared baseline by 47% in diagnosis report quality and is preferred by human multilingual experts in 87.9% of pairwise comparisons. Applied with multilingual experts, MADE further surfaces four actionable findings on deployment, iteration, and cross-cultural pitfalls, turning benchmark score tables into model-selection and remediation guidance.

cs.CL

The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models

Evaluating the multilingual and multicultural capabilities of Large Language Models (LLMs) is essential for their global utility. However, current benchmarks face three critical limitations: (1) fragmented evaluation dimensions that often neglect deep cultural nuances; (2) insufficient language coverage in subjective tasks relying on low-quality machine translation; and (3) shallow analysis that lacks diagnostic depth beyond simple rankings. To address these, we introduce GaoYao, a comprehensive benchmark with 182.3k samples, 26 languages and 51 nations/areas. First, GaoYao proposes a unified framework categorizing evaluation tasks into three cultural layers (General Multilingual, Cross-cultural, Monocultural) and nine cognitive sub-layers. Second, we achieve native-quality expansion by leveraging experts to rigorously localize subjective benchmarks into 19 languages and synthesizing cross-cultural test sets for 34 cultures, surpassing prior coverage by up to 111%. Third, we conduct an in-depth diagnostic analysis on 20+ flagship and compact LLMs. Our findings reveal significant geographical performance disparities and distinct gaps between tasks, offering a reliable map for future work. We release the benchmark (https://github.com/lunyiliu/GaoYao).

cs.CL

C-Mining: Unsupervised Discovery of Seeds for Cultural Data Synthesis via Geometric Misalignment

Achieving cultural alignment in Large Language Models (LLMs) increasingly depends on synthetic data generation. For such synthesis, the most vital initial step is seed curation; however, current methods lack quantifiable standards for selecting these seeds. Existing approaches rely on unscalable manual curation or bias-prone LLM extraction, treating cultural specificity as an abstract concept rather than a measurable signal. In this paper, we address this "quantification gap" by proposing C-Mining, an unsupervised framework that transforms the discovery of cultural seeds from a subjective selection process into a computable data mining formulation. Our approach exploits a novel geometric insight, leveraging the cross-lingual misalignment of cultural concepts within pre-trained embedding spaces as a quantifiable discovery signal. By systematically identifying these regions characterized by pronounced linguistic exclusivity and geometric isolation, while actively filtering out noise, C-Mining automatically extracts high-fidelity Culture Points (CPs) from raw multilingual corpora without reliance on human or LLM supervision, reducing preparation costs by more than 150-fold. We further leverage the mined knowledge to steer the synthesis of diverse instruction-tuning datasets. Extensive experiments demonstrate that this seed-centric approach significantly enhances cultural understanding and reasoning capabilities, achieving a +6.03 point improvement on CulturalBench-Hard and surpassing state-of-the-art baselines, providing a scalable, quantifiable solution for high-quality cultural data synthesis.

cs.CL

Revealing spatio-temporal interaction patterns behind complex cities

Cities are typical dynamic complex systems that connect people and facilitate interactions. Revealing universal collective patterns behind spatio-temporal interactions between residents is crucial for various urban studies, of which we are still lacking a comprehensive understanding. Massive cellphone data enable us to construct interaction networks based on spatio-temporal co-occurrence of individuals. The rank-size distributions of hourly dynamic population of locations are stable, although people are almost constantly moving in cities and hotspots that attract people are changing over time in a day. A larger city is of a stronger heterogeneity as indicated by a larger scaling exponent. After aggregating spatio-temporal interaction networks over consecutive time windows, we reveal a switching behavior of cities between two states. During the "active" state, the whole city is concentrated in fewer larger communities; while in the "sleeping" state, people are scattered in more smaller communities. Above discoveries are universal over diversified cities across continents. In addition, a city sleeps less, when its population grows larger. And spatio-temporal interaction segregation can be well approximated by residential segregation in smaller cities, but not in larger ones. We propose a temporal-population-weighted-opportunity model by integrating time-dependent departure probability to make dynamic predictions on human mobility, which can reasonably well explain observed patterns of spatio-temporal interactions in cities.

physics.soc-ph

Quantifying relation between mobility patterns and socioeconomic status of dockless sharing-bike users

Bikes are among the healthiest, greenest, and most affordable means of transportation for a better future city, but mobility patterns of riders with different income were rarely studied due to limitations on collecting data. Newly emergent dockless bike-sharing platforms that record detailed information regarding each trip provide us a unique opportunity. Attribute to its better usage flexibility and accessibility, dockless bike-sharing platforms are booming over the past a few years worldwide and reviving the riding fashion in cities. In this work, by exploiting massive riding records in two megacities from a dockless bike-sharing platform, we reveal that individual mobility patterns, including radius of gyration and average travel distance, are similar among users with different income, which indicates that human beings all follow similar physical rules. However, collective mobility patterns, including average range and diversity of visitation, and commuting directions, all exhibit different behaviors and spatial patterns across income categories. Hotspot locations that attract more cycling activities are quite different over groups, and locations where users reside are of a low user ratio for both higher and lower income groups. Lower income groups are inclined to visit less flourishing locations, and commute towards the direction to the city center in both cities, and of a smaller mobility diversity in Beijing but a larger diversity in Shanghai. In addition, differences on mobility patterns among socioeconomic categories are more evident in Beijing than in Shanghai. Our findings would be helpful on designing better promotion strategies for dockless bike-sharing platforms and towards the transition to a more sustainable green transportation.

physics.soc-ph