SearcharxivSearch

arXiv subjects

Yutong Wu

Publications and source records attributed to Yutong Wu.

At least 19 recordsLinked to original sources

GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction

Whole-body motion tracking policies turn a humanoid into a robust control interface: the teleoperator---or an upstream model---only supplies a coarse movement intent, while the low-level policy keeps the robot balanced and physically feasible. Existing trackers deliver this interface only on flat ground: trained in empty scenes, they never learn how contact with terrain and objects reshapes their dynamics, and they attempt to teach the policy to balance under any command by continually enlarging the reference-motion corpus, which stops working once feasible behaviors become environment-dependent. We present GigaBrain-WBC-0.5, the first Behavior World Model (BWM) for humanoid whole-body control. Rather than a purely reactive tracker, we train a causal Transformer to jointly predict its next action, next state, and the distribution over its next latent behavior command, so the network that acts also models how the environment shapes what it can do next. An automatic terrain-annotation pipeline recovers full 3D contact geometry from retargeted motion, enabling terrain annotation at the scale of existing motion datasets. The predicted distribution is reused at deployment to detect implausible commands online and retract them onto learned behaviors, so the robot attempts tasks in a "best-effort" manner. The result is a unified policy that takes real-time command, interacts with environment, and stays robust to implausible commands, falls, and disturbances. GigaBrain-WBC-0.5 achieves the highest success rate across all four regimes among three large-scale tracker baselines: 81.3% on terrain interaction (4.3x the strongest baseline), 83.1% under implausible commands, and 99.3% fall recovery (16.8x the strongest baseline). Hardware trials show robust interaction under missing supports and disturbances; the Unitree G1 checkpoint transfers to the Maker L01 robot with simple fine-tuning.

cs.RO

Targeted Counterfactual Fingerprinting for Black-Box LLM Ownership Verification

Large language models (LLMs) are high-value assets that can be derived through redeployment, fine-tuning, quantization, or further alignment. Because deployed LLMs are commonly exposed only through query APIs, ownership verification must often rely on black-box text responses. This setting is difficult: generations are open-ended and can vary across repeated queries, while existing black-box fingerprints rely on signals that are fragile under a final-response interface, including full-text matching, soft behavioral features, or model-specific prompts designed not to transfer. We propose TCF (Targeted Counterfactual Fingerprinting), a black-box LLM fingerprinting framework that converts open-ended generation comparison into constrained-answer targeted counterfactual transfer. TCF restricts each verification query to a finite answer space, reducing the surface-form ambiguity that enters the verification score, and optimizes a prompt perturbation toward a counterfactual target different from the protected model's clean answer on the original prompt. Verification reduces to checking whether the suspect model's parsed final answer matches the recorded target. We introduce the source-model counterfactual margin (SCM), a protected-model-only quantity that certifies the target is unlikely before the perturbation and likely after it; SCM controls target selection, perturbation stopping, and fingerprint filtering. Under explicit derived-preservation and independent-transfer budgets motivated by local behavioral closeness, we derive a target-accuracy gap between derived and independent models. Across four LLM families, TCF achieves an average AUC of 0.9861, improving over TRAP, ProFLingo, and ZeroPrint by 0.07 to 0.19.

cs.CR

Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents

LLM-based web agents are increasingly deployed in real-world settings such as e-commerce, where they interact extensively with untrusted web content while executing actions that carry direct financial consequences. This makes them vulnerable to prompt-injection attacks, in which seemingly benign web content conceals adversarial instructions that manipulate the agent's behavior. Existing security benchmarks adopt an \textit{attack-centric} perspective, focusing on the technical feasibility of injections while overlooking the nuanced distribution of resulting harms. In practice, however, prompt-injection risk is victim-dependent: a single exploit can produce asymmetric consequences for different stakeholders, and the same attack pattern may exhibit substantially different effectiveness depending on whom it targets. To capture these properties, we introduce StakeBench, a stakeholder-centric benchmark that systematically categorizes and attributes harm in real-world web agent systems for online shopping. In general, StakeBench decomposes prompt-injection risk into 12 concrete attack objectives across three stakeholder classes, realized by 22 reusable templates and instantiated into 264 executable adversarial cases spanning 12 product categories, with each case evaluated along complementary outcome- and process-level metrics. Evaluating four deployable agent-backbone configurations across 3,168 attacked runs, we find substantial and heterogeneous vulnerabilities: no attack objective is reliably resisted by current LLM-based web agents, and outcomes span four qualitatively distinct modes. These patterns are missed by conventional attack-centric, single-metric evaluation, underscoring the need for stakeholder-aware assessment of LLM-based agents in real-world deployments.

cs.CR

How Millions Coordinate at Scale: Engagement, Collaboration, and Conflict in Three Editions of Reddit r/place

Mass peer-production environments are shaped by a complex interplay between decentralized coordination, platform design, and potential conflict over resources. While online infrastructures enable large-scale collaboration, they can also introduce coordination challenges and contested interactions among participants. The Reddit r/place experiment provides a unique socio-technical setting for studying these dynamics across three distinct editions (2017, 2022, and 2023). By allowing millions of participants to collaborate and compete as they update pixels on a finite shared digital canvas, the events generated fine-grained traces comprising hundreds of millions of actions taken over multiple days. In this paper, we examine how participation, collaboration, and conflict evolve within r/place. First, we conduct a longitudinal, cross-edition analysis of engagement and collaboration patterns across the 2017, 2022, and 2023 events. Second, to investigate collaborative activity beyond surviving final artifacts, we introduce a scalable graph-based dynamic clustering framework and apply it to the 2017 event to reconstruct coalition trajectories from behavioral interaction logs, enabling the analysis of both persistent and transient efforts. Our findings reveal recurring organizational patterns across editions: participation remains highly concentrated to few participants despite individual rate-limiting constraints, larger coalitions exhibit both greater coordination inefficiencies and lower median per-participant activity, while still success increasingly concentrates within large collaborative groups over the course of an event. Analysis of recovered coalition trajectories further shows that coalition outcomes are difficult to predict based on state characteristics during much of the event, highlighting the dynamic and contested nature of collaborative production in r/place.

cs.SI

Learning Roller-Skating Motions of Humanoid Robots Based on Adversarial Motion Priors

Humanoid roller-skating is difficult because the robot must coordinate whole-body balance, rolling contacts, and velocity-dependent posture regulation. This paper presents an adversarial motion prior based reinforcement learning framework for two humanoid roller-skating gaits: Pump Glide skating and Push Glide skating. The two gait datasets are collected independently through motion capture and retargeted to the humanoid robot separately. The retargeted data are then smoothed and resampled into reference motion states for AMP training. The two gaits are learned by independent AMP training pipelines with separate reference datasets, separate policies, and independent reward architectures. Simulation experiments are designed to evaluate gait quality, velocity tracking, turning, and gait-specific reward ablations.

cs.RO

On soliton clusters and collision blow up for the $L^2$-critical Hartree equation

We consider the $L^2$-critical nonlinear Hartree equation in $\mathbb{R}^{1+4}$ and multisoliton solutions for which the trajectories are approximated to leading order by an $m$-body law. We obtain soliton clusters asymptotically following hyperbolic-parabolic trajectories of the corresponding $m$-body problem. By pseudo-conformal invariance, we then conclude finite-time collision blow-up with any number of clusters, each consisting of an arbitrary number of solitons, colliding simultaneously at distinct prescribed points.

math.AP

Multisoliton solutions and blow up for the $L^2$-critical Hartree equation

We construct multisoliton solutions for the $L^2$-critical Hartree equation with trajectories asymptotically obeying a many-body law for an inverse square potential. Precisely, we consider the $m$-body hyperbolic and parabolic non-trapped dynamics. The pseudo-conformal symmetry then implies finite-time collision blow up in the latter case and a solution blowing up at $m$ distinct points in the former case. The approach we take is based on the ideas of [Krieger-Martel-Raphaël, 2009] and the third author's recent extension [Wu, 2026]. The approximation scheme requires new aspects in order to deal with a certain degeneracy for generalized root space elements.

math.AP

SwitchPatch: Physical Adversarial Attack Strategy with Switchable Adversarial Objectives

Physical adversarial patch (PAP) attacks attach carefully crafted patches to physical objects to manipulate a deployed model. However, existing PAP attacks suffer from several limitations. First, existing patches remain continuously active, which prevents selective targeting of specific attack objectives and compromises stealth. Second, these approaches require target device access or hardware configuration knowledge, and often rely on costly external equipment. To address these limitations, this paper introduces SwitchPatch, a novel physical adversarial attack strategy that employs a physically static adversarial patch yet can be triggered to produce dynamic and controllable attack effects. Unlike existing approaches, SwitchPatch can transition between states through predefined triggers, enabling adaptation to dynamic environments. Moreover, to improve stealth, we design two trigger patterns: one overlapping with the patch and another spatially separated from it. These triggers can be implemented at low cost without target device access or hardware configuration knowledge. We make three contributions. First, we provide theoretical and empirical analysis to establish the feasibility of SwitchPatch and characterize the number of attack objectives it can support. Second, we develop a gradient-based framework for static yet switchable attacks through diverse trigger patterns. Third, we conduct extensive Unmanned Ground Vehicle (UGV) experiments to validate the effectiveness, transferability, and robustness of SwitchPatch.

cs.CR

Wavefront super-resolution for Adaptive Optics systems on ground-based telescopes

In ground-based astronomy, Adaptive Optics (AO) is a pivotal technique, engineered to correct wavefront phase distortions and thereby enhance the quality of the observed images. Integral to an AO system is the wavefront sensor (WFS), which is crucial for detecting wavefront aberrations from guide stars, essential for phase calculations. Many models based on a single-WFS model have been proposed to obtain the high-resolution phase of the incoming wavefront. In this paper, we delve into the realm of multiple WFSs within the framework of state-of-the-art telescope setups for high-resolution phase reconstruction. We propose a model for reconstructing a high-resolution wavefront from a sequence of wavefront gradient data from multiple WFSs in a multi-frame post-processing setting. Our model is based on the turbulence statistics and the Taylor frozen flow hypothesis, incorporating knowledge of the wind velocities in atmospheric turbulence layers. We also introduce an $H_2$ regularization term, especially for atmospheric characteristics under von Karman statistics, and provide a theoretical analysis for $H^2$ space within $H^{11/6}$. Numerical simulations are conducted to demonstrate the robustness and effectiveness of our regularization term and multi-WFS reconstruction strategy under identical experimental conditions.

astro-ph.IM

QiMeng-CodeV-SVA: Training Specialized LLMs for Hardware Assertion Generation via RTL-Grounded Bidirectional Data Synthesis

SystemVerilog Assertions (SVAs) are crucial for hardware verification. Recent studies leverage general-purpose LLMs to translate natural language properties to SVAs (NL2SVA), but they perform poorly due to limited data. We propose a data synthesis framework to tackle two challenges: the scarcity of high-quality real-world SVA corpora and the lack of reliable methods to determine NL-SVA semantic equivalence. For the former, large-scale open-source RTLs are used to guide LLMs to generate real-world SVAs; for the latter, bidirectional translation serves as a data selection method. With the synthesized data, we train CodeV-SVA, a series of SVA generation models. Notably, CodeV-SVA-14B achieves 75.8% on NL2SVA-Human and 84.0% on NL2SVA-Machine in Func.@1, matching or exceeding advanced LLMs like GPT-5 and DeepSeek-R1.

cs.CL

QiMeng-CodeV-R1: Reasoning-Enhanced Verilog Generation

Large language models (LLMs) trained via reinforcement learning with verifiable reward (RLVR) have achieved breakthroughs on tasks with explicit, automatable verification, such as software programming and mathematical problems. Extending RLVR to electronic design automation (EDA), especially automatically generating hardware description languages (HDLs) like Verilog from natural-language (NL) specifications, however, poses three key challenges: the lack of automated and accurate verification environments, the scarcity of high-quality NL-code pairs, and the prohibitive computation cost of RLVR. To this end, we introduce CodeV-R1, an RLVR framework for training Verilog generation LLMs. First, we develop a rule-based testbench generator that performs robust equivalence checking against golden references. Second, we propose a round-trip data synthesis method that pairs open-source Verilog snippets with LLM-generated NL descriptions, verifies code-NL-code consistency via the generated testbench, and filters out inequivalent examples to yield a high-quality dataset. Third, we employ a two-stage "distill-then-RL" training pipeline: distillation for the cold start of reasoning abilities, followed by adaptive DAPO, our novel RLVR algorithm that can reduce training cost by adaptively adjusting sampling rate. The resulting model, CodeV-R1-7B, achieves 68.6% and 72.9% pass@1 on VerilogEval v2 and RTLLM v1.1, respectively, surpassing prior state-of-the-art by 12~20%, while even exceeding the performance of 671B DeepSeek-R1 on RTLLM. We have released our model, training code, and dataset to facilitate research in EDA and LLM communities.

cs.LG

Nanodiamond-Enabled Torsion Microscopy Uncovers Multidimensional Cell-Matrix Mechanical Interactions

Traditional cellular force-sensing techniques, such as traction force microscopy (TFM), are predominantly limited to measuring linear tractions, overlooking and technically unable to capture the nanoscale torsional forces that are critical in cell-matrix interactions. Here, we introduce a nanodiamond-enabled torsion microscopy (DTM) that integrates nitrogen-vacancy (NV) centers as orientation markers with micropillar arrays to decouple and quantify nanoscale rotational and translational motions induced by cells. This approach achieves high precision (~1.47 degree rotational accuracy and ~3.13*10-15 Nm torque sensitivity), enabling reconstruction of cellular torsional force fields and twisting energy distributions previously underestimated. Our findings reveal the widespread presence of torsional forces in cell-matrix interactions, introducing "cellular mechanical modes" where different adhesion patterns dictate the balance between traction- and torque- mediated mechanical energy transferred to the substrate. Notably, in immune cells like macrophages that generally exert low linear tractions, torque overwhelmingly dominates traction, highlighting a unique mechanical output for specific cellular functions. By uncovering these differential modes, DTM provides a versatile tool to advance biomechanical investigations, with potential applications in disease diagnostics and therapeutics.

physics.bio-ph

StepFun-Formalizer: Unlocking the Autoformalization Potential of LLMs through Knowledge-Reasoning Fusion

Autoformalization aims to translate natural-language mathematical statements into a formal language. While LLMs have accelerated progress in this area, existing methods still suffer from low accuracy. We identify two key abilities for effective autoformalization: comprehensive mastery of formal-language domain knowledge, and reasoning capability of natural language problem understanding and informal-formal alignment. Without the former, a model cannot identify the correct formal objects; without the latter, it struggles to interpret real-world contexts and map them precisely into formal expressions. To address these gaps, we introduce ThinkingF, a data synthesis and training pipeline that improves both abilities. First, we construct two datasets: one by distilling and selecting large-scale examples rich in formal knowledge, and another by generating informal-to-formal reasoning trajectories guided by expert-designed templates. We then apply SFT and RLVR with these datasets to further fuse and refine the two abilities. The resulting 7B and 32B models exhibit both comprehensive formal knowledge and strong informal-to-formal reasoning. Notably, StepFun-Formalizer-32B achieves SOTA BEq@1 scores of 40.5% on FormalMATH-Lite and 26.7% on ProverBench, surpassing all prior general-purpose and specialized models.

cs.CL

Early Lung Cancer Diagnosis from Virtual Follow-up LDCT Generation via Correlational Autoencoder and Latent Flow Matching

Lung cancer is one of the most commonly diagnosed cancers, and early diagnosis is critical because the survival rate declines sharply once the disease progresses to advanced stages. However, achieving an early diagnosis remains challenging, particularly in distinguishing subtle early signals of malignancy from those of benign conditions. In clinical practice, a patient with a high risk may need to undergo an initial baseline and several annual follow-up examinations (e.g., CT scans) before receiving a definitive diagnosis, which can result in missing the optimal treatment. Recently, Artificial Intelligence (AI) methods have been increasingly used for early diagnosis of lung cancer, but most existing algorithms focus on radiomic features extraction from single early-stage CT scans. Inspired by recent advances in diffusion models for image generation, this paper proposes a generative method, named CorrFlowNet, which creates a virtual, one-year follow-up CT scan after the initial baseline scan. This virtual follow-up would allow for an early detection of malignant/benign nodules, reducing the need to wait for clinical follow-ups. During training, our approach employs a correlational autoencoder to encode both early baseline and follow-up CT images into a latent space that captures the dynamics of nodule progression as well as the correlations between them, followed by a flow matching algorithm on the latent space with a neural ordinary differential equation. An auxiliary classifier is used to further enhance the diagnostic accuracy. Evaluations on a real clinical dataset show our method can significantly improve downstream lung nodule risk assessment compared with existing baseline models. Moreover, its diagnostic accuracy is comparable with real clinical CT follow-ups, highlighting its potential to improve cancer diagnosis.

cs.CV

Cost-Optimal Grouped-Query Attention for Long-Context Modeling

Grouped-Query Attention (GQA) is a widely adopted strategy for reducing the computational cost of attention layers in large language models (LLMs). However, current GQA configurations are often suboptimal because they overlook how context length influences inference cost. Since inference cost grows with context length, the most cost-efficient GQA configuration should also vary accordingly. In this work, we analyze the relationship among context length, model size, GQA configuration, and model loss, and introduce two innovations: (1) we decouple the total head size from the hidden size, enabling more flexible control over attention FLOPs; and (2) we jointly optimize the model size and the GQA configuration to arrive at a better allocation of inference resources between attention layers and other components. Our analysis reveals that commonly used GQA configurations are highly suboptimal for long-context scenarios. More importantly, we propose a recipe for deriving cost-optimal GQA configurations. Our results show that for long-context scenarios, one should use fewer attention heads while scaling up model size. Configurations selected by our recipe can reduce both memory usage and FLOPs by more than 50% compared to Llama-3's GQA, with *no degradation in model capabilities*. Our findings offer valuable insights for designing efficient long-context LLMs. The code is available at https://www.github.com/THUNLP/cost-optimal-gqa .

cs.CL