SearcharxivSearch

arXiv subjects

Yifei Yao

Publications and source records attributed to Yifei Yao.

12 recordsLinked to original sources

CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification

Anthropic proposes the concept of skills for LLM agents to tackle multi-step professional tasks that simple tool invocations cannot address. A tool is a single, self-contained function, whereas a skill is a structured bundle of interdependent multi-file artifacts. Currently, skill generation is not only label-intensive due to manual authoring, but also may suffer from human--machine cognitive misalignment, which can lead to degraded agent performance, as evidenced by evaluations on SkillsBench. Therefore, we aim to enable agents to autonomously generate skills. However, existing self-evolving methods designed for tools cannot be directly applied to skills due to their increased complexity. To address these issues, we propose CoEvoSkills, a self-evolving skills framework that enables agents to autonomously construct complex, multi-file skill packages. Specifically, CoEvoSkills couples a Skill Generator that iteratively refines skills with a Surrogate Verifier that co-evolves to provide informative and actionable feedback without access to ground-truth test content. On SkillsBench, CoEvoSkills outperforms five baselines on both Claude Code and Codex, and generalizes strongly to six additional LLMs. The code is publicly available at https://github.com/Zhang-Henry/CoEvoSkills.

cs.AI

Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation

Real-world deployment of embodied agents requires active exploration, visual grounding, and interactive intent disambiguation. However, existing frameworks often rely on privileged simulator states or assume complete instructions, bypassing realistic deployment challenges. To bridge this gap, we present REAL, an agentic framework for open-world mobile manipulation. REAL establishes sim-to-real-consistent environment APIs without oracle perception and integrates a simulated user to enable human-in-the-loop interaction. Within this environment, we design diverse task compositions to drive data collection, supervised fine-tuning, and online reinforcement learning, systematically optimizing agent performance. To comprehensively evaluate this approach, we introduce REAL-Bench, a benchmark spanning 241 tasks across active exploration, visual distraction, articulated manipulation, and interactive disambiguation. Experimental results demonstrate that our trained agent outperforms leading commercial closed-source VLMs on interactive tasks with a 56.9% success rate. Further empirical analysis reveals that our hierarchical training pipeline successfully aligns the model's tool-use capabilities while maintaining robust open-vocabulary reasoning under extended exploration horizons. Finally, we deploy and evaluate our framework on a physical dual-arm mobile robot, where it achieves a 78.3% end-to-end success rate over 60 real-world episodes. These physical trials demonstrate robust zero-shot transferability to unseen household scenarios, validating that our sim-to-real-consistent design successfully bridges the reality gap for long-horizon mobile manipulation. Code is available at https://github.com/InternRobotics/REAL.

cs.CV

AdaptiveK: Complexity-Driven Sparse Autoencoders for Interpretable Language Model Representations

Understanding the internal representations of large language models (LLMs) remains a central challenge for interpretability research. Sparse autoencoders (SAEs) offer a promising solution by decomposing activations into interpretable features, but existing approaches rely on fixed sparsity constraints that fail to account for input complexity. We propose AdaptiveK SAE (Adaptive Top K Sparse Autoencoders), a novel framework that dynamically adjusts sparsity levels based on the semantic complexity of each input. Leveraging linear probes, we demonstrate that context complexity is linearly encoded in LLM representations, and we use this signal to guide feature allocation during training. Experiments across ten language models demonstrate that this complexity-driven adaptation outperforms fixed-sparsity approaches on reconstruction fidelity, explained variance, cosine similarity and interpretability metrics while eliminating the burden of extensive hyperparameter tuning. Our code is available at: https://github.com/hiyukie/adaptiveK.

cs.LG

GBC: Generalized Behavior-Cloning Framework for Whole-Body Humanoid Imitation

The creation of human-like humanoid robots is hindered by a fundamental fragmentation: data processing and learning algorithms are rarely universal across different robot morphologies. This paper introduces the Generalized Behavior Cloning (GBC) framework, a comprehensive and unified solution designed to solve this end-to-end challenge. GBC establishes a complete pathway from human motion to robot action through three synergistic innovations. First, an adaptive data pipeline leverages a differentiable IK network to automatically retarget any human MoCap data to any humanoid. Building on this foundation, our novel DAgger-MMPPO algorithm with its MMTransformer architecture learns robust, high-fidelity imitation policies. To complete the ecosystem, the entire framework is delivered as an efficient, open-source platform based on Isaac Lab, empowering the community to deploy the full workflow via simple configuration scripts. We validate the power and generality of GBC by training policies on multiple heterogeneous humanoids, demonstrating excellent performance and transfer to novel motions. This work establishes the first practical and unified pathway for creating truly generalized humanoid controllers.

cs.RO

Visible Brillouin-quadratic microlaser in a high-Q thin-film lithium niobate microdisk

Narrow-linewidth lasers at short/visible wavelengths are crucial for quantum and atomic applications, such as atomic clocks, quantum computing, atomic and molecular spectroscopy, and quantum sensing. However, such lasers are often only accessible in bulky tabletop systems and remain scarce in integrated photonic platform. Here, we report an on-chip visible Brillouin-quadratic microlaser in a 117-um-diameter thin-film lithium niobate (TFLN) microdisk via dispersion engineering. Enabled by the ultra-high Q factor of 4.0X10(6) and small mode volume, strong photon-phonon interaction and high second-order nonlinearity of the TFLN microdisk, narrow-linewidth Stokes Brillouin lasing (SBL) is demonstrated with 10.17 GHz Brillouin shift under a 1560-nm pump, exhibiting a short-term narrow linewidth of 254 Hz and a low threshold of only 1.81 mW. Meanwhile, efficient second harmonic generation (SHG) of the SBL signal is also observed at 780 nm, with a normalized conversion efficiency of 3.61%/mW, made possible by simultaneous phase matching fulfillments for both narrow-linewidth SBL and its SHG. This demonstration of an integrated ultra-narrow linewidth visible wavelength Brillouin-quadratic lasers opens new avenues toward chip-scale quantum information processing and precise metrology.

physics.optics

Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents

Although LLM-based agents, powered by Large Language Models (LLMs), can use external tools and memory mechanisms to solve complex real-world tasks, they may also introduce critical security vulnerabilities. However, the existing literature does not comprehensively evaluate attacks and defenses against LLM-based agents. To address this, we introduce Agent Security Bench (ASB), a comprehensive framework designed to formalize, benchmark, and evaluate the attacks and defenses of LLM-based agents, including 10 scenarios (e.g., e-commerce, autonomous driving, finance), 10 agents targeting the scenarios, over 400 tools, 27 different types of attack/defense methods, and 7 evaluation metrics. Based on ASB, we benchmark 10 prompt injection attacks, a memory poisoning attack, a novel Plan-of-Thought backdoor attack, 4 mixed attacks, and 11 corresponding defenses across 13 LLM backbones. Our benchmark results reveal critical vulnerabilities in different stages of agent operation, including system prompt, user prompt handling, tool usage, and memory retrieval, with the highest average attack success rate of 84.30\%, but limited effectiveness shown in current defenses, unveiling important works to be done in terms of agent security for the community. We also introduce a new metric to evaluate the agents' capability to balance utility and security. Our code can be found at https://github.com/agiresearch/ASB.

cs.CR

Simultaneously generating Brillouin microlaser and second harmonic within a lithium niobate microdisk

We report the simultaneous generation of second-harmonic generation (SHG) and Brillouin microlaser in a high-quality thin-film lithium niobate (TFLN) microdisk resonator. The microdisk is fabricated with ultrahigh-Q factor of 4X10(6) by photolithography-assisted chemo-mechanical etching, enabling significant cavity-enhancement effect for boosting nonlinear frequency conversion. Under 1559.632 nm pumping, Brillouin microlaser is demonstrated in the microdisk with Stokes Brillouin shift of 10 GHz, a low threshold of 1.81 mW, and a fundamental linewidth of 254.365 Hz. Meanwhile, efficient SHG is observed at 779.816 nm with an absolute conversion efficiency of 3.8% at pump level of 3.028 mW. The coexistence of these two nonlinear processes is enabled by the simultaneous confinement of the light and acoustic fileds for effect coupling in the microdisk, which enhances both optomechanical and second-order nonlinear interactions. This research provides new possibilities for integrated multi-frequency laser sources and multifunctional nonlinear photonic devices.

physics.optics

AnyBipe: An End-to-End Framework for Training and Deploying Bipedal Robots Guided by Large Language Models

Training and deploying reinforcement learning (RL) policies for robots, especially in accomplishing specific tasks, presents substantial challenges. Recent advancements have explored diverse reward function designs, training techniques, simulation-to-reality (sim-to-real) transfers, and performance analysis methodologies, yet these still require significant human intervention. This paper introduces an end-to-end framework for training and deploying RL policies, guided by Large Language Models (LLMs), and evaluates its effectiveness on bipedal robots. The framework consists of three interconnected modules: an LLM-guided reward function design module, an RL training module leveraging prior work, and a sim-to-real homomorphic evaluation module. This design significantly reduces the need for human input by utilizing only essential simulation and deployment platforms, with the option to incorporate human-engineered strategies and historical data. We detail the construction of these modules, their advantages over traditional approaches, and demonstrate the framework's capability to autonomously develop and refine controlling strategies for bipedal robot locomotion, showcasing its potential to operate independently of human intervention.

cs.RO

Class Incremental Fault Diagnosis under Limited Fault Data via Supervised Contrastive Knowledge Distillation

Class-incremental fault diagnosis requires a model to adapt to new fault classes while retaining previous knowledge. However, limited research exists for imbalanced and long-tailed data. Extracting discriminative features from few-shot fault data is challenging, and adding new fault classes often demands costly model retraining. Moreover, incremental training of existing methods risks catastrophic forgetting, and severe class imbalance can bias the model's decisions toward normal classes. To tackle these issues, we introduce a Supervised Contrastive knowledge distiLlation for class Incremental Fault Diagnosis (SCLIFD) framework proposing supervised contrastive knowledge distillation for improved representation learning capability and less forgetting, a novel prioritized exemplar selection method for sample replay to alleviate catastrophic forgetting, and the Random Forest Classifier to address the class imbalance. Extensive experimentation on simulated and real-world industrial datasets across various imbalance ratios demonstrates the superiority of SCLIFD over existing approaches. Our code can be found at https://github.com/Zhang-Henry/SCLIFD_TII.

cs.LG

From Uncertainty to Clarity: Uncertainty-Guided Class-Incremental Learning for Limited Biomedical Samples via Semantic Expansion

In real-world clinical settings, data distributions evolve over time, with a continuous influx of new, limited disease cases. Therefore, class incremental learning is of great significance, i.e., deep learning models are required to learn new class knowledge while maintaining accurate recognition of previous diseases. However, traditional deep neural networks often suffer from severe forgetting of prior knowledge when adapting to new data unless trained from scratch, which undesirably costs much time and computational burden. Additionally, the sample sizes for different diseases can be highly imbalanced, with newly emerging diseases typically having much fewer instances, consequently causing the classification bias. To tackle these challenges, we are the first to propose a class-incremental learning method under limited samples in the biomedical field. First, we propose a novel cumulative entropy prediction module to measure the uncertainty of the samples, of which the most uncertain samples are stored in a memory bank as exemplars for the model's later review. Furthermore, we theoretically demonstrate its effectiveness in measuring uncertainty. Second, we developed a fine-grained semantic expansion module through various augmentations, leading to more compact distributions within the feature space and creating sufficient room for generalization to new classes. Besides, a cosine classifier is utilized to mitigate classification bias caused by imbalanced datasets. Across four imbalanced data distributions over two datasets, our method achieves optimal performance, surpassing state-of-the-art methods by as much as 53.54% in accuracy.

cs.CV

Structural and Dynamical Mechanisms of a Naturally Occurring Variant of the Human Prion Protein in Preventing Prion Conversion

Prion diseases are associated with the misfolding of the normal helical cellular form of prion protein (PrPC) into the beta-sheet-rich scrapie form (PrPSc) and the subsequent aggregation of PrPSc into amyloid fibrils. Recent studies demonstrated that a naturally occurring variant V127 of human PrPC is intrinsically resistant to prion conversion and aggregation, and can completely prevent prion diseases. However, the underlying molecular mechanism remains elusive. Herein we perform multiple microsecond molecular dynamics simulations on both wildtype (WT) and V127 variant of human PrPC to understand at atomic level the protective effect of V127 variant. Our simulations show that G127V mutation not only increases the rigidity of the S2-H2 loop between strand-2 (S2) and helix-2 (H2), but also allosterically enhances the stability of the H2 C-terminal region. Interestingly, previous studies reported that animals with rigid S2-H2 loop usually do not develop prion diseases, and the increase in H2 C-terminal stability can prevent misfolding and oligomerization of prion protein. The allosteric paths from G/V127 to H2 C-terminal region are identified using dynamical network analyses. Moreover, community network analyses illustrate that G127V mutation enhances the global correlations and intra-molecular interactions of PrP, thus stabilizing the overall PrPC structure and inhibiting its conversion into PrPSc. This study provides mechanistic understanding of human V127 variant in preventing prion conversion which may be helpful for the rational design of potent anti-prion compounds.

physics.bio-ph

Enhanced optical nonlinearities in air-cladding silicon pedestal waveguides

The third-order optical nonlinearity in optical waveguides has found applications in optical switching, optical wavelength conversion, optical frequency comb generation, and ultrafast optical signal processing. The development of an integrated waveguide platform with a high nonlinearity is therefore important for nonlinear integrated photonics. Here, we report the observation of an enhancement in the nonlinearity of an air-cladding silicon pedestal waveguide. We observe enhanced nonlinear spectral broadening compared to a conventional silicon-on-insulator waveguide. At the center wavelength of 1555 nm, the nonlinear-index coefficient of air-cladding silicon pedestal waveguide is measured to be about 5% larger than that of a conventional silicon-on-insulator waveguide. We observe enhanced spectral broadening from self-phase modulation of an optical pulse in the pedestal waveguide. The interaction of light with the confined acoustic phonons in the pedestal structure gives rise to a larger nonlinear-index coefficient. The experimental results agree well with the theoretical models.

physics.optics