SearcharxivSearch

arXiv subjects

Yuzhao Peng

Publications and source records attributed to Yuzhao Peng.

3 recordsLinked to original sources

Scalable Behaviour Cloning on Browser Using via Skill Distillation

Internet users collectively perform an enormous range of skilled work through web browsers, from software development and document editing to search, forms, and enterprise workflows, making human browsing a highly scalable but under-exploited source of reusable browser skills. We argue that the bottleneck for browser agents is decision-making under incomplete information rather than low-level operation, and that the priors agents lack are already implicit in human interaction traces. We therefore study scalable behavior cloning for browser agents via skill distillation, converting user interaction trajectories into compact natural-language skills that agents can read, retrieve, reuse, and compose directly. We further organize the distilled skills into a skill graph so that growth proceeds through consolidation rather than unbounded accumulation. This suggests that the scalability of browser agents may come less from manually designed tasks and more from the collective skills already expressed by internet users. Our project is available at: https://lab.einsia.ai/browserbc/.

cs.CL

PaperX: A Unified Framework for Multimodal Academic Presentation Generation with Scholar DAG

Transforming scientific papers into multimodal presentation content is essential for research dissemination but remains labor intensive. Existing automated solutions typically treat each format as an isolated downstream task, leading to redundant processing and semantic inconsistency. We introduce PaperX, a unified framework that models academic presentation generation as a structural transformation and rendering process. Central to our approach is the Scholar DAG, an intermediate representation that decouples the paper's logical structure from its final presentation syntax. By applying adaptive graph traversal strategies, PaperX generates diverse, high quality outputs from a single source. Comprehensive evaluations demonstrate that our framework achieves the state of the art performance in content fidelity and aesthetic quality while significantly improving cost efficiency compared to specialized single task agents.

cs.DL

Why Does Weak-OOD Help? A Further Step Towards Understanding Jailbreaking VLMs

Large Vision-Language Models (VLMs) are susceptible to jailbreak attacks: researchers have developed various attack strategies that bypass the safety mechanisms of VLMs. Among these approaches, jailbreak methods based on the Out-of-Distribution (OOD) strategy have garnered widespread attention due to their simplicity and effectiveness. This paper further advances the understanding of OOD-based VLM jailbreak methods. We show that mild OOD manipulations can achieve stronger jailbreak performance than both clean inputs and overly strong perturbations, a non-monotonic pattern we define as "weak-OOD". We explain this phenomenon through a trade-off between two dominant factors: input intent perception and model refusal triggering. Our evidence suggests that these two factors respond differently to OOD manipulations, which is consistent with a discrepancy between broad pre-training robustness and narrower safety alignment. Building on this insight, we draw inspiration from optical character recognition (OCR) capability enhancement---a core task in the pre-training phase of mainstream VLMs. Leveraging this capability, we design JOCR (Jailbreak via OCR-Aware Embedded Text Perturbation), a practical OCR-readable extension of embedded-text jailbreaks that achieves the best average ASR among evaluated baselines. Code is available at GitHub: https://github.com/Yuxuan2003/weak-ood-jailbreak.

cs.CR