Searcharxiv⌕ Search

arXiv subjects

Bo Peng

Publications and source records attributed to Bo Peng.

At least 73 records · Page 4Linked to original sources

Topology-Enhanced Alignment for Large Language Models: Trajectory Topology Loss and Topological Preference Optimization

Alignment of large language models (LLMs) via SFT and RLHF/DPO typically ignores the global geometry of the representation space, relying instead on local token likelihoods or scalar scores. We view generation as tracing a semantic trajectory in hidden space and propose a topology-enhanced alignment framework that regularizes these trajectories using 0-dimensional persistent homology. First, for SFT, we introduce Trajectory Topology Loss (TTL). Treating prompt and gold-answer embeddings as a mixed point cloud, we use a 0D persistent homology algorithm to extract "prompt-answer bridges." TTL aligns the model's actual update direction with these topological bridges rather than arbitrary directions. Second, for DPO, we propose Topological Preference Optimization (TPO). TPO constructs topic-specific semantic preference vectors and aligns the improvement direction between rejected and chosen responses with these vectors in an intermediate hidden layer. We also introduce a dynamic weighting scheme to balance DPO and TPO losses. Evaluating on Qwen2.5-7B-Instruct using UltraChat and Anthropic HH-RLHF, our topology-enhanced objectives consistently outperform strong non-topological baselines (e.g., per-example, nearest-neighbor, random regularizers) on automatic preference metrics and LLM-judge evaluations, while maintaining or improving toxicity. Results show persistent homology and trajectory geometry offer a promising direction for controllable alignment.

cs.CL↗

PACEvolve++: Improving Test-time Learning for Evolutionary Search Agents

Large language models have become drivers of evolutionary search, but most systems rely on a fixed, prompt-elicited policy to sample next candidates. This limits adaptation in practical engineering and research tasks, where evaluations are expensive, and progress depends on learning task-specific search dynamics. We introduce PACEvolve++, an advisor-model reinforcement learning framework for test-time policy adaptation in evolutionary search agents. PACEvolve++ decouples strategic search decisions from implementation: a trainable advisor generates, assesses, and selects hypotheses, while a stronger frontier model translates selected hypotheses into executable candidates. To train the advisor under non-stationary feedback, we propose a phase-adaptive approach that adapts its optimization strategy to different phases of the evolutionary process. Early in evolution, it uses group-relative feedback to learn broad search preferences; later, as reward gaps compress, it emphasizes best-of-$k$ frontier contribution to support stable refinement. Across expert-parallel load balancing, sequential recommendation, and protein fitness extrapolation, PACEvolve++ outperforms the state-of-the-art evolutionary search framework with frontier models, achieving faster convergence and stabilizing test-time training during evolutionary search.

cs.LG↗

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

Language-model agents are increasingly used as persistent coworkers that assist users across multiple working days. During such workflows, the surrounding environment may change independently of the agent: new emails arrive, calendar entries shift, knowledge-base records are updated, and evidence appears across images, scanned PDFs, audio, video, and spreadsheets. Existing benchmarks do not adequately evaluate this setting because they typically run within a single static episode and remain largely text-centric. We introduce \bench{}, a benchmark for coworker agents built around multi-turn multi-day tasks, a stateful sandboxed service environment whose state evolves between turns, and rule-based verification. The current release contains 100 tasks across 13 professional scenarios, executed against five stateful sandboxed services (filesystem, email, calendar, knowledge base, spreadsheet) and scored by 1537 deterministic Python checkers over post-execution service state; no LLM-as-judge is invoked during scoring. We benchmark seven frontier agent systems. The strongest model reaches 75.8 weighted score, but the best strict Task Success is only 20.0\%, indicating that partial progress is common while complete end-to-end workflow completion remains rare. Turn-level analysis shows that performance drops after the first exogenous environment update, highlighting adaptation to changing state as a key open challenge. We release the benchmark, evaluation harness, and construction pipeline to support reproducible coworker-agent evaluation.

cs.CV↗

Magnetononlinear Hall effect from multigap topology in metal-organic frameworks

We unveil that non-Abelian multigap band topology characterized by nontrivial Euler class invariants induces observable magnetononlinear Hall transport phenomena. We demonstrate these effects in a highly-tunable class of recently synthesized two-dimensional kagome N-heterocyclic carbene (NHC) metal-organic frameworks. We showcase the controllability of the nonlinear effect upon applying external voltage, changing temperature, and chemical substitutions that preserve the bulk topology and associated edge states. Our findings therefore reveal an uncharted presence of Euler class topology in metal-organic materials that can be experimentally deduced through measurable magnetotransport.

cond-mat.mes-hall↗

Towards High Fidelity Face Swapping: A Comprehensive Survey and New Benchmark

Face swapping has witnessed significant progress in recent years, largely driven by advances in deep generative models such as GANs and diffusion models.Despite these advances, existing methods remain fragmented across different paradigms, and their evaluation is highly inconsistent due to the lack of standardized datasets and protocols. Moreover, prior surveys primarily focus on broader deepfake generation or detection, leaving face swapping insufficiently studied as a standalone problem. In this paper, we present a comprehensive survey and benchmark for face swapping. We provide a structured review of existing methods, organizing them into five major paradigms and systematically analyzing their design principles, strengths, and limitations. To enable fair and controlled evaluation, we introduce CASIA FaceSwapping, a high-quality benchmark with balanced demographic distributions and explicit attribute variations, and establish standardized protocols to assess the robustness of different face swapping methods. Extensive experiments on representative approaches yield new insights into the performance characteristics and limitations of current techniques. Overall, our work provides a unified perspective and a principled evaluation framework to facilitate the development of more robust and controllable face swapping methods. More results can be found at https://github.com/CASIA-NLPRAI/face-swapping-survey.

cs.CV↗

A transformable slender microrobot inspired by nematode parasites for interventional endovascular surgery

Cardiovascular diseases account for around 17.9 million deaths per year globally, the treatment of which is challenging considering the confined space and complex topology of the vascular network and high risks during operations. Robots, although promising, still face the dilemma of possessing versatility or maneuverability after decades of development. Inspired by nematodes, the parasites living, feeding, and moving in the human body's vascular system, this work develops a transformable slender magnetic microrobot. Based on the experiments and analyses, we optimize the fabrication and geometry of the robot and finally create a slender prototype with an aspect ratio larger than 100 (smaller than 200 microns in diameter and longer than 20 mm in length), which possesses uniformly distributed magnetic beads on the body of an ultrathin polymer string and a big bead on the head. This prototype shows great flexibility (largest curvature 0.904 mm-1) and locomotion capability (the maximum speed: 125 mm/s). Moreover, the nematode-inspired robot can pass through sharp turns with a radius of 0.84 mm and holes distributed in three-dimensional (3D) space. We also display the potential application in interventional surgery of the microrobot by navigating it through a narrow blood vessel mold to wrap and transport a drug (95 times heavier than the robot) by deforming the robot's slender body and releasing the drug to the aim position finally. Moreover, the robot also demonstrates the possible applications in embolization by transforming and winding itself into an aneurysms phantom and exhibits its outstanding injectability by being successfully withdrawn and injected through a medical needle (diameter: 1.2 mm) of a syringe.

cs.RO↗

Evidence-Based Actor-Verifier Reasoning for Echocardiographic Agents

Echocardiography plays an important role in the screening and diagnosis of cardiovascular diseases. However, automated intelligent analysis of echocardiographic data remains challenging due to complex cardiac dynamics and strong view heterogeneity. In recent years, visual language models (VLM) have opened a new avenue for building ultrasound understanding systems for clinical decision support. Nevertheless, most existing methods formulate this task as a direct mapping from video and question to answer, making them vulnerable to template shortcuts and spurious explanations. To address these issues, we propose EchoTrust, an evidence-driven Actor-Verifier framework for trustworthy reasoning in echocardiography VLM-based agents. EchoTrust produces a structured intermediate representation that is subsequently analyzed by distinct roles, enabling more reliable and interpretable decision-making for high-stakes clinical applications.

cs.CV↗

Emergence and transition of incompressible phases in decorated Landau levels

A single Landau level (LL) dressed with periodic electrostatic potentials can realize a plethora of interacting topological phases where the Hall conductivity generally does not equal to the LL filling factor. Their physics can be captured by a new family of flat topological bands: decorated Landau levels (dLL) from imposing an electrostatic delta potential lattice within a single LL. With $p/q$ magnetic fluxes per unit cell, there are $q$ dispersive bands and $p-q$ zero energy bands forming the dLL. When the electrostatic potential strength dominates the electron-electron interaction, band mixing is suppressed and the dispersion bands consist of ``localized states" with vanishing total Chern number. Nevertheless these dispersive bands can have highly nontrivial Berry curvature distribution, and even non-zero Chern numbers when $q>1$. Interestingly even in the limit of large short range interaction, band mixing between dLL and dispersion bands can be strongly suppressed at low filling factor, leading to robust topological phases within the dLL stabilized by the one-body potential. The dLL and the associated dispersive bands can serve as minimal theoretical models for correlated physics in lattice or moiré systems; they are also highly tunable experimental platforms for realizing rich phase diagrams of exotic 2D quantum fluids.

cond-mat.str-el↗

A 44-minute periodic radio transient in a supernova remnant

Long-period radio transients (LPTs) are a newly discovered class of radio emitters with periods ranging from minutes to hours. The astrophysical nature remains undetermined, particularly of LPTs with no detectable companions. We report the first evidence for a plausible supernova remnant (SNR) association with an LPT (DART J1832-0911, 2656.23+-0.15 s period), which supports a neutron star origin of such objects. The dispersion measure of this LPT, SNR's CO emission and HI absorption, and low probability of chance of alignment with field pulsars are all consistent with such an association. The source displays either phase-locked circular or nearly 100\% linear polarization, indicating its strong and geometrically stable magnetic field. No detectable optical counterpart was found, even with a 10m-class telescope. The SNR association and the stable polarization suggest that DART J1832-0911 most likely originates from a young neutron star, whose spin could have been braked by supernova's fallback materials. This discovery provides critical insights into the nature of ultra-long period transients and their link to stellar remnants.

astro-ph.HE↗

Coupled atmospHere Interior modeL Intercomparison (CHILI) Protocol Version 1.0: A CUISINES Intercomparison Project of Magma Ocean Models

Spectroscopic characterization of rocky exoplanets with the James Webb Space Telescope has brought the origin and evolution of their atmospheres into the focus of exoplanet science. Time-evolved models of the feedback between interior and atmosphere are critical to predict and interpret these observations and link them to the Solar System terrestrial planets. However, models differ in methodologies and input data, which can lead to significant differences in interpretation. In this paper, we present the experimental protocol of the Coupled atmospHere Interior modeL Intercomparison (CHILI) project. CHILI is an (exo-)planet model intercomparison project within the Climates Using Interactive Suites of Intercomparisons Nested for Exoplanet Studies (CUISINES) framework, which aims to support a diverse set of multi-model intercomparison projects in the exoplanet community. The present protocol includes the initial set of participating magma ocean models, divided into evolutionary and static models, and two types of test categories, one focused on Solar System planets (Earth & Venus) and the other on exoplanets orbiting low-mass M-dwarfs. Both test categories aim to quantify the evolution of key markers of the links between planetary atmospheres and interiors over geological timescales. The proposed tests would allow us to quantify and compare the differences between coupled atmosphere-interior models used by the exoplanet and planetary science communities. Results from the proposed tests will be published in dedicated follow-up papers. To encourage the community to join this comparison effort and as an example, we present initial test results for the early Earth and TRAPPIST-1 b, conducted with models differing in the treatment of energy transport in the planetary interior and atmosphere, surface boundary layer, geochemistry, and the in- and outgassing of volatile compounds.

astro-ph.EP↗

DREAM: A Benchmark Study for Deepfake photoREalism AssessMent

Deep learning based face-swap videos, widely known as deepfakes, have drawn wide attention due to their threat to information credibility. Recent works mainly focus on the problem of deepfake detection that aims to reliably tell deepfakes apart from real ones, in an objective way. On the other hand, the subjective perception of deepfakes, especially its computational modeling and imitation, is also a significant problem but lacks adequate study. In this paper, we focus on the photorealism assessment of deepfakes, which is defined as the automatic assessment of deepfake photorealism that approximates human perception of deepfakes. It is important for evaluating the quality and deceptiveness of deepfakes which can be used for predicting the influence of deepfakes on Internet, and it also has potentials in improving the deepfake generation process by serving as a critic. This paper promotes this new direction by presenting a comprehensive benchmark called DREAM, which stands for Deepfake photoREalism AssessMent. It is comprised of a deepfake video dataset of diverse quality, a large scale annotation that includes 140,000 photorealism scores and textual descriptions obtained from 3,500 human annotators, and a comprehensive evaluation and analysis of 18 representative photorealism assessment methods, including recent large vision language model based methods and a newly proposed description-aligned CLIP method. The benchmark and insights included in this study can lay the foundation for future research in this direction and other related areas. We make the dataset available to the research community at https://github.com/bomb2peng/DREAM-A-Benchmark-Study-for-Deepfake-photoREalism-AssessMent.

cs.CV↗

Delving into Spectral Clustering with Vision-Language Representations

Spectral clustering is known as a powerful technique in unsupervised data analysis. The vast majority of approaches to spectral clustering are driven by a single modality, leaving the rich information in multi-modal representations untapped. Inspired by the recent success of vision-language pre-training, this paper enriches the landscape of spectral clustering from a single-modal to a multi-modal regime. Particularly, we propose Neural Tangent Kernel Spectral Clustering that leverages cross-modal alignment in pre-trained vision-language models. By anchoring the neural tangent kernel with positive nouns, i.e., those semantically close to the images of interest, we arrive at formulating the affinity between images as a coupling of their visual proximity and semantic overlap. We show that this formulation amplifies within-cluster connections while suppressing spurious ones across clusters, hence encouraging block-diagonal structures. In addition, we present a regularized affinity diffusion mechanism that adaptively ensembles affinity matrices induced by different prompts. Extensive experiments on \textbf{16} benchmarks -- including classical, large-scale, fine-grained and domain-shifted datasets -- manifest that our method consistently outperforms the state-of-the-art by a large margin.

cs.CV↗

Near-limit quantum control beyond analytic tractability in many-body spin systems

As quantum control approaches hardware-imposed performance limits, weak effects omitted by reduced models become consequential. Assumptions required for analytic tractability then cease to guide control design and instead constrain further improvement. Here, we relax such assumptions and use simulation-guided stochastic tree search to navigate combinatorially large, discrete pulse-sequence spaces for robust many-body spin control. Experimentally, in a solid-state spin ensemble, the resulting computationally discovered pulse sequences substantially outperform analytically optimized baselines, despite being excluded by construction from analytic design criteria. Importantly, these unconventional sequences expose predictive structural features that enable rapid neural network--based performance evaluation. This efficiency gain makes the combinatorial scaling tractable and expands the control alphabet from 8 symmetry-restricted pulses to over 26,000 hardware-resolved options. The resulting fine-grained design freedom provides the control resolution required to reliably address weak, performance-limiting effects, unlocking qualitatively different spin-control capabilities beyond decades of traditional sequence design. Together, these results show that near performance limits, simplifying assumptions can become a primary constraint on quantum control in realistic hardware, and must be repurposed to guide computational discovery.

quant-ph↗

AgenticTagger: Structured Item Representation for Recommendation with LLM Agents

High-quality representations are a core requirement for effective recommendation. In this work, we study the problem of LLM-based descriptor generation, i.e., keyphrase-like natural language item representation generation frameworks with minimal constraints on downstream applications. We propose AgenticTagger, a framework that queries LLMs for representing items with sequences of text descriptors. However, open-ended generation provides little control over the generation space, leading to high cardinality, low-performance descriptors that render downstream modeling challenging. To this end, AgenticTagger features two core stages: (1) a vocabulary-building stage in which a set of hierarchical, low-cardinality, and high-quality descriptors is identified, and (2) a vocabulary-assignment stage in which LLMs assign in-vocabulary descriptors to items. To effectively and efficiently ground vocabulary in the item corpus of interest, we design a multi-agent reflection mechanism in which an architect LLM iteratively refines the vocabulary guided by parallelized feedback from annotator LLMs that validate the vocabulary against item data. Experiments on public and private data show AgenticTagger brings consistent improvements across diverse recommendation scenarios, including generative and term-based retrieval, ranking, and controllability-oriented, critique-based recommendation.

cs.IR↗

Compile-once block encodings for masked similarity-transformed effective Hamiltonians

We present COMPOSER, a compile-once modular parametric oracle for similarity-encoded effective reduction of electronic-structure operators (e.g., Schrieffer-Wolff-type constructions). Low-rank factorizations compress Hamiltonians and anti-Hermitian generators into rank-one bilinear and projected-quadratic ladders with near-linear scaling at fixed thresholds; each ladder admits deterministic, number-conserving preparation and a block encoding using constant number of signal ancillas. A fixed PREP-SELECT-PREP template multiplexes these ladders, and one QSP polynomial performs the spectral transformation with degree set by operator norms. For a fixed orbital pool and qubit register, the two-qubit fabric is compiled once; geometry, active-space (mask) updates, and truncations are absorbed by re-dialed single-qubit rotations. We introduce a mask-aware similarity-sandwich effective-Hamiltonian construction and benchmark stability under low-rank and second-order-perturation-guided screening. COMPOSER is an execution architecture: algorithmic errors (block-encoding and QSP approximation) are tunable for any supplied parameters, while physical accuracy depends on how those parameters are obtained if not refined.

quant-ph↗

Seeking Necessary and Sufficient Information from Multimodal Medical Data

Learning multimodal representations from medical images and other data sources can provide richer information for decision-making. While various multimodal models have been developed for this, they overlook learning features that are both necessary (must be present for the outcome to occur) and sufficient (enough to determine the outcome). We argue learning such features is crucial as they can improve model performance by capturing essential predictive information, and enhance model robustness to missing modalities as each modality can provide adequate predictive signals. Such features can be learned by leveraging the Probability of Necessity and Sufficiency (PNS) as a learning objective, an approach that has proven effective in unimodal settings. However, extending PNS to multimodal scenarios remains underexplored and is non-trivial as key conditions of PNS estimation are violated. We address this by decomposing multimodal representations into modality-invariant and modality-specific components, then deriving tractable PNS objectives for each. Experiments on synthetic and real-world medical datasets demonstrate our method's effectiveness. Code will be available on GitHub.

cs.CV↗

Uncovering the absorbed atomic Universe with the [OI]63um line

We report the discovery of strongly absorbed [OI]63um in a sample of 12 DSFGs at 4.2 4. Using ALMA Bands 9 and 10, we obtain spatially and spectrally resolved observations that probe the interstellar medium on sub-kpc scales. Despite reaching sensitivities 10-100x deeper than most previous studies, we detect [OI]63um in emission in only 2 sources at low significance, with the remaining galaxies yielding stringent non-detections over the full velocity range covered by robust detections of other far-infrared lines, including [CII] and [NII]205um. We identify several compact (0.05-0.2") regions having [OI]63um absorption against the far-infrared dust continuum, some of which are possibly reaching below rest-frame CMB radiation level. We also detect narrow, spatially localised [OI]63um emission "escape channels" preferentially detected in regions with weak or absent dust continuum emission. We predict that similar absorption effects may appear in the [CII] line, particularly when concentrating on the regions with the densest foreground material along the line of sight. The [OI]63um line appears to be originate from a mix of compact, high optical depth [OI]63um emitting regions and sub-thermally excited, oxygen-rich molecular clouds dispersed throughout high-redshift starbursts that are capable of absorbing the ground-state line emission. Combined with a comparison to cosmological radiation hydrodynamical simulations, this supports the interpretation that regions with higher gas and dust column densities may lead to weakening an intrinsically strong [OI]63um line emission. We argue that the high [OI]63um optical depth is the dominant effect causing the strong absorption, limiting the diagnostic power of this line to trace regions of massive star formation in high-redshift DSFGs.

astro-ph.GA↗

How Foundational Skills Influence VLM-based Embodied Agents:A Native Perspective

Recent advances in vision-language models (VLMs) have shown promise for human-level embodied intelligence. However, existing benchmarks for VLM-driven embodied agents often rely on high-level commands or discretized action spaces, which are non-native settings that differ markedly from real-world control. In addition, current benchmarks focus primarily on high-level tasks and lack joint evaluation and analysis at both low and high levels. To address these limitations, we present NativeEmbodied, a challenging benchmark for VLM-driven embodied agents that uses a unified, native low-level action space. Built on diverse simulated scenes, NativeEmbodied includes three representative high-level tasks in complex scenarios to evaluate overall performance. For more detailed analysis, we further decouple the skills required by complex tasks and construct four types of low-level tasks, each targeting a fundamental embodied skill. This joint evaluation across task and skill granularities enables fine-grained assessment of embodied agents. Experiments with state-of-the-art VLMs reveal clear deficiencies in several fundamental embodied skills, and further analysis shows that these bottlenecks significantly limit performance on high-level tasks. NativeEmbodied highlights key challenges for current VLM-driven embodied agents and provides insights to guide future research.

cs.AI↗