SearcharxivSearch

arXiv subjects

Ping Jiang

Publications and source records attributed to Ping Jiang.

16 recordsLinked to original sources

CoEvoKG: Co-Evolving Knowledge Graphs with Self-Evolving Search Agents

Large language models can improve with reinforcement learning for search agents, yet existing self play agents repeatedly generate tasks while discarding the knowledge gained during successful searches. We introduce CoEvoKG, a framework that turns a knowledge graph into both a source of verifiable training tasks and a persistent evidence memory for agent evolution. CoEvoKG jointly trains a task generator and a search agent: the generator creates multihop questions from entity chains sampled from the knowledge graph, while the agent learns from rewards for answer correctness and search trajectories whose entity paths are supported by graph evidence. When a search succeeds, CoEvoKG verifies and deduplicates the retrieved evidence, then writes it back to the corresponding graph nodes and edges. Future rounds reuse this enriched graph for task generation and reward computation, closing the loop between model self evolution and knowledge accumulation. Experiments on six QA benchmarks (NQ, TriviaQA, PopQA, HotpotQA, 2WikiMultiHopQA, and Bamboogle) with three backbone models show that CoEvoKG improves macro average accuracy over the corresponding base models by +11.2, +10.1, and +11.6 points on Qwen2.5-3B-Instruct, Qwen2.5-7B-Instruct, and Llama-3.1-8B-Instruct, respectively. Under matched training budgets, CoEvoKG further improves over competitive self play baselines and RL baselines for search agents by +2.6 to +3.7 macro average points across the three backbones. Code is available at https://github.com/lazzy1225/CoEvoKG.

cs.AI

Multi-Catheter Digitization in Brachytherapy via Few-Shot Synthetic-to-Real Learning and Structure-Aware Tracking

Accurate catheter digitization in CT-guided interstitial brachytherapy is a critical but time-consuming task, especially for complex implant configurations. We developed a data-efficient, physics-guided framework for automated multi-catheter digitization with minimal clinical annotation. The pipeline consists of two stages. First, an implant region-aware network was pretrained on synthetic CT volumes with simulated metallic signatures and then fine-tuned using only 10 clinical cases. Second, a structure-aware reconstruction module combined a direction-constrained 3D Hough transform with synchronous physics-constrained inward tracking to separate adherent catheter trajectories. The method was evaluated by patient-level five-fold cross-validation on 203 treatment fractions from 38 patients. The fine-tuned network achieved an HD95 of 0.853 +/- 0.362 mm. End-to-end evaluation yielded an F1 score of 0.891 +/- 0.178, with shaft and tip errors of 0.334 +/- 0.367 mm and 0.896 +/- 0.680 mm, respectively. In cases with severe catheter adhesion, the tracking F1 score remained 0.843 +/- 0.190. The complete workflow required approximately 11.6 s per case. These results indicate that combining few-shot synthetic-to-real learning with physics-guided structural tracking can provide robust and efficient multi-catheter digitization for time-sensitive clinical workflows.

physics.med-ph

Dive Into the Implicit Biases of Low-rank Vision-language Alignment

Vision-language alignment, the stage that bridges pretrained vision encoders and large language models, is widely treated as a form of pretraining requiring full-parameter updates. We challenge this view and investigate what happens when low-rank adaptation is applied to the LLM during this stage instead. We find that low-rank alignment not only reduces computational costs but also outperforms full-parameter alignment on most benchmarks. To understand this phenomenon, we systematically characterize the implicit biases introduced by low-rank adaptation during alignment. Empirically, we find that low-rank alignment shifts model behavior from hallucinatory to conservative and preserves per-token linear separability of visual features that full-parameter alignment disrupts, a phenomenon we term LS-curse. Geometrically, low rank aligned models exhibit more homogeneous and structurally stable visual representations, maintaining modality-specific knowledge rather than prematurely fusing entity-level semantics. Theoretically, we establish two theorems showing that low-rank alignment induces preferences for parameter subspaces with flat gradients and feature subspaces robust to perturbations, providing a principled explanation for the observed structure-preserving behavior. Extensive experiments cover ablation over 100 alignment configurations, three families of low-rank operators, and various rank, encoder, and other settings.

cs.CV

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation

As web agents increasingly demonstrate capabilities in automated task execution, the development of robust evaluation frameworks for assessing their navigation and task completion performance has emerged as a critical research priority. However, existing benchmarks exhibit fundamental limitations. First, they suffer from insufficient scale and limited domain diversity, constraining comprehensive evaluation of cross-domain generalization. Second, prevailing LLM-as-Judge evaluation methodologies inadequately capture fine-grained interaction semantics, particularly regarding precise query formulation and filtering operations. Third, current benchmarks predominantly emphasize navigation success metrics while neglecting critical requirements for real-world deployment scenarios. To address these limitations, we introduce WebRetriever, a large-scale benchmark encompassing 800 websites and 1,550 tasks across diverse domains, including consumer, professional, and enterprise sectors, with comprehensive coverage of user intent patterns. We propose NavEval (Navigation Evaluation), a novel LLM-as-Judge framework that leverages rich interaction context beyond visual screenshots, achieving state-of-the-art alignment with human judgment across multiple evaluation datasets. Furthermore, we establish three complementary evaluation protocols that collectively provide holistic assessment of web agent capabilities: navigation proficiency, knowledge-assisted interaction, and end-to-end task completion with information extraction. Extensive experimental analysis reveals substantial performance disparities across evaluation protocols, demonstrating that navigation success alone is an insufficient predictor of real-world application effectiveness. WebRetriever delivers fine-grained diagnostic insights into agent capabilities and establishes a rigorous foundation for advancing web agent research and development.

cs.CV

A Novel Method for Differential-Algebraic Dynamic Model Discovery in Power Systems: An LLM-Based Multi-Agent Collaborative Framework

With large-scale integration of emerging power electronic devices represented by grid-forming inverters, power system dynamics increasingly exhibit strong nonlinearity, multi-timescale coupling, and black-box control logic. These features hinder conventional parameter identification requiring known model structures and structure identification based on predefined function libraries, making complete differential-algebraic dynamic model recovery difficult under weak prior information. To address this challenge, this paper proposes an LLM-based multi-agent collaborative framework for differential-algebraic dynamic model discovery in power systems. It integrates heterogeneous exploratory agents, individual candidate model memories, parameter fitting and evaluation, and a coordinator agent. Under unified measurement-data constraints, agents generate candidate equation structures in parallel, while candidates are optimized, evaluated, retained, and summarized to provide closed-loop search guidance. The task is decomposed into differential equation structure discovery and algebraic closure discovery, enabling joint recovery of state dynamics, algebraic constraints, and key intermediate variables with incomplete prior information. Case studies on synchronous generators and grid-forming inverters show that the proposed method outperforms single-agent LLM-based discovery and conventional symbolic regression in reconstruction accuracy, generalization, search efficiency, and noise robustness. In the generator case, OOD MAPE reaches 0.19\%; in the inverter case, discovery time is reduced by 25.7\% compared with the single-agent LLM baseline.

eess.SY

Strain-released epitaxy of GaN enabled by compliant single-crystalline metal foils

Heteroepitaxy conventionally relies on rigid crystalline substrates, implicitly assuming that lattice and thermal mismatch must be accommodated within the epitaxial layer, leading to residual strain and defects that worsen with increasing substrate size. Here we demonstrate a substrate-mediated strain-partitioning regime in which lattice and thermal mismatch are preferentially partitioned into the substrate rather than stored in the epitaxial layer. We report the epitaxial growth of single-crystalline GaN on mechanically compliant yet crystallographically ordered single-crystalline copper foils. Atomic-resolution microscopy, geometric phase analysis and density functional theory reveal that mismatch-induced stress is primarily screened by elastic deformation of the Cu lattice, accompanied by localized interfacial slip confined to a few atomic layers, leaving the AlN and GaN epilayers nearly strain-free despite large nominal mismatch. Leveraging this strain-released epitaxial platform, we further demonstrate dense GaN micro-light-emitting diode arrays that benefit from efficient vertical electrical conduction and thermal dissipation enabled by the metallic substrate. By establishing compliant single-crystal metal foils as a new substrate class, this work identifies mechanical contrast as an underexplored governing parameter in heteroepitaxial design, with implications extending beyond GaN.

cond-mat.mtrl-sci

Mano Technical Report

Graphical user interfaces (GUIs) are the primary medium for human-computer interaction, yet automating GUI interactions remains challenging due to the complexity of visual elements, dynamic environments, and the need for multi-step reasoning. Existing methods based on vision-language models (VLMs) often suffer from limited resolution, domain mismatch, and insufficient sequential decisionmaking capability. To address these issues, we propose Mano, a robust GUI agent built upon a multi-modal foundation model pre-trained on extensive web and computer system data. Our approach integrates a novel simulated environment for high-fidelity data generation, a three-stage training pipeline (supervised fine-tuning, offline reinforcement learning, and online reinforcement learning), and a verification module for error recovery. Mano demonstrates state-of-the-art performance on multiple GUI benchmarks, including Mind2Web and OSWorld, achieving significant improvements in success rate and operational accuracy. Our work provides new insights into the effective integration of reinforcement learning with VLMs for practical GUI agent deployment, highlighting the importance of domain-specific data, iterative training, and holistic reward design.

cs.MM

PRE-MAP: Personalized Reinforced Eye-tracking Multimodal LLM for High-Resolution Multi-Attribute Point Prediction

Visual selective attention, driven by individual preferences, regulates human prioritization of visual stimuli by bridging subjective cognitive mechanisms with objective visual elements, thereby steering the semantic interpretation and hierarchical processing of dynamic visual scenes. However, existing models and datasets predominantly neglect the influence of subjective cognitive diversity on fixation behavior. Conventional saliency prediction models, typically employing segmentation approaches, rely on low-resolution imagery to generate saliency heatmaps, subsequently upscaled to native resolutions, which limiting their capacity to capture personalized attention patterns. Furthermore, MLLMs are constrained by factors such as hallucinations, making it very costly to strictly adhere to the expected format in tasks involving multiple point predictions, and achieving precise point positioning is challenging. To address these limitations, we present Subjective Personalized Attention for Advertisement Videos, namely SPA-ADV, a large-scale multimodal dataset capturing gaze behaviors from over 4,500 participants varying in age and gender with 486 videos. Furthermore, we propose PRE-MAP, a novel eye-tracking saliency model that characterizes Personalized visual disparities through Reinforcement learning-optimized Eye-tracking, built upon MLLMs and guided by Multi-Attribute user profiles to predict Points. To ensure MLLMs produce prediction points that are both format-correct and spatially accurate, we introduce Consistency Group Relative Policy Optimization (C-GRPO), inspired by the variability in eye movement points and Multi-Attribute profiles. Extensive experiments on SPA-ADV and other benchmarks demonstrate the effectiveness of our approach. The code and dataset are available at \href{https://github.com/mininglamp-MLLM/PRE-MAP}{this URL}.

cs.CV

COM Adjustment Mechanism Control for Multi-Configuration Motion Stability of Unmanned Deformable Vehicle

An unmanned deformable vehicle is a wheel-legged robot transforming between two configurations: vehicular and humanoid states, with different motion modes and stability characteristics. To address motion stability in multiple configurations, a center-of-mass adjustment mechanism was designed. Further, a motion stability hierarchical control algorithm was proposed, and an electromechanical model based on a two-degree-of-freedom center-of-mass adjustment mechanism was established. An unmanned-deformable-vehicle vehicular-state steady-state steering dynamics model and a gait planning kinematic model of humanoid state walking were established. A stability hierarchical control strategy was designed to realize the stability control. The results showed that the steady-state steering stability in vehicular state and the walking stability in humanoid state could be significantly improved by controlling the slider motion.

cs.RO

Phase-field modeling of dendritic growth with gas bubbles in the solidification of binary alloys

In this work, a phase-field model is developed for the dendritic growth with gas bubbles in the solidification of binary alloys. In this model, a total free energy for the complex gas-liquid-dendrite system is proposed through considering the interactions of gas bubbles, liquid melt and solid dendrites, and it can reduce to the energy for gas-liquid flows in the region far from the solid phase, while degenerate to the energy for thermosolutal dendritic growth when the gas bubble disappears. The governing equations are usually obtained by minimizing the total free energy, but here some modifications are made to improve the capacity of the conservative phase-field equation for gas bubbles and convection-diffusion equation for solute transfer. Additionally, through the asymptotic analysis of the thin-interface limit, the present general phase-field model for alloy solidification can match the corresponding free boundary problem, and it is identical to the commonly used models under a specific choice of model parameters. Furthermore, to describe the fluid flow, the incompressible Navier-Stokes equations are adopted in the entire domain including gas, liquid, and solid regions, where the fluid-structure interaction is considered by a simple diffuse-interface method. To test the present phase-field model, the lattice Boltzmann method is used to study several problems of gas-liquid flows, dendritic growth as well as the solidification in presence of gas bubbles, and a good performance of the present model for such complex problems is observed.

physics.flu-dyn

A novel directly energy-preserving method for charged particle dynamics

In this paper, we apply the coordinate increment discrete gradient (CIDG) method to solve the Lorentz force system which can be written as a non-canonical Hamiltonian system. Then we can obtain a new energy-preserving CIDG-I method for the system. The CIDG-I method can combine with its adjoint method CIDG-II which is also a energy-preserving method to form a new method, namely CIDG-C method. The CIDG-C method is symmetrical and can conserve the Hamiltonian energy directly and exactly. With comparison to the well-used Boris method, numerical experiments indicate that the CIDG-C method holds advantage over the Boris method in terms of energy-conserving.

math.NA

Multiple-object Grasping Using a Multiple-suction-cup Vacuum Gripper in Cluttered Scenes

Multiple-suction-cup grasping can improve the efficiency of bin picking in cluttered scenes. In this paper, we propose a grasp planner for a vacuum gripper to use multiple suction cups to simultaneously grasp multiple objects or an object with a large surface. To take on the challenge of determining where to grasp and which cups to activate when grasping, we used 3D convolution to convolve the affordable areas inferred by neural network with the gripper kernel in order to find graspable positions of sampled gripper orientations. The kernel used for 3D convolution in this work was encoded including cup ID information, which helps to directly determine which cups to activate by decoding the convolution results. Furthermore, a sorting algorithm is proposed to find the optimal grasp among the candidates. Our planner exhibited good generality and successfully found multiple-cup grasps in previous affordance map datasets. Our planner also exhibited improved picking efficiency using multiple suction cups in physical robot picking experiments. Compared with single-object (single-cup) grasping, multiple-cup grasping contributed to 1.45x, 1.65x, and 1.16x increases in efficiency for picking boxes, fruits, and daily necessities, respectively.

cs.RO

A diffuse-interface lattice Boltzmann method for the dendritic growth with thermosolutal convection

In this work, we proposed a diffuse interface model for the dendritic growth with thermosolutal convection. In this model, the sharp boundary between the fluid and solid dendrite is replaced by a thin but nonzero thickness diffuse interface, which is described by the order parameter governed by the phase-field equation for the dendritic growth. The governing equations for solute and heat transfer are modified such that the previous special treatments for source term can be avoided. To solve the model for the dendritic growth with thermosolutal convection, we also developed a diffuse-interface multi-relaxation-time lattice Boltzmann (LB) method. In this method, the order parameter in the phase-field equation is combined into the force caused by the fluid-solid interaction, and the treatment on the complex fluid-solid interface can be avoided. In addition, four LB models are developed for the phase-field equation, concentration equation, temperature equation and the Navier-Stokes equations in a unified framework. Finally, to test the present diffuse-interface LB method, we performed some simulations of the dendritic growth, and found that the numerical results are in good agreements with some previous works.

physics.flu-dyn

Learning suction graspability considering grasp quality and robot reachability for bin-picking

Deep learning has been widely used for inferring robust grasps. Although human-labeled RGB-D datasets were initially used to learn grasp configurations, preparation of this kind of large dataset is expensive. To address this problem, images were generated by a physical simulator, and a physically inspired model (e.g., a contact model between a suction vacuum cup and object) was used as a grasp quality evaluation metric to annotate the synthesized images. However, this kind of contact model is complicated and requires parameter identification by experiments to ensure real world performance. In addition, previous studies have not considered manipulator reachability such as when a grasp configuration with high grasp quality is unable to reach the target due to collisions or the physical limitations of the robot. In this study, we propose an intuitive geometric analytic-based grasp quality evaluation metric. We further incorporate a reachability evaluation metric. We annotate the pixel-wise grasp quality and reachability by the proposed evaluation metric on synthesized images in a simulator to train an auto-encoder--decoder called suction graspability U-Net++ (SG-U-Net++). Experiment results show that our intuitive grasp quality evaluation metric is competitive with a physically-inspired metric. Learning the reachability helps to reduce motion planning computation time by removing obviously unreachable candidates. The system achieves an overall picking speed of 560 PPH (pieces per hour).

cs.RO

Credit Card Fraud Detection Using Autoencoder Neural Network

Imbalanced data classification problem has always been a popular topic in the field of machine learning research. In order to balance the samples between majority and minority class. Oversampling algorithm is used to synthesize new minority class samples, but it could bring in noise. Pointing to the noise problems, this paper proposed a denoising autoencoder neural network (DAE) algorithm which can not only oversample minority class sample through misclassification cost, but it can denoise and classify the sampled dataset. Through experiments, compared with the denoising autoencoder neural network (DAE) with oversampling process and traditional fully connected neural networks, the results showed the proposed algorithm improves the classification accuracy of minority class of imbalanced datasets.

cs.LG

A quantum plasmonic nanocircuit on a semiconductor platform

Quantum photonics holds great promise for future technologies such as secure communication, quantum computation, quantum simulation, and quantum metrology. An outstanding challenge for quantum photonics is to develop scalable miniature circuits that integrate single-photon sources, linear optical components, and detectors on a chip. Plasmonic nanocircuits will play essential roles in such developments. Plasmonic components feature ultracompact geometries and can be controlled more flexibly and more energy-efficiently compared to conventional dielectric components due to strong field confinement and enhancement. Moreover, plasmonic components are compatible with electronic circuits, thanks to their deep subwavelength sizes as well as their electrically conducting materials. However, for quantum plasmonic circuits, integration of stable, bright, and narrow-band single photon sources in the structure has so far not been reported. Here we present a quantum plasmonic nanocircuit driven by a self-assembled GaAs quantum dot. The quantum dot efficiently excites narrow-band single plasmons that are guided in a two-wire transmission line until they are converted into single photons by an optical antenna. Our work demonstrates the feasibility of fully on-chip plasmonic nanocircuits for quantum optical applications.

physics.optics