Searcharxiv⌕ Search

arXiv subjects

Yan Peng

Publications and source records attributed to Yan Peng.

At least 37 records · Page 2Linked to original sources

Raw Data Matters: Enhancing Prompt Tuning by Internal Augmentation on Vision-Language Models

For CLIP-based prompt tuning, introducing more data as additional knowledge for enhancing fine-tuning process is proved to be an effective approach. Existing data amplification strategies for prompt tuning typically rely on external knowledge (e.g., large language models or pre-structured knowledge bases), resulting in higher costs for data collection and processing, while generally ignoring further utilization of features in image modality. To address this, we propose Augmentation-driven Prompt Tuning (AugPT), a self-contained distillation-based prompt tuning approach using only internal augmentation on raw dataset to better exploit known features. Specifically, AugPT employs self-supervised augmentation on unlabeled images in the training set, and introduces a novel gating mechanism based on consensus test, reusing the pre-trained prompt tuning backbone model to spontaneously filter noisy samples, further enhancing the quality of augmented views. Extensive experiments validate that AugPT simultaneously enhances model performance and generalization capability without using appended external knowledge. The code of AugPT is available at: https://github.com/JREion/AugPT .

cs.CV↗

Non-linearly scalarized supermassive black holes

In this study, we investigate a nonlinear mechanism driving the formation of scalarized rotating black holes within a scalar-Gauss-Bonnet gravity framework that includes an additional squared Gauss-Bonnet term. With the specific coupling function, Kerr metric is a solution to this modified gravity. In linear level Kerr black holes are stable against the scalar perturbation, while nonlinearly they suffer the so-called ``nonlinear scalarization" and are unstable. By employing a pseudo-spectral method, we derive the spectrum of nonlinearly scalarized rotating black hole solutions, revealing multiple scalarized branches. Our analysis demonstrates that both the black hole's spin and the additional squared Gauss-Bonnet term significantly influence the existence and properties of these solutions. Furthermore, we explore the thermodynamic properties of nonlinearly scalarized rotating black holes, and find that the scalarized black holes are entropically favored over Kerr black holes of the same mass and spin across a wide range of parameters.

gr-qc↗

Spontaneous excitation of a centripetally accelerated atom coupled to electromagnetic vacuum fluctuations near a reflecting boundary

We investigate the rate of change of the mean atomic energy for centripetally accelerated atoms interacting with electromagnetic vacuum fluctuations near a reflecting boundary, using the Dalibard-Dupont-Roc-Cohen-Tannoudji formalism. The distinct contributions from vacuum fluctuations and radiation reaction are analyzed separately. Our results reveal that, when the centripetal acceleration significantly exceeds the characteristic acceleration set by the atomic transition frequency, vacuum fluctuations dominates over radiation reaction, irrespective of the atom-boundary distance and the atomic polarization. In the near-zone regime, where the atom-boundary distance is much smaller than both the characteristic length associated with the acceleration and the transition wavelength of the atom, the boundary introduces substantial corrections to the rate of change of the mean atomic energy. These corrections are comparable in magnitude to those in free space and exhibit strong dependence on the atomic polarization. Remarkably, in the intermediate and far regions, contributions stemming from the combined effects of the boundary and acceleration can become the leading and subleading terms, respectively. An acceleration-independent term also arises from their interplay. These findings highlight the significant interplay between acceleration and the presence of a boundary in shaping atomic radiative properties and may have potential implications for experimentally probing the circular Unruh effect.

gr-qc↗

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework

Visual reasoning refers to the task of solving questions about visual information. Current visual reasoning methods typically employ pre-trained vision-language model (VLM) strategies or deep neural network approaches. However, existing efforts are constrained by limited reasoning interpretability, while hindering by the phenomenon of underspecification in the question text. Additionally, the absence of fine-grained visual knowledge limits the precise understanding of subject behavior in visual reasoning tasks. To address these issues, we propose VIKSER (Visual Knowledge-Driven Self-Reinforcing Reasoning Framework). Specifically, VIKSER, trained using knowledge distilled from large language models, extracts fine-grained visual knowledge with the assistance of visual relationship detection techniques. Subsequently, VIKSER utilizes fine-grained visual knowledge to paraphrase the question with underspecification. Additionally, we design a novel prompting method called Chain-of-Evidence (CoE), which leverages the power of "evidence for reasoning" to endow VIKSER with interpretable reasoning capabilities. Meanwhile, the integration of self-reflection technology empowers VIKSER with the ability to learn and improve from its mistakes. Experiments conducted on widely used datasets demonstrate that VIKSER achieves new state-of-the-art (SOTA) results in relevant tasks. Moreover, VIKSER achieves performance on par with leading proprietary models, such as the latest ChatGPT-5.

cs.CV↗

Rotating scalarized supermassive black holes

In this study, we investigate rotating black hole solutions within a scalar Gauss-Bonnet gravity framework that incorporates a squared Gauss-Bonnet term. By employing a quadratic exponential coupling function between the scalar field and the Gauss-Bonnet invariant, we derive both the standard General Relativity solutions and novel scalarized black hole configurations. Utilizing a pseudo spectral method to solve the coupled field equations, we examine how black hole spin and coupling constants influence the existence and properties of these solutions. Our findings reveal that both the rotation of the black hole and the squared coupling term effectively constrain the parameter space available for scalarization. Moreover, we demonstrate that, over a wide range of parameters, scalarized black holes exhibit higher entropy than Kerr black holes of equivalent mass and spin, indicating that they are thermodynamically favored. These results significantly expand the phase space of black holes in modified gravity theories.

gr-qc↗

CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Robot foundation models, particularly Vision-Language-Action (VLA) models, have garnered significant attention for their ability to enhance robot policy learning, greatly improving robot's generalization and robustness. OpenAI's recent model, O1, showcased impressive capabilities in solving complex problems by utilizing extensive reasoning chains. This prompts an important question: can robot models achieve better performance in multi-task , complex environments by reviewing prior observations and then providing task-specific reasoning to guide action prediction? In this paper, we introduce Chain-of-Affordance (CoA-VLA) , a novel approach to scaling robot models by incorporating reasoning in the format of sequential robot affordances to facilitate task completion. Specifically, we prompt the model to consider the following four types of affordances before taking action: (1) object affordance - what object to manipulate and where it is ; (2) grasp affordance - the specific object part to grasp ; (3) spatial affordance - the optimal space to place the object ; and (4) movement affordance-the collision - free path for movement. We further transform each affordance into two prompting formats: visual affordance and textual affordance. We introduce a novel vision-language co-injection module that integrates this knowledge into the policy network. This allows the robot to leverage essential contextual information during action inference, resulting in improved precision and robustness. Our experiments demonstrate that CoA-VLA outperforms state-of-the-art robot foundation models, including OpenVLA and Octo, on a variety of tasks. Furthermore, CoA-VLA exhibits strong generalization capabilities, including recognizing unseen object poses, identifying free space, and avoiding obstacles in novel environments.

cs.RO↗

Lower bound on the orbital period of Kerr-Newman black holes

Based on the orbital period of Kerr black holes, Hod proposed a conjecture that a general lower bound on the orbital period may exist. In this work, we examined this bound by exploring the orbital period of Kerr-Newman black holes using analytical and numerical methods. By choosing different charge and spin of Kerr-Newman black holes, we found a lower bound for the orbital period of Kerr-Newman black holes as $T(r)\geqslant 4πM$, where $r$ is the orbital radius, $T(r)$ is the orbital period observed from infinity and $M$ is the black hole mass. This bound is just the same as Hod's conjectured lower bound. So our results further demonstrated that Hod's lower bound may be a general property in black hole spacetimes.

gr-qc↗

Hybrid Physics-ML Modeling for Marine Vehicle Maneuvering Motions in the Presence of Environmental Disturbances

A hybrid physics-machine learning modeling framework is proposed for the surface vehicles' maneuvering motions to address the modeling capability and stability in the presence of environmental disturbances. From a deep learning perspective, the framework is based on a variant version of residual networks with additional feature extraction. Initially, an imperfect physical model is derived and identified to capture the fundamental hydrodynamic characteristics of marine vehicles. This model is then integrated with a feedforward network through a residual block. Additionally, feature extraction from trigonometric transformations is employed in the machine learning component to account for the periodic influence of currents and waves. The proposed method is evaluated using real navigational data from the 'JH7500' unmanned surface vehicle. The results demonstrate the robust generalizability and accurate long-term prediction capabilities of the nonlinear dynamic model in specific environmental conditions. This approach has the potential to be extended and applied to develop a comprehensive high-fidelity simulator.

cs.RO↗

DPC: Dual-Prompt Collaboration for Tuning Vision-Language Models

The Base-New Trade-off (BNT) problem universally exists during the optimization of CLIP-based prompt tuning, where continuous fine-tuning on base (target) classes leads to a simultaneous decrease of generalization ability on new (unseen) classes. Existing approaches attempt to regulate the prompt tuning process to balance BNT by appending constraints. However, imposed on the same target prompt, these constraints fail to fully avert the mutual exclusivity between the optimization directions for base and new. As a novel solution to this challenge, we propose the plug-and-play Dual-Prompt Collaboration (DPC) framework, the first that decoupling the optimization processes of base and new tasks at the prompt level. Specifically, we clone a learnable parallel prompt based on the backbone prompt, and introduce a variable Weighting-Decoupling framework to independently control the optimization directions of dual prompts specific to base or new tasks, thus avoiding the conflict in generalization. Meanwhile, we propose a Dynamic Hard Negative Optimizer, utilizing dual prompts to construct a more challenging optimization task on base classes for enhancement. For interpretability, we prove the feature channel invariance of the prompt vector during the optimization process, providing theoretical support for the Weighting-Decoupling of DPC. Extensive experiments on multiple backbones demonstrate that DPC can significantly improve base performance without introducing any external knowledge beyond the base classes, while maintaining generalization to new classes. Code is available at: https://github.com/JREion/DPC.

cs.CV↗

PointVLA: Injecting the 3D World into Vision-Language-Action Models

Vision-Language-Action (VLA) models excel at robotic tasks by leveraging large-scale 2D vision-language pretraining, but their reliance on RGB images limits spatial reasoning critical for real-world interaction. Retraining these models with 3D data is computationally prohibitive, while discarding existing 2D datasets wastes valuable resources. To bridge this gap, we propose PointVLA, a framework that enhances pre-trained VLAs with point cloud inputs without requiring retraining. Our method freezes the vanilla action expert and injects 3D features via a lightweight modular block. To identify the most effective way of integrating point cloud representations, we conduct a skip-block analysis to pinpoint less useful blocks in the vanilla action expert, ensuring that 3D features are injected only into these blocks--minimizing disruption to pre-trained representations. Extensive experiments demonstrate that PointVLA outperforms state-of-the-art 2D imitation learning methods, such as OpenVLA, Diffusion Policy and DexVLA, across both simulated and real-world robotic tasks. Specifically, we highlight several key advantages of PointVLA enabled by point cloud integration: (1) Few-shot multi-tasking, where PointVLA successfully performs four different tasks using only 20 demonstrations each; (2) Real-vs-photo discrimination, where PointVLA distinguishes real objects from their images, leveraging 3D world knowledge to improve safety and reliability; (3) Height adaptability, Unlike conventional 2D imitation learning methods, PointVLA enables robots to adapt to objects at varying table height that unseen in train data. Furthermore, PointVLA achieves strong performance in long-horizon tasks, such as picking and packing objects from a moving conveyor belt, showcasing its ability to generalize across complex, dynamic environments.

cs.RO↗

Unity RL Playground: A Versatile Reinforcement Learning Framework for Mobile Robots

This paper introduces Unity RL Playground, an open-source reinforcement learning framework built on top of Unity ML-Agents. Unity RL Playground automates the process of training mobile robots to perform various locomotion tasks such as walking, running, and jumping in simulation, with the potential for seamless transfer to real hardware. Key features include one-click training for imported robot models, universal compatibility with diverse robot configurations, multi-mode motion learning capabilities, and extreme performance testing to aid in robot design optimization and morphological evolution. The attached video can be found at https://linqi-ye.github.io/video/iros25.mp4 and the code is coming soon.

cs.RO↗

PhiP-G: Physics-Guided Text-to-3D Compositional Scene Generation

Text-to-3D asset generation has achieved significant optimization under the supervision of 2D diffusion priors. However, when dealing with compositional scenes, existing methods encounter several challenges: 1). failure to ensure that composite scene layouts comply with physical laws; 2). difficulty in accurately capturing the assets and relationships described in complex scene descriptions; 3). limited autonomous asset generation capabilities among layout approaches leveraging large language models (LLMs). To avoid these compromises, we propose a novel framework for compositional scene generation, PhiP-G, which seamlessly integrates generation techniques with layout guidance based on a world model. Leveraging LLM-based agents, PhiP-G analyzes the complex scene description to generate a scene graph, and integrating a multimodal 2D generation agent and a 3D Gaussian generation method for targeted assets creation. For the stage of layout, PhiP-G employs a physical pool with adhesion capabilities and a visual supervision agent, forming a world model for layout prediction and planning. Extensive experiments demonstrate that PhiP-G significantly enhances the generation quality and physical rationality of the compositional scenes. Notably, PhiP-G attains state-of-the-art (SOTA) performance in CLIP scores, achieves parity with the leading methods in generation quality as measured by the T$^3$Bench, and improves efficiency by 24x.

cs.CV↗

Anisotropic multi-orbital Hubbard model simulated with impurity approximation

Motivated by the recent experimental findings on the orbital ordering of cuprate SC, we have investigated the multi-orbital Hubbard model in the framework of Cu impurity approximation embedded in the O lattice by incorporating the 3d$^{8}$ multiplet structure coupled to a full O-2p band.Our systematic investigation on the impact of anisotropy of various parameters reveal rich phenomena in terms of the ground state (GS) weight asymmetry between $\hat{x}$ and $\hat{y}$ directions.The numerical evidence demonstrate that the GS weight of Zhang-Rice singlet (ZRS) can be affected by the asymmetry of these parameters to distinct extent. Although the experimentally motivated asymmetric charge transfer energy only induces tiny weight difference, the asymmetric $d$-$p$ hybridization can result in considerable change of the weight. Besides, the nearest-neighbor $V_{pd}$ has much stronger impact than the local $U_{pp}$, which stems from the nature of ZRS consisting of nearest-neighbor two holes. Our systematic exploration provide valuable knowledge on the role of the artificial symmetry breaking on the two-hole GS nature and serves as the starting point of more sophisticated many-body simulations to uncover more interesting physics of multi-orbital Hubbard model within the symmetry breaking setup.

cond-mat.str-el↗

Upper bound on the radius of the innermost photonsphere in the regular compact star spacetime

We study properties of the innermost photonsphere in the regular compact star background. We take the traceless energy-momentum tensor and dominant energy conditions. In the regular compact star background, we analytically obtain an upper bound on the radius of the innermost photonsphere as $r_γ^{in}\leqslant \frac{12}{5}M$, where $r_γ^{in}$ is the radius of the innermost photonsphere and $M$ is the total ADM mass of the asymptotically flat compact star spacetime.

gr-qc↗

Analytical investigations on the Maxwell electromagnetic invariant in the spinning and charged horizonless star background

We study properties of the Maxwell electromagnetic invariant in the external region of spinning and charged horizonless stars. We analytically find that the minimum negative value of the Maxwell electromagnetic invariant is obtained on the equator of the star surface. We are interested in scalar fields non-minimally coupled to the Maxwell electromagnetic invariant. The negative enough Maxwell electromagnetic invariant can lead to a negative effective mass term, which forms a binding potential well for the scalar field. It means that the scalar field coupled to the Maxwell electromagnetic invariant may mostly exist around the surface of the star on the equator.

gr-qc↗

Reducing Hallucinations: Enhancing VQA for Flood Disaster Damage Assessment with Visual Contexts

The zero-shot performance of visual question answering (VQA) models relies heavily on prompts. For example, a zero-shot VQA for disaster scenarios could leverage well-designed Chain of Thought (CoT) prompts to stimulate the model's potential. However, using CoT prompts has some problems, such as causing an incorrect answer in the end due to the hallucination in the thought process. In this paper, we propose a zero-shot VQA named Flood Disaster VQA with Two-Stage Prompt (VQA-TSP). The model generates the thought process in the first stage and then uses the thought process to generate the final answer in the second stage. In particular, visual context is added in the second stage to relieve the hallucination problem that exists in the thought process. Experimental results show that our method exceeds the performance of state-of-the-art zero-shot VQA models for flood disaster scenarios in total. Our study provides a research basis for improving the performance of CoT-based zero-shot VQA.

cs.CV↗

Unleashing the Potential of Large Language Model: Zero-shot VQA for Flood Disaster Scenario

Visual question answering (VQA) is a fundamental and essential AI task, and VQA-based disaster scenario understanding is a hot research topic. For instance, we can ask questions about a disaster image by the VQA model and the answer can help identify whether anyone or anything is affected by the disaster. However, previous VQA models for disaster damage assessment have some shortcomings, such as limited candidate answer space, monotonous question types, and limited answering capability of existing models. In this paper, we propose a zero-shot VQA model named Zero-shot VQA for Flood Disaster Damage Assessment (ZFDDA). It is a VQA model for damage assessment without pre-training. Also, with flood disaster as the main research object, we build a Freestyle Flood Disaster Image Question Answering dataset (FFD-IQA) to evaluate our VQA model. This new dataset expands the question types to include free-form, multiple-choice, and yes-no questions. At the same time, we expand the size of the previous dataset to contain a total of 2,058 images and 22,422 question-meta ground truth pairs. Most importantly, our model uses well-designed chain of thought (CoT) demonstrations to unlock the potential of the large language model, allowing zero-shot VQA to show better performance in disaster scenarios. The experimental results show that the accuracy in answering complex questions is greatly improved with CoT prompts. Our study provides a research basis for subsequent research of VQA for other disaster scenarios.

cs.CV↗

From Knowing to Doing: Learning Diverse Motor Skills through Instruction Learning

Recent years have witnessed many successful trials in the robot learning field. For contact-rich robotic tasks, it is challenging to learn coordinated motor skills by reinforcement learning. Imitation learning solves this problem by using a mimic reward to encourage the robot to track a given reference trajectory. However, imitation learning is not so efficient and may constrain the learned motion. In this paper, we propose instruction learning, which is inspired by the human learning process and is highly efficient, flexible, and versatile for robot motion learning. Instead of using a reference signal in the reward, instruction learning applies a reference signal directly as a feedforward action, and it is combined with a feedback action learned by reinforcement learning to control the robot. Besides, we propose the action bounding technique and remove the mimic reward, which is shown to be crucial for efficient and flexible learning. We compare the performance of instruction learning with imitation learning, indicating that instruction learning can greatly speed up the training process and guarantee learning the desired motion correctly. The effectiveness of instruction learning is validated through a bunch of motion learning examples for a biped robot and a quadruped robot, where skills can be learned typically within several million steps. Besides, we also conduct sim-to-real transfer and online learning experiments on a real quadruped robot. Instruction learning has shown great merits and potential, making it a promising alternative for imitation learning.

cs.RO↗