SearcharxivSearch

arXiv subjects

Gang Liu

Publications and source records attributed to Gang Liu.

At least 19 recordsLinked to original sources

Harness Robotic OS: A Unified Embodied-Agent Runtime for Closed-Loop Quadruped Inspection

Autonomous property inspection requires more than robust robot navigation: a deployable system must connect heterogeneous sensing, reusable autonomy capabilities, multimodal scene understanding, human interaction, and enterprise response within a traceable operational loop. Existing quadruped inspection systems commonly integrate these functions through task-specific interfaces, making contextual coordination, knowledge reuse, and controlled adaptation difficult. This paper presents \textit{Harness Robotic OS} (HROS), a unified embodied-agent runtime, and Argos, its realization for residential-community inspection. HROS organizes the system into robot runtime, embodied autonomy skills, cognitive agent runtime, and interaction and operations planes. A shared context connects physical state with agent reasoning; streaming ASR/TTS supports voice-based mission interaction; hierarchical working, episodic, and semantic memory preserves operational knowledge; and a safety-gated self-evolution loop converts execution traces into versioned candidate updates without permitting unconstrained online modification. The Argos prototype integrates a Vbot quadruped, Fast-LIO2 localization and mapping, Hobot-Stereo depth perception, PCT-Planner global planning, EGO-Planner local motion generation, and OpenClaw-orchestrated Qwen3-VL inspection analysis. Experiments in a residential property environment achieved 100\% waypoint reachability, outdoor localization error below 10~cm, local obstacle-response latency below 200~ms, representative hazard-detection rates of 85--95\%, and 99\% success in alarm delivery and structured-report generation. These results validate the deployed navigation and inspection closed loop, while HROS provides an extensible software foundation for memory-augmented, voice-aware, and continuously improvable embodied inspection agents.

cs.RO

Symplectic Dirac operators on homogeneous spaces

We define symplectic Dirac operators on homogeneous spaces and study their representation-theoretic role. For an invariant polarization, the symplectic Dirac operator decomposes into two symplectic Dolbeault operators. We compute their commutator as the natural symplectic analogue of the square of the classical Dirac operator. Our first main result gives a necessary and sufficient condition for this commutator to satisfy a Parthasarathy-type formula. We further prove that, whenever this condition fails, no cubic perturbation of the symplectic Dolbeault operators can yield such a formula, in contrast with Kostant's cubic Dirac operator in the orthogonal setting. As applications, we establish an ${\mathfrak s}{\mathfrak l}_2$-structure generated by the symplectic Dolbeault operators and derive Dirac-type inequalities for unitary representations of Hermitian symmetric spaces labelled by the levels of the symmetric powers of the antiholomorphic tangent space at the identity. The level-zero inequality recovers the standard Parthasarathy-Dirac inequality, while the higher levels inequalities yield new constraints. For $SU(1,n)$, we show that, for representations with a specific Kraljevi\'c corner, the level-one inequality strengthens all basic Parthasarathy inequalities of the first kind for particular $K$-types, precisely those satisfying an explicit highest-weight condition.

math.RT

Zonal dislocations in Laves phases: A coupled synchro-shear slip mechanism

Synchro-shear is the primary plastic deformation mechanism in Laves phases at elevated temperatures, mediated by synchro-Shockley partial dislocations-zonal dislocations that proceed via localized events such as kink-pair nucleation and propagation. Using atomistic simulations, we identified a novel slip mechanism in Laves phases, namely coupled synchro-shear slip, involving the synchronized glide of two synchro-Shockley partial dislocations on adjacent slip planes, leading to the formation of extrinsic stacking faults. High-resolution scanning transmission electron microscopy revealed the extended core structures consistent with coupled synchro-Shockley partial dislocations bounded by extrinsic stacking faults in C15 NbCr2 and their involvement in twinning. These results highlight the critical role of coupled synchro-shear slip in enabling phase transformations between Laves polytypes and in governing twinning behavior, providing new atomistic insight into the kinetic nature of plasticity in topologically close-packed intermetallic phases.

cond-mat.mtrl-sci

DNative-Twin: Decision Graphs and Digital Twins for Reconstructable Agentic Decisions

AI agents increasingly gather evidence, invoke tools, apply constraints, and produce decisions that people or software may commit to action. A final output alone cannot show which evidence, tool state, rule, authorization, or action path produced it. We present DNative-Twin, a graph-native digital twin that records a committed agentic decision as a typed trajectory and re-executes its decision mechanism under declared conditions. The graph links the state observed by the agent, the path it followed, and the authority behind the resulting action. The twin synchronizes this information, replays the mechanism in isolation, and compares it under controlled changes. We instantiate the framework in enterprise decision processes using three public process logs and controlled replay suites. The experiments identify a specific failure: graph structure localizes represented changes but cannot determine the consequence of an unobserved tool state. In a three-condition controlled experiment with 300 injected instances, unresolved-divergence recall increased from 0 to 0.667 when replay-contract state was added and to 1.0 when verification results were also available; the held-out set contained no critical-class instance. Across 500--5,000 BPI 2020 cases, median end-to-end time increased from 0.794 to 8.889 seconds on the reported platform. These results separate the roles of graph structure, replay context, and verification evidence in reviewing a decision mechanism.

cs.AI

Knowledge-Distilled End-to-End Reinforcement Learning for Smooth 6-DOF Thrust Control and Rapid Adaptation to Ocean Currents in Remotely Operated Vehicles

With the continuous improvement of computational capabilities, end-to-end reinforcement learning has been rapidly developed for remotely operated vehicles control. Nevertheless, existing end-to-end reinforcement-learningbased methods still face challenges in achieving optimal control under oceancurrent disturbances. In particular, there remains a lack of a unified control framework that can simultaneously achieve low steady-state tracking error, rapid transient response, energy-efficient operation, and smooth controlforce outputs under disturbances. To address the issue, this paper proposes the thrust smoothness rapid current adaptation proximal policy optimization (TSRCA-PPO) method which learns a near-optimal strategy by a twostage distillation learning framework. The core innovations of this work lie in the reward-function design and the privileged multi-encoder architecture. Ablation studies validate the effectiveness of each module. Simulation results demonstrate that the proposed TSRCA-PPO method consistently outperforms the conventional cascaded P-PID controller across all evaluation metrics. Specifically, TSRCA-PPO reduces the steady-state position error, steady-state attitude error, settling time, energy index, and thrustsmoothness index to 42.7%, 76.5%, 10.6%, 93.5%, and 15.9% of the corresponding P-PID values, respectively.

cs.RO

Quantum Magnonics: Quantum States Generation and Applications

Hybrid systems based on magnons in ferromagnetic materials, such as yttrium iron garnet, have achieved remarkable development in the last decade. These include the coupling of magnons to microwave and optical photons, superconducting qubits, phonons, spins, the center-of-mass motion of a ferromagnet, etc. Here, we review both the experimental and theoretical progress in this field, focusing on the generation of magnonic quantum states and their applications in a broad range of fields. Since the strong coupling is a prerequisite for achieving coherent quantum control of magnons and preparing magnonic quantum states, we start by introducing representative strong-coupling experiments in cavity magnonics, then review a series of protocols for creating various magnonic quantum states, such as Fock, cat, squeezed, and entangled states, and discuss their potential applications in macroscopic quantum studies, quantum information science, quantum sensing, magnonic quantum devices, dark matter detection, and so on. Finally, we summarize the review and give an outlook for the future study of quantum magnonics.

quant-ph

Preparing two-mode magnonic Schr\"odinger cat states in a cavity-magnon-qubit system

The cavity-magnon-qubit system has recently been demonstrated as a new platform for preparing macroscopic quantum states in magnonic systems. Here, we propose to prepare a two-mode magnonic cat state, which is also a non-Gaussian entangled state, based on this practical system involving two yttrium-iron-garnet (YIG) spheres and a superconducting qubit coupled to a common microwave cavity. By adiabatically eliminating the cavity and resonantly driving the qubit, an effective magnon-qubit conditional-displacement interaction is achieved. Further working in the magnon-magnon strong-coupling regime and considering two identical magnon frequencies and coupling strengths to the cavity, two hybridized magnon modes are formed, of which the bright mode is prepared in a cat state after a projective measurement on the qubit, while the dark mode remains in its initial vacuum state. Such a state corresponds to a two-mode cat state of two original magnon modes, which share strong non-Gaussian entanglement. We also discuss practical dissipation and dephasing effects on the cat state. The results indicate that strong nonclassicality and non-Gaussian entanglement are present in the two-mode cat state using fully feasible parameters.

quant-ph

Y-BotFrame: An Extensible Embodied Agent Framework for Quadruped Robot Assistants

Quadruped robots are capable of traversing a wide range of complex terrains with high flexibility. As highly mobile ground-based intelligent platforms, they can be equipped with modules for navigation control, environmental perception, and intelligent interaction, thereby serving as real-world mobile deployment platforms for various algorithms. In this paper, we introduce Y-BotFrame, an extensible embodied platform that turns a robot into an intelligent ground assistant. Y-BotFrame integrates multimodal perception capabilities, including speech, vision, and LiDAR, and employs a large language model as the cognitive core for environmental understanding, contextual reasoning, and task planning. The system maps user natural-language instructions into executable embodied task units that can be carried out by the robot. Y-BotFrame supports natural interaction through voice commands and visual feedback, removing the need for a remote controller and enabling efficient human-robot collaboration. With a highly extensible framework, Y-BotFrame supports plug-and-play integration of new functional modules as well as modular upgrades and iterative development, offering a reference implementation for the real-world deployment of general-purpose, instruction-driven embodied agents.The supplementary video is available at https://xdei-group.github.io/Y-BotFrame/.

cs.RO

AerialClaw: An Open-Source Framework for LLM-Driven Autonomous Aerial Agents

Unmanned aerial vehicles (UAVs) are increasingly used in inspection, search and rescue, environmental monitoring, and emergency response. However, most UAV applications still rely on pre-defined command sequences or task-specific pipelines, where developers manually connect perception, planning, flight control, simulation, logging, and safety modules. This limits the flexibility, reproducibility, and extensibility of autonomous aerial systems. This paper presents AerialClaw, an open-source software framework that enables UAVs to operate as decision-making aerial agents rather than merely command-following platforms. Given a natural-language mission, AerialClaw allows an LLM-based agent to understand the task, maintain context, invoke executable aerial skills, observe perception and runtime feedback, and iteratively update its decisions in a closed loop. The framework adopts a modular brain-skill-runtime architecture, combining hard skills for atomic UAV operations, Markdown-based soft skills for reusable task strategies, document-driven agent state and capability boundaries, memory-driven reflection, safety-oriented runtime validation, and platform-agnostic execution adapters. AerialClaw supports lightweight mock execution, PX4 SITL with Gazebo, and AirSim-based simulation, together with a web console, pluggable model backends, example missions, simulation assets, and staged deployment scripts. By combining standardized aerial skills, document-driven agent state, memory, and closed-loop LLM decision-making, AerialClaw provides a reproducible and extensible open-source framework for building UAV systems that can interpret missions, make decisions, execute skills, and adapt their behavior from feedback.

cs.RO

Hessian-matching Based Weighting for Attitude Determination Using Short-Range DoA Measurements with IMU Assistance

Accurate and reliable attitude determination (AD) is essential for unmanned vehicles operating in Global Navigation Satellite System (GNSS)-denied environments. Short-range wireless arrays can provide direction-of-arrival (DoA) measurements from multiple anchors, enabling AD by aligning corresponding direction vectors (DVs) expressed in the body and navigation frames. In short-range scenarios, navigation-frame DVs inherit non-negligible uncertainty induced by anchor/vehicle position errors in addition to DoA-induced errors in body-frame DVs. Moreover, due to projection and unit-norm normalization, the DV errors are generally anisotropic, which motivates a total least squares (TLS) viewpoint. This paper identifies the key modeling distinction in short-range AD, develops a TLS-consistent formulation based on the total DV error and solves the resulting covariance-weighted orthogonal Procrustes problem via a manifold Gauss--Newton method. To retain the efficiency and numerical robustness of the closed-form weighted Wahba solution, we further propose Hessian-matching based scalar weighting strategies that approximate the Hessian of Wahba formulation to the TLS formulation, including a full-attitude strategy for overall accuracy and a direction-of-interest (DOI) strategy for prioritizing a selected attitude component. Finally, we incorporate IMU-derived gravity as an additional DV pair for static initialization, leading to extended Wahba and extended TLS formulations. Simulation results demonstrate that the proposed Hessian-matching weighting improves accuracy and robustness compared with existing baselines, and that gravity-DV augmentation further reduces attitude errors and improves solution availability under limited anchor availability.

eess.SP

LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations?

Advanced Large Multimodal Models (LMMs) have demonstrated impressive performance in K-12 reasoning tasks, exhibiting great promise as intelligent tutors. Realizing this potential requires models to navigate real-world examinations effectively, yet most existing benchmarks fail to capture the complexity of authentic testing environments. Specifically, most datasets are static, prone to data contamination, and are often confined to restricted modalities, disciplines, and evaluation criteria. To address these issues, we introduce LiveK12Bench, a dynamic, holistic, multi-disciplinary benchmark designed to evaluate the reasoning abilities of LMMs in realistic examination scenarios. LiveK12Bench comprises 2K+ verified questions spanning Mathematics, Physics, Chemistry, and Biology, sourced from the latest real-world exam papers and designed to grow over time. Our framework features several core innovations: 1) featuring an automated pipeline that continuously ingests and parses the latest examination papers to mitigate data leakage; and 2) proposing a novel `Mock Exam' evaluation scheme, which assesses the ability to complete end-to-end exams autonomously with accurate and efficient reasoning paths. Extensive experiments on 12 LMMs reveal that advanced models suffer substantial performance degradation under exam-realistic constraints: GPT-5's score drops from 79 to 53 (out of 100) when process rigor and efficiency are jointly evaluated. Our findings expose critical vulnerabilities, such as sensitivity to complex visual layouts, highlighting the gap between idealized reasoning capabilities and true educational readiness. Both code and dataset are publicly available.

cs.AI

Macroscopic entanglement between two magnon modes via two-tone driving of a superconducting qubit

The cavity-mediated coupling between magnons in an yttrium-iron-garnet (YIG) sphere and a superconducting qubit has recently been demonstrated as a new platform for preparing macroscopic quantum states. Here, based on this system, we propose to entangle two magnon modes in two YIG spheres by driving the qubit with a two-tone field and by appropriately choosing the frequencies and strengths of the two driving fields. We show that strong entanglement can be achieved with fully feasible parameters. We further provide a detection scheme for experimentally verifying the entanglement. Our results indicate that macroscopic entanglement between two magnon modes in two millimeter-sized YIG spheres, involving more than $10^{18}$ spins, can be realized using currently available parameters, which finds promising applications in fundamental studies, such as macroscopic quantum mechanics and the test of unconventional decoherence theories.

quant-ph

Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning

In-context learning (ICL) allows large models to adapt to tasks using a few examples, yet its extension to vision-language models (VLMs) remains fragile. Our analysis reveals that the fundamental limitation lies in an inductive gap, models often produce correct answers from flawed reasoning, while struggling to extract consistent rules across demonstrations. This gap is further exacerbated by two visual-level obstacles: an overwhelming proportion of redundant visual tokens that obscure textual cues, and a skewed attention distribution that favors the initial image at the expense of subsequent context. To address these issues, we introduce a framework that restructures multimodal ICL as a principled inductive-deductive process. The framework incorporates a similarity-based visual token compression module to filter out redundant patches, a dynamic attention rebalancing mechanism to distribute focus equitably across all images, and a chain-of-thought paradigm that explicitly guides the model to analyze individual examples, derive a generalizable rule, and then apply it to the query. An auxiliary learning pipeline combines supervised fine-tuning with reinforcement learning using verifiable rewards to reinforce faithful citation and noise filtering. Evaluations across eight benchmarks covering visual perception, logical reasoning, STEM problems, and sarcasm detection demonstrate consistent and significant improvements over standard ICL baselines for multiple open-source VLMs, highlighting the potential of equipping models with genuine inductive capabilities in multimodal settings.

cs.CV

Magnonic Gottesman-Kitaev-Preskill states

Bosonic quantum error correction encodes a logical qubit in an oscillator, avoiding the hardware overhead of large qubit arrays. Among such encodings, Gottesman-Kitaev-Preskill (GKP) states are paticularly powerful because their phase-space grid structure protects against small displacement errors simultaneously in both conjugate quadratures. Here we provide the first protocol for preparing magnonic GKP states, which involves an ellipsoidal magnetic crystal effectively coupled to a superconducting qubit via a microwave cavity. The geometric anisotropy intrinsically squeezes the magnon mode, while the cavity-mediated qubit control realizes an effective conditional-displacement interaction. We show that two rounds of a conditional-displacement interaction and a qubit projective measurement yield three- and four-component magnonic GKP-like states. We also show how to realize single logical qubit gate operations, such as Pauli, Hadamard and phase gates, completing the logical Pauli basis of the approximate GKP code. Our results establish hybrid magnon-qubit systems as a promising platform for preparing bosonic code states, with applications in magnonic fault-tolerant quantum computation and quantum sensing.

quant-ph

Generation of magnonic squeezed state and its superposition in a hybrid qubit-magnon system

We propose a protocol for generating magnonic squeezed states (MSS) and their superpositions (SMSS) in a hybrid system comprising a superconducting flux qubit magnetically coupled to the Kittel mode of a yttrium iron garnet (YIG) sphere. The flux qubit provides an intrinsic longitudinal interaction with the magnon mode, which, under resonant microwave driving, gives rise to an effective qubit-state-dependent squeezing Hamiltonian. Numerical simulations incorporating realistic dissipation demonstrate that magnon quadrature noise reduction exceeding $8~\mathrm{dB}$ is achievable with experimentally accessible parameters.~By preparing the qubit in a superposition state followed by projective measurement, we further obtain symmetric and antisymmetric superpositions of orthogonally squeezed magnon states exhibiting clear phase-space interference fringes.~We discuss how the fourfold rotational symmetry of these states supports a bosonic logical encoding with potential for protecting against dominant error channels in magnonic platforms.

quant-ph

ADEPT-PolyGraphMT: Automated Molecular Simulation and Multi-Task Multi-Fidelity Machine Learning for Polymer Property Generation and Prediction

The discovery of polymers with targeted properties is challenged by the vast chemical design space and the limited availability of consistent, high-quality data across multiple properties. In this work, an integrated polymer informatics framework is presented that combines the Automated molecular Dynamics Engine for Polymer simulaTions (ADEPT) workflow with multi-task and multi-fidelity machine learning (PolyGraphMT). Polymer repeat units are represented as molecular graphs and processed using a graph neural network to learn structure-property relationships. Starting from SMILES representations for monomers, ADEPT automates the construction of atomistic models and the evaluation of their properties using molecular dynamics simulations and density functional theory calculations. The simulation data are combined with curated experimental data and group contribution theory estimates to construct a unified dataset of approximately 62,000 polymer property values spanning 28 properties. Using this dataset, inter-property correlations are analyzed, and multi-task learning strategies are evaluated for joint property prediction. The results show that multi-task models achieve performance comparable to single-task models in data-rich regimes and exhibit superior accuracy as training data become limited. In addition, fidelity-aware training improves predictive accuracy when combining experimental and computational data sources. The trained models are further applied to large-scale property prediction for polymers in the PolyInfo database and the PI1M virtual polymer library, producing physically consistent property distributions across a broad chemical space. Overall, the proposed framework provides a structured approach for scalable prediction and screening of polymer properties across multiple property types and data fidelity levels.

physics.chem-ph

Engineering near-unitary one-axis twisting evolution via a driven Tavis-Cummings model

One-axis twisting (OAT) interaction is a pivotal resource for manipulating quantum states of atomic ensembles, enabling spin squeezing, atomic-cat-state generation, and weak-phase amplification. Current implementations of OAT dynamics predominantly rely on the Tavis-Cummings model of light-atoms coupling; however, this approach inevitably introduces an additional Stark term that entangles the light with the atoms, which compromises the unitarity of OAT evolution and thereby degrades the OAT-based control precision. Here we propose a scheme based on a driven Tavis-Cummings model to achieve near-unitary OAT evolution. We demonstrate that both constant and time-varying driving of an atoms-cavity hybrid system can realize near-unitary OAT evolution, albeit with distinct coupling strength. Furthermore, when atomic dissipation is taken into account, we find that the time-varying-driving scheme exhibits superior resistance to decoherence. Our approach is broadly applicable to a variety of atomic platforms, including cold atoms, trapped ions, and nitrogen-vacancy centers.

quant-ph

FireRed-OCR Technical Report

We present FireRed-OCR, a systematic framework to specialize general VLMs into high-performance OCR models. Large Vision-Language Models (VLMs) have demonstrated impressive general capabilities but frequently suffer from ``structural hallucination'' when processing complex documents, limiting their utility in industrial OCR applications. In this paper, we introduce FireRed-OCR, a novel framework designed to transform general-purpose VLMs (based on Qwen3-VL) into pixel-precise structural document parsing experts. To address the scarcity of high-quality structured data, we construct a ``Geometry + Semantics'' Data Factory. Unlike traditional random sampling, our pipeline leverages geometric feature clustering and multi-dimensional tagging to synthesize and curate a highly balanced dataset, effectively handling long-tail layouts and rare document types. Furthermore, we propose a Three-Stage Progressive Training strategy that guides the model from pixel-level perception to logical structure generation. This curriculum includes: (1) Multi-task Pre-alignment to ground the model's understanding of document structure; (2) Specialized SFT for standardizing full-image Markdown output; and (3) Format-Constrained Group Relative Policy Optimization (GRPO), which utilizes reinforcement learning to enforce strict syntactic validity and structural integrity (e.g., table closure, formula syntax). Extensive evaluations on OmniDocBench v1.5 demonstrate that FireRed-OCR achieves state-of-the-art performance with an overall score of 92.94\%, significantly outperforming strong baselines such as DeepSeek-OCR 2 and OCRVerse across text, formula, table, and reading order metrics. We open-source our code and model weights to facilitate the ``General VLM to Specialized Structural Expert'' paradigm.

cs.CV