SearcharxivSearch

arXiv subjects

Jingjing Ma

Publications and source records attributed to Jingjing Ma.

7 recordsLinked to original sources

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory

Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification, and cross-embodiment execution. We present ABot-AgentOS, a general robotic Agent Operating System that sits above low-level controllers and provides a deliberative agent layer for scene-conditioned planning, context-isolated skill execution, multi-stage verification, multi-modal memory, and edge-cloud collaboration. To evaluate such systems, we introduce EmbodiedWorldBench, an executable benchmark with 16 indoor, outdoor, and hybrid scenes, four difficulty levels, and over 200 tasks involving navigation, object search, NPC dialogue, dynamic events, and trace-grounded scoring. ABot-AgentOS further introduces Universal Multi-modal Graph Memory, a persistent source-grounded substrate that converts dialogue, visual observations, spatial context, temporal relations, and task traces into typed nodes and edges. A failure-driven self-evolution loop converts diagnosed memory failures into gated runtime evo-assets that are promoted only to later evaluation splits, preventing current-split ground-truth leakage while enabling continual improvement. On an initial EmbodiedWorldBench subset, ABot-AgentOS improves over a single-controller baseline in both task success and goal completion. Across memory benchmarks, ABot-AgentOS Static achieves 87.5 on LoCoMo, 59.9 on OpenEQA EM-EQA, 88.6 on Mem-Gallery, and 76.5 Acc@All on NExT-QA; self-evolution further improves LoCoMo to 88.7, OpenEQA to 60.4, and Mem-Gallery to 89.0. These results suggest that a general Agent OS layer can improve long-horizon embodied execution while providing persistent, auditable memory for continual interaction.

cs.AI

ABot-N1: Toward a General Visual Language Navigation Foundation Model

Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this integration via monolithic policies that map observations directly to actions, yet they often suffer from coordinate drift and poor handling of long-tail semantics. Furthermore, these black-box mappings lack interpretability, hindering the simultaneous achievement of generality, robustness, and transparency. We present ABot-N1, a step toward a general Visual Language Navigation foundation model, that addresses these challenges by decoupling cognition from control via a slow-fast architecture guided by dual visual-language signals. More specifically, a slow vision-language reasoner performs explicit Chain-of-Thought reasoning while producing a pixel goal. This compact set of image-space anchor points serves as a universal interface for diverse tasks, including point-goal, object-goal, poi-goal, instruction-following, and person-following. Subsequently, a fast action expert leverages both the textual cues and the pixel guidance to generate continuous waypoints at the native control frequency. By bridging high-level intents and low-level control through pixel-grounded anchors paired with explicit linguistic traces, our approach ensures robust, generalizable, and interpretable navigation across simulation and real-world benchmarks. ABot-N1 establishes new state-of-the-art records, delivering massive gains specifically in urban-scale navigation: boosting POI arrival by 35.0% (to 77.3%) and achieving 95.4%/92.9% SR in complex indoor and outdoor scenes. It also maintains superior robustness across object-reaching, person-following, and instruction-following tasks. New Point-Goal/POI-Goal benchmarks are released as open source to advance the field of urban-scale navigation.

cs.CV

Exploration on the Two-stream Instability in the Polar Cusp Under Solar Storm Disturbances and its Potential Impacts on Spacecraft

During solar storms, the polar cusp often exhibits electron populations with distinct velocity distributions, which may be associated with the two-stream instability. This study reveals the evolution of the two-stream instability associated with electron velocities and the interaction between the growth phase of the two-stream instability and the electrostatic solitary waves (ESWs). The results from particle-in-cell (PIC) simulations are compared with satellite observational data and computational outcomes. The potential risks associated with two-stream instability, including surface charge accumulation and communication system interference on spacecraft, are also explored. The findings show that, in the high-latitude polar cusp region, the interaction between the solar wind plasma propagating along magnetic field lines and the upward-moving ionospheric plasma could drive two-stream instability, leading to the formation of electron hole structures in phase space and triggering a bipolar distribution of ESWs. When the spatial magnetic field and wave vector meet specific conditions, the enhanced electron cyclotron motion could suppress the formation of two-stream instability and electron hole structures, leading to a reduction in the amplitude of the ESWs. The results offer valuable insights for a deeper understanding of the impact of solar storms on the polar cusp environment, as well as for monitoring electromagnetic environment and ensuring the stable operation of spacecraft.

physics.plasm-ph

Semantic-aware DropSplat: Adaptive Pruning of Redundant Gaussians for 3D Aerial-View Segmentation

In the task of 3D Aerial-view Scene Semantic Segmentation (3D-AVS-SS), traditional methods struggle to address semantic ambiguity caused by scale variations and structural occlusions in aerial images. This limits their segmentation accuracy and consistency. To tackle these challenges, we propose a novel 3D-AVS-SS approach named SAD-Splat. Our method introduces a Gaussian point drop module, which integrates semantic confidence estimation with a learnable sparsity mechanism based on the Hard Concrete distribution. This module effectively eliminates redundant and semantically ambiguous Gaussian points, enhancing both segmentation performance and representation compactness. Furthermore, SAD-Splat incorporates a high-confidence pseudo-label generation pipeline. It leverages 2D foundation models to enhance supervision when ground-truth labels are limited, thereby further improving segmentation accuracy. To advance research in this domain, we introduce a challenging benchmark dataset: 3D Aerial Semantic (3D-AS), which encompasses diverse real-world aerial scenes with sparse annotations. Experimental results demonstrate that SAD-Splat achieves an excellent balance between segmentation accuracy and representation compactness. It offers an efficient and scalable solution for 3D aerial scene understanding.

cs.CV

Stein-Weiss type inequalities with partial variable weight on the upper half space and related weighted inequalities

In this paper, we establish a class of Stein-Weiss type inequality with partial variable weight functions on the upper half space using a weighted Hardy type inequality. Overcoming the impact of weighted functions, the existence of extremal functions is proved via the concentration compactness principle, whereas Riesz rearrangement inequality is not available. Moreover, the cylindrical symmetry with respect to $t$-axis and the explicit forms on the boundary of all nonnegative extremal functions are discussed via the method of moving planes and method of moving spheres, as well as, regularity results are obtained by the regularity lift lemma and bootstrap technique. As applications, we obtain some weighted Sobolev inequalities with partial variable weight function for Laplacian and fractional Laplacian.

math.AP

Reversed Hardy-Littlewood-Sobolev inequalities with vertical weights on the upper half space

In this paper, we obtain the reversed Hardy-Littlewood-Sobolev inequality with vertical weights on the upper half space and discuss the extremal functions. We show that the sharp constants in this inequality are attained by introducing a renormalization method. The classification of corresponding extremal functions is discussed via the method of moving spheres. Moreover, we prove the sufficient and necessary conditions of existence for positive solutions to the Euler-Lagrange equations by using Pohozaev identities in weak sense. This renormalization method is rearrangement free, which can be also applied to prove the existence of extremal functions for sharp (reversed) Hardy-Littlewood-Sobolev inequality with extended kernels and other similar inequalities.

math.AP

SoftMatch Distance: A Novel Distance for Weakly-Supervised Trend Change Detection in Bi-Temporal Images

General change detection (GCD) and semantic change detection (SCD) are common methods for identifying changes and distinguishing object categories involved in those changes, respectively. However, the binary changes provided by GCD is often not practical enough, while annotating semantic labels for training SCD models is very expensive. Therefore, there is a novel solution that intuitively dividing changes into three trends (``appear'', ``disappear'' and ``transform'') instead of semantic categories, named it trend change detection (TCD) in this paper. It offers more detailed change information than GCD, while requiring less manual annotation cost than SCD. However, there are limited public data sets with specific trend labels to support TCD application. To address this issue, we propose a softmatch distance which is used to construct a weakly-supervised TCD branch in a simple GCD model, using GCD labels instead of TCD label for training. Furthermore, a strategic approach is presented to successfully explore and extract background information, which is crucial for the weakly-supervised TCD task. The experiment results on four public data sets are highly encouraging, which demonstrates the effectiveness of our proposed model.

cs.CV