SearcharxivSearch

arXiv subjects

Li Ji

Publications and source records attributed to Li Ji.

At least 19 recordsLinked to original sources

ContextWeave: A Real-World Workflow Benchmark

Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering. We introduce ContextWeave, a longitudinal benchmark that evaluates whether recalled experience improves downstream agent performance in realistic office-work streams. ContextWeave reconstructs privacy-preserved, multi-month workflows of 14 participants into 1,005 executable tasks, including 568 core evaluation tasks, with instructions, containerized environments, trajectories, and task-specific rubrics. It measures workspace quality and alignment with participant-specific preferences, complemented by diagnostics of relevance, continuity, solvability, and robustness to misleading recall. Across six memory components under a fixed model, the strongest configuration raises Workspace Score from 68.08 to 78.20 and Preference Score from 41.50 to 70.60. With a fixed memory component, recall improves both outcomes for all five tested base models, although gains vary substantially. Our analysis shows that actionable, experience-rich memory supports workflow continuation and reduces redundant exploration more effectively than compact summaries, while it can also be more susceptible to misleading recall. These findings motivate memory systems that optimize not only retrieval relevance but also reliable use during execution.

cs.AI

The XMM-Newton Line Emission Analysis Program (X-LEAP) III: Earth's Magnetospheric X-ray Emission Revealed by 22-Year XMM-Newton Observations

The magnetosphere, protecting the Earth from intense solar activity, is also shaped by the solar wind, while its structure is still uncertain in observation. In this study, we map the X-ray emission in the magnetosphere, which is induced by the charge exchange between the highly-ionized solar wind and the neutral gas around the Earth, known as the magnetospheric solar wind charge exchange (SWCX). In particular, we extract the magnetospheric SWCX in the O VII line emission data adopted from the XMM Line Emission Analysis Program (X-LEAP). The observed magnetospheric SWCX shows an enhanced emission of approximately $I_{\rm OVII}^{\rm mag}\approx 2$ photons $\rm cm^{-2}~ s^{-1}~sr^{-1}$ toward the Sun, showing a consistent shape predicted by numerical simulations. Furthermore, this magnetospheric SWCX exhibits a dependence on the XMM pointing direction, which traces the path length of SWCX emission in the Earth's magnetosphere. Building on this directional dependence, we model the 3D magnetosheath structure using soft X-ray observations for the first time, constraining the averaged boundary geometry and SWCX emissivity distribution over 22 years. Finally, utilizing the XMM data, we derive an empirical O VII emission efficiency of $\alpha_{\rm OVII}=(2.1\pm0.4) \times 10^{-16}\ {\rm eV\,cm^{2}}$.

astro-ph.HE

HiMe: Hierarchical Embodied Memory for Long-Horizon Vision-Language-Action Control

Current Vision-Language-Action (VLA) models excel at robotic manipulation but often struggle with non-Markovian tasks requiring long-term memory and reasoning due to their reliance on immediate observations. Existing solutions face a ''frequency-competence paradox,'' where stronger reasoning models are too slow for real-time control, while faster models lack sufficient reasoning capabilities. To resolve this architectural misalignment, we propose HiMe, a Hierarchical Embodied Memory framework that decouples embodied intelligence into a high-frequency Executor for execution, a Sentry for working memory, and a Planner for long-term strategy. We also introduce a dynamic knowledge system based on cross-modal semantic schemas and active management mechanisms, allowing robots to maintain memory plasticity through ''Add, Update, and Delete'' operations. This hierarchical design effectively balances the conflict between real-time execution and slow thinking planning, significantly improving success rates in long-horizon tasks. Experiments demonstrate that this approach not only outperforms flat memory baselines but also exhibits the novel ability to self-correct its internal knowledge based on human preferences.

cs.RO

Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs

Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- triplets of observations, instructions, and actions that are costly to collect at scale. We argue that this bottleneck stems from conflating two distinct learning objectives: acquiring physical competence (how to move) and acquiring semantic alignment (what to do). Crucially, only the latter requires language supervision. Building on this Decomposition Hypothesis, we propose Task-Agnostic Pretraining (TAP), a two-stage framework that first learns transferable motor priors from cheap, unlabeled interaction data -- including discarded off-task trajectories and autonomous robot play -- via a self-supervised Inverse Dynamics objective. A lightweight second stage then grounds these priors in language using minimal expert data. On the SIMPLER benchmark, TAP matches models trained on over 1M expert trajectories while using orders of magnitude less labeled data, yielding a 10% absolute gain over standard behavior cloning. On a real-world WidowX platform, TAP retains 25% success under camera perturbations where internet-scale baselines collapse to 0%, demonstrating that task-agnostic pretraining produces robust, transferable physical representations and offers a scalable path forward for Embodied AI.

cs.RO

In-Context World Modeling for Robotic Control

Modern Vision-Language-Action (VLA) models often fail to generalize to novel setups, such as altered camera viewpoints or robot morphologies, because they are typically conditioned only on current observations and language instructions. By ignoring the underlying system configuration as a variable, these models implicitly assume a fixed execution context encountered during training, necessitating data-intensive fine-tuning for any new environment. In this work, we introduce In-Context World Modeling (ICWM), a framework that treats system identification as an in-context adaptation problem. ICWM enables robot policies to autonomously infer essential system variables from a short history of self-generated, task-agnostic interactions. Unlike traditional In-Context Learning that uses demonstrations to specify what task to perform, ICWM leverages the context window to understand how the system operates. By processing these interactions before task execution, the model implicitly captures the world dynamics of the current system, enabling adaptation to novel configurations without parameter updates. Extensive experiments in simulation and on real-world robot platforms demonstrate that ICWM significantly outperforms standard VLA baselines on novel camera viewpoints.

cs.RO

Twisted-pair unilateral reconnection: A unifying driver for magnetically powered astrophysical bursts

Magnetic reconnection in twisted loops has long been invoked as an engine powering energetic transients from black hole accretion to neutron star mergers, yet never directly observed. Here we report the first direct observation of the complete reconnection of this type in a solar flare. We find a magnetic loop twisted to about 540 degrees, far exceeding the 180 degrees twist assumed in existing simulations. This extreme twist inherently enables efficient multiple X-line reconnection, akin to the role of turbulence in contemporary theory. Remarkably, the intertwined end breaks unilaterally after reconnection (unlike symmetric breaking in simulations), forming open field lines that release hot plasma -- providing a promising mechanism for coronal generation or heating. We first detect hard X-ray emission from the current sheet, directly proving it as a particle accelerator. Moreover, we discover a power-law relationship between quasi-periodic oscillation frequency and magnetic field strength across solar flares, black hole binaries, active galactic nuclei, magnetars, and gamma-ray bursts. This relation identifies twisted-pair unilateral reconnection as a common burst mechanism and provides a natural ruler for cosmic magnetic fields. These findings establish an observational foundation for future reconnection theory and simulations, offering a unified framework for magnetically powered bursts.

astro-ph.HE

Wafer-scale Demonstration of High-voltage beta-Ga2O3 MOSFETs with Excellent Uniformity and over 3kV Breakdown Voltages

This study demonstrates a wafer-scale growth of a 2-inch Si-doped $\beta$-Ga2O3 (100) epitaxial wafer and the realization of uniform, high-voltage lateral $\beta$-Ga2O3 MOSFET arrays. The 2-inch homoepitaxial $\beta$-Ga2O3 (100) film grown by MOCVD exhibit excellent crystalline uniformity with an average rocking curve FWHM of ~27.0 arcsec and a low surface roughness less than 1 nm, alongside a uniform net doping concentration on the value of 4.60 $\times$ 1E17 cm-3. The fabricated MOSFETs deliver a threshold voltage of -31.75 V, a drain-current on/off ratio over 1E9, a specific on-resistance of 126.52 mohm$\cdot$cm2 and breakdown voltage exceeding 3 kV. Statistical analysis across the entire wafer presents good device uniformity, with threshold voltages ranging from -28 V to -36 V, output current densities of 60-75 mA/mm, and a breakdown voltage over 3 kV. These results provide the demonstration using the 2-inch $\beta$-Ga2O3 epitaxial wafer to realize high-voltage $\beta$-Ga2O3 MOSFETs with wafer-scale performance uniformity for next-generation power device application.

cond-mat.mtrl-sci

Isolated neutron star candidates from the fourth generation XMM-Newton catalogues

X-ray thermally emitting isolated neutron stars (XINSs) are a rare population that provides insights into neutron star cooling, magnetic-field evolution, and Galactic demographics. Using more than two decades of observations from the European Space Agency's XMM-Newton Observatory, we searched the 4XMM-DR9 and 4XMM-DR12 catalogues for absorbed XINS candidates down to a flux of $10^{-14}$ erg cm$^{-2}$ s$^{-1}$ in the 0.5--1 keV band. Candidates were selected based on soft X-ray spectra and the absence of catalogued optical, ultraviolet, or infrared counterparts. Follow-up observations with XMM-Newton and FAST were complemented by data from the SRG/eROSITA All-Sky Survey, Chandra, and optical surveys. Of ten sources analysed, five are compelling XINS candidates, one is the known XINS 4XMM J022141.5-735632, two are extragalactic contaminants, and two remain ambiguous because of limited photon statistics. The five candidates exhibit soft ($kT\sim80-100$ eV), moderately absorbed, and stable X-ray emission consistent with distant XINSs. They are located primarily in the Galactic plane, with possible associations at distances of $\sim$1.8 and $\sim$6 kpc. Population-synthesis simulations predict $20\pm5$ XINSs within the 4XMM-DR12 footprint, of which $6^{+2}_{-3}$ exceed our flux threshold, consistent with the observed sample if additional candidates are confirmed. The model further predicts that $\sim$70% of the population remains below the detection threshold. Deep optical and additional X-ray observations are required to establish the nature of the candidates. Future missions such as NewAthena will enable more detailed studies of these distant populations and improve constraints on the Galactic XINS population.

astro-ph.HE

Persistent incommensurate amorphous/crystalline meta-interfaces enable engineering-grade superlubricity

Friction dissipates a substantial portion of global energy, motivating the pursuit of superlubricity, a state of near-zero friction, in real-world systems. Conventional approaches rely on crystalline lattice mismatch to suppress periodic energy barriers, but real interfaces invariably contain defects, edges and grain boundaries that restore high-friction states. Here we introduce a materials-agnostic strategy based on amorphous/crystalline heterointerfaces to achieve robust superlubricity under engineering-relevant conditions. Using diamond-like carbon (DLC) and crystalline MoS2 as a model system, we show through experiments and atomistic simulations that their interface remains incommensurate at all orientations and exhibits vanishing energy barriers during friction. In contrast, twisted MoS2 bilayers readily reorient into commensurate, high-friction states. We scale this effect by fabricating laser-patterned arrays of DLC/MoS2 meta-contacts reinforced with Ti3C2Tx MXene, forming hierarchical interfaces that sustain a friction coefficient of ~0.008 over 100000 cycles under combined extreme conditions: millimetre-scale contact size, 12.7 GPa contact pressure and RH 40% air. This unprecedented performance arises from four synergistic factors: intrinsic incommensurability at amorphous/crystalline interface, the rigidity of DLC support, MXene-based mechanical reinforcement and normalized load distribution by geometric patterning. These findings establish a general design paradigm that extends structural superlubricity from nanoscale model systems to practical technologies for sustainable engineering.

cond-mat.mtrl-sci

A biomimetic feedback loop for sustaining self-lubrication and wear resistance

Intelligent materials that self-sense and self-regulate are an emerging frontier in sustainable technology. Here we introduce Cu(Au)/C nanocomposite films that act as bioinspired self-adjusting lubricants. In these films, frictional heating triggers melting and migration of soft metal nanoparticles (NPs) such as Cu or Au along nano-pores to the friction interface, where the metal catalyzes the in-situ formation of ordered carbon nano-structures. Real-time monitoring of friction coefficient, electrical resistance(R), and metal release confirms an autonomous cycle: high friction coefficient generates heat, melting the metal NPs; the migrating metal then lowers friction coefficent by creating low-friction nanostructures, which reduces heat and arrests further migration until friction rises again. This self-limiting feedback enables stable ultra-low friction (~0.04) and an exceptional wear life (>40 km) even in high vacuum. By utilizing friction-derived heat as an intrinsic activation signal, our system establishes a general paradigm for intelligent, self-regulating materials with applications extending beyond tribology.

cond-mat.mtrl-sci

ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development

The evolution of Large Language Models (LLMs) into autonomous agents has expanded the scope of AI coding from localized code generation to complex, repository-level, and execution-driven problem solving. However, current benchmarks predominantly evaluate code logic in static contexts, neglecting the dynamic, full-process requirements of real-world engineering, particularly in backend development which demands rigorous environment configuration and service deployment. To address this gap, we introduce ABC-Bench, a benchmark explicitly designed to evaluate agentic backend coding within a realistic, executable workflow. Using a scalable automated pipeline, we curated 224 practical tasks spanning 8 languages and 19 frameworks from open-source repositories. Distinct from previous evaluations, ABC-Bench require the agents to manage the entire development lifecycle from repository exploration to instantiating containerized services and pass the external end-to-end API tests. Our extensive evaluation reveals that even state-of-the-art models struggle to deliver reliable performance on these holistic tasks, highlighting a substantial disparity between current model capabilities and the demands of practical backend engineering. Our code is available at https://github.com/OpenMOSS/ABC-Bench.

cs.SE

Nex-N1: Agentic Models Trained via a Unified Ecosystem for Large-Scale Environment Construction

The evolution of Large Language Models (LLMs) from passive responders to autonomous agents necessitates a fundamental shift in learning paradigms -- from static imitation to incentive-driven decision making. However, this transition is significantly impeded by the lack of scalable infrastructure capable of constructing high-quality interaction signals for effective policy learning. To address this, we introduce a comprehensive method designed to systematically scale the diversity and complexity of interactive environments. Our method realizes this scaling by addressing three orthogonal dimensions: (1) Complexity: NexAU, a flexible agent framework that supports building complex agent hierarchies via simple configurations; (2) Diversity: NexA4A automatically generates diverse agent hierarchies from natural language to cover infinite domains; and (3) Fidelity: NexGAP bridges the simulation-reality gap by integrating dynamic real-world environment for grounded trajectories synthesis. We train Nex-N1 upon the diverse and complex interactive environments established by our infrastructure. Empirical results on benchmarks such as SWE-bench and tau2 demonstrate that Nex-N1 consistently outperforms SOTA open-source models and achieves competitive performance against frontier proprietary models on complex agentic tasks. We open-source the Nex ecosystem and model weights to facilitate further research.

cs.CL

SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models

Vision-Language-Action (VLA) models excel in robotic manipulation but are constrained by their heavy reliance on expert demonstrations, leading to demonstration bias and limiting performance. Reinforcement learning (RL) is a vital post-training strategy to overcome these limits, yet current VLA-RL methods, including group-based optimization approaches, are crippled by severe reward sparsity. Relying on binary success indicators wastes valuable information in failed trajectories, resulting in low training efficiency. To solve this, we propose Self-Referential Policy Optimization (SRPO), a novel VLA-RL framework. SRPO eliminates the need for external demonstrations or manual reward engineering by leveraging the model's own successful trajectories, generated within the current training batch, as a self-reference. This allows us to assign a progress-wise reward to failed attempts. A core innovation is the use of latent world representations to measure behavioral progress robustly. Instead of relying on raw pixels or requiring domain-specific fine-tuning, we utilize the compressed, transferable encodings from a world model's latent space. These representations naturally capture progress patterns across environments, enabling accurate, generalized trajectory comparison. Empirical evaluations on the LIBERO benchmark demonstrate SRPO's efficiency and effectiveness. Starting from a supervised baseline with 48.9% success, SRPO achieves a new state-of-the-art success rate of 99.2% in just 200 RL steps, representing a 103% relative improvement without any extra supervision. Furthermore, SRPO shows substantial robustness, achieving a 167% performance improvement on the LIBERO-Plus benchmark.

cs.RO

LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

Visual-Language-Action (VLA) models report impressive success rates on robotic manipulation benchmarks, yet these results may mask fundamental weaknesses in robustness. We perform a systematic vulnerability analysis by introducing controlled perturbations across seven dimensions: objects layout, camera viewpoints, robot initial states, language instructions, light conditions, background textures and sensor noise. We comprehensively analyzed multiple state-of-the-art models and revealed consistent brittleness beneath apparent competence. Our analysis exposes critical weaknesses: models exhibit extreme sensitivity to perturbation factors, including camera viewpoints and robot initial states, with performance dropping from 95% to below 30% under modest perturbations. Surprisingly, models are largely insensitive to language variations, with further experiments revealing that models tend to ignore language instructions completely. Our findings challenge the assumption that high benchmark scores equate to true competency and highlight the need for evaluation practices that assess reliability under realistic variation.

cs.RO

Probing the \ion{He}{2} re-Ionization ERa via Absorbing \ion{C}{4} Historical Yield (HIERACHY) IV: A complex redshifted absorption system intrinsic to quasar

High-resolution spectra provide a powerful tool in studying the associated absorption lines (AALs) in quasars. We present a case study of the quasar J014741-030247 at $z \sim$ 4.75, which hosts complex intrinsic absorption lines revealed by the high-resolution Magellan/MIKE spectrum obtained from the HIERACHY program. We focus on one of the strongest absorption systems ($z$ $\sim$ 4.7804) and determine the column densities of multiple ionization species. We find that the Apparent Optical Depth method may significantly underestimate the column densities of high ions. Decomposing the absorption into multiple components yields a better fit and reveals clear evidence of partial coverage. The variation in covering fractions among different ions suggests that high ions are distributed more extensively in this system. We estimate electron densities of different components ($630 - 4070 \ \mathrm{cm}^{-3}$), these are based on the column densities of \ion{Si}{2}* and \ion{C}{2}*. By combining these with the hydrogen number density and ionization parameter derived from photoionization modeling, we infer that the different components are located at distances of 2.3 to 9.5 kpc from the quasar. The derived $N_{\mathrm H} / n_{\mathrm e}$ and the partial coverage observed in low ions all require cloud sizes smaller than 1 pc, even down to 0.01 pc. Finally, the low kinetic luminosity of the gas ($< 0.5\% L_\mathrm{bol}$) indicates that it is insufficient to drive significant AGN feedback and may only suppress star formation via `multistage' mechanism.

astro-ph.GA

Magnetic Reconnection as a Potential Driver of X-ray Variability in Active Galactic Nuclei

We present a systematic analysis on the X-ray variability in 13 bright quasars at z > 4.5, combining recent Swift observations from 2021 to 2023 and archival multi-epoch observations. Upper limits of the luminosity measurements were included in the analysis by using the Kaplan-Meier estimator method. It is found that the high-z quasars exhibit X-ray variability on both short-term (hours-to-days) and intermediate-term (weeks-to-months) timescales, with short-term variability dominating the overall variation. A linear correlation exists between the global mean ($\mu_{\mathrm{L_{2-10\,keV}}}$) and standard deviation ($\sigma_{\mathrm{L_{2-10\,keV}}}$) of X-ray luminosities, which is independent of the X-ray photon index and optical-to-X-ray spectral slope. The localized stochastic magnetic reconnection mechanism is strongly favored, which can naturally lead to a scale-invariant power-law energy distribution and satisfactorily explain the correlation. The $\sigma$-$\mu$ correlation parallels with the well-documented rms-flux relation of low-z active galactic nuclei (AGNs), implying the magnetic reconnection mechanism could drive short-timescale X-ray variability in both high- and low-z AGNs. The highest-z quasar in our sample, J142952+544717 (z = 6.18), shows a luminosity distribution extending to ${10}^{47}\ \rm{erg\ {s}^{-1}}$ with a not conspicuous median luminosity. On the other hand, J143023+420436 (z = 4.7), which hosts the most relativistic jet among known high-z blazars, is dominated in the high-luminosity regime (${10}^{47}\ \rm{erg\ {s}^{-1}}$ ), making it an ideal target for multi-wavelength follow-up observations. J090630+693030 is found to have a rest-frame period of 182.46 days and J143023+420436 has a period of 16.89 days, both could be explained by the global evolution of plasmoid chains, in which magnetic islands formed during reconnection may merge successively.

astro-ph.HE

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning

Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks, yet they face significant challenges in embodied task planning scenarios that require continuous environmental understanding and action generation. Existing approaches generate open-loop action scripts based on static knowledge, making it difficult to learn causal relationships between actions and environmental feedback, particularly in partially observable environments. We introduce Embodied Planner-R1, a novel outcome-driven reinforcement learning framework that enables LLMs to develop interactive capabilities through autonomous exploration with minimal supervision. Our framework incorporates three key innovations: (1) Without human annotations, we employ pure reinforcement learning with group rollout, incorporating in-environment interaction through parallel exploration; (2) completion-driven sparse reward; and (3) Interactive Policy Optimization (IPO) for efficient learning from grouped trajectories. Across two challenging text-based Embodied planning benchmarks, Embodied Planner-R1 achieves impressive completion rates of 97.78% on ALFWorld and 79.92% on ScienceWorld, surpassing prior methods by a large margin, and suffers only a -3.66% drop in previously unseen environments, evidencing strong generalization.

cs.CL

The MALATANG survey: Dense gas distribution on sub-kiloparsec scales across the disk of M82

We present observations of HCN J=4-3 and HCO^+ J=4-3 lines obtained with the James Clerk Maxwell Telescope as part of the MALATANG survey, combined with archival HCN J=1-0 and HCO^+ J=1-0 data from the Green Bank Telescope, to study the spatial distribution and excitation conditions of dense molecular gas in the disk of M82. We detect HCN J=4-3 and HCO^+ J=4-3 emission within the central region (< 500 pc) of the galaxy, while the J=1-0 emission lines exhibit a more extended spatial distribution (> 700 pc). The dense gas shows a clear double-lobed structure in both spatial distribution and kinematics, with the HCN and HCO^+ J=4-3 lines in the southwest lobe blueshifted by ~ 40 km/s relative to the J=1-0 lines. The HCN J=4-3/1-0 and HCO^+ J=4-3/1-0 line-luminosity ratios range from 0.09 to 0.53 and from 0.14 to 0.87, respectively, with mean values of 0.18 +/- 0.04 and 0.36 +/- 0.06. The HCN ratio is lower than the typical average observed in nearby star-forming galaxies, whereas the HCO^+ ratio is comparatively higher, suggesting that the high-J HCN emission in M82 is significantly sub-thermally excited. Spatially, the peak values of the J=4-3/1-0 ratios are found in the northwest region of M82, coinciding with the galaxy-scale outflow. Elevated HCN/HCO^+ ratios are also detected in roughly the same area, potentially tracing local excitation enhancements driven by the outflow. The HCN/HCO^+ J=4-3 ratio across all detected regions ranges from 0.19 to 1.07 with a mean value of 0.41 +/- 0.11, which is significantly lower than the average J=1-0 ratio of 0.76 +/- 0.08. Both ratios are significantly lower than the average values observed in nearby star-forming galaxies, which could be related to the relatively low gas density and the presence of an extended photo-dissociation region in M82.

astro-ph.GA