SearcharxivSearch

arXiv subjects

Liang Ma

Publications and source records attributed to Liang Ma.

At least 19 recordsLinked to original sources

AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment

Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but existing aviation evaluations centered on static knowledge do not support systematic testing of procedural execution and safety compliance in interactive environments. This paper presents the AeroCopilot Operational Environment (ACOE), a reproducible interactive virtual-cockpit test environment, and AeroCopilotBench, a two-tier aviation agent evaluation benchmark. Tier-1 evaluates aviation knowledge using 1,200 multiple-choice questions, while Tier-2 comprises 73 emergency and abnormal tasks derived from the manufacturers' Pilot's Operating Handbooks (POHs) and instantiated in ACOE. ACOE converts natural-language procedures into executable state transitions, final-state goal conditions, and hard safety constraints, enabling models to interpret cockpit state, diagnose faults, and operate aircraft systems through standardized tool interfaces. We establish a safety-gated evaluation framework in which a trajectory succeeds only when all task goals are achieved without violating any hard safety constraint, while safe goal progress and trajectory safety are measured separately. Across 12 models, the highest Tier-2 success rate is 72.6%, while static knowledge performance does not consistently translate into procedural execution. Analysis of 451 failed episodes from 3 representative models identifies recurring failures in procedural completeness, use of state feedback, and long-horizon execution management. These findings motivate state-aware agent orchestration, joint assessment of task completion and trajectory safety, and repeated regression testing. ACOE and AeroCopilotBench provide a reproducible foundation for testing knowledge application, interactive execution, and operational safety in aviation agents.

cs.AI

Global Distinctions Between New Electrovacuum and Kundt Class

It was observed [2606.30426] that our recently constructed electrovacuum [2606.23782], supported by electromagnetic fields, can be locally transformed to a special Petrov type D solution within the general Kundt class, which is not an electrovacuum but instead describes a spacetime generated by (accelerating) electric and magnetic charges. In this short note, which serves as a supplemental to [2606.23782], we analyse the major global differences between the two locally-equivalent solutions.

gr-qc

New Rotating Black Hole in Electromagnetic Fields: Cosmological Horizon without Cosmological Constant

We obtain a new electrovacuum and related background spacetime that contains a cosmological horizon supported entirely by the electromagnetic field. The local solution belongs to the Kundt class of type D, but is globally distinct. We further construct an exact solution describing a Kerr black hole in this background. We study the global structure including horizons and singularities, and derive the first law of black hole thermodynamics. The emergence of a cosmological horizon in Einstein-Maxwell gravity without invoking a positive cosmological constant or dark energy is tantalizing, and may provide a new avenue for exploring cosmological and astrophysical phenomena related to black holes and the late-time cosmology.

gr-qc

Demagnetizing KBR and New Ricci-flat Rotating Metric

We construct a new Ricci-flat metric by demagnetizing the recently reported Kerr-Bertotti-Robinson (KBR) solution. The metric is a deformation of the Kerr metric characterized by a parameter $B$, so that the asymptotic Kerr becomes a regular dome of spindle shape with north and south poles. Despite lacking an asymptotically-flat region, we find that the first law of black hole thermodynamics can be established. Some thermodynamic relations are identical to those of the Kerr black hole, as if the constant $B$ is absent. Our Ricci-flat rotating metric serves a neutral seed for a variety of inequivalent schemes of magnetizing the Schwarzschild and Kerr black holes.

gr-qc

Fatigue-Related Reaction Time Forecasting via EEG Functional Connectivity in Sustained Attention Task

Mental fatigue related behavioral performance decline precipitates catastrophic accidents in sustained attention tasks. While existing neurophysiological systems effectively detect current behavioral performance, they often lack the capability to forecast behavioral lapses with sufficient temporal lead time for intervention. This study proposes a novel model for the reaction time (RT) forecasting using EEG functional connectivity features. Thirty participants engaged in a sustained Psychomotor Vigilance Test (PVT) with concurrent 30-channel EEG recording. Mutual information (MI) between electrodes was calculated as functional connectivity features. Random Forest regression model (RF) was trained to predict single-trial RTs across forecasting horizons ranging from 0 to 20 seconds. The model demonstrated robust predictive validity, achieving a Root Mean Square Error (RMSE) of 23.75 ms for immediate detection and maintaining high accuracy (RMSE = 24.07 ms) across different forecasting horizons. Interpretability analysis via SHAP and Linear Mixed Effects model further support the validity of the proposed model and revealed distinct temporal biomarkers. This study validates the feasibility of forecasting behavioral performance 20 seconds in advance, offering a promising methodology for proactive fatigue management in safety-critical systems.

cs.HC

Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE

We present Mamoda2.5, a unified AR-Diffusion framework that seamlessly integrates multimodal understanding and generation within a single architecture. To efficiently enhance the model's generation capability, we equip the Diffusion Transformer backbone with a fine-grained Mixture-of-Experts (MoE) design (128 experts, Top-8 routing), yielding a 25B-parameter model that activates only 3B parameters, significantly reducing training costs while scaling up the model capacity. Mamoda2.5 achieves top-tier generation performance on VBench 2.0 and sets a new record in video editing quality, surpassing evaluated open-source models and matching the performance of current top-tier proprietary models, including the Kling O1 on OpenVE-Bench. Furthermore, we introduce a joint few-step distillation and reinforcement learning framework that compresses the 30-step editing model into a 4-step model and greatly accelerates model inference. Compared to open-source baselines, Mamoda2.5 achieves up to $95.9\times$ faster video editing inference. In real-world applications, Mamoda2.5 has been successfully deployed for content moderation and creative restoration tasks in advertising scenarios, achieving a 98% success rate in internal advertising video editing scenario.

cs.CV

Robust Flat Magnetoresistivity in D0$_3$-Fe$_3$Ga Driven by Chiral Anomaly

Topologically non-trivial nodes emerging from flat-band crossings not only enhance unconventional topological responses but also play a fundamental role in exploring correlation-driven topological physics. Here, we report the exceptionally robust chiral-anomaly-dominated transport in D0_3-Fe_3Ga. First, we observe a combination of positive and negative magnetoresistance, ideal planar longitudinal magnetoresistance (PLMR), and the planar Hall effect (PHE). Second, ultra-low-temperature resistivity exhibits pronounced non-Fermi-liquid (NFL) behavior, accompanied by the emergence of giant intrinsic anomalous Hall conductivity (AHC), in excellent agreement with our DFT calculations, which confirm the existence of tilted Weyl points arising from crossings of nearly three-dimensional (3D) flat bands. Most remarkably, we detect an exceptionally robust flat magnetoresistance (flat-MR) that persists without decay up to 33 T. This set of phenomena provides strong evidence that the Fermi level intersects the flattened Weyl crossings, offering confirmation of a topological flat-band semimetal. D0_3-Fe_3Ga presents a promising magnetic platform for quantum device innovations.

cond-mat.mtrl-sci

ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation

Vision-Language-Action (VLA) models and world-action models have emerged as central paradigms for general-purpose robotic intelligence, yet their empirical progress remains constrained by the absence of evaluation protocols that are both physically realistic and diagnostically controlled. Simulator-centric benchmarks provide scale and reproducibility, but cannot fully capture the reality gap induced by perception noise, contact dynamics, latency, calibration error, and hardware constraints. Conversely, real-robot evaluations are often fragmented across platforms, scenes, objects, and scoring rules, making fair comparison and failure attribution difficult. We introduce ManipArena, a standardized real-robot evaluation framework for studying manipulation generalization under matched physical conditions. ManipArena comprises 20 tasks, 10,812 expert trajectories, 13.5M frames, and approximately 188 robot hours across tabletop and mobile manipulation. The framework combines schema-defined task variation, stratified in-domain, visualshift, and semantic-OOD trials, subtask-level partial-credit scoring, three-level language annotations, low-level motor signals, and paired real-to-sim environments reconstructed from physical scenes. Using ManipArena, we evaluate seven tabletop configurations spanning VLA and world-action-model policies. The results show that real-robot conclusions depend not only on architecture, but also on model provenance, fine-tuning regime, data sampling, and annotation granularity. ManipArena thus provides a reproducible and interpretable foundation for diagnosing capability boundaries and failure modes in embodied generalization.

cs.RO

UniFluids: Unified Neural Operator Learning with Conditional Flow-matching

Partial differential equation (PDE) simulation holds extensive significance in scientific research. Currently, the integration of deep neural networks to learn solution operators of PDEs has introduced great potential. In this paper, we present UniFluids, a conditional flow-matching framework that harnesses the scalability of diffusion Transformer to unify learning of solution operators across diverse PDEs with varying dimensionality and physical variables. Unlike the autoregressive PDE foundation models, UniFluids adopts flow-matching to achieve parallel sequence generation, making it the first such approach for unified operator learning. Specifically, the introduction of a unified four-dimensional spatiotemporal representation for the heterogeneous PDE datasets enables joint training and conditional encoding. Furthermore, we find the effective dimension of the PDE dataset is much lower than its patch dimension. We thus employ $x$-prediction in the flow-matching operator learning, which is verified to significantly improve prediction accuracy. We conduct a large-scale evaluation of UniFluids on several PDE datasets covering spatial dimensions 1D, 2D and 3D. Experimental results show that UniFluids achieves strong prediction accuracy and demonstrates good scalability and cross-scenario generalization capability. The code will be released later.

cs.LG

World2Act: Latent Action Post-Training from World Model Dynamics

World Models (WMs) offer a promising mechanism for post-training Vision-Language-Action (VLA) policies by providing dynamics priors that improve generalization under task and scene variation. However, most WM-based post-training methods rely on pixel-space supervision, making policies sensitive to visual artifacts introduced by imperfect WM rollouts. We present World2Act, a latent-space post-training framework that transfers WM dynamics to the VLA policy without pixel-space supervision. World2Act operates in two stages: 1) it induces a shared video-action latent space by contrastively aligning WM-dynamics latents with action embeddings, and 2) it post-trains the VLA by guiding policy action representations toward WM-imagined dynamics rather than decoded pixels. Built on GR00T-N1.6, World2Act delivers absolute success-rate gains of up to +2.5% on simulation benchmarks (RoboCasa, LIBERO, Bridge-SIMPLER) and +6.7% on a real robot over finetuned VLA baselines. Notably, it outperforms pixel-space WM supervision by up to +6.0%, including on LIBERO where pixel supervision degrades the baseline, suggesting that latent WM dynamics offer a more stable WM-based post-training alternative to pixel-space transfer.

cs.CV

Implicit Geometry Representations for Vision-and-Language Navigation from Web Videos

Vision-and-Language Navigation (VLN) has long been constrained by the limited diversity and scalability of simulator-curated datasets, which fail to capture the complexity of real-world environments. To overcome this limitation, we introduce a large-scale video-instruction framework derived from web-based room tour videos, enabling agents to learn from natural human walking demonstrations in diverse, realistic indoor settings. Unlike existing datasets, our framework integrates both open-ended description-enriched trajectories and action-enriched trajectories reconstructed in 3D, providing richer spatial and semantic supervision. A key extension in this work is the incorporation of implicit geometry representations, which extract spatial cues directly from RGB frames without requiring fragile 3D reconstruction. This approach substantially improves data utilization, alleviates reconstruction failures, and unlocks large portions of previously unusable video data. Comprehensive experiments across multiple VLN benchmarks (CVDN, SOON, R2R, and REVERIE) demonstrate that our method not only sets new state-of-the-art performance but also enables the development of robust zero-shot navigation agents. By bridging large-scale web videos with implicit spatial reasoning, this work advances embodied navigation towards more scalable, generalizable, and real-world applicable solutions.

cs.CV

Choose What to Observe: Task-Aware Semantic-Geometric Representations for Visuomotor Policy

Visuomotor policies learned from demonstrations often overfit to nuisance visual factors in raw RGB observations, resulting in brittle behavior under appearance shifts such as background changes and object recoloring. We propose a task-aware observation interface that canonicalizes visual input into a shared representation, improving robustness to out-of-distribution (OOD) appearance changes without modifying or fine-tuning the policy. Given an RGB image and an open-vocabulary specification of task-relevant entities, we use SAM3 to segment the target object and robot/gripper. We construct an L0 observation by repainting segmented entities with predefined semantic colors on a constant background. For tasks requiring stronger geometric cues, we further inject monocular depth from Depth Anything 3 into the segmented regions via depth-guided overwrite, yielding a unified semantic--geometric observation (L1) that remains a standard 3-channel, image-like input. We evaluate on RoboMimic (Lift), ManiSkill YCB grasping under clutter, four RLBench tasks under controlled appearance shifts, and two real-world Franka tasks (ReachX and CloseCabinet). Across benchmarks and policy backbones (Flow Matching Policy and SmolVLA), our interface preserves in-distribution performance while substantially improving robustness under OOD visual shifts.

cs.RO

Pressure-Induced Metal-Insulator and Paramagnet-Altermagnet Transitions in Rutile OsO2 Single Crystals

Altermagnets with compensated spin structures and nonrelativistic spin splitting have emerged as a new class of magnetic materials. Rutile OsO2 has been theoretically predicted to be altermagnetic, but experimental studies have been limited by synthesis challenges. We have succeeded in synthesizing high-quality single crystals of rutile OsO2. Electrical transport studies reveal that OsO2 is highly conductive and exhibits clear Fermi liquid behavior, indicating strong electron-electron scattering. Magnetic measurements show that the crystals are isotropically paramagnetic. Density-functional theory calculations indicate that bulk OsO2 is semimetallic with coexisting electron and hole pockets, with its magnetic ground state strongly dependent on the on-site Coulomb correlation U. Angle-resolved photoemission spectroscopy studies unveil that the bulk bands do not yet show altermagnetic spin splitting. Interestingly, resistivity is rather pressure sensitive: at 44 GPa, a clear metal-insulator transition occurs. Hybrid functional calculations reveal that applying pressure significantly increases the Hubbard U value, driving a phase transition from a paramagnetic metal to an altermagnetic metal, and eventually to an altermagnetic insulator. These findings suggest that tuning external pressure effectively modulates the magnetic ground state of OsO2, providing a pathway to realize altermagnetism in this material.

cond-mat.mes-hall

Reply to "Threefold error in the reported zero-field cooled magnetic moment of single crystal $La_2SmNi_2O_7$ (arXiv: 2602.23240)"

We respond to the critique by Aleksandr V. Korolev and Evgeny F. Talantsev on the superconducting phase fraction ($f$) calculations in Li et al. Nature 649, 871-878 (2026). First, the weak upturn in the low-temperature tail of our data has been confirmed to originate from the background, and the paramagnetic Meissner effect is absent in our case; thus, field-cooled (FC) data can be used for superconducting phase fraction calculations. Second, demagnetization effect must be calculated based on the actual measured moment as a function of $f$, which has been well-established and routinely employed in the superconductivity community. In contrast, Korolev and Talantsev treated the demagnetization field as a constant; thus, their calculation underestimates $f$ by a factor of $(1-N\chi_{meas})(1-N)$. This factor is close to 1/3, given $N$ = 0.849, $\chi_{meas}$ = -1.313 in our study, which explains the origin of their deviated result (nearly three times smaller than our results). Third, our sample is a homogeneous high-quality bulk single crystal, evidenced by various techniques, making the existence of multiple discrete superconducting regions highly unlikely. We conclude that the superconducting phase fraction calculations reported in Li et al. Nature 649, 871-878 (2026) are not invalidated by the analyses presented in Korolev et al. arXiv: 2602.23240 (2026).

cond-mat.supr-con

Black Hole Thermodynamic Ensembles, Euclidean Action and Legendre Transformation

In thermodynamics, a Legendre transformation of the free energy provides a mapping between different statistical ensembles. In this work, we demonstrate that performing a Legendre transformation of the black hole on-shell action is equivalent to imposing different boundary conditions on the fields. Consequently, the choice of ensemble must be consistent with, and cannot contradict, the imposed boundary conditions. From this perspective, it follows that for four-dimensional dyonic black holes, the on-shell action can only be expressed either as a function of the electric charge and the magnetic potential, or alternatively as a function of the magnetic charge and the electric potential. Inspired by the Legendre transformation of the Maxwell field, we argue that for purely gravitational theories whose metric geometries admit a \(U(1)\) fiber bundle structure, i.e.\ rotating, boosted, or Kaluza-Klein monopole configurations, one can similarly introduce appropriate Legendre terms, in the sense of dimensional reduction, to modify the thermodynamic ensemble of the black hole. Within the dimensional reduction framework, we study the on-shell action of black holes in five-dimensional minimal supergravity with a Chern-Simons term, analyze the corresponding Legendre transformation procedure, and show how the resulting formulation remains consistent with the Wald formalism.

hep-th

Deep Deterministic Nonlinear ICA via Total Correlation Minimization with Matrix-Based Entropy Functional

Blind source separation, particularly through independent component analysis (ICA), is widely utilized across various signal processing domains for disentangling underlying components from observed mixed signals, owing to its fully data-driven nature that minimizes reliance on prior assumptions. However, conventional ICA methods rely on an assumption of linear mixing, limiting their ability to capture complex nonlinear relationships and to maintain robustness in noisy environments. In this work, we present deep deterministic nonlinear independent component analysis (DDICA), a novel deep neural network-based framework designed to address these limitations. DDICA leverages a matrix-based entropy function to directly optimize the independence criterion via stochastic gradient descent, bypassing the need for variational approximations or adversarial schemes. This results in a streamlined training process and improved resilience to noise. We validated the effectiveness and generalizability of DDICA across a range of applications, including simulated signal mixtures, hyperspectral image unmixing, modeling of primary visual receptive fields, and resting-state functional magnetic resonance imaging (fMRI) data analysis. Experimental results demonstrate that DDICA effectively separates independent components with high accuracy across a range of applications. These findings suggest that DDICA offers a robust and versatile solution for blind source separation in diverse signal processing tasks.

stat.ME

Quadratic Curvature Correction to 5D Myers-Perry Metric

We consider quadratic curvature perturbation to the Myers-Perry black hole in five dimensions at the linear level in the coupling constant. The solution can then be solved order by order in terms of two dimensionless angular momentum parameters up to an arbitrary order. We present the results up to tenth order. The perturbed solution allows us to obtain the higher-derivative correction to the black hole thermodynamics, which we find is in complete agreement with the Reall-Santos method.

hep-th

Full spectrum of Love numbers of Reissner-Nordstrom black hole in D-dimensions

We present a comprehensive analysis of the full spectrum of tidal Love numbers for Reissner-Nordstr\"om (RN) black holes in general spacetime dimensions. By perturbing the Einstein-Maxwell theory around the $D$-dimensional RN background, we derive an effective two dimensional quadratic action encompassing tensor, vector, and scalar-type perturbation sectors. Through diagonalization, we obtain master equations governing each sector and extract the corresponding Love numbers from the asymptotic behavior of the solutions. Our results confirm that all Love numbers vanish for four-dimensional RN black holes. In higher dimensions, the tensor and vector Love numbers reproduce previously known results. For the previously unknown scalar-type Love numbers, we show also they vanish for integer valued effective multipolar indices and display logarithmic running behavior when the corresponding indices are half integers.

hep-th