SearcharxivSearch

arXiv subjects

Xinqi Li

Publications and source records attributed to Xinqi Li.

16 recordsLinked to original sources

MILD: Tractable Terrain Modeling for Learning Improved Bipedal Locomotion on Deformable Surfaces

Enabling robots to walk on yielding terrain is vital for applications ranging from disaster response to planetary exploration. While bipedal robots hold immense potential, their locomotion on deformable surfaces remains limited as current simulators fail to capture the spatiotemporal heterogeneity of such yielding substrates. We present MILD, featuring a physics-grounded discrete-element contact solver that accurately simulates spatially varying foot-terrain interactions. Complementing this model, we train a terrain-aware locomotion controller via deep reinforcement learning with latent modulation and proprioceptive estimation. Quantitative comparisons against state-of-the-art methods show our approach generates more diverse and realistic contact scenarios during training, resulting in controllers that exhibit natural adaptation on real deformable surfaces. Through hardware experiments, we demonstrate the system's capability for online terrain identification and adaptation across a wide range of surface stiffness.

cs.RO

PassNet: Scaling Large Language Models for Graph Compiler Pass Generation

Modern tensor compilers such as TorchInductor deliver substantial speedups on mainstream models, yet face a systematic performance ceiling on long-tail workloads -- our profiling shows that 43% of real-world subgraphs experience end-to-end slowdowns under default compilation. While LLMs offer a path toward automated optimization, existing efforts focus on standalone kernel generation. We argue that pass generation -- where LLMs author structured graph transformations that integrate directly into compiler pipelines -- is the more appropriate abstraction. We propose PassNet, the first large-scale ecosystem for LLM-based compiler pass generation, comprising: (1) PassNet-Dataset, over 18K unique computational graphs from 100K real-world models; and (2) PassBench, 200 curated long-tail fusible tasks (comprising 2,060 subgraphs in total) evaluated under the Error-aware Speedup Score (ES_t) -- a metric unifying correctness, stability, and performance -- with layered integrity defenses against systematic LLM exploitation. Experiments reveal that PassBench is both highly discriminative and genuinely unsaturated: the best frontier model trails TorchInductor by 37% in aggregate, yet on individual subgraphs LLMs achieve up to 3x speedup over the same compiler -- indicating that the bottleneck is consistency, not capability. Fine-tuning a small model on merely ~4K PassNet trajectories yields a 2.67x improvement approaching frontier-model performance, demonstrating substantial headroom and validating PassNet as live training infrastructure for advancing LLM-driven compiler optimization. All data, benchmarks, and tooling are publicly available.

cs.AI

BCER Agent: Reliable Long-Horizon MRI Workflow Execution via Compilation, Artifact Binding, and Bounded Local Recovery

Many recent medical VLM and agent studies are benchmarked on 2D images or comparatively short tool-calling exchanges, whereas real MRI analysis typically demands long, interdependent pipelines that operate on 3D/4D volumetric data. Under these conditions, reactive tool-calling agents are prone to cascading breakdowns triggered by faulty intermediate references, mismatched tool arguments, and limited control over cross-step dependencies. To address this, we introduce BCER (Brain-Cerebellum-Extremity-Reflector), a controller architecture aimed at dependable long-horizon MRI workflow execution. BCER decouples high-level planning from execution and provides bounded local recovery. We assess BCER on a multi-organ MRI benchmark covering brain, prostate, and cardiac tasks with both short- and long-chain workflows, using matched task contracts across controller variants and several backbone models. Relative to reactive baselines, BCER yields consistent improvements in end-to-end execution, with the most pronounced gains observed on long-chain workflows. BCER additionally enables auditability by maintaining explicit links between final outputs and intermediate artifacts and measurements. Code and benchmark are released at https://github.com/Albertlongzi/BCER.

eess.IV

PSD: Pushing the Pareto Frontier of Diffusion LLMs via Parallel Speculative Decoding

Diffusion large language models (dLLMs) generate text by iteratively denoising masked token sequences. Although dLLMs can predict all masked positions in parallel within each step, the large number of denoising iterations still makes inference expensive. This cost can be reduced spatially by unmasking multiple tokens per step, or temporally by collapsing multiple denoising steps into one verification call. We propose Parallel Speculative Decoding (PSD), a training-free framework that jointly improves inference along both axes. Using the confidence scores from a single forward pass, PSD selects positions to unmask via a configurable, adaptive unmasking policy and constructs multi-depth speculative drafts without extra model calls. A final batched verification pass then applies hierarchical acceptance, keeping the deepest draft that remains consistent with the updated predictions. Experiments on three dLLMs across reasoning and code generation tasks show that PSD achieves favorable trade-offs between inference efficiency and generation quality, reaching up to $5.5\times$ tokens per forward pass with accuracy comparable to greedy decoding.

cs.CL

Beyond Static Artifacts: A Forensic Benchmark for Video Deepfake Reasoning in Vision Language Models

Current Vision-Language Models (VLMs) for deepfake detection excel at identifying spatial artifacts but overlook a critical dimension: temporal inconsistencies in video forgeries. Adapting VLMs to reason about these dynamic cues remains a distinct challenge. To bridge this gap, we propose Forensic Answer-Questioning (FAQ), a large-scale benchmark that formulates temporal deepfake analysis as a multiple-choice task. FAQ introduces a three-level hierarchy to progressively evaluate and equip VLMs with forensic capabilities: (1) Facial Perception, testing the ability to identify static visual artifacts; (2) Temporal Deepfake Grounding, requiring the localization of dynamic forgery artifacts across frames; and (3) Forensic Reasoning, challenging models to synthesize evidence for final authenticity verdicts. We evaluate a range of VLMs on FAQ and generate a corresponding instruction-tuning set, FAQ-IT. Extensive experiments show that models fine-tuned on FAQ-IT achieve advanced performance on both in-domain and cross-dataset detection benchmarks. Ablation studies further validate the impact of our key design choices, confirming that FAQ is the driving force behind the temporal reasoning capabilities of these VLMs.

cs.CV

Groupwise Registration with Physics-Informed Test-Time Adaptation on Multi-parametric Cardiac MRI

Multiparametric mapping MRI has become a viable tool for myocardial tissue characterization. However, misalignment between multiparametric maps makes pixel-wise analysis challenging. To address this challenge, we developed a generalizable physics-informed deep-learning model using test-time adaptation to enable group image registration across contrast weighted images acquired from multiple physical models (e.g., a T1 mapping model and T2 mapping model). The physics-informed adaptation utilized the synthetic images from specific physics model as registration reference, allows for transductive learning for various tissue contrast. We validated the model in healthy volunteers with various MRI sequences, demonstrating its improvement for multi-modal registration with a wide range of image contrast variability.

eess.IV

GraphNet: A Large-Scale Computational Graph Dataset for Tensor Compiler Research

We introduce GraphNet, a dataset of 2.7K real-world deep learning computational graphs with rich metadata, spanning six major task categories across multiple deep learning frameworks. To evaluate tensor compiler performance on these samples, we propose the benchmark metric Speedup Score S(t), which jointly considers runtime speedup and execution correctness under tunable tolerance levels, offering a reliable measure of general optimization capability. Furthermore, we extend S(t) to the Error-aware Speedup Score ES(t), which incorporates error information and helps compiler developers identify key performance bottlenecks. In this report, we benchmark the default tensor compilers, CINN for PaddlePaddle and TorchInductor for PyTorch, on computer vision (CV) and natural language processing (NLP) samples to demonstrate the practicality of GraphNet. The full construction pipeline with graph extraction and compiler evaluation tools is available at https://github.com/PaddlePaddle/GraphNet .

cs.LG

Charge-polarized superconducting state emerging in a superatomic antipolar metal

The simultaneous presence of polarity and metallicity or superconductivity in a material signifies the exotic polar metallic or superconducting (SC) state, while such materials are extremely rare due to their exclusive nature. Recently, the interweaved CDW and antipolar charge orders have been discovered in a metallic superatomic crystal of Au6Te12Se8 (ATS), while their interplay and competition with the following emergent SC state remains elusive. Here, we report a further experimental investigation of the SC state emerged from the preformed CDW and antipolar order states using scanning tunneling microscopy/spectroscopy in combination with transport and Raman measurements. The temperature-dependent pre-formation and condensation of Cooper pairs are experimentally identified. The pre-existent CDW is gradually suppressed by the preformed Cooper pairs, and then the antipolar charge order is spatially suppressed into a ferrielectric-like polar order by the condensed Cooper pairs of SC state. The exotic charge-polarized superconducting state is discovered in the polar metal of ATS, suggesting a valuable platform for the exploration of intriguing polar superconducting properties.

cond-mat.supr-con

Contrast-Agnostic Groupwise Registration by Robust PCA for Quantitative Cardiac MRI

Quantitative cardiac magnetic resonance imaging (MRI) is an increasingly important diagnostic tool for cardiovascular diseases. Yet, co-registration of all baseline images within the quantitative MRI sequence is essential for the accuracy and precision of quantitative maps. However, co-registering all baseline images from a quantitative cardiac MRI sequence remains a nontrivial task because of the simultaneous changes in intensity and contrast, in combination with cardiac and respiratory motion. To address the challenge, we propose a novel motion correction framework based on robust principle component analysis (rPCA) that decomposes quantitative cardiac MRI into low-rank and sparse components, and we integrate the groupwise CNN-based registration backbone within the rPCA framework. The low-rank component of rPCA corresponds to the quantitative mapping (i.e. limited degree of freedom in variation), while the sparse component corresponds to the residual motion, making it easier to formulate and solve the groupwise registration problem. We evaluated our proposed method on cardiac T1 mapping by the modified Look-Locker inversion recovery (MOLLI) sequence, both before and after the Gadolinium contrast agent administration. Our experiments showed that our method effectively improved registration performance over baseline methods without introducing rPCA, and reduced quantitative mapping error in both in-domain (pre-contrast MOLLI) and out-of-domain (post-contrast MOLLI) inference. The proposed rPCA framework is generic and can be integrated with other registration backbones.

eess.IV

OneFlow: Redesign the Distributed Deep Learning Framework from Scratch

Deep learning frameworks such as TensorFlow and PyTorch provide a productive interface for expressing and training a deep neural network (DNN) model on a single device or using data parallelism. Still, they may not be flexible or efficient enough in training emerging large models on distributed devices, which require more sophisticated parallelism beyond data parallelism. Plugins or wrappers have been developed to strengthen these frameworks for model or pipeline parallelism, but they complicate the usage and implementation of distributed deep learning. Aiming at a simple, neat redesign of distributed deep learning frameworks for various parallelism paradigms, we present OneFlow, a novel distributed training framework based on an SBP (split, broadcast and partial-value) abstraction and the actor model. SBP enables much easier programming of data parallelism and model parallelism than existing frameworks, and the actor model provides a succinct runtime mechanism to manage the complex dependencies imposed by resource constraints, data movement and computation in distributed deep learning. We demonstrate the general applicability and efficiency of OneFlow for training various large DNN models with case studies and extensive experiments. The results show that OneFlow outperforms many well-known customized libraries built on top of the state-of-the-art frameworks. The code of OneFlow is available at: https://github.com/Oneflow-Inc/oneflow.

cs.DC

Facile and fast growth of high mobility nanoribbons of ZrTe$_5$

Recently, ZrTe$_5$ has received a lot of attention as it exhibits various topological phases, such as weak and strong topological insulators, a Dirac semimetal, and a quantum spin Hall insulator in the monolayer limit. While most of studies have been focused on the three-dimensional bulk material, it is highly desired to obtain nanostructured materials due to their advantages in device applications. We report the synthesis and characterizations of ZrTe$_5$ nanoribbons. Via a silicon-assisted chemical vapor transport method, long nanoribbons with thickness as thin as 20 nm can be grown. The growth rate is over an order of magnitude faster than the previous method for growth of bulk crystals. Moreover, transport studies show that nanoribbons are of low unintentional doping and high carrier mobility, over 30,000 cm$^2$/Vs, which enable reliable determination of the Berry phase of $π$ in the $ac$ plane from quantum oscillations. Our method holds great potential in growth of high quality ultra-thin nanostructures of ZrTe$_5$.

cond-mat.mtrl-sci

Scanning tunneling microscopy and spectroscopy of nanoscale twisted bilayer graphene

Nanoscale twisted bilayer graphene (TBG) is quite instable and will change its structure to Bernal (or AB-stacking) bilayer with a much lower energy. Therefore, the lack of nanoscale TBG makes its electronic properties not accessible in experiment up to now. In this work, a special confined TBG is obtained in the overlaid area of two continuous misoriented graphene sheets. The width of the confined region of the TBG changes gradually from about 22 nm to 0 nm. By using scanning tunnelling microscopy, we studied carefully the structure and the electronic properties of the nanoscale TBG. Our results indicate that the low-energy electronic properties, including twist-induced van Hove singularities (VHSs) and spatial modulation of local density-of-state, are strongly affected by the translational symmetry breaking of the nanoscale TBG. Whereas, the electronic properties above the energy of the VHSs are almost not influenced by the quantum confinement even when the width of the TBG is reduced to only a single moire spot.

cond-mat.mes-hall

Zero-Shot Fine-Grained Classification by Deep Feature Learning with Semantics

Fine-grained image classification, which aims to distinguish images with subtle distinctions, is a challenging task due to two main issues: lack of sufficient training data for every class and difficulty in learning discriminative features for representation. In this paper, to address the two issues, we propose a two-phase framework for recognizing images from unseen fine-grained classes, i.e. zero-shot fine-grained classification. In the first feature learning phase, we finetune deep convolutional neural networks using hierarchical semantic structure among fine-grained classes to extract discriminative deep visual features. Meanwhile, a domain adaptation structure is induced into deep convolutional neural networks to avoid domain shift from training data to test data. In the second label inference phase, a semantic directed graph is constructed over attributes of fine-grained classes. Based on this graph, we develop a label propagation algorithm to infer the labels of images in the unseen classes. Experimental results on two benchmark datasets demonstrate that our model outperforms the state-of-the-art zero-shot learning models. In addition, the features obtained by our feature learning model also yield significant gains when they are used by other zero-shot learning models, which shows the flexility of our model in zero-shot fine-grained classification.

cs.CV

Imaging the dynamics of individual hydrogen atom intercalated between two graphene sheets

The interlayer gallery between two adjacent sheets of van der Waals materials is expected to modify properties of atoms and molecules confined at the atomic interfaces. Here, we directly image individual hydrogen atom intercalated between two graphene sheets and investigate its dynamics by scanning tunnelling microscope (STM). The intercalated hydrogen atom is found to be remarkably different from atomic hydrogen chemisorbed on external surface of graphene. Our STM measurements, complemented by first-principles calculations, show that the hydrogen atom intercalated between two graphene sheets has dramatically reduced potential barriers for elementary migration steps. Especially, the confined atomic hydrogen dissociation energy from graphene is reduced to 0.34 eV, which is only about a third of a hydrogen atom chemisorbed on graphene. This offers a unique platform for direct imaging of the atomic dynamics of confined atoms. Our results suggest that the atomic interfaces of van der Waals materials may provide a confined environment to tune the interfacial chemical reactions.

cond-mat.mes-hall

Electrical transport in nano-thick ZrTe$_5$ sheets: from three to two dimensions

ZrTe$_5$ is a newly discovered topological material. Shortly after a single layer ZrTe$_5$ had been predicted to be a two-dimensional topological insulator, a handful of experiments have been carried out on bulk ZrTe$_5$ crystals, which however suggest that its bulk form may be a three-dimensional topological Dirac semimetal. We report the first transport study on ultra thin ZrTe$_5$ flakes down to 10 nm. A significant modulation of the characteristic resistivity maximum in the temperature dependence by thickness has been observed. Remarkably, the metallic behavior, occurring only below about 150 K in bulk, persists to over 320 K for flakes less than 20 nm thick. Furthermore, the resistivity maximum can be greatly tuned by ionic gating. Combined with the Hall resistance, we identify contributions from a semiconducting and a semimetallic bands. The enhancement of the metallic state in thin flakes are consequence of shifting of the energy bands. Our results suggest that the band structure sensitively depends on the film thickness, which may explain the divergent experimental observations on bulk materials.

cond-mat.mtrl-sci

Energy gaps of atomically precise armchair graphene nanoribbons

Graphene nanoribbons (GNRs) are one-dimensional (1D) structures that exhibit a rich variety of electronic properties1-17. Therefore, they are predicted to be the building blocks in next-generation nanoelectronic devices. Theoretically, it has been demonstrated that armchair GNRs can be divided into three families, i.e., Na = 3p, Na = 3p + 1, and Na = 3p + 2 (here Na is the number of dimer lines across the ribbon width and p is an integer), according to their electronic structures, and the energy gaps for the three families are quite different even with the same p1,3-6. However, a systematic experimental verification of this fundamental prediction is still lacking, owing to very limited atomic-level control of the width of the armchair GNRs investigated7,9,10,13,17. Here, we studied electronic structures of the armchair GNRs with atomically well-defined widths ranging from Na = 6 to Na = 26 by using scanning tunnelling microscope (STM). Our result demonstrated explicitly that all the studied armchair GNRs exhibit semiconducting gaps due to quantum confinement and, more importantly, the observed gaps as a function of Na are well grouped into the three categories, as predicted by density-functional theory calculations3. Such a result indicated that we can tune the electronic properties of the armchair GNRs dramatically by simply adding or cutting one carbon dimer line along the ribbon width.

cond-mat.mtrl-sci