SearcharxivSearch

arXiv subjects

He Zheng

Publications and source records attributed to He Zheng.

11 recordsLinked to original sources

One-Step Epitaxial Access to Rhombohedral Graphene Flat-Band States on Step-Bunched SiC

Rhombohedral graphene multilayers provide a moir\'e-free platform for correlated and topological flat-band physics, but direct, transfer-free epitaxial access to thickness-tunable multilayers remains limited. Here we report a one-step graphitization route on 4$^\circ$ off-axis 4H-SiC, in which high-temperature flash annealing simultaneously drives self-organized step bunching and multilayer graphene formation. Atomic-resolution cross-sectional scanning transmission electron microscopy identify local ABC registry and distinguish rhombohedral from Bernal stacking. The thickness is tuned from bilayer to more than twenty layers by varying single parameter, the annealing temperature. Angle-resolved photoemission spectroscopy directly tracks the thickness-dependent evolution from interface-dominated low-energy states toward pronounced near-Fermi-level flat-band spectral weight in thick multilayers. Low-temperature scanning tunneling microscopy and spectroscopy on a 17-layer film further reveal a 13.4 meV low-energy spectral reconstruction and a $\sqrt{3} \times \sqrt{3}$ Kekul\'e-like modulation, providing microscopic signatures consistent with an intervalley-mixed electronic texture. This one-step, transfer-free approach establishes step-bunched SiC as an epitaxial platform that links stacking engineering with moir\'e-free correlated flat-band electronic states.

cond-mat.mtrl-sci

PerceptDrive: Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous Driving

Frozen perception foundation models encode rich geometric, semantic, and dynamic knowledge. Yet narrow conditioning interfaces may attenuate task-relevant cues, while static fusion cannot adjust expert contributions to each scene. We cast this challenge as the prior-to-plan transfer problem and introduce PerceptDrive, a perception prior world-action modeling framework with adaptive expert routing. PerceptDrive feeds teacher-distilled priors from a frozen, driving-adapted provider and dense observation latents from a frozen self-supervised video encoder into a trainable expert-routed world-action model. Expert-specific query branches process these signals, while a prior-retention objective anchors each branch to its prior. A router predicts soft gates from a shared scene representation and combines the expert conditions before trajectory generation. During training, privileged rule-based sub-metric estimates for branch-specific trajectory drafts provide soft-gate distillation targets. The predicted action-free future latent conditions a flow-matching actor. At inference, privileged components are absent; with one front-facing camera, PerceptDrive generates one trajectory per planning step without test-time scoring, reranking, or search. Experiments show that PerceptDrive achieves state-of-the-art performance with 90.4 PDMS on NAVSIM v1 and 90.2 EPDMS on NAVSIM v2, outperforming existing methods. Ablations confirm complementary gains from prior retention and scene-conditioned routing, alongside differential reliance on the three priors. These results demonstrate that preserving and adaptively routing perception priors improves direct planning without test-time candidate selection.

cs.CV

ForgeDrive: Bidirectional Cross-Conditioning for Unified Visual-Action Generation in Autonomous Driving

World-model-based autonomous driving endows the model with the ability to understand scene evolution. Yet this promise is undermined by the prevailing imagine-then-act paradigm, which allows errors from the more challenging visual generation stage to cascade into action planning. We introduce ForgeDrive, a unified autoregressive diffusion framework with visual-action cross-conditioning that closes this gap through act-then-imagine paradigm. ForgeDrive factorizes the future as a sequence of per-timestep frame-action pairs, intertwining each action with its corresponding visual observation. During training, we decouple the diffusion timesteps of the two modalities and introduce a UniDiffuser-style noise scheduler to get the ability to infer either modality from its counterpart and deepen understanding of relationships between images and actions. At inference, we propose a novel act-then-imagine inference paradigm, and find that at each step, action generation is a capability internalized during training, requiring no clean future frame as a prerequisite at inference time; instead, the generated action can improve the accuracy of future frame generation, which in turn enhances the quality of the next action. Additionally, we augment each step with future ego-status prediction, further sharpening planning ability. Extensive experiments on NAVSIM demonstrate that ForgeDrive not only unifies driving simulation, planning, and visual odometry into a single model, but also outperforms existing strong planners without any post-training strategy.

cs.CV

Untangling 3D atomic reconstruction in twisted bilayer 2D crystals via dark field transmission electron microscopy

Reconstruction of the atomic crystal structure in twisted 2D materials has been demonstrated to be responsible for multiple exciting phenomena in van der Waals heterostructures, from the appearance of flat bands in twisted bilayer graphene to Wigner crystallization in transition metal dichalcogenides (TMDs). However, there are still no experimental methods for accessing the 3D atomic distributions nor models that describe the exact atomic shifts in such reconstructed structures, which significantly impedes the development of the field. Dark field (DF) transmission electron microscopy (TEM) has been conventionally employed to visualize the local in-plane atomic displacements. Here we expand this method to obtain a full description of the reconstructed atomic systems and demonstrate the quantitative relations between the local stacking and the intensity in the DF image. We show how local 3D atomic displacements and the interlayer distance can be extracted from a DF image.

cond-mat.mes-hall

Atoms, Worldlines, and the Scalar Approximation

The worldline path-integral method, developed thus far for scalar fields, offers promising computational efficiency in general geometries, However, it relies so far on the scalar approximation that decomposes electromagnetic waves into two independent polarizations. In this work, we investigate different theoretical frameworks of fluctuation-induced effects and analyze the limitations of the worldline path-integral method in modeling multiple-atom Casimir-Polder interactions. In particular, we ask the question: how accurate is the scalar approximation? Using the worldline approach, it appears that a simple sum of the contributions from the two polarizations agrees with the exact Casimir-Polder force for two-atom systems. However, it turns out that this agreement is fortuitous. To enable calculations beyond two atoms via worldlines, we develop general N-atom expressions for the Casimir-Polder force within the scalar approximation. For three-body systems, the scalar worldline method fails drastically, predicting significant discrepancies in both magnitude and sign due to strong polarization mixing. Furthermore, we show that the TE/TM decomposition in the worldline method differs from that of the Green-tensor formalism, and we discuss why this is. This study highlights the inadequacy of scalar worldline models that rely on the polarization-decomposition approximation in general geometries.

quant-ph

Pathwise Differentiation of Worldline Path Integrals

The worldline method is a powerful numerical path-integral framework for computing Casimir and Casimir-Polder energies. An important challenge arises when one desires derivatives of path-integral quantities--standard finite-difference techniques, for example, yield results of poor accuracy. In this work we present methods for computing derivatives of worldline-type path integrals of scalar fields to calculate forces, energy curvatures, and torques. In Casimir-Polder-type path integrals, which require derivatives with respect to the source point of the paths, the derivatives can be computed by a simple reweighting of the path integral. However, a partial-averaging technique is necessary to render the differentiated path integral computationally efficient. We also discuss the computation of Casimir forces, curvatures, and torques between macroscopic bodies. Here a different method is used, involving summing over the derivatives of all the intersections with a body; again, a different partial-averaging method makes the path integral efficient. To demonstrate the efficiency of the techniques, we give the results of numerical implementations of these worldline methods in atomplane and plane-plane geometries. Being quite general, the methods here should apply to path integrals outside the worldline context (e.g., financial mathematics).

quant-ph

NTIRE 2021 Challenge on Quality Enhancement of Compressed Video: Methods and Results

This paper reviews the first NTIRE challenge on quality enhancement of compressed video, with a focus on the proposed methods and results. In this challenge, the new Large-scale Diverse Video (LDV) dataset is employed. The challenge has three tracks. Tracks 1 and 2 aim at enhancing the videos compressed by HEVC at a fixed QP, while Track 3 is designed for enhancing the videos compressed by x265 at a fixed bit-rate. Besides, the quality enhancement of Tracks 1 and 3 targets at improving the fidelity (PSNR), and Track 2 targets at enhancing the perceptual quality. The three tracks totally attract 482 registrations. In the test phase, 12 teams, 8 teams and 11 teams submitted the final results of Tracks 1, 2 and 3, respectively. The proposed methods and solutions gauge the state-of-the-art of video quality enhancement. The homepage of the challenge: https://github.com/RenYang-home/NTIRE21_VEnh

eess.IV

AIM 2022 Challenge on Super-Resolution of Compressed Image and Video: Dataset, Methods and Results

This paper reviews the Challenge on Super-Resolution of Compressed Image and Video at AIM 2022. This challenge includes two tracks. Track 1 aims at the super-resolution of compressed image, and Track~2 targets the super-resolution of compressed video. In Track 1, we use the popular dataset DIV2K as the training, validation and test sets. In Track 2, we propose the LDV 3.0 dataset, which contains 365 videos, including the LDV 2.0 dataset (335 videos) and 30 additional videos. In this challenge, there are 12 teams and 2 teams that submitted the final results to Track 1 and Track 2, respectively. The proposed methods and solutions gauge the state-of-the-art of super-resolution on compressed image and video. The proposed LDV 3.0 dataset is available at https://github.com/RenYang-home/LDV_dataset. The homepage of this challenge is at https://github.com/RenYang-home/AIM22_CompressSR.

eess.IV

Atomistic manipulation of reversible oxidation and reduction in Ag by electron beam

Employing electrons for direct control of nanoscale reaction is highly desirable since it provides fabrication of nanostructures with different properties at atomic resolution and with flexibility of dimension and location. Here, applying in situ transmission electron microscopy, we show the reversible oxidation and reduction kinetics in Ag, well controlled by changing the dose rate of electron beam. Aberration-corrected high-resolution transmission electron microscopy observation reveals that O atoms are preferably inserted and extracted along the {111} close-packed planes of Ag, leading to the nucleation and decomposition of nanoscale Ag2O islands on the Ag substrate. By controlling electron beam size and dose rate, we demonstrated fabrication of an array of 3 nm Ag2O nanodots in an Ag matrix. Our results open up a new pathway to manipulate atomistic reaction with electron beam towards the precise fabrication of nanostructures for device applications.

cond-mat.mtrl-sci

AI Challenger : A Large-scale Dataset for Going Deeper in Image Understanding

Significant progress has been achieved in Computer Vision by leveraging large-scale image datasets. However, large-scale datasets for complex Computer Vision tasks beyond classification are still limited. This paper proposed a large-scale dataset named AIC (AI Challenger) with three sub-datasets, human keypoint detection (HKD), large-scale attribute dataset (LAD) and image Chinese captioning (ICC). In this dataset, we annotate class labels (LAD), keypoint coordinate (HKD), bounding box (HKD and LAD), attribute (LAD) and caption (ICC). These rich annotations bridge the semantic gap between low-level images and high-level concepts. The proposed dataset is an effective benchmark to evaluate and improve different computational methods. In addition, for related tasks, others can also use our dataset as a new resource to pre-train their models.

cs.CV

A Fusion Method Based on Decision Reliability Ratio for Finger Vein Verification

Finger vein verification has developed a lot since its first proposal, but there is still not a perfect algorithm. It is proved that algorithms with the same overall accuracy may have different misclassified patterns. We could make use of this complementation to fuse individual algorithms together for more precise result. According to our observation, algorithm has different confidence on its decisions but it is seldom considered in fusion methods. Our work is first to define decision reliability ratio to quantify this confidence, and then propose the Maximum Decision Reliability Ratio (MDRR) fusion method incorporating Weighted Voting. Experiment conducted on a data set of 1000 fingers and 5 images per finger proves the effectiveness of the method. The classifier obtained by MDRR method gets an accuracy of 99.42% while the maximum accuracy of the original individual classifiers is 97.77%. The experiment results also show the MDRR outperforms the traditional fusion methods as Voting, Weighted Voting, Sum and Weighted Sum.

cs.CV