SearcharxivSearch

arXiv subjects

Feng Ding

Publications and source records attributed to Feng Ding.

At least 19 recordsLinked to original sources

When Diffusion Models Forget Who You Are: Identity Preservation in Face Inpainting under Large Occlusions

Face inpainting with diffusion models has recently achieved impressive visual quality, yet preserving identity fidelity under significant occlusion and conflicting text guidance remains a major challenge. To address this issue, we present Reference Semantic Inpainting for Face (ReSem-Face), a cascaded diffusion framework that introduces an explicit identity-conditioned semantic prior for multi-reference face inpainting. Our approach distills representative identity features from multiple references to reconstruct missing semantic regions, which then guide the diffusion process through a multi-stream conditioning architecture. This design provides strong semantic constraints when pixels are absent and stabilizes identity reconstruction while remaining compatible with prompt-driven edits. Experiments on CelebAHQ-IDI-5 and VGGFace2 demonstrate that ReSem-Face yields more reliable identity-preserving completion under severe semantic masks and improves text-controlled editing quality compared with representative baselines.

cs.CV

Active rejection enables reliable generalization of universal machine-learning interatomic potentials

Universal machine learning interatomic potentials (uMLIPs) bridge quantum-mechanical accuracy and large-scale molecular dynamics, but the cost of high-accuracy calculations such as r$^2$SCAN limits training to datasets that remain small relative to the open materials space. Strong average benchmark performance also does not guarantee reliable energy--force predictions for every structure. We propose Adaptive Multi-Teacher Routing (ATR), which reformulates high-fidelity data construction as a structure-wise decision problem under uncertainty. Using a small set of real r$^2$SCAN labels, ATR calibrates multiple pretrained uMLIP teachers and combines structural descriptors, teacher identity, and inter-teacher disagreement to estimate the reliability of each structure--teacher pair. It selects high-confidence predictions for pseudo-label generation and rejects structures for which no teacher is sufficiently reliable. With real r$^2$SCAN labels for only 0.2\% of candidate structures, ATR distils 2.89 million traceable r$^2$SCAN-level pseudo-labels for pretraining. On held-out r$^2$SCAN structures and the MP-r$^2$SCAN benchmark, a lightweight CHGNet trained on the ATR-generated dataset consistently outperforms the baseline and non-routed controls. Finite-temperature molecular dynamics further shows that ATR improves dynamical robustness across multiple material systems, maintaining stable trajectories where baseline simulations undergo catastrophic structural collapse. These results establish active rejection as an effective mechanism for converting multiple pretrained uMLIPs into a scalable and reliable data-construction system for high-fidelity uMLIPs.

cs.LG

ChWDTA: Channel-wise Wavelet-Domain Transformer Attention and Entropy Modeling for Learned Image Compression

State-of-the-art learned image compression (LIC) schemes are increasingly based on hybrid CNN-transformer architectures. To further improve rate-distortion performance, we introduce channel-wise wavelet transforms into both the transformer and entropy-coding components. First, we propose a channel-wise wavelet-domain transformer attention (ChWDTA) mechanism. ChWDTA keeps the efficient windowed spatial self-attention used in modern LIC backbones, but computes the Q/K/V projections on channel-wise wavelet-transformed features before mapping the attention output back with the inverse transform. The resulting Channel-wise Wavelet-Domain Transformer Block (ChWDTB) therefore preserves the spatial tokenization pattern of windowed attention while sparsifying the channel covariance seen by the attention projections. Second, in the entropy-coding stage, we introduce a channel-wise wavelet packet (ChWP) decomposition that produces four equal-sized subbands, which better fit channel-wise slice-based autoregressive entropy modeling. When each channel-wise subband is divided into two slices, we use eight slices for entropy coding. With this configuration, the proposed scheme obtains BD-rate reductions of -17.82%, -19.15%, and -22.56% on the Kodak, CLIC Professional Validation, and Tecnick test sets, respectively. Even when each channel-wise subband is coded as a single slice, the scheme still retains most of the coding gains with lower complexity. The results confirm the advantage of introducing wavelet transform in CNN-transformer-based LIC schemes.

eess.IV

ML-CLIPSim: Multi-Layer CLIP Similarity for Machine-Oriented Image Quality

We study full-reference image quality assessment from a machine-centric perspective, where images are evaluated by how well they preserve information for downstream models. We formulate machine-oriented quality as a latent machine utility and approximate it through pairwise predictive-consistency comparisons. To this end, we construct PCMP, a dataset of PSNR-matched distortion pairs labeled by consistency votes from multiple pretrained models. We further propose ML-CLIPSim, a differentiable quality metric built on a frozen CLIP visual encoder, which aggregates intermediate patch-token similarities and global image embeddings. Experiments on machine-preference benchmarks, human-IQA datasets, and learned image compression show that ML-CLIPSim better aligns with machine-oriented preferences than conventional fidelity and perceptual metrics, while remaining competitive for human quality prediction. Used as a compression distortion term, it improves rate--task trade-offs across multiple downstream tasks.

eess.IV

Phase transformation kinetics in MoS2 governed by S-S repulsive interactions and defect-interface compatibility

The metastable T' phase in monolayer MoS2 exhibits remarkable persistence despite a strong thermodynamic driving force toward the stable H phase. Using machine learning-accelerated molecular dynamics and first-principles calculations, we reveal that this kinetic arrest originates from repulsive S-S interactions, which impose high energy barriers during both nucleation and grain boundary propagation. While sulfur vacancies can alleviate these barriers in certain interfaces, they fail to accelerate transformation at the most stable interface, ZZ-Mo|-, due to their thermodynamic instability there. Instead, vacancies migrate into the T' phase, leaving the advancing front defect-free. Direct simulations of nanostructures confirm that H-phase nucleation initiates at corners or edges, and all observed growth fronts adopt the ZZ-Mo|- configuration, consistent with its low interfacial energy but slow kinetics. Our work establishes that phase transformation in 2D materials is governed not by global defect concentration, but by the local compatibility between defects and moving interfaces, offering a new paradigm for controlling structural transitions through interface-specific design.

cond-mat.mtrl-sci

Trillion-atom molecular dynamics simulations with ab initio accuracy

Material properties are fundamentally dictated by multiscale phenomena, which often reach mesoscale in size. The {\mu}m mesoscale is also the size which can be observed directly under an optical microscope, bridging the atomistic microscopic description with the continuous model macroscopic world. In this work, we report an unprecedented molecular dynamics (MD) simulation comprising 1.62 trillion atoms. Utilizing the neuroevolution potential (NEP) framework, we attained ab initio accuracy on China's New-generation Intelligent Supercomputer. Our implementation achieves a time-to-solution (s/step/atom) 100 times faster than previous state-of-the-art machine learning force field simulations, and 1,000 times faster than the Gordon Bell Prize-winning application from six years ago. Furthermore, we demonstrate an 86.9% weak scaling efficiency from a single GPGPU to 45,000 GPGPUs. These results redefine atomistic simulation boundaries, enabling direct mesoscopic modeling with quantum-level precision.

cond-mat.mtrl-sci

Facet-dependent Chemical Kinetics Governed Growth of Twisted Graphene Layers with Pre-designed Angles

Twisted graphene layers (TGLs) provide a powerful platform for investigating multiple quantum phenomena, yet their scalable deployment is hindered by the lack of reliable synthesis with precise angle. Here, benefited from a deeper understanding of the interplay between grain index and graphene growth kinetics, we report a scalable strategy to grow TGLs with pre-designed twist angles on platinum (Pt) via chemical vapor deposition (CVD), Through a combination of complementary in situ methods, we identified the activity sequence of different Pt grains and attributed it to the area ratio of exposed (110) facets during graphene-induced surface reconstruction. Moreover, we revealed that CVD-grown graphene orientation is determined by the grain-orientation-dependent surface morphology. By leveraging the so-established correlations between grain index with both graphene growth priority and its orientation, we achieve controlled folding and tearing of graphene overlayer using a pair of adjacent grains with dramatically different catalytical activity and kink-free atomic steps. We reveal that overlayer-induced step bunching and terrace reconfiguration critically govern the domain morphology and folding direction. Building on this mechanistic insight, we demonstrate a substrate-engineering framework where specific platinum grains are rationally selected to yield TGLs with pre-designed twist angles, including magic angle with flat band dispersion. This work not only highlights fundamental kinetics of Pt catalyzed graphene CVD growth, but also offers a generalizable methodology for manipulating foldable two-dimensional materials via dynamic substrate reconstruction, exampled by programmable growth of high-quality TGLs on open surfaces.

cond-mat.mtrl-sci

SCALE: Semantic- and Confidence-Aware Conditional Variational Autoencoder for Zero-shot Skeleton-based Action Recognition

Zero-shot skeleton-based action recognition (ZSAR) aims to recognize action classes without any training skeletons from those classes, relying instead on auxiliary semantics from text. Existing approaches frequently depend on explicit skeleton-text alignment, which can be brittle when action names underspecify fine-grained dynamics and when unseen classes are semantically confusable. We propose SCALE, a lightweight and deterministic Semantic- and Confidence-Aware Listwise Energy-based framework that formulates ZSAR as class-conditional energy ranking. SCALE builds a text-conditioned Conditional Variational Autoencoder where frozen text representations parameterize both the latent prior and the decoder, enabling likelihood-based evaluation for unseen classes without generating samples at test time. To separate competing hypotheses, we introduce a semantic- and confidence-aware listwise energy loss that emphasizes semantically similar hard negatives and incorporates posterior uncertainty to adapt decision margins and reweight ambiguous training instances. Additionally, we utilize a latent prototype contrast objective to align posterior means with text-derived latent prototypes, improving semantic organization and class separability without direct feature matching. Experiments on NTU-60 and NTU-120 datasets show that SCALE consistently improves over prior VAE- and alignment-based baselines while remaining competitive with diffusion-based methods.

cs.CV

Simplicity Prevails: The Emergence of Generalizable AIGI Detection in Visual Foundation Models

While specialized detectors for AI-Generated Images (AIGI) achieve near-perfect accuracy on curated benchmarks, they suffer from a dramatic performance collapse in realistic, in-the-wild scenarios. In this work, we demonstrate that simplicity prevails over complex architectural designs. A simple linear classifier trained on the frozen features of modern Vision Foundation Models , including Perception Encoder, MetaCLIP 2, and DINOv3, establishes a new state-of-the-art. Through a comprehensive evaluation spanning traditional benchmarks, unseen generators, and challenging in-the-wild distributions, we show that this baseline not only matches specialized detectors on standard benchmarks but also decisively outperforms them on in-the-wild datasets, boosting accuracy by striking margins of over 30\%. We posit that this superior capability is an emergent property driven by the massive scale of pre-training data containing synthetic content. We trace the source of this capability to two distinct manifestations of data exposure: Vision-Language Models internalize an explicit semantic concept of forgery, while Self-Supervised Learning models implicitly acquire discriminative forensic features from the pretraining data. However, we also reveal persistent limitations: these models suffer from performance degradation under recapture and transmission, remain blind to VAE reconstruction and localized editing. We conclude by advocating for a paradigm shift in AI forensics, moving from overfitting on static benchmarks to harnessing the evolving world knowledge of foundation models for real-world reliability.

cs.CV

Turning Insulators into Accelerators: Deciphering the Interfacial Conductivity Boost in ZrO2-Li2ZrCl6 Composites through Machine Learning Molecular Dynamics Simulations

Halide solid-state electrolytes have emerged as promising candidates for all-solid-state lithium batteries due to their high oxidative stability and deformability, yet their moderate ionic conductivity remains a bottleneck. While incorporating ionically insulating ZrO2 nanoparticles (Nat. Commun. 2023, 14, 2459) has been experimentally shown to enhance the ionic conductivity of Li2ZrCl6, the atomistic origin governing this interfacial phenomenon remains unclear. Here, we bridge the spatiotemporal gap in modeling complex heterostructures by developing an accurate machine-learned force fields based on neuroevolution potential, enabling large-scale molecular dynamics simulations of ZrO2/Li2ZrCl6 heterostructures. By systematically investigating four representative low-lattice-mismatch ZrO2/Li2ZrCl6 interfaces, we identify spontaneous interfacial amorphization driven by space-charge effects upon surface cleavage, trapping Li+ and leading to under-coordinated Li+ polyhedrons with pronounced geometric distortion. These distorted amorphous interfacial regions exhibit markedly enhanced Li+ hopping activity, significantly outperforming the bulk lattice, provided that local mobile Li+ inventory is not depleted by surface charge redistribution. This work establishes a computational framework for training validated machine-learned force fields for interfaces and provides mechanistic understandings of the interfacial conductivity boost in the insulator-conductor composites, guiding the rational design of electrolytes toward next-generation solid-state batteries.

cond-mat.mtrl-sci

MPF-Net: Exposing High-Fidelity AI-Generated Video Forgeries via Hierarchical Manifold Deviation and Micro-Temporal Fluctuations

With the rapid advancement of video generation models such as Veo and Wan, the visual quality of synthetic content has reached a level where macro-level semantic errors and temporal inconsistencies are no longer prominent. However, this does not imply that the distinction between real and cutting-edge high-fidelity fake is untraceable. We argue that AI-generated videos are essentially products of a manifold-fitting process rather than a physical recording. Consequently, the pixel composition logic of consecutive adjacent frames residual in AI videos exhibits a structured and homogenous characteristic. We term this phenomenon `Manifold Projection Fluctuations' (MPF). Driven by this insight, we propose a hierarchical dual-path framework that operates as a sequential filtering process. The first, the Static Manifold Deviation Branch, leverages the refined perceptual boundaries of Large-Scale Vision Foundation Models (VFMs) to capture residual spatial anomalies or physical violations that deviate from the natural real-world manifold (off-manifold). For the remaining high-fidelity videos that successfully reside on-manifold and evade spatial detection, we introduce the Micro-Temporal Fluctuation Branch as a secondary, fine-grained filter. By analyzing the structured MPF that persists even in visually perfect sequences, our framework ensures that forgeries are exposed regardless of whether they manifest as global real-world manifold deviations or subtle computational fingerprints.

cs.CV

DiffFace-Edit: A Diffusion-Based Facial Dataset for Forgery-Semantic Driven Deepfake Detection Analysis

Generative models now produce imperceptible, fine-grained manipulated faces, posing significant privacy risks. However, existing AI-generated face datasets generally lack focus on samples with fine-grained regional manipulations. Furthermore, no researchers have yet studied the real impact of splice attacks, which occur between real and manipulated samples, on detectors. We refer to these as detector-evasive samples. Based on this, we introduce the DiffFace-Edit dataset, which has the following advantages: 1) It contains over two million AI-generated fake images. 2) It features edits across eight facial regions (e.g., eyes, nose) and includes a richer variety of editing combinations, such as single-region and multi-region edits. Additionally, we specifically analyze the impact of detector-evasive samples on detection models. We conduct a comprehensive analysis of the dataset and propose a cross-domain evaluation that combines IMDL methods. Dataset will be available at https://github.com/ywh1093/DiffFace-Edit.

cs.CV

An Overlay Multicast Routing Method Based on Network Situational Awareness and Hierarchical Multi-Agent Reinforcement Learning

Compared with IP multicast, Overlay Multicast (OM) offers better compatibility and flexible deployment in heterogeneous, cross-domain networks. However, traditional OM struggles to adapt to dynamic traffic due to unawareness of physical resource states, and existing reinforcement learning methods fail to decouple OM's tightly coupled multi-objective nature, leading to high complexity, slow convergence, and instability. To address this, we propose MA-DHRL-OM, a multi-agent deep hierarchical reinforcement learning approach. Using SDN's global view, it builds a traffic-aware model for OM path planning. The method decomposes OM tree construction into two stages via hierarchical agents, reducing action space and improving convergence stability. Multi-agent collaboration balances multi-objective optimization while enhancing scalability and adaptability. Experiments show MA-DHRL-OM outperforms existing methods in delay, bandwidth utilization, and packet loss, with more stable convergence and flexible routing.

cs.NI

Observation of robust macroscale structural superlubricity

Structural superlubricity (SSL) promises nearly frictionless and wearless sliding, but has until now been considered a special and extreme interfacial phenomenon limited to micro- and nanoscale contacts. Here, we demonstrate robust macroscale SSL within a single sub-millimeter graphite contact. Previously reported near-zero friction coefficients, where friction is nearly independent of normal load, have only been observed at microscale contacts under low loads. Our system expands both contact size and load into the macroscopic regime, exhibiting friction coefficients that fluctuate around zero and reach values as low as $10^{-6}$ across a broad load range from 1 mN to 0.5 N. Negative friction coefficients are also observed. Similar behavior is observed at graphite/MoS$_2$ interfaces, indicating that macroscale SSL is a generalizable phenomenon across flat layered materials. These findings overturn long-standing scaling limitations and establish macroscale SSL as a paradigm-shifting platform for next-generation mechanical and electromechanical systems.

cond-mat.mtrl-sci

Non-Euclidean interfaces decode the continuous landscape of graphene-induced surface reconstructions

Interfacial reconstruction between two-dimensional (2D) materials and metal substrates fundamentally governs heterostructure properties, yet conventional flat substrates fail to capture the continuous crystallographic landscape. Here, we overcome this topological limitation using non-Euclidean interfaces-curved 2D graphene-copper surfaces as a model system-to traverse the infinite spectrum of lattice orientations. By integrating multimodal microscopy with a deep-learning-enhanced dimensional upscaling framework, we translate 2D scanning electron microscopy (SEM) contrast into quantitative three-dimensional (3D) morphologies with accurate facet identification. Coupling these observations with machine-learning-assisted density functional theory, we demonstrate that reconstruction is governed by a unified thermodynamic mechanism where high-index facets correspond to specific local minima in the surface energy landscape. This work resolves the long-standing complexity of graphene-copper faceting and establishes non-Euclidean surface topologies as a generalizable paradigm for decoding and controlling interfacial reconstruction in diverse metal-2D material systems.

cond-mat.mes-hall

From cluster to nanocrystal: the continuous evolution and critical size of copper clusters revealed by machine learning

The evolution of cluster structure with size and the critical size for the transition from cluster to nanocrystal have long been fundamental problems in nanoscience. Due to limitations of experimental technology and computational methods, the exploration of the continuous evolution of clusters towards nanocrystal is still a big challenge. Here, we proposed a machine learning force field (MLFF) that can generalize well to various copper systems ranging from small clusters to large clusters and bulk. The continuous evolution of copper clusters CuN towards nanocrystal was revealed by investigating clusters in a wide size range (7 <= N <= 17885) based on MLFF simulated annealing. For small CuN (N < 40), electron counting rule plays a major role in stability. For large CuN (N > 80), geometric magic number rule plays a dominant role and the evolution of clusters is based on the formation of more and more icosahedral shells. For medium size CuN (40 <= N <= 80), both rules contribute. The critical size from cluster to nanocrystal was calculated to be around 8000 atoms (about 6 nm in diameter). Our work terminates the long-term challenge in nanoscience, and lay the methodological foundation for subsequent research on other cluster systems.

cond-mat.mtrl-sci

Artificial Intelligence-Enabled Holistic Design of Catalysts Tailored for Semiconducting Carbon Nanotube Growth

Catalyst design is crucial for materials synthesis, especially for complex reaction networks. Strategies like collaborative catalytic systems and multifunctional catalysts are effective but face challenges at the nanoscale. Carbon nanotube synthesis contains complicated nanoscale catalytic reactions, thus achieving high-density, high-quality semiconducting CNTs demands innovative catalyst design. In this work, we present a holistic framework integrating machine learning into traditional catalyst design for semiconducting CNT synthesis. It combines knowledge-based insights with data-driven techniques. Three key components, including open-access electronic structure databases for precise physicochemical descriptors, pre-trained natural language processing-based embedding model for higher-level abstractions, and physical - driven predictive models based on experiment data, are utilized. Through this framework, a new method for selective semiconducting CNT synthesis via catalyst - mediated electron injection, tuned by light during growth, is proposed. 54 candidate catalysts are screened, and three with high potential are identified. High-throughput experiments validate the predictions, with semiconducting selectivity exceeding 91% and the FeTiO3 catalyst reaching 98.6%. This approach not only addresses semiconducting CNT synthesis but also offers a generalizable methodology for global catalyst design and nanomaterials synthesis, advancing materials science in precise control.

cond-mat.mtrl-sci

Atomically-precise synthesis and simultaneous integration of 2D transition metal dichalcogenides enabled by nano-confinement

Two-dimensional (2D) materials, such as graphene, transition metal dichalcogenides (TMDs), and hBN, exhibit intriguing properties that are sensitive to their atomic-scale structures and can be further enriched through van der Waals (vdW) integration. However, the precise synthesis and clean integration of 2D materials remain challenging. Here, using graphene or hBN as a vdW capping layer, we create a nano-confined environment that directs the growth kinetics of 2D TMDs (e.g., NbSe2 and MoS2), enabling precise formation of TMD monolayers with tailored morphologies, from isolated monolayer domains to large-scale continuous films and intrinsically-patterned rings. Moreover, Janus S-Mo-Se monolayers are synthesized with atomic precision via vdW-protected bottom-plane chalcogen substitution. Importantly, our approach simultaneously produces ultraclean vdW interfaces. This in situ encapsulation reliably preserves air-sensitive materials, as evidenced by the enhanced superconductivity of nano-confined NbSe2 monolayers. Altogether, our study establishes a versatile platform for the controlled synthesis and integration of 2D TMDs for advanced applications.

cond-mat.mtrl-sci