SearcharxivSearch

arXiv subjects

Tongtong Wang

Publications and source records attributed to Tongtong Wang.

11 recordsLinked to original sources

HairCS: Reconstructing Strand-Based Hair from Hair Cards

We present an automated pipeline that converts hair-card models into high-quality strand-based hairstyles. Given a collection of textured triangular or quad strips as input, our method produces a strand-based representation that preserves the original hairstyle while enriching it with fine-scale geometric detail and adhering to standard production requirements: strands originate from the scalp, roots are uniformly distributed, and the hair volume is plausibly filled. The resulting assets are directly compatible with strand-based rendering, physics-based simulation, and common grooming modifiers (e.g., clumping, curling, noise) for enhanced realism and artistic control. We validate our approach on a large and diverse set of hairstyles, including short and long hair, curly styles, and complex styles such as buns and ponytails.

cs.GR

ADGNet: Asymmetric Dual-text Guided Network for Infrared Small Target Detection

InfRared Small Target Detection (IRSTD) is a challenging task. Relying solely on pixel-level information, vision-only methods struggle to distinguish targets from clutter. Current multimodal methods typically describe both targets and backgrounds with a single textual prompt. Such an approach lacks dedicated regional guidance and ignores infrared semantic asymmetry. Consequently, it provides insufficient background suppression information and introduces severe feature optimization conflicts, overwhelming small targets with noise. To address these issues, we propose a novel Asymmetric Dual-text Guided Network (ADGNet). Specifically, accounting for the infrared semantic asymmetry, we first design the Asymmetric Dual-text Prompt (ADP), comprising an image-agnostic abstract target prompt and an image-specific detailed background prompt. To leverage these prompts, we introduce an Asymmetric Dual-Branch Interaction (ADBI) module to separately guide visual features with their respective text priors, protecting targets from noise while suppressing background clutter. Subsequently, we introduce an Adaptive Feature Aggregation (AFA) module to dynamically fuse features from the two branches. Furthermore, we construct a multimodal Asymmetric Image-Text Infrared (AITIR) dataset by providing asymmetric text annotations for three public datasets (IRSTD-1K, NUDT-SIRST, and SIRST). Extensive experiments demonstrate that ADGNet outperforms 21 state-of-the-art (SOTA) methods. Code is available at https://github.com/iLearn-Lab/MM26-ADGNet.

cs.CV

DGNet: Dual-knowledge Guided Network for Infrared Small Target Detection

InfRared Small Target Detection (IRSTD) is a prominent and challenging task in computer vision. In recent years, text-guided methods have significantly improved detection performance. However, they still suffer from two key limitations. First, a single text description simultaneously modeling both background and target leads to semantic entanglement, which contradicts the objective of background suppression and target enhancement. Second, reliance on image-specific textual prompts (requiring additional external models such as CLIP during inference) results in deployment constraints. To address these issues, we propose a novel Dual-knowledge Guided Network (DGNet) based on multiple generalizable texts. Specifically, we design a Prior-knowledge Wavelet Modulation (PWM) module, which leverages dual textual priors that separately characterize large-scale backgrounds and sparse targets to effectively disentangle and modulate entangled semantics in the frequency domain. Furthermore, we introduce a Consensus-knowledge Directional Alignment (CDA) loss, which models the initial state and the ideal target across samples as `complex background' and `bright target', respectively, thereby constructing a clear and unified directional optimization trajectory for the model. Extensive experiments on three public datasets demonstrate the superior performance of DGNet and the effectiveness of each component. The source code is available at https://github.com/iLearn-Lab/MM26-DGNet.

cs.CV

Strand-based Hairstyle Generation via Large Reconstruction and Multimodal Models

Creating high-quality strand-based hairstyles in current production pipelines remains heavily dependent on skilled artists and time-consuming manual authoring, making it costly and difficult to scale. Existing learning-based methods have advanced image-driven hair reconstruction, but typically require large, diverse training datasets, struggle to generalize to complex styles such as buns and ponytails, and often operate in representations that are not directly compatible with strand-based modeling, editing, and simulation. We present a novel automatic pipeline that combines the capabilities of Large Reconstruction Models (LRMs), Large Multimodal Models (LMMs), and classical geometry processing to generate high-quality strand-based hairstyles from single-view images. Our approach produces detailed, production-ready strand geometry without task-specific training or data collection and can handle a wide variety of hairstyles, including straight and curly hair, short and long styles, and challenging structured configurations such as ponytails and buns. Across this diverse set of examples, our method generates visually compelling strand-level reconstructions within only a few minutes, making it well-suited for integration into modern digital human workflows.

cs.GR

Generalized Phase Diagrams for Graphene CVD growth on Copper

Understanding the competition between first-layer lateral expansion and second-layer nucleation is essential for layer-controlled graphene growth via chemical vapor deposition (CVD). Building on our previous phase diagram framework based on the dimensionless parameters $α$ and $Γ$, we develop an enhanced model incorporating two previously neglected effects: thermal-expansion-induced substrate strain and chemical desorption of carbon monomers via reverse dehydrogenation. First-principles calculations are employed to determine the strain-dependent diffusion and attachment barriers on both exposed and graphene-covered Cu(111) surfaces. By mapping the multi-step CVD process into an effective quasi-physical vapor deposition, we construct a generalized phase diagram characterized by the coupled effects of $α$, $Γ$, and a newly introduced desorption parameter $Z$. Our results show that tensile strain expands the bilayer graphene (BLG) growth window for critical nucleus sizes $i^*>1$. In contrast, chemical desorption suppresses BLG formation in the high-$Γ$ regime via $Z$-dependent monomer depletion. This unified framework provides a predictive guide for the rational synthesis of high-quality bilayer graphene by linking macroscopic growth parameters to microscopic layer-selection mechanisms.

cond-mat.mtrl-sci

Real-time Neural Six-way Lightmaps

Participating media are a pervasive and intriguing visual effect in virtual environments. Unfortunately, rendering such phenomena in real-time is notoriously difficult due to the computational expense of estimating the volume rendering equation. While the six-way lightmaps technique has been widely used in video games to render smoke with a camera-oriented billboard and approximate lighting effects using six precomputed lightmaps, achieving a balance between realism and efficiency, it is limited to pre-simulated animation sequences and is ignorant of camera movement. In this work, we propose a neural six-way lightmaps method to strike a long-sought balance between dynamics and visual realism. Our approach first generates a guiding map from the camera view using ray marching with a large sampling distance to approximate smoke scattering and silhouette. Then, given a guiding map, we train a neural network to predict the corresponding six-way lightmaps. The resulting lightmaps can be seamlessly used in existing game engine pipelines. This approach supports visually appealing rendering effects while enabling real-time user interactivity, including smoke-obstacle interaction, camera movement, and light change. By conducting a series of comprehensive benchmarks, we demonstrate that our method is well-suited for real-time applications, such as games and VR/AR.

cs.GR

Anisotropic Kinetics of Ion-Irradiation-Induced Phase Transition in Gallium Oxide

Radiation-tolerant semiconductors have traditionally been engineered by the principle of suppressing defect accumulation and amorphization, based on the assumption that radiation damage is inherently stochastic. Here we show that, in monoclinic $β$-\ce{Ga2O3}, a promising ultrawide-bandgap semiconductor, surface crystallographic orientation deterministically governs radiation tolerance through highly anisotropic kinetics of the $β$-to-$γ$ phase transition. Using machine-learning molecular dynamics coupled with a local configurational-entropy descriptor, we quantitatively map anisotropic $β$-to-$γ$ transition kinetics, showing that the critical dose, transition-layer depth, and kinetic stability of the $γ$-phase are fundamentally governed by surface orientation. Under ion irradiation, non-channeling surfaces such as (100), (001), and (-201) undergo severe surface amorphization, whereas the strongly channeling (010) surface resists damage accumulation and promotes subsurface $γ$-phase nucleation. During thermal annealing recovery process, these initial states follow two distinct recovery pathways: the channeling (010) surface reverts directly from $γ$-to-$β$, whereas non-channeling surfaces follow a sequential amorphous-to-$γ$-to-$β$ transition pathway. This work establishes surface orientation as a fundamental design principle for achieving radiation tolerance through controlled polymorphic transitions, providing a universal framework for engineering functional materials capable of withstanding extreme irradiation environments.

cond-mat.mtrl-sci

Out of Distribution Detection in Self-adaptive Robots with AI-powered Digital Twins

Self-adaptive robots (SARs) in complex, uncertain environments must proactively detect and address abnormal behaviors, including out-of-distribution (OOD) cases. To this end, digital twins offer a valuable solution for OOD detection. Thus, we present a digital twin-based approach for OOD detection (ODiSAR) in SARs. ODiSAR uses a Transformer-based digital twin to forecast SAR states and employs reconstruction error and Monte Carlo dropout for uncertainty quantification. By combining reconstruction error with predictive variance, the digital twin effectively detects OOD behaviors, even in previously unseen conditions. The digital twin also includes an explainability layer that links potential OOD to specific SAR states, offering insights for self-adaptation. We evaluated ODiSAR by creating digital twins of two industrial robots: one navigating an office environment, and another performing maritime ship navigation. In both cases, ODiSAR forecasts SAR behaviors (i.e., robot trajectories and vessel motion) and proactively detects OOD events. Our results showed that ODiSAR achieved high detection performance -- up to 98\% AUROC, 96\% TNR@TPR95, and 95\% F1-score -- while providing interpretable insights to support self-adaptation.

cs.RO

Auto Hair Card Extraction for Smooth Hair with Differentiable Rendering

Hair cards remain a widely used representation for hair modeling in real-time applications, offering a practical trade-off between visual fidelity, memory usage, and performance. However, generating high-quality hair card models remains a challenging and labor-intensive task. This work presents an automated pipeline for converting strand-based hair models into hair card models with a limited number of cards and textures while preserving the hairstyle appearance. Our key idea is a novel differentiable representation where each strand is encoded as a projected 2D spline in the texture space, which enables efficient optimization with differentiable rendering and structured results respecting the hair geometry. Based on this representation, we develop a novel algorithm pipeline, where we first cluster hair strands into initial hair cards and project the strands into the texture space. We then conduct a two-stage optimization where our first stage optimizes the texture and geometry of each hair card separately, and after texture reduction, our second stage conducts joint optimization of all the cards for fine-tuning. Put together, our method is evaluated on a wide range of hairstyles, including straight, wavy, curly, and coily hairs. To better capture the appearance of short or coily hair, we additionally support hair cap and cross-card. Furthermore, our framework supports seamless LoD transitions via texture sharing, balancing texture memory efficiency and visual quality.

cs.GR

Phase Diagram of growth modes in Graphene Growth on Cooper by Vapor Deposition

Understanding the atomistic mechanism in graphene growth is crucial for controlling the number of layers or domain sizes to meet practical needs. In this work, focusing on the growth of graphene by chemical vapor deposition on copper substrates, the surface kinetics in the growth are systematically investigated by first-principles calculations. The phase diagram, predicting whether the growth mode is monolayer graphene or bilayer graphene under various experimental conditions, is constructed based on classical nucleation theory. Our phase diagram well illustrates the effect of high hydrogen pressure on bilayer graphene growth and clarifies the mechanism of the most widely used experimental growth approaches. The phase diagram can provide guidance and predictions for experiments and inspires the study of other two-dimensional materials with graphene-like growth mechanisms.

cond-mat.mtrl-sci

CAT-DM: Controllable Accelerated Virtual Try-on with Diffusion Model

Generative Adversarial Networks (GANs) dominate the research field in image-based virtual try-on, but have not resolved problems such as unnatural deformation of garments and the blurry generation quality. While the generative quality of diffusion models is impressive, achieving controllability poses a significant challenge when applying it to virtual try-on and multiple denoising iterations limit its potential for real-time applications. In this paper, we propose Controllable Accelerated virtual Try-on with Diffusion Model (CAT-DM). To enhance the controllability, a basic diffusion-based virtual try-on network is designed, which utilizes ControlNet to introduce additional control conditions and improves the feature extraction of garment images. In terms of acceleration, CAT-DM initiates a reverse denoising process with an implicit distribution generated by a pre-trained GAN-based model. Compared with previous try-on methods based on diffusion models, CAT-DM not only retains the pattern and texture details of the inshop garment but also reduces the sampling steps without compromising generation quality. Extensive experiments demonstrate the superiority of CAT-DM against both GANbased and diffusion-based methods in producing more realistic images and accurately reproducing garment patterns.

cs.CV