SearcharxivSearch

arXiv subjects

Xujie Zhu

Publications and source records attributed to Xujie Zhu.

3 recordsLinked to original sources

WaveInst: A Frequency-Domain Enhanced Network for Fine-Grained Thin Tree Trunk Extraction in Forest Scenes

Analyzing tree morphology, particularly trunk and branch extraction, is valuable for genetic breeding and forestry management. Existing image-based deep learning methods tend to misidentify overlapping trunks as a single trunk when structural discontinuities occur due to front-back overlap, while low contrast between trunk textures and the background further complicates segmentation. Moreover, limited juvenile tree data, coupled with substantial variations in trunk diameter across growth stages, restricts model performance in extracting thin trunks and branches. Based on this, we propose an instance segmentation network leveraging frequency-domain features, WaveInst. Its core is a Frequency-domain Feature Compensation branch, consisting of Discrete Wavelet Transform block and High-Frequency Enhancement block. The former performs high- and low-frequency decomposition and aggregates high-frequency responses along multiple directions, while the latter further refines the high-frequency features through multi-path processing. An Adaptive Gated Fusion Module is then applied to effectively integrate spatial-domain convolutional features with frequency-domain representations, allowing the decoder to utilize embedded frequency-domain features to enhance its fine-grained detail representation. We conduct experiments on public datasets including SynthTree43k, CaneTree100, and UrbanStreet, as well as PoplarDataset, which contains both mature and juvenile poplar trees. On the public datasets, WaveInst demonstrates strong performance and stable robustness across diverse scenarios. On PoplarDataset, it achieves a mean average precision of 53.1 for mature and 29.1 for juvenile, outperforming existing state-of-the-art methods on juvenile by 6.6 points, demonstrating its effectiveness in extracting thin tree trunk.

cs.CV

Pluggable Pruning with Contiguous Layer Distillation for Diffusion Transformers

Diffusion Transformers (DiTs) have shown exceptional performance in image generation, yet their large parameter counts incur high computational costs, impeding deployment in resource-constrained settings. To address this, we propose Pluggable Pruning with Contiguous Layer Distillation (PPCL), a flexible structured pruning framework specifically designed for DiT architectures. First, we identify redundant layer intervals through a linear probing mechanism combined with the first-order differential trend analysis of similarity metrics. Subsequently, we propose a plug-and-play teacher-student alternating distillation scheme tailored to integrate depth-wise and width-wise pruning within a single training phase. This distillation framework enables flexible knowledge transfer across diverse pruning ratios, eliminating the need for per-configuration retraining. Extensive experiments on multiple Multi-Modal Diffusion Transformer architecture models demonstrate that PPCL achieves a 50\% reduction in parameter count compared to the full model, with less than 3\% degradation in key objective metrics. Notably, our method maintains high-quality image generation capabilities while achieving higher compression ratios, rendering it well-suited for resource-constrained environments. The open-source code, checkpoints for PPCL can be found at the following link: https://github.com/OPPO-Mente-Lab/Qwen-Image-Pruning.

cs.CV

X2Edit: Revisiting Arbitrary-Instruction Image Editing through Self-Constructed Data and Task-Aware Representation Learning

Existing open-source datasets for arbitrary-instruction image editing remain suboptimal, while a plug-and-play editing module compatible with community-prevalent generative models is notably absent. In this paper, we first introduce the X2Edit Dataset, a comprehensive dataset covering 14 diverse editing tasks, including subject-driven generation. We utilize the industry-leading unified image generation models and expert models to construct the data. Meanwhile, we design reasonable editing instructions with the VLM and implement various scoring mechanisms to filter the data. As a result, we construct 3.7 million high-quality data with balanced categories. Second, to better integrate seamlessly with community image generation models, we design task-aware MoE-LoRA training based on FLUX.1, with only 8\% of the parameters of the full model. To further improve the final performance, we utilize the internal representations of the diffusion model and define positive/negative samples based on image editing types to introduce contrastive learning. Extensive experiments demonstrate that the model's editing performance is competitive among many excellent models. Additionally, the constructed dataset exhibits substantial advantages over existing open-source datasets. The open-source code, checkpoints, and datasets for X2Edit can be found at the following link: https://github.com/OPPO-Mente-Lab/X2Edit.

cs.CV