Searcharxiv⌕ Search

arXiv subjects

Ping Huang

Publications and source records attributed to Ping Huang.

At least 37 records · Page 2Linked to original sources

VISION Datasets: A Benchmark for Vision-based InduStrial InspectiON

Despite progress in vision-based inspection algorithms, real-world industrial challenges -- specifically in data availability, quality, and complex production requirements -- often remain under-addressed. We introduce the VISION Datasets, a diverse collection of 14 industrial inspection datasets, uniquely poised to meet these challenges. Unlike previous datasets, VISION brings versatility to defect detection, offering annotation masks across all splits and catering to various detection methodologies. Our datasets also feature instance-segmentation annotation, enabling precise defect identification. With a total of 18k images encompassing 44 defect types, VISION strives to mirror a wide range of real-world production scenarios. By supporting two ongoing challenge competitions on the VISION Datasets, we hope to foster further advancements in vision-based industrial inspection.

cs.CV↗

Railway Network Delay Evolution: A Heterogeneous Graph Neural Network Approach

Railway operations involve different types of entities (stations, trains, etc.), making the existing graph/network models with homogenous nodes (i.e., the same kind of nodes) incapable of capturing the interactions between the entities. This paper aims to develop a heterogeneous graph neural network (HetGNN) model, which can address different types of nodes (i.e., heterogeneous nodes), to investigate the train delay evolution on railway networks. To this end, a graph architecture combining the HetGNN model and the GraphSAGE homogeneous GNN (HomoGNN), called SAGE-Het, is proposed. The aim is to capture the interactions between trains, trains and stations, and stations and other stations on delay evolution based on different edges. In contrast to the traditional methods that require the inputs to have constant dimensions (e.g., in rectangular or grid-like arrays) or only allow homogeneous nodes in the graph, SAGE-Het allows for flexible inputs and heterogeneous nodes. The data from two sub-networks of the China railway network are applied to test the performance and robustness of the proposed SAGE-Het model. The experimental results show that SAGE-Het exhibits better performance than the existing delay prediction methods and some advanced HetGNNs used for other prediction tasks; the predictive performances of SAGE-Het under different prediction time horizons (10/20/30 min ahead) all outperform other baseline methods; Specifically, the influences of train interactions on delay propagation are investigated based on the proposed model. The results show that train interactions become subtle when the train headways increase . This finding directly contributes to decision-making in the situation where conflict-resolution or train-canceling actions are needed.

cs.LG↗

DeSTSeg: Segmentation Guided Denoising Student-Teacher for Anomaly Detection

Visual anomaly detection, an important problem in computer vision, is usually formulated as a one-class classification and segmentation task. The student-teacher (S-T) framework has proved to be effective in solving this challenge. However, previous works based on S-T only empirically applied constraints on normal data and fused multi-level information. In this study, we propose an improved model called DeSTSeg, which integrates a pre-trained teacher network, a denoising student encoder-decoder, and a segmentation network into one framework. First, to strengthen the constraints on anomalous data, we introduce a denoising procedure that allows the student network to learn more robust representations. From synthetically corrupted normal images, we train the student network to match the teacher network feature of the same images without corruption. Second, to fuse the multi-level S-T features adaptively, we train a segmentation network with rich supervision from synthetic anomaly masks, achieving a substantial performance improvement. Experiments on the industrial inspection benchmark dataset demonstrate that our method achieves state-of-the-art performance, 98.6% on image-level AUC, 75.8% on pixel-level average precision, and 76.4% on instance-level average precision.

cs.CV↗

RGI: robust GAN-inversion for mask-free image inpainting and unsupervised pixel-wise anomaly detection

Generative adversarial networks (GANs), trained on a large-scale image dataset, can be a good approximator of the natural image manifold. GAN-inversion, using a pre-trained generator as a deep generative prior, is a promising tool for image restoration under corruptions. However, the performance of GAN-inversion can be limited by a lack of robustness to unknown gross corruptions, i.e., the restored image might easily deviate from the ground truth. In this paper, we propose a Robust GAN-inversion (RGI) method with a provable robustness guarantee to achieve image restoration under unknown \textit{gross} corruptions, where a small fraction of pixels are completely corrupted. Under mild assumptions, we show that the restored image and the identified corrupted region mask converge asymptotically to the ground truth. Moreover, we extend RGI to Relaxed-RGI (R-RGI) for generator fine-tuning to mitigate the gap between the GAN learned manifold and the true image manifold while avoiding trivial overfitting to the corrupted input image, which further improves the image restoration and corrupted region mask identification performance. The proposed RGI/R-RGI method unifies two important applications with state-of-the-art (SOTA) performance: (i) mask-free semantic inpainting, where the corruptions are unknown missing regions, the restored background can be used to restore the missing content; (ii) unsupervised pixel-wise anomaly detection, where the corruptions are unknown anomalous regions, the retrieved mask can be used as the anomalous region's segmentation mask.

cs.CV↗

PAEDID: Patch Autoencoder Based Deep Image Decomposition For Pixel-level Defective Region Segmentation

Unsupervised pixel-level defective region segmentation is an important task in image-based anomaly detection for various industrial applications. The state-of-the-art methods have their own advantages and limitations: matrix-decomposition-based methods are robust to noise but lack complex background image modeling capability; representation-based methods are good at defective region localization but lack accuracy in defective region shape contour extraction; reconstruction-based methods detected defective region match well with the ground truth defective region shape contour but are noisy. To combine the best of both worlds, we present an unsupervised patch autoencoder based deep image decomposition (PAEDID) method for defective region segmentation. In the training stage, we learn the common background as a deep image prior by a patch autoencoder (PAE) network. In the inference stage, we formulate anomaly detection as an image decomposition problem with the deep image prior and domain-specific regularizations. By adopting the proposed approach, the defective regions in the image can be accurately extracted in an unsupervised fashion. We demonstrate the effectiveness of the PAEDID method in simulation studies and an industrial dataset in the case study.

cs.CV↗

Synthetic Defect Generation for Display Front-of-Screen Quality Inspection: A Survey

Display front-of-screen (FOS) quality inspection is essential for the mass production of displays in the manufacturing process. However, the severe imbalanced data, especially the limited number of defect samples, has been a long-standing problem that hinders the successful application of deep learning algorithms. Synthetic defect data generation can help address this issue. This paper reviews the state-of-the-art synthetic data generation methods and the evaluation metrics that can potentially be applied to display FOS quality inspection tasks.

cs.LG↗

Information Gain Propagation: a new way to Graph Active Learning with Soft Labels

Graph Neural Networks (GNNs) have achieved great success in various tasks, but their performance highly relies on a large number of labeled nodes, which typically requires considerable human effort. GNN-based Active Learning (AL) methods are proposed to improve the labeling efficiency by selecting the most valuable nodes to label. Existing methods assume an oracle can correctly categorize all the selected nodes and thus just focus on the node selection. However, such an exact labeling task is costly, especially when the categorization is out of the domain of individual expert (oracle). The paper goes further, presenting a soft-label approach to AL on GNNs. Our key innovations are: i) relaxed queries where a domain expert (oracle) only judges the correctness of the predicted labels (a binary question) rather than identifying the exact class (a multi-class question), and ii) new criteria of maximizing information gain propagation for active learner with relaxed queries and soft labels. Empirical studies on public datasets demonstrate that our method significantly outperforms the state-of-the-art GNN-based AL methods in terms of both accuracy and labeling cost.

cs.LG↗

Self-supervised Semi-supervised Learning for Data Labeling and Quality Evaluation

As the adoption of deep learning techniques in industrial applications grows with increasing speed and scale, successful deployment of deep learning models often hinges on the availability, volume, and quality of annotated data. In this paper, we tackle the problems of efficient data labeling and annotation verification under the human-in-the-loop setting. We showcase that the latest advancements in the field of self-supervised visual representation learning can lead to tools and methods that benefit the curation and engineering of natural image datasets, reducing annotation cost and increasing annotation quality. We propose a unifying framework by leveraging self-supervised semi-supervised learning and use it to construct workflows for data labeling and annotation verification tasks. We demonstrate the effectiveness of our workflows over existing methodologies. On active learning task, our method achieves 97.0% Top-1 Accuracy on CIFAR10 with 0.1% annotated data, and 83.9% Top-1 Accuracy on CIFAR100 with 10% annotated data. When learning with 50% of wrong labels, our method achieves 97.4% Top-1 Accuracy on CIFAR10 and 85.5% Top-1 Accuracy on CIFAR100.

cs.CV↗

RIM: Reliable Influence-based Active Learning on Graphs

Message passing is the core of most graph models such as Graph Convolutional Network (GCN) and Label Propagation (LP), which usually require a large number of clean labeled data to smooth out the neighborhood over the graph. However, the labeling process can be tedious, costly, and error-prone in practice. In this paper, we propose to unify active learning (AL) and message passing towards minimizing labeling costs, e.g., making use of few and unreliable labels that can be obtained cheaply. We make two contributions towards that end. First, we open up a perspective by drawing a connection between AL enforcing message passing and social influence maximization, ensuring that the selected samples effectively improve the model performance. Second, we propose an extension to the influence model that incorporates an explicit quality factor to model label noise. In this way, we derive a fundamentally new AL selection criterion for GCN and LP--reliable influence maximization (RIM)--by considering quantity and quality of influence simultaneously. Empirical studies on public datasets show that RIM significantly outperforms current AL methods in terms of accuracy and efficiency.

cs.LG↗

BatchQuant: Quantized-for-all Architecture Search with Robust Quantizer

As the applications of deep learning models on edge devices increase at an accelerating pace, fast adaptation to various scenarios with varying resource constraints has become a crucial aspect of model deployment. As a result, model optimization strategies with adaptive configuration are becoming increasingly popular. While single-shot quantized neural architecture search enjoys flexibility in both model architecture and quantization policy, the combined search space comes with many challenges, including instability when training the weight-sharing supernet and difficulty in navigating the exponentially growing search space. Existing methods tend to either limit the architecture search space to a small set of options or limit the quantization policy search space to fixed precision policies. To this end, we propose BatchQuant, a robust quantizer formulation that allows fast and stable training of a compact, single-shot, mixed-precision, weight-sharing supernet. We employ BatchQuant to train a compact supernet (offering over $10^{76}$ quantized subnets) within substantially fewer GPU hours than previous methods. Our approach, Quantized-for-all (QFA), is the first to seamlessly extend one-shot weight-sharing NAS supernet to support subnets with arbitrary ultra-low bitwidth mixed-precision quantization policies without retraining. QFA opens up new possibilities in joint hardware-aware neural architecture search and quantization. We demonstrate the effectiveness of our method on ImageNet and achieve SOTA Top-1 accuracy under a low complexity constraint ($<20$ MFLOPs). The code and models will be made publicly available at https://github.com/bhpfelix/QFA.

cs.CV↗

SSNE: Effective Node Representation for Link Prediction in Sparse Networks

Graph embedding is gaining its popularity for link prediction in complex networks and achieving excellent performance. However, limited work has been done in sparse networks that represent most of real networks. In this paper, we propose a model, Sparse Structural Network Embedding (SSNE), to obtain node representation for link predication in sparse networks. The SSNE first transforms the adjacency matrix into the Sum of Normalized $H$-order Adjacency Matrix (SNHAM), and then maps the SNHAM matrix into a $d$-dimensional feature matrix for node representation via a neural network model. The mapping operation is proved to be an equivalent variation of singular value decomposition. Finally, we calculate nodal similarities for link prediction based on such feature matrix. By extensive testing experiments bases on synthetic and real sparse network, we show that the proposed method presents better link prediction performance in comparison of those of structural similarity indexes, matrix optimization and other graph embedding models.

cs.SI↗

Improved simulation of El Niño and its influence on the climate anomalies of the East Asia-western North Pacific in the ICM Version 2

This study introduces the second version of the Integrated Climate Model (ICM). ICM is developed by the Center for Monsoon System Research, Institute of Atmospheric Physics to improve the short-term climate prediction of the East Asia-western North Pacific (EA-WNP). The main update of the second version of ICM (ICM.V2) relative to the first version (ICM.V1) is the improvement of the horizontal resolution of the atmospheric model from T31 spectral resolution (3.75°*3.75°) to T63 (1.875°*1.875°). As a result, some important factors for the short-term climate prediction of the EA-WNP is apparently improved from ICM.V1 to ICM.V2, including the climatological SST, the rainfall and circulation of the East Asian summer monsoon, and the variability and spatial pattern of ENSO. The impact of El Niño on the EA-WNP climate simulated in ICM.V2 is also improved with more realistic anticyclonic anomalies and precipitation pattern over the EA-WNP. The tropical Indian ocean capacitor effect and the WNP local air-sea interaction feedback, two popular mechanisms to explain the impact of El Niño on the EA-WNP climate is also realistically reproduced in ICM.V2, much improved relative to that in ICM.V1.

physics.ao-ph↗

LUDA: Boost LSM Key Value Store Compactions with GPUs

Log-Structured-Merge (LSM) tree-based key value stores are facing critical challenges of fully leveraging the dramatic performance improvements of the underlying storage devices, which makes the compaction operations of LSM key value stores become CPU-bound, and slow compactions significantly degrade key value store performance. To address this issue, we propose LUDA, an LSM key value store with CUDA, which uses a GPU to accelerate compaction operations of LSM key value stores. How to efficiently parallelize compaction procedures as well as accommodate the optimal performance contract of the GPU architecture challenge LUDA. Specifically, LUDA overcomes these challenges by exploiting the data independence between compaction procedures and using cooperative sort mechanism and judicious data movements. Running on a commodity GPU under different levels of CPU overhead, evaluation results show that LUDA provides up to 2x higher throughput and 2x data processing speed, and achieves more stable 99th percentile latencies than LevelDB and RocksDB.

cs.DC↗

Synthesis of MAX Phases Nb2CuC and Ti2(Al0.1Cu0.9)N by A-site Replacement Reaction in Molten Salts

New MAX phases Ti2(AlxCu1-x)N and Nb2CuC were synthesized by A-site replacement by reacting Ti2AlN and Nb2AlC, respectively, with CuCl2 or CuI molten salt. X-ray diffraction, scanning electron microscopy, and atomically-resolved scanning transmission electron microscopy showed complete A-site replacement in Nb2AlC, which lead to formation of Nb2CuC. However, the replacement of Al in Ti2AlN phase was only close to complete at Ti2(Al0.1Cu0.9)N. Density-functional theory calculations corroborated the structural stability of Nb2CuC and Ti2CuN phases. Moreover, the calculated cleavage energy in these Cu-containing MAX phases are weaker than in their Al-containing counterparts, indicating that they are precursor candidates for MXene derivation.

cond-mat.mtrl-sci↗

Melting a skyrmion lattice topologically: through the hexatic phase to a skyrmion liquid

Skyrmions are twirling magnetic textures whose non-trivial topology leads to particle-like properties promising for information technology applications. Perhaps the most important aspect of interacting particles is their ability to form thermodynamically distinct phases from gases and liquids to crystalline solids. Dilute gases of skyrmions have been realized in artificial multilayers, and solid crystalline skyrmion lattices have been observed in bulk skyrmion hosting materials. Yet, to date melting of the skyrmion lattice into a skyrmion liquid has not been reported experimentally. Through direct imaging with cryo-Lorentz transmission electron microscopy, we demonstrate that the skyrmion lattice in the material Cu$_2$OSeO$_3$ can be dynamically melted. Remarkably, we discover this melting process to be a topological defects mediated two-step transition via a theoretically hypothesized hexatic phase to the liquid phase. The existence of hexatic and liquid phases instead of a simple fading of the local magnetic moments upon thermal excitations implies that even in bulk materials skyrmions possess considerable particle nature, which is a pre-requisite for application schemes.

cond-mat.str-el↗

In situ Electric Field Skyrmion Creation in Magnetoelectric Cu$_2$OSeO$_3$

Magnetic skyrmions are localized nanometric spin textures with quantized winding numbers as the topological invariant. Rapidly increasing attention has been paid to the investigations of skyrmions since their experimental discovery in 2009, due both to the fundamental properties and the promising potential in spintronics based applications. However, controlled creation of skyrmions remains a pivotal challenge towards technological applications. Here, we report that skyrmions can be created locally by electric field in the magnetoelectric helimagnet Cu$\mathsf{_2}$OSeO$\mathsf{_3}$. Using Lorentz transmission electron microscopy, we successfully write skyrmions in situ from a helical spin background. Our discovery is highly coveted since it implies that skyrmionics can be integrated into contemporary field effect transistor based electronic technology, where very low energy dissipation can be achieved, and hence realizes a large step forward to its practical applications.

cond-mat.str-el↗

Magnetic skyrmions and skyrmion clusters in the helical phase of Cu$_2$OSeO$_3$

Skyrmions are nanometric spin whirls that can be stabilized in magnets lacking inversion symmetry. The properties of isolated skyrmions embedded in a ferromagnetic background have been intensively studied. We show that single skyrmions and clusters of skyrmions can also form in the helical phase and investigate theoretically their energetics and dynamics. The helical background provides natural one-dimensional channels along which a skyrmion can move rapidly. In contrast to skyrmions in ferromagnets, the skymion-skyrmion interaction has a strong attractive component and thus skyrmions tend to form clusters with characteristic shapes. These clusters are directly observed in transmission electron microscopy measurements in thin films of Cu$_2$OSeO$_3$. Topological quantization, high mobility and the confinement of skyrmions in channels provided by the helical background may be useful for future spintronics devices.

cond-mat.str-el↗