SearcharxivSearch

arXiv subjects

Xiaomin Wang

Publications and source records attributed to Xiaomin Wang.

At least 19 recordsLinked to original sources

High-efficiency integrated laser on erbium-doped lithium niobate-on-insulator

Lithium niobate on insulator (LNOI) combines the outstanding optical properties of lithium niobate (LN) with strong optical confinement, scalable fabrication and high-density integration, making it a leading platform for integrated photonic chips. Recent advances in LNOI photonics have mainly centred on passive and electro-optic components, including couplers, waveguides, microcavities and modulators, whereas efficient on-chip laser sources remain insufficiently developed, limiting the realization of fully integrated LN photonic systems. Because LN is an indirect-bandgap material, lasing on LNOI generally relies on photoluminescence from rare-earth-ion doping, yet the conversion efficiency of doped LNOI lasers has remained low. By comparing LNOI microcavity lasers with fibre lasers and waveguide amplifiers, we identify the limited number of rare-earth ions participating in stimulated emission as a key factor responsible for inefficient pump utilization. Here we demonstrate an integrated Er-doped LNOI laser that combines high-quality, highly Er-doped LN, a large-diameter wide-microring resonator, a low-loss waveguide amplifier and bidirectional pumping. This architecture enables a slope efficiency of 16.91% at 1562 nm, exceeding 10% on the LNOI platform for the first time. Our results provide a route towards high-efficiency LNOI lasers for fully integrated photonic systems.

physics.optics

One- and two-dimensional cluster states for topological phase simulation and measurement-based quantum computation

Quantum entanglement is a fundamental resource for quantum information processing and serves as a critical benchmark for quantum hardware performance. Cluster states are a special class of entangled states that serve as universal resources for measurement-based quantum computation and possess an intrinsic symmetry-protected topological order, which confers robustness against symmetry-respecting noise. Here we report the scalable preparation and verification of genuine multipartite cluster states on the 105-qubit Zuchongzhi 3.1 superconducting processor. We achieve one-dimensional cluster states of up to 95 qubits and two-dimensional cluster states of up to 72 qubits. The symmetry-protected topological cluster states exhibit input-state-dependent robustness under symmetry-breaking perturbations due to an operational parity structure that enhances the performance of measurement-based quantum computation. Furthermore, we use our two-dimensional cluster states to implement the Deutsch-Jozsa algorithm within the measurement-based quantum computation framework, achieving higher output-state fidelity compared with traditional circuit-based models and a query efficiency advantage over classical approaches. Our work establishes a scalable platform that combines large-scale entanglement generation, symmetry-protected topological order and practical quantum algorithms to enable robust, fault-tolerant measurement-based quantum computation.

quant-ph

A Transformer-Based Conditional GAN with Multiple Instance Learning for UAV Signal Detection and Classification

Unmanned Aerial Vehicles (UAVs) are increasingly used in surveillance, logistics, agriculture, disaster management, and military operations. Accurate detection and classification of UAV flight states, such as hovering, cruising, ascending, or transitioning, which are essential for safe and effective operations. However, conventional time series classification (TSC) methods often lack robustness and generalization for dynamic UAV environments, while state of the art(SOTA) models like Transformers and LSTM based architectures typically require large datasets and entail high computational costs, especially with high-dimensional data streams. This paper proposes a novel framework that integrates a Transformer-based Generative Adversarial Network (GAN) with Multiple Instance Locally Explainable Learning (MILET) to address these challenges in UAV flight state classification. The Transformer encoder captures long-range temporal dependencies and complex telemetry dynamics, while the GAN module augments limited datasets with realistic synthetic samples. MIL is incorporated to focus attention on the most discriminative input segments, reducing noise and computational overhead. Experimental results show that the proposed method achieves superior accuracy 96.5% on the DroneDetect dataset and 98.6% on the DroneRF dataset that outperforming other SOTA approaches. The framework also demonstrates strong computational efficiency and robust generalization across diverse UAV platforms and flight states, highlighting its potential for real-time deployment in resource constrained environments.

cs.LG

Self-Powered, Ultra-thin, Flexible, and Scalable Ultraviolet Detector Utilizing Diamond-MoS$_2$ Heterojunction

The escalating demand for ultraviolet (UV) sensing in space exploration, environmental monitoring, and agricultural productivity necessitates detectors that are both environmentally and mechanically resilient. Diamond, featuring its high bandgap and UV absorption, exceptional mechanical/chemical robustness, and excellent thermal stability, emerges as a highly promising material for next-generation UV detection in various scenarios. However, conventional diamond-based UV detectors are constrained by rigid bulk architectures and reliance on external power supplies, hindering their integration with curved and flexible platforms and complicating device scalability due to auxiliary power requirements. To tackle these challenges, herein, we firstly demonstrated a large-scale, self-powered, and flexible diamond UV detector by heterogeneously integrating a MoS$_2$ monolayer with an ultrathin, freestanding diamond membrane. The fabricated device operates at zero external bias, and simultaneously exhibits a high responsivity of 94 mA W$^{-1}$ at 220 nm, and detectivity of 5.88 x 109 Jones. Notably, mechanical bending enables strain-induced bandgap modulation of the diamond membrane, allowing dynamically tunable photoresponse-a capability absent in rigid diamond counterparts. To validate its practicality and scalability, a proof-of-concept UV imager with 3x3 pixels was demonstrated. This newly developed configuration will undoubtedly open up new routes toward scalable, integrable, flexible, and cost-effective UV sensing solutions for emerging technologies

physics.ins-det

Boundary-enhanced time series data imputation with long-term dependency diffusion models

Data imputation is crucial for addressing challenges posed by missing values in multivariate time series data across various fields, such as healthcare, traffic, and economics, and has garnered significant attention. Among various methods, diffusion model-based approaches show notable performance improvements. However, existing methods often cause disharmonious boundaries between missing and known regions and overlook long-range dependencies in missing data estimation, leading to suboptimal results. To address these issues, we propose a Diffusion-based time Series Data Imputation (DSDI) framework. We develop a weight-reducing injection strategy that incorporates the predicted values of missing points with reducing weights into the reverse diffusion process to mitigate boundary inconsistencies. Further, we introduce a multi-scale S4-based U-Net, which combines hierarchical information from different levels via multi-resolution integration to capture long-term dependencies. Experimental results demonstrate that our model outperforms existing imputation methods.

cs.LG

Temporally Resolution Decrement: Utilizing the Shape Consistency for Higher Computational Efficiency

Image resolution that has close relations with accuracy and computational cost plays a pivotal role in network training. In this paper, we observe that the reduced image retains relatively complete shape semantics but loses extensive texture information. Inspired by the consistency of the shape semantics as well as the fragility of the texture information, we propose a novel training strategy named Temporally Resolution Decrement. Wherein, we randomly reduce the training images to a smaller resolution in the time domain. During the alternate training with the reduced images and the original images, the unstable texture information in the images results in a weaker correlation between the texture-related patterns and the correct label, naturally enforcing the model to rely more on shape properties that are robust and conform to the human decision rule. Surprisingly, our approach greatly improves both the training and inference efficiency of convolutional neural networks. On ImageNet classification, using only 33\% calculation quantity (randomly reducing the training image to 112$\times$112 within 90\% epochs) can still improve ResNet-50 from 76.32\% to 77.71\%. Superimposed with the strong training procedure of ResNet-50 on ImageNet, our method achieves 80.42\% top-1 accuracy with saving 37.5\% calculation overhead. To the best of our knowledge this is the highest ImageNet single-crop accuracy on ResNet-50 under 224$\times$224 without extra data or distillation.

cs.CV

White Paper Assistance: A Step Forward Beyond the Shortcut Learning

The promising performances of CNNs often overshadow the need to examine whether they are doing in the way we are actually interested. We show through experiments that even over-parameterized models would still solve a dataset by recklessly leveraging spurious correlations, or so-called 'shortcuts'. To combat with this unintended propensity, we borrow the idea of printer test page and propose a novel approach called White Paper Assistance. Our proposed method involves the white paper to detect the extent to which the model has preference for certain characterized patterns and alleviates it by forcing the model to make a random guess on the white paper. We show the consistent accuracy improvements that are manifest in various architectures, datasets and combinations with other techniques. Experiments have also demonstrated the versatility of our approach on fine-grained recognition, imbalanced classification and robustness to corruptions.

cs.CV

Selective Output Smoothing Regularization: Regularize Neural Networks by Softening Output Distributions

In this paper, we propose Selective Output Smoothing Regularization, a novel regularization method for training the Convolutional Neural Networks (CNNs). Inspired by the diverse effects on training from different samples, Selective Output Smoothing Regularization improves the performance by encouraging the model to produce equal logits on incorrect classes when dealing with samples that the model classifies correctly and over-confidently. This plug-and-play regularization method can be conveniently incorporated into almost any CNN-based project without extra hassle. Extensive experiments have shown that Selective Output Smoothing Regularization consistently achieves significant improvement in image classification benchmarks, such as CIFAR-100, Tiny ImageNet, ImageNet, and CUB-200-2011. Particularly, our method obtains 77.30% accuracy on ImageNet with ResNet-50, which gains 1.1% than baseline (76.2%). We also empirically demonstrate the ability of our method to make further improvements when combining with other widely used regularization techniques. On Pascal detection, using the SOSR-trained ImageNet classifier as the pretrained model leads to better detection performances.

cs.CV

Channel Self-Supervision for Online Knowledge Distillation

Recently, researchers have shown an increased interest in the online knowledge distillation. Adopting an one-stage and end-to-end training fashion, online knowledge distillation uses aggregated intermediated predictions of multiple peer models for training. However, the absence of a powerful teacher model may result in the homogeneity problem between group peers, affecting the effectiveness of group distillation adversely. In this paper, we propose a novel online knowledge distillation method, \textbf{C}hannel \textbf{S}elf-\textbf{S}upervision for Online Knowledge Distillation (CSS), which structures diversity in terms of input, target, and network to alleviate the homogenization problem. Specifically, we construct a dual-network multi-branch structure and enhance inter-branch diversity through self-supervised learning, adopting the feature-level transformation and augmenting the corresponding labels. Meanwhile, the dual network structure has a larger space of independent parameters to resist the homogenization problem during distillation. Extensive quantitative experiments on CIFAR-100 illustrate that our method provides greater diversity than OKDDip and we also give pretty performance improvement, even over the state-of-the-art such as PCL. The results on three fine-grained datasets (StanfordDogs, StanfordCars, CUB-200-211) also show the significant generalization capability of our approach.

cs.CV

Cut-Thumbnail: A Novel Data Augmentation for Convolutional Neural Network

In this paper, we propose a novel data augmentation strategy named Cut-Thumbnail, that aims to improve the shape bias of the network. We reduce an image to a certain size and replace the random region of the original image with the reduced image. The generated image not only retains most of the original image information but also has global information in the reduced image. We call the reduced image as thumbnail. Furthermore, we find that the idea of thumbnail can be perfectly integrated with Mixed Sample Data Augmentation, so we put one image's thumbnail on another image while the ground truth labels are also mixed, making great achievements on various computer vision tasks. Extensive experiments show that Cut-Thumbnail works better than state-of-the-art augmentation strategies across classification, fine-grained image classification, and object detection. On ImageNet classification, ResNet-50 architecture with our method achieves 79.21\% accuracy, which is more than 2.8\% improvement on the baseline.

cs.CV

Feature Mining: A Novel Training Strategy for Convolutional Neural Network

In this paper, we propose a novel training strategy for convolutional neural network(CNN) named Feature Mining, that aims to strengthen the network's learning of the local feature. Through experiments, we find that semantic contained in different parts of the feature is different, while the network will inevitably lose the local information during feedforward propagation. In order to enhance the learning of local feature, Feature Mining divides the complete feature into two complementary parts and reuse these divided feature to make the network learn more local information, we call the two steps as feature segmentation and feature reusing. Feature Mining is a parameter-free method and has plug-and-play nature, and can be applied to any CNN models. Extensive experiments demonstrate the wide applicability, versatility, and compatibility of our method.

cs.CV

Go Small and Similar: A Simple Output Decay Brings Better Performance

Regularization and data augmentation methods have been widely used and become increasingly indispensable in deep learning training. Researchers who devote themselves to this have considered various possibilities. But so far, there has been little discussion about regularizing outputs of the model. This paper begins with empirical observations that better performances are significantly associated with output distributions, that have smaller average values and variances. By audaciously assuming there is causality involved, we propose a novel regularization term, called Output Decay, that enforces the model to assign smaller and similar output values on each class. Though being counter-intuitive, such a small modification result in a remarkable improvement on performance. Extensive experiments demonstrate the wide applicability, versatility, and compatibility of Output Decay.

cs.CV

Self-supervised Feature Enhancement: Applying Internal Pretext Task to Supervised Learning

Traditional self-supervised learning requires CNNs using external pretext tasks (i.e., image- or video-based tasks) to encode high-level semantic visual representations. In this paper, we show that feature transformations within CNNs can also be regarded as supervisory signals to construct the self-supervised task, called \emph{internal pretext task}. And such a task can be applied for the enhancement of supervised learning. Specifically, we first transform the internal feature maps by discarding different channels, and then define an additional internal pretext task to identify the discarded channels. CNNs are trained to predict the joint labels generated by the combination of self-supervised labels and original labels. By doing so, we let CNNs know which channels are missing while classifying in the hope to mine richer feature information. Extensive experiments show that our approach is effective on various models and datasets. And it's worth noting that we only incur negligible computational overhead. Furthermore, our approach can also be compatible with other methods to get better results.

cs.CV

Self-supervision of Feature Transformation for Further Improving Supervised Learning

Self-supervised learning, which benefits from automatically constructing labels through pre-designed pretext task, has recently been applied for strengthen supervised learning. Since previous self-supervised pretext tasks are based on input, they may incur huge additional training overhead. In this paper we find that features in CNNs can be also used for self-supervision. Thus we creatively design the \emph{feature-based pretext task} which requires only a small amount of additional training overhead. In our task we discard different particular regions of features, and then train the model to distinguish these different features. In order to fully apply our feature-based pretext task in supervised learning, we also propose a novel learning framework containing multi-classifiers for further improvement. Original labels will be expanded to joint labels via self-supervision of feature transformations. With more semantic information provided by our self-supervised tasks, this approach can train CNNs more effectively. Extensive experiments on various supervised learning tasks demonstrate the accuracy improvement and wide applicability of our method.

cs.CV

FocusedDropout for Convolutional Neural Network

In convolutional neural network (CNN), dropout cannot work well because dropped information is not entirely obscured in convolutional layers where features are correlated spatially. Except randomly discarding regions or channels, many approaches try to overcome this defect by dropping influential units. In this paper, we propose a non-random dropout method named FocusedDropout, aiming to make the network focus more on the target. In FocusedDropout, we use a simple but effective way to search for the target-related features, retain these features and discard others, which is contrary to the existing methods. We found that this novel method can improve network performance by making the network more target-focused. Besides, increasing the weight decay while using FocusedDropout can avoid the overfitting and increase accuracy. Experimental results show that even a slight cost, 10\% of batches employing FocusedDropout, can produce a nice performance boost over the baselines on multiple datasets of classification, including CIFAR10, CIFAR100, Tiny Imagenet, and has a good versatility for different CNN models.

cs.CV

Exact mean first-passage time on generalized Vicsek fractal

Fractal phenomena may be widely observed in a great number of complex systems. In this paper, we revisit the well-known Vicsek fractal, and study some of its structural properties for purpose of understanding how the underlying topology influences its dynamic behaviors. For instance, we analytically determine the exact solution to mean first-passage time for random walks on Vicsek fractal in a more light mapping-based manner than previous other methods, including typical spectral technique. More importantly, our method can be quite efficient to precisely calculate the solutions to mean first-passage time on all generalized versions of Vicsek fractal generated based on an arbitrary allowed seed, while other previous methods suitable for typical Vicsek fractal will become prohibitively complicated and even fail. Lastly, this analytic results suggest that the scaling relation between mean first-passage time and vertex number in generalized versions of Vicsek fractal keeps unchanged in the large graph size limit no matter what seed is selected.

math.PR

Two Cumulative Distributions For Scale-freeness of Dynamic Networks

It is well-known that the scale-free networks are ubiquitous in nature and society and have been one of the hotspot topic in complex networks. Recently, scholars presented a large quantity of scale-free networks by calculating cumulative distribution. The purpose of this paper is to discuss the relationship between two cumulative distributions, namely, cumulative distribution, edge-cumulative distribution. Here, firstly, we introduce a relationship between degree distribution and cumulative distribution. Secondly, we introduce the definition of cumulative distribution and edge-cumulative distribution, and compare the relationship between them. Thirdly, we apply algorithmic techniques to construct three deterministic networks, calculate their cumulative distribution and edge-cumulative distribution, and analyze the relationship between cumulative distribution and edge-cumulative distribution. Finally, we offer some open problems for future research in order to understand the interplay between the degree distribution, cumulative distribution and edge-cumulative distribution.

cs.SI

Constructions and properties of a class of random scale-free networks

Complex networks have abundant and extensive applications in real life. Recently, researchers have proposed a number of complex networks, in which some are deterministic and others are random. Compared with deterministic networks, random network is not only interesting and typical but also practical to illustrate and study many real-world complex networks, especially for random scale-free networks. Here, we introduce three types of operations, i.e., type-A operation, type-B operation and type-C operation, for generating random scale-free networks $N(p,q,r,t)$. On the basis of our operations, we put forward the concrete process of producing networks, which constitute the network space $\mathcal{N}(p,q,r,t)$, and then discuss their topological properties. Firstly, we calculate the range of the average degree of each member in our network space and discover that each member is a sparse network. Secondly, we prove that each member in our space obeys the power-law distribution with degree exponent $γ=1+\frac{\ln(4-r)}{\ln2}$, which implies that each member is scale-free. Next, we analyze the diameter, and find that the diameter may abruptly transform from small to large due to type-B operation. Afterwards, we study the clustering coefficient of network and discover that its value is only determined by type-C operation. Ultimately, we make an elaborate conclusion. \\ \textbf{Keywords:} Random network; degree distribution; diameter; clustering coefficient.

physics.soc-ph