Searcharxiv⌕ Search

arXiv subjects

Hanwei Zhang

Publications and source records attributed to Hanwei Zhang.

32 records · Page 2Linked to original sources

A Learning Paradigm for Interpretable Gradients

This paper studies interpretability of convolutional networks by means of saliency maps. Most approaches based on Class Activation Maps (CAM) combine information from fully connected layers and gradient through variants of backpropagation. However, it is well understood that gradients are noisy and alternatives like guided backpropagation have been proposed to obtain better visualization at inference. In this work, we present a novel training approach to improve the quality of gradients for interpretability. In particular, we introduce a regularization loss such that the gradient with respect to the input image obtained by standard backpropagation is similar to the gradient obtained by guided backpropagation. We find that the resulting gradient is qualitatively less noisy and improves quantitatively the interpretability properties of different networks, using several interpretability methods.

cs.CV↗

DP-Net: Learning Discriminative Parts for image recognition

This paper presents Discriminative Part Network (DP-Net), a deep architecture with strong interpretation capabilities, which exploits a pretrained Convolutional Neural Network (CNN) combined with a part-based recognition module. This system learns and detects parts in the images that are discriminative among categories, without the need for fine-tuning the CNN, making it more scalable than other part-based models. While part-based approaches naturally offer interpretable representations, we propose explanations at image and category levels and introduce specific constraints on the part learning process to make them more discrimative.

cs.CV↗

Opti-CAM: Optimizing saliency maps for interpretability

Methods based on class activation maps (CAM) provide a simple mechanism to interpret predictions of convolutional neural networks by using linear combinations of feature maps as saliency maps. By contrast, masking-based methods optimize a saliency map directly in the image space or learn it by training another network on additional data. In this work we introduce Opti-CAM, combining ideas from CAM-based and masking-based approaches. Our saliency map is a linear combination of feature maps, where weights are optimized per image such that the logit of the masked image for a given class is maximized. We also fix a fundamental flaw in two of the most common evaluation metrics of attribution methods. On several datasets, Opti-CAM largely outperforms other CAM-based approaches according to the most relevant classification metrics. We provide empirical evidence supporting that localization and classifier interpretability are not necessarily aligned.

cs.CV↗

Towards Good Practices in Evaluating Transfer Adversarial Attacks

Transfer adversarial attacks raise critical security concerns in real-world, black-box scenarios. However, the actual progress of this field is difficult to assess due to two common limitations in existing evaluations. First, different methods are often not systematically and fairly evaluated in a one-to-one comparison. Second, only transferability is evaluated but another key attack property, stealthiness, is largely overlooked. In this work, we design good practices to address these limitations, and we present the first comprehensive evaluation of transfer attacks, covering 23 representative attacks against 9 defenses on ImageNet. In particular, we propose to categorize existing attacks into five categories, which enables our systematic category-wise analyses. These analyses lead to new findings that even challenge existing knowledge and also help determine the optimal attack hyperparameters for our attack-wise comprehensive evaluation. We also pay particular attention to stealthiness, by adopting diverse imperceptibility metrics and looking into new, finer-grained characteristics. Overall, our new insights into transferability and stealthiness lead to actionable good practices for future evaluations.

cs.CR↗

A Stochastic Online Forecast-and-Optimize Framework for Real-Time Energy Dispatch in Virtual Power Plants under Uncertainty

Aggregating distributed energy resources in power systems significantly increases uncertainties, in particular caused by the fluctuation of renewable energy generation. This issue has driven the necessity of widely exploiting advanced predictive control techniques under uncertainty to ensure long-term economics and decarbonization. In this paper, we propose a real-time uncertainty-aware energy dispatch framework, which is composed of two key elements: (i) A hybrid forecast-and-optimize sequential task, integrating deep learning-based forecasting and stochastic optimization, where these two stages are connected by the uncertainty estimation at multiple temporal resolutions; (ii) An efficient online data augmentation scheme, jointly involving model pre-training and online fine-tuning stages. In this way, the proposed framework is capable to rapidly adapt to the real-time data distribution, as well as to target on uncertainties caused by data drift, model discrepancy and environment perturbations in the control process, and finally to realize an optimal and robust dispatch solution. The proposed framework won the championship in CityLearn Challenge 2022, which provided an influential opportunity to investigate the potential of AI application in the energy domain. In addition, comprehensive experiments are conducted to interpret its effectiveness in the real-life scenario of smart building energy management.

eess.SY↗

MOTSLAM: MOT-assisted monocular dynamic SLAM using single-view depth estimation

Visual SLAM systems targeting static scenes have been developed with satisfactory accuracy and robustness. Dynamic 3D object tracking has then become a significant capability in visual SLAM with the requirement of understanding dynamic surroundings in various scenarios including autonomous driving, augmented and virtual reality. However, performing dynamic SLAM solely with monocular images remains a challenging problem due to the difficulty of associating dynamic features and estimating their positions. In this paper, we present MOTSLAM, a dynamic visual SLAM system with the monocular configuration that tracks both poses and bounding boxes of dynamic objects. MOTSLAM first performs multiple object tracking (MOT) with associated both 2D and 3D bounding box detection to create initial 3D objects. Then, neural-network-based monocular depth estimation is applied to fetch the depth of dynamic features. Finally, camera poses, object poses, and both static, as well as dynamic map points, are jointly optimized using a novel bundle adjustment. Our experiments on the KITTI dataset demonstrate that our system has reached best performance on both camera ego-motion and object tracking on monocular dynamic SLAM.

cs.CV↗

KDD CUP 2022 Wind Power Forecasting Team 88VIP Solution

KDD CUP 2022 proposes a time-series forecasting task on spatial dynamic wind power dataset, in which the participants are required to predict the future generation given the historical context factors. The evaluation metrics contain RMSE and MAE. This paper describes the solution of Team 88VIP, which mainly comprises two types of models: a gradient boosting decision tree to memorize the basic data patterns and a recurrent neural network to capture the deep and latent probabilistic transitions. Ensembling these models contributes to tackle the fluctuation of wind power, and training submodels targets on the distinguished properties in heterogeneous timescales of forecasting, from minutes to days. In addition, feature engineering, imputation techniques and the design of offline evaluation are also described in details. The proposed solution achieves an overall online score of -45.213 in Phase 3.

cs.LG↗

Ensemble Defense with Data Diversity: Weak Correlation Implies Strong Robustness

In this paper, we propose a framework of filter-based ensemble of deep neuralnetworks (DNNs) to defend against adversarial attacks. The framework builds an ensemble of sub-models -- DNNs with differentiated preprocessing filters. From the theoretical perspective of DNN robustness, we argue that under the assumption of high quality of the filters, the weaker the correlations of the sensitivity of the filters are, the more robust the ensemble model tends to be, and this is corroborated by the experiments of transfer-based attacks. Correspondingly, we propose a principle that chooses the specific filters with smaller Pearson correlation coefficients, which ensures the diversity of the inputs received by DNNs, as well as the effectiveness of the entire framework against attacks. Our ensemble models are more robust than those constructed by previous defense methods like adversarial training, and even competitive with the classical ensemble of adversarial trained DNNs under adversarial attacks when the attacking radius is large.

cs.LG↗

Walking on the Edge: Fast, Low-Distortion Adversarial Examples

Adversarial examples of deep neural networks are receiving ever increasing attention because they help in understanding and reducing the sensitivity to their input. This is natural given the increasing applications of deep neural networks in our everyday lives. When white-box attacks are almost always successful, it is typically only the distortion of the perturbations that matters in their evaluation. In this work, we argue that speed is important as well, especially when considering that fast attacks are required by adversarial training. Given more time, iterative methods can always find better solutions. We investigate this speed-distortion trade-off in some depth and introduce a new attack called boundary projection (BP) that improves upon existing methods by a large margin. Our key idea is that the classification boundary is a manifold in the image space: we therefore quickly reach the boundary and then optimize distortion on this manifold.

cs.CV↗

Smooth Adversarial Examples

This paper investigates the visual quality of the adversarial examples. Recent papers propose to smooth the perturbations to get rid of high frequency artefacts. In this work, smoothing has a different meaning as it perceptually shapes the perturbation according to the visual content of the image to be attacked. The perturbation becomes locally smooth on the flat areas of the input image, but it may be noisy on its textured areas and sharp across its edges. This operation relies on Laplacian smoothing, well-known in graph signal processing, which we integrate in the attack pipeline. We benchmark several attacks with and without smoothing under a white-box scenario and evaluate their transferability. Despite the additional constraint of smoothness, our attack has the same probability of success at lower distortion.

cs.CV↗

Optical rogue wave in random distributed feedback fiber laser

The famous demonstration of optical rogue wave (RW)-rarely and unexpectedly event with extremely high intensity-had opened a flourishing time for temporal statistic investigation as a powerful tool to reveal the fundamental physics in different laser scenarios. However, up to now, optical RW behavior with temporally localized structure has yet not been presented in random fiber laser (RFL) characterized with mirrorless open cavity, whose feedback arises from distinctive distributed multiple scattering. Here, thanks to the participation of sustained and crucial stimulated Brillouin scattering (SBS) process, experimental explorations of optical RW are done in the highly-skewed transient intensity of an incoherently-pumped standard-telecom-fiber-constructed RFL. Furthermore, threshold-like beating peak behavior can also been resolved in the radiofrequency spectroscopy. Bringing the concept of optical RW to RFL domain without fixed cavity may greatly extend our comprehension of the rich and complex kinetics such as photon propagation and localization in disordered amplifying media with multiple scattering.

physics.optics↗

Spectral adjustment of a powerful random fiber laser

Random fiber laser (RFL) based on random distributed feedback and Raman gain has earned much attention in recent years. In this presentation, we demonstrate a powerful linearly polarized RFL with spectral adjustment, in which the central wavelength and the linewidth of the spectrum can be tuned independently through a bandwidth-adjustable tunable optical filter (BA-TOF). As a result, the central wavelength can be continuously tuned from 1095 to 1115 nm, while the full width at half-maximum (FWHM) linewidth has a maximal tuning range from ~0.6 to more than 2 nm. This laser provides a flexible tool for scenes where the temporal coherence property accounts, such as coherent sensing/communication and nonlinear frequency conversion. To the best of our knowledge, this is the first demonstration of a high power linearly polarized RFL with both wavelength and linewidth tunability.

physics.optics↗

Self-started stable pulsing operation of random fiber laser

Unlike traditional fiber laser with defined resonant cavity, random fiber laser (RFL), whose operation is based on distributed gain and feedback via Rayleigh scattering and stimulated Raman scattering in long passive fiber, has fundamental scientific challenges in pulsing operation for its remarkable cavity-free feature. Here, we propose and experimentally realize the passively spatiotemporal gain modulation induced self-started stable pulsing operation of counter-pumped RFL. Thanks to the good temporal stability of employed pumping amplified spontaneous emission source and the superiority of this pulse generation scheme, stable and regular pulse train can be obtained. Furthermore, the pump hysteresis and bistability phenomena with the generation of high order Stokes light is presented and the dynamics of pulsing operation is discussed. This work extends our comprehension of temporal property of RFL and provides an effective novel avenue for the exploration of pulsed RFL with structural simplicity, low cost and stable output.

physics.optics↗

Incoherently pumped high-power linearly-polarized single-mode random fiber laser: experimental investigations and theoretical prospects

We present a hundred-watt-level linearly-polarized random fiber laser (RFL) pumped by incoherent broadband amplified spontaneous emission (ASE) source and prospect the power scaling potential theoretically. The RFL employs half-opened cavity structure which is composed by a section of 330 m polarization maintained (PM) passive fiber and two PM high reflectivity fiber Bragg gratings. The 2nd order Stokes light centered at 1178 nm reaches the pump limited maximal power of 100.7 W with a full width at half-maximum linewidth of 2.58 nm and polarization extinction ratio of 23.5 dB. The corresponding ultimate quantum efficiency of pump to 2nd order Stokes light is 89.01%. To the best of our knowledge, this is the first demonstration of linearly-polarized high-order RFL with hundred-watt output power. Furthermore, the theoretical investigation indicates that 300 W-level linearly-polarized single-mode 1st order Stokes light can be obtained from incoherently pumped RFL with 100 m PM passive fiber.

physics.optics↗