Searcharxiv⌕ Search

arXiv subjects

Yan Feng

Publications and source records attributed to Yan Feng.

At least 91 records · Page 5Linked to original sources

Cross-Layer Distillation with Semantic Calibration

Knowledge distillation is a technique to enhance the generalization ability of a student model by exploiting outputs from a teacher model. Recently, feature-map based variants explore knowledge transfer between manually assigned teacher-student pairs in intermediate layers for further improvement. However, layer semantics may vary in different neural networks and semantic mismatch in manual layer associations will lead to performance degeneration due to negative regularization. To address this issue, we propose Semantic Calibration for cross-layer Knowledge Distillation (SemCKD), which automatically assigns proper target layers of the teacher model for each student layer with an attention mechanism. With a learned attention distribution, each student layer distills knowledge contained in multiple teacher layers rather than a specific intermediate layer for appropriate cross-layer supervision. We further provide theoretical analysis of the association weights and conduct extensive experiments to demonstrate the effectiveness of our approach. Code is avaliable at \url{https://github.com/DefangChen/SemCKD}.

cs.CV↗

Spectral compression by phase doubling in second harmonic generation

In second harmonic generation, the phase of the optical field is doubled which has important implication. Here the phase doubling effect is utilized to solve a long-standing challenge in power scaling of single frequency laser. When a (-π/2, π/2) binary phase modulation is applied to a single frequency seed laser to broaden the spectrum and suppress the stimulated Brillouin scattering in high power fiber amplifier, the second harmonic of the phase-modulated laser will return to single frequency, because the (-π/2, π/2) modulation is doubled to (-π, π) for the second harmonic. A compression rate as high as 95% is demonstrated in the experiment limited by the electronic bandwidth of the setup, which can be improved with optimized devices.

physics.optics↗

An Accuracy-Lossless Perturbation Method for Defending Privacy Attacks in Federated Learning

Although federated learning improves privacy of training data by exchanging local gradients or parameters rather than raw data, the adversary still can leverage local gradients and parameters to obtain local training data by launching reconstruction and membership inference attacks. To defend such privacy attacks, many noises perturbation methods (like differential privacy or CountSketch matrix) have been widely designed. However, the strong defence ability and high learning accuracy of these schemes cannot be ensured at the same time, which will impede the wide application of FL in practice (especially for medical or financial institutions that require both high accuracy and strong privacy guarantee). To overcome this issue, in this paper, we propose \emph{an efficient model perturbation method for federated learning} to defend reconstruction and membership inference attacks launched by curious clients. On the one hand, similar to the differential privacy, our method also selects random numbers as perturbed noises added to the global model parameters, and thus it is very efficient and easy to be integrated in practice. Meanwhile, the random selected noises are positive real numbers and the corresponding value can be arbitrarily large, and thus the strong defence ability can be ensured. On the other hand, unlike differential privacy or other perturbation methods that cannot eliminate the added noises, our method allows the server to recover the true gradients by eliminating the added noises. Therefore, our method does not hinder learning accuracy at all.

cs.LG↗

SamWalker++: recommendation with informative sampling strategy

Recommendation from implicit feedback is a highly challenging task due to the lack of reliable negative feedback data. Existing methods address this challenge by treating all the un-observed data as negative (dislike) but downweight the confidence of these data. However, this treatment causes two problems: (1) Confidence weights of the unobserved data are usually assigned manually, which lack flexibility and may create empirical bias on evaluating user's preference. (2) To handle massive volume of the unobserved feedback data, most of the existing methods rely on stochastic inference and data sampling strategies. However, since a user is only aware of a very small fraction of items in a large dataset, it is difficult for existing samplers to select informative training instances in which the user really dislikes the item rather than does not know it. To address the above two problems, we propose two novel recommendation methods SamWalker and SamWalker++ that support both adaptive confidence assignment and efficient model learning. SamWalker models data confidence with a social network-aware function, which can adaptively specify different weights to different data according to users' social contexts. However, the social network information may not be available in many recommender systems, which hinders application of SamWalker. Thus, we further propose SamWalker++, which does not require any side information and models data confidence with a constructed pseudo-social network. We also develop fast random-walk-based sampling strategies for our SamWalker and SamWalker++ to adaptively draw informative training instances, which can speed up gradient estimation and reduce sampling variance. Extensive experiments on five real-world datasets demonstrate the superiority of the proposed SamWalker and SamWalker++.

cs.IR↗

Boosting Black-Box Attack with Partially Transferred Conditional Adversarial Distribution

This work studies black-box adversarial attacks against deep neural networks (DNNs), where the attacker can only access the query feedback returned by the attacked DNN model, while other information such as model parameters or the training datasets are unknown. One promising approach to improve attack performance is utilizing the adversarial transferability between some white-box surrogate models and the target model (i.e., the attacked model). However, due to the possible differences on model architectures and training datasets between surrogate and target models, dubbed "surrogate biases", the contribution of adversarial transferability to improving the attack performance may be weakened. To tackle this issue, we innovatively propose a black-box attack method by developing a novel mechanism of adversarial transferability, which is robust to the surrogate biases. The general idea is transferring partial parameters of the conditional adversarial distribution (CAD) of surrogate models, while learning the untransferred parameters based on queries to the target model, to keep the flexibility to adjust the CAD of the target model on any new benign sample. Extensive experiments on benchmark datasets and attacking against real-world API demonstrate the superior attack performance of the proposed method.

cs.CR↗

Development of a VR tool to study pedestrian route and exit choice behaviour in a multi-story building

Although route and exit choice in complex buildings are important aspects of pedestrian behaviour, studies predominantly investigated pedestrian movement in a single level. This paper presents an innovative VR tool that was designed to investigate pedestrian route and exit choice in a multi-story building. This tool supports free navigation and collects pedestrian walking trajectories, head movements and gaze points automatically. An experiment was conducted to evaluate the VR tool from objective standpoints (i.e., pedestrian behaviour) and subjective standpoints (i.e., the feeling of presence, system usability, simulation sickness). The results show that the VR tool allows for accurate collection of pedestrian behavioural data in the complex building. Moreover, the results of the questionnaire report high realism of the virtual environment, high immersive feeling, high usability, and low simulator sickness. This paper contributes by showcasing an innovative approach of applying VR technologies to study pedestrian behaviour in complex and realistic environments.

cs.HC↗

CoSam: An Efficient Collaborative Adaptive Sampler for Recommendation

Sampling strategies have been widely applied in many recommendation systems to accelerate model learning from implicit feedback data. A typical strategy is to draw negative instances with uniform distribution, which however will severely affect model's convergency, stability, and even recommendation accuracy. A promising solution for this problem is to over-sample the ``difficult'' (a.k.a informative) instances that contribute more on training. But this will increase the risk of biasing the model and leading to non-optimal results. Moreover, existing samplers are either heuristic, which require domain knowledge and often fail to capture real ``difficult'' instances; or rely on a sampler model that suffers from low efficiency. To deal with these problems, we propose an efficient and effective collaborative sampling method CoSam, which consists of: (1) a collaborative sampler model that explicitly leverages user-item interaction information in sampling probability and exhibits good properties of normalization, adaption, interaction information awareness, and sampling efficiency; and (2) an integrated sampler-recommender framework, leveraging the sampler model in prediction to offset the bias caused by uneven sampling. Correspondingly, we derive a fast reinforced training algorithm of our framework to boost the sampler performance and sampler-recommender collaboration. Extensive experiments on four real-world datasets demonstrate the superiority of the proposed collaborative sampler model and integrated sampler-recommender framework.

cs.IR↗

Continuous and Discontinuous Transitions in the Depinning of Two-Dimensional Dusty Plasmas on a One-Dimensional Periodic Substrate

Langevin dynamical simulations are performed to study the depinning dynamics of two-dimensional dusty plasmas on a one-dimensional periodic substrate. From the diagnostics of the sixfold coordinated particles $P_6$ and the collective drift velocity $V_x$, three different states appear, which are the pinning, disordered plastic flow, and moving ordered states. It is found that the depth of the substrate is able to modulate the properties of the depinning phase transition, based on the results of $P_6$ and $V_x$, as well as the observation of hysteresis of $V_x$ while increasing and decreasing the driving force monotonically. When the depth of the substrate is shallow, there are two continuous phase transitions. When the potential well depth slightly increases, the phase transition from the pinned to the disordered plastic flow states is continuous; however, the phase transition from the disordered plastic flow to the moving ordered states is discontinuous. When the substrate is even deeper, the phase transition from the pinned to the disordered plastic flow states changes to discontinuous. When the substrate further increases, as the driving force increases, the pinned state changes to the moving ordered state directly, so that the disordered plastic flow state disappears completely.

physics.plasm-ph↗

High order tensor moments of random vectors

A random vector $\bx\in \R^n$ is a vector whose coordinates are all random variables. A random vector is called a Gaussian vector if it follows Gaussian distribution. These terminology can also be extended to a random (Gaussian) matrix and random (Gaussian) tensor. The classical form of an $k$-order moment (for any positive integer $k$) of a random vector $\bx\in \R^n$ is usually expressed in a matrix form of size $n\times n^{k-1}$ generated from the $k$th derivative of the characteristic function or the moment generating function of $\bx$ , and the expression of an $k$-order moment is very complicate even for a standard normal distributed vector. With the tensor form, we can simplify all the expressions related to high order moments. The main purpose of this paper is to introduce the high order moments of a random vector in tensor forms and the high order moments of a standard normal distributed vector. Finally we present an expression of high order moments of a random vector that follows a Gaussian distribution.

math.PR↗

Toward Adversarial Robustness via Semi-supervised Robust Training

Adversarial examples have been shown to be the severe threat to deep neural networks (DNNs). One of the most effective adversarial defense methods is adversarial training (AT) through minimizing the adversarial risk $R_{adv}$, which encourages both the benign example $x$ and its adversarially perturbed neighborhoods within the $\ell_{p}$-ball to be predicted as the ground-truth label. In this work, we propose a novel defense method, the robust training (RT), by jointly minimizing two separated risks ($R_{stand}$ and $R_{rob}$), which is with respect to the benign example and its neighborhoods respectively. The motivation is to explicitly and jointly enhance the accuracy and the adversarial robustness. We prove that $R_{adv}$ is upper-bounded by $R_{stand} + R_{rob}$, which implies that RT has similar effect as AT. Intuitively, minimizing the standard risk enforces the benign example to be correctly predicted, and the robust risk minimization encourages the predictions of the neighbor examples to be consistent with the prediction of the benign example. Besides, since $R_{rob}$ is independent of the ground-truth label, RT is naturally extended to the semi-supervised mode ($i.e.$, SRT), to further enhance the adversarial robustness. Moreover, we extend the $\ell_{p}$-bounded neighborhood to a general case, which covers different types of perturbations, such as the pixel-wise ($i.e.$, $x + δ$) or the spatial perturbation ($i.e.$, $ AX + b$). Extensive experiments on benchmark datasets not only verify the superiority of the proposed SRT method to state-of-the-art methods for defensing pixel-wise or spatial perturbations separately, but also demonstrate its robustness to both perturbations simultaneously. The code for reproducing main results is available at \url{https://github.com/THUYimingLi/Semi-supervised_Robust_Training}.

cs.LG↗

Fast Adaptively Weighted Matrix Factorization for Recommendation with Implicit Feedback

Recommendation from implicit feedback is a highly challenging task due to the lack of the reliable observed negative data. A popular and effective approach for implicit recommendation is to treat unobserved data as negative but downweight their confidence. Naturally, how to assign confidence weights and how to handle the large number of the unobserved data are two key problems for implicit recommendation models. However, existing methods either pursuit fast learning by manually assigning simple confidence weights, which lacks flexibility and may create empirical bias in evaluating user's preference; or adaptively infer personalized confidence weights but suffer from low efficiency. To achieve both adaptive weights assignment and efficient model learning, we propose a fast adaptively weighted matrix factorization (FAWMF) based on variational auto-encoder. The personalized data confidence weights are adaptively assigned with a parameterized neural network (function) and the network can be inferred from the data. Further, to support fast and stable learning of FAWMF, a new specific batch-based learning algorithm fBGD has been developed, which trains on all feedback data but its complexity is linear to the number of observed data. Extensive experiments on real-world datasets demonstrate the superiority of the proposed FAWMF and its learning algorithm fBGD.

cs.IR↗

Adversarial Attack on Deep Product Quantization Network for Image Retrieval

Deep product quantization network (DPQN) has recently received much attention in fast image retrieval tasks due to its efficiency of encoding high-dimensional visual features especially when dealing with large-scale datasets. Recent studies show that deep neural networks (DNNs) are vulnerable to input with small and maliciously designed perturbations (a.k.a., adversarial examples). This phenomenon raises the concern of security issues for DPQN in the testing/deploying stage as well. However, little effort has been devoted to investigating how adversarial examples affect DPQN. To this end, we propose product quantization adversarial generation (PQ-AG), a simple yet effective method to generate adversarial examples for product quantization based retrieval systems. PQ-AG aims to generate imperceptible adversarial perturbations for query images to form adversarial queries, whose nearest neighbors from a targeted product quantizaiton model are not semantically related to those from the original queries. Extensive experiments show that our PQ-AQ successfully creates adversarial examples to mislead targeted product quantization retrieval models. Besides, we found that our PQ-AG significantly degrades retrieval performance in both white-box and black-box settings.

cs.CV↗

Experimental demonstration of a dusty plasma ratchet rectification and its reversal

The naturally persistent flow of hundreds of dust particles is experimentally achieved in a dusty plasma system with the asymmetric sawteeth of gears on the electrode. It is also demonstrated that the direction of the dust particle flowcan be controlled by changing the plasma conditions of the gas pressure or the plasma power. Numerical simulations of dust particles with the ion drag inside the asymmetric sawteeth verify the experimental observations of the flow rectification of dust particles. Both experiments and simulations suggest that the asymmetric potential and the collective effect are the twokeys in this dusty plasma ratchet.With the nonequilibrium ion drag, the dust flow along the asymmetric orientation of this electric potential of the ratchet can be reversed by changing the balance height of dust particles using different plasma conditions.

physics.plasm-ph↗

Oscillation-like diffusion of two-dimensional liquid dusty plasmas on one-dimensional periodic substrates with varied widths

The long-time diffusion of two-dimensional dusty plasmas on a one-dimensional periodic substrate with varied widths is investigated using Langevin dynamical simulations. When the substrate is narrow and the dust particles form a single row, the diffusion is the smallest in both directions. We find that as the substrate width gradually increases to twice its initial value, the long-time diffusion of the two-dimensional dusty plasmas first increases, then decreases, and finally increases again, giving an oscillation-like diffusion with varied substrate width. When the width increases to a specific value, the dust particles within each potential well arrange themselves in a stable zigzag pattern, greatly reducing the diffusion, and leading to the observed oscillation in the diffusion with the increasing width. In addition, the long-time oscillation-like diffusion is consistent with the number of dust particles that are hopping across the potential wells of the substrate.

physics.plasm-ph↗

Online Knowledge Distillation with Diverse Peers

Distillation is an effective knowledge-transfer technique that uses predicted distributions of a powerful teacher model as soft targets to train a less-parameterized student model. A pre-trained high capacity teacher, however, is not always available. Recently proposed online variants use the aggregated intermediate predictions of multiple student models as targets to train each student model. Although group-derived targets give a good recipe for teacher-free distillation, group members are homogenized quickly with simple aggregation functions, leading to early saturated solutions. In this work, we propose Online Knowledge Distillation with Diverse peers (OKDDip), which performs two-level distillation during training with multiple auxiliary peers and one group leader. In the first-level distillation, each auxiliary peer holds an individual set of aggregation weights generated with an attention-based mechanism to derive its own targets from predictions of other auxiliary peers. Learning from distinct target distributions helps to boost peer diversity for effectiveness of group-based distillation. The second-level distillation is performed to transfer the knowledge in the ensemble of auxiliary peers further to the group leader, i.e., the model used for inference. Experimental results show that the proposed framework consistently gives better performance than state-of-the-art approaches without sacrificing training or inference complexity, demonstrating the effectiveness of the proposed two-level distillation framework.

cs.LG↗

Structures and Diffusion of Two Dimensional Dusty Plasmas on One Dimensional Periodic Substrates

Using numerical simulations, we examine the structure and diffusion of a two-dimensional dusty plasma (2DDP) in the presence of a one-dimensional periodic substrate (1DPS) as a function of increasing substrate strength. Both the pair correlation function perpendicular to the substrate modulation and the mean squared displacement (MSD) of dust particles are calculated. It is found that both the structure and dynamics of 2DDP exhibit strong anisotropic effects, due to the applied 1DPS. As the substrate strength increases from 0, the structure order of dusty plasma along each potential well of 1DPS increases first probably due to the competition between the inter-particle interactions and the particle-substrate interactions, and then decreases gradually, which may be due to the reduced dimensionality and the enhanced fluctuations. The obtained MSD along potential wells of 1DPS clearly shows three processes of diffusion in our studied 2DDP. Between the initial ballistic and finally diffusive motion, there is the intermediate sub-diffusion discovered here, which may result from the substrate-induced distortion of the caging dynamics.

physics.plasm-ph↗

Dissipative solitary wave at the interface of a binary complex plasma

The propagation of a dissipative solitary wave across an interface is studied in a binary complex plasma. The experiments were performed under microgravity conditions in the PK-3 Plus Laboratory on board the International Space Station using microparticles with diameters of 1.55 micrometre and 2.55 micrometre immersed in a low-temperature plasma. The solitary wave was excited at the edge of a particle-free region and propagated from the sub-cloud of small particles into that of big particles. The interfacial effect was observed by measuring the deceleration of particles in the wave crest. The results are compared with a Langevin dynamics simulation, where the waves were excited by a gentle push on the edge of the sub-cloud of small particles. Reflection of the wave at the interface is induced by increasing the strength of the push. By tuning the ion drag force exerted on big particles in the simulation, the effective width of the interface is adjusted. We show that the strength of reflection increases with narrower interfaces.

physics.plasm-ph↗

Fraunhofer diffraction at the two-dimensional quadratically distorted (QD) Grating

A two-dimensional (2D) mathematical model of quadratically distorted (QD) grating is established with the principles of Fraunhofer diffraction and Fourier optics. Discrete sampling and bisection algorithm are applied for finding numerical solution of the diffraction pattern of QD grating. This 2D mathematical model allows the precise design of QD grating and improves the optical performance of simultaneous multiplane imaging system.

physics.optics↗