Searcharxiv⌕ Search

arXiv subjects

Deng Pan

Publications and source records attributed to Deng Pan.

At least 37 records · Page 2Linked to original sources

Jamming is a first-order transition with quenched disorder in amorphous materials sheared by cyclic quasistatic deformations

Jamming is an athermal transition between flowing and rigid states in amorphous systems such as granular matter, colloidal suspensions, complex fluids and cells. The jamming transition seems to display mixed aspects of a first-order transition, evidenced by a discontinuity in the coordination number, and a second-order transition, indicated by power-law scalings and diverging lengths. Here we demonstrate that jamming is a first-order transition with quenched disorder in cyclically sheared systems with quasistatic deformations, in two and three dimensions. Based on scaling analyses, we show that fluctuations of the jamming density in finite-sized systems have important consequences on the finite-size effects of various quantities, resulting in a square relationship between disconnected and connected susceptibilities, a key signature of the first-order transition with quenched disorder. This study puts the jamming transition into the category of a broad class of transitions in disordered systems where sample-to-sample fluctuations dominate over thermal fluctuations, suggesting that the nature and behavior of the jamming transition might be better understood within the developed theoretical framework of the athermally driven random-field Ising model.

cond-mat.soft↗

Criticality in the Luria-Delbrück model with an arbitrary mutation rate

The Luria-Delbrück model is a classic model of population dynamics with random mutations, that has been used historically to prove that random mutations drive evolution. In typical scenarios, the relevant mutation rate is exceedingly small, and mutants are counted only at the final time point. Here, inspired by recent experiments on DNA repair, we study a mathematical model that is formally equivalent to the Luria-Delbrück setup, with the repair rate $p$ playing the role of mutation rate, albeit taking on large values, of order unity per cell division. We find that although at large times the fraction of repaired cells approaches one, the variance of the number of repaired cells undergoes a phase transition: when $p>1/2$ the variance decreases with time, but, intriguingly, for $p<1/2$ even though the fraction of repaired cells approaches 1, the variance in number of repaired cells increases with time. Analyzing DNA-repair experiments, we find that in order to explain the data the model should also take into account the probability of a successful repair process once it is initiated. Taken together, our work shows how the study of variability can lead to surprising phase-transitions as well as provide biological insights into the process of DNA-repair.

physics.bio-ph↗

A review on shear jamming

Jamming is a ubiquitous phenomenon that appears in many soft matter systems, including granular materials, foams, colloidal suspensions, emulsions, polymers, and cells -- when jamming occurs, the system undergoes a transition from flow-like to solid-like states. Conventionally, the jamming transition occurs when the system reaches a threshold jamming density under isotropic compression, but recent studies reveal that jamming can also be induced by shear. Shear jamming has attracted much interest in the context of non-equilibrium phase transitions, mechanics and rheology of amorphous materials. Here we review the phenomenology of shear jamming and its related physics. We first describe basic observations obtained in experiments and simulations, and results from theories. Shear jamming is then demonstrated as a "bridge" that connects the rheology of athermal soft spheres and thermal hard spheres. Based on a generalized jamming phase diagram, a universal description is provided for shear jamming in frictionless and frictional systems. We further review the isostaticity and criticality of the shear jamming transition, and the elasticity of shear jammed solids. The broader relevance of shear jamming is discussed, including its relation to other phenomena such as shear hardening, dilatancy, fragility, and discrete shear thickening.

cond-mat.soft↗

Negative Flux Aggregation to Estimate Feature Attributions

There are increasing demands for understanding deep neural networks' (DNNs) behavior spurred by growing security and/or transparency concerns. Due to multi-layer nonlinearity of the deep neural network architectures, explaining DNN predictions still remains as an open problem, preventing us from gaining a deeper understanding of the mechanisms. To enhance the explainability of DNNs, we estimate the input feature's attributions to the prediction task using divergence and flux. Inspired by the divergence theorem in vector analysis, we develop a novel Negative Flux Aggregation (NeFLAG) formulation and an efficient approximation algorithm to estimate attribution map. Unlike the previous techniques, ours doesn't rely on fitting a surrogate model nor need any path integration of gradients. Both qualitative and quantitative experiments demonstrate a superior performance of NeFLAG in generating more faithful attribution maps than the competing methods. Our code is available at \url{https://github.com/xinli0928/NeFLAG}

cs.LG↗

Decentralized federated learning methods for reducing communication cost and energy consumption in UAV networks

Unmanned aerial vehicles (UAV) or drones play many roles in a modern smart city such as the delivery of goods, mapping real-time road traffic and monitoring pollution. The ability of drones to perform these functions often requires the support of machine learning technology. However, traditional machine learning models for drones encounter data privacy problems, communication costs and energy limitations. Federated Learning, an emerging distributed machine learning approach, is an excellent solution to address these issues. Federated learning (FL) allows drones to train local models without transmitting raw data. However, existing FL requires a central server to aggregate the trained model parameters of the UAV. A failure of the central server can significantly impact the overall training. In this paper, we propose two aggregation methods: Commutative FL and Alternate FL, based on the existing architecture of decentralised Federated Learning for UAV Networks (DFL-UN) by adding a unique aggregation method of decentralised FL. Those two methods can effectively control energy consumption and communication cost by controlling the number of local training epochs, local communication, and global communication. The simulation results of the proposed training methods are also presented to verify the feasibility and efficiency of the architecture compared with two benchmark methods (e.g. standard machine learning training and standard single aggregation server training). The simulation results show that the proposed methods outperform the benchmark methods in terms of operational stability, energy consumption and communication cost.

cs.LG↗

Shear hardening in frictionless amorphous solids near the jamming transition

The jamming transition, generally manifested by a rapid increase of rigidity under compression (i.e., compression hardening), is ubiquitous in amorphous materials. Here we study shear hardening in deeply annealed frictionless packings generated by numerical simulations, reporting critical scalings absent in compression hardening. We demonstrate that hardening is a natural consequence of shear-induced memory destruction. Based on an elasticity theory, we reveal two independent microscopic origins of shear hardening: (i) the increase of the interaction bond number and (ii) the emergence of anisotropy and long-range correlations in the orientations of bonds, the latter highlights the essential difference between compression and shear hardening. Through the establishment of physical laws specific to anisotropy, our work completes the criticality and universality of jamming transition, and the elasticity theory of amorphous solids.

cond-mat.soft↗

Learning Compact Features via In-Training Representation Alignment

Deep neural networks (DNNs) for supervised learning can be viewed as a pipeline of the feature extractor (i.e., last hidden layer) and a linear classifier (i.e., output layer) that are trained jointly with stochastic gradient descent (SGD) on the loss function (e.g., cross-entropy). In each epoch, the true gradient of the loss function is estimated using a mini-batch sampled from the training set and model parameters are then updated with the mini-batch gradients. Although the latter provides an unbiased estimation of the former, they are subject to substantial variances derived from the size and number of sampled mini-batches, leading to noisy and jumpy updates. To stabilize such undesirable variance in estimating the true gradients, we propose In-Training Representation Alignment (ITRA) that explicitly aligns feature distributions of two different mini-batches with a matching loss in the SGD training process. We also provide a rigorous analysis of the desirable effects of the matching loss on feature representation learning: (1) extracting compact feature representation; (2) reducing over-adaption on mini-batches via an adaptive weighting mechanism; and (3) accommodating to multi-modalities. Finally, we conduct large-scale experiments on both image and text classifications to demonstrate its superior performance to the strong baselines.

cs.LG↗

Spin flips of electron beams in optical near fields

Manipulating the spin polarization of electron beams using light is highly desirable but exceedingly challenging, as the approaches proposed in previous studies using free-space light usually require enormous laser intensities. Here, we propose the use of a transverse electric optical near field, extended on nanostructures, to efficiently induce spin flips of an adjacent electron beam by exploiting the strong inelastic electron scattering in phase-matched optical near fields. Our calculations show that the use of a dramatically reduced laser intensity ($\sim 10^{12}\,$W/cm$^2$) with a short interaction length ($16\,μ$m) achieves an electron spin-flip probability of approximately $12\%$. Intriguingly, the two spin components of an unpolarized incident electron beam -- parallel and antiparallel to the electric field -- are spin-flipped and inelastically scattered to different energy states, providing an analog of the Stern--Gerlach experiment in the energy dimension. Our findings are important for optical control of free-electron spins, preparation of spin-polarized electron beams, and applications as varied as in material science and high-energy physics.

physics.optics↗

Non-linear elasticity, yielding and entropy in amorphous solids

The holographic duality has proven successful in linking seemingly unrelated problems in physics.Recently, intriguing correspondences between the physics of soft matter and gravity are emerging,including strong similarities between the rheology of amorphous solids, effective field theories for elasticity and the physics of black holes. However, direct comparisons between theoretical predictions and experimental/simulation observations remain limited. Here, we study the effects of non-linear elasticity on the mechanical and thermodynamic properties of amorphous materials responding to shear, using effective field and gravitational theories. The predicted correlations among the non-linear elastic exponent, the yielding strain/stress and the entropy change due to shear are supported qualitatively by simulations of granular matter models. Our approach opens a path towards understanding complex mechanical responses of amorphous solids, such as mixed effects of shear softening and shear hardening, and offers the possibility to study the rheology of solid states and black holes in a unified framework.

cond-mat.soft↗

First-Generation Inference Accelerator Deployment at Facebook

In this paper, we provide a deep dive into the deployment of inference accelerators at Facebook. Many of our ML workloads have unique characteristics, such as sparse memory accesses, large model sizes, as well as high compute, memory and network bandwidth requirements. We co-designed a high-performance, energy-efficient inference accelerator platform based on these requirements. We describe the inference accelerator platform ecosystem we developed and deployed at Facebook: both hardware, through Open Compute Platform (OCP), and software framework and tooling, through Pytorch/Caffe2/Glow. A characteristic of this ecosystem from the start is its openness to enable a variety of AI accelerators from different vendors. This platform, with six low-power accelerator cards alongside a single-socket host CPU, allows us to serve models of high complexity that cannot be easily or efficiently run on CPUs. We describe various performance optimizations, at both platform and accelerator level, which enables this platform to serve production traffic at Facebook. We also share deployment challenges, lessons learned during performance optimization, as well as provide guidance for future inference hardware co-design.

cs.AR↗

Temporal Evolution of Flow in Pore-Networks: From Homogenization to Instability

We study the dynamics of flow-networks in porous media using a pore-network model. First, we consider a class of erosion dynamics assuming a constitutive law depending on flow rate, local velocities, or shear stress at the walls. We show that depending on the erosion law, the flow may become uniform and homogenized or become unstable and develop channels. By defining an order parameter capturing these different behaviors we show that a phase transition occurs depending on the erosion dynamics. Using a simple model, we identify quantitative criteria to distinguish these regimes and correctly predict the fate of the network, and discuss the experimental relevance of our result.

physics.flu-dyn↗

Improving Adversarial Robustness via Probabilistically Compact Loss with Logit Constraints

Convolutional neural networks (CNNs) have achieved state-of-the-art performance on various tasks in computer vision. However, recent studies demonstrate that these models are vulnerable to carefully crafted adversarial samples and suffer from a significant performance drop when predicting them. Many methods have been proposed to improve adversarial robustness (e.g., adversarial training and new loss functions to learn adversarially robust feature representations). Here we offer a unique insight into the predictive behavior of CNNs that they tend to misclassify adversarial samples into the most probable false classes. This inspires us to propose a new Probabilistically Compact (PC) loss with logit constraints which can be used as a drop-in replacement for cross-entropy (CE) loss to improve CNN's adversarial robustness. Specifically, PC loss enlarges the probability gaps between true class and false classes meanwhile the logit constraints prevent the gaps from being melted by a small perturbation. We extensively compare our method with the state-of-the-art using large scale datasets under both white-box and black-box attacks to demonstrate its effectiveness. The source codes are available from the following url: https://github.com/xinli0928/PC-LC.

cs.LG↗

Emergence of material momentum in optical media

Understanding the momentum of light when propagating through optical media is not only fundamental for studies as varied as classical electrodynamics and polaritonics in condensed matter physics, but also for important applications such as optical-force manipulations and photovoltaics. From a microscopic perspective, an optical medium is in fact a complex system that can split the light momentum into the electromagnetic field, as well as the material electrons and the ionic lattice. Here, we disentangle the partition of momentum associated with light propagation in optical media, and develop a quantum theory to explicitly calculate its distribution. The material momentum here revealed, which is distributed among electrons and ionic lattice, leads to the prediction of unexpected phenomena. In particular, the electron momentum manifests through an intrinsic DC current, and strikingly, we find that under certain conditions this current can be along the photonic wave vector, implying an optical pulling effect on the electrons. Likewise, an optical pulling effect on the lattice can also be observed, such as in graphene during plasmon propagation. We also predict the emergence of boundary electric dipoles associated with light transmission through finite media, offering a microscopic explanation of optical pressure on material boundaries.

physics.optics↗

Explainable Recommendation via Interpretable Feature Mapping and Evaluation of Explainability

Latent factor collaborative filtering (CF) has been a widely used technique for recommender system by learning the semantic representations of users and items. Recently, explainable recommendation has attracted much attention from research community. However, trade-off exists between explainability and performance of the recommendation where metadata is often needed to alleviate the dilemma. We present a novel feature mapping approach that maps the uninterpretable general features onto the interpretable aspect features, achieving both satisfactory accuracy and explainability in the recommendations by simultaneous minimization of rating prediction loss and interpretation loss. To evaluate the explainability, we propose two new evaluation metrics specifically designed for aspect-level explanation using surrogate ground truth. Experimental results demonstrate a strong performance in both recommendation and explaining explanation, eliminating the need for metadata. Code is available from https://github.com/pd90506/AMCF.

cs.LG↗

Defending against adversarial attacks on medical imaging AI system, classification or detection?

Medical imaging AI systems such as disease classification and segmentation are increasingly inspired and transformed from computer vision based AI systems. Although an array of adversarial training and/or loss function based defense techniques have been developed and proved to be effective in computer vision, defending against adversarial attacks on medical images remains largely an uncharted territory due to the following unique challenges: 1) label scarcity in medical images significantly limits adversarial generalizability of the AI system; 2) vastly similar and dominant fore- and background in medical images make it hard samples for learning the discriminating features between different disease classes; and 3) crafted adversarial noises added to the entire medical image as opposed to the focused organ target can make clean and adversarial examples more discriminate than that between different disease classes. In this paper, we propose a novel robust medical imaging AI framework based on Semi-Supervised Adversarial Training (SSAT) and Unsupervised Adversarial Detection (UAD), followed by designing a new measure for assessing systems adversarial risk. We systematically demonstrate the advantages of our robust medical imaging AI system over the existing adversarial defense techniques under diverse real-world settings of adversarial attacks using a benchmark OCT imaging data set.

cs.LG↗

Anomalous thermodiffusion of electrons in graphene

We reveal a dramatic departure of electron thermodiffusion in solids relative to the commonly accepted picture of the ideal free-electron gas model. In particular, we show that the interaction with the lattice and impurities, combined with a strong material dependence of the electron dispersion relation, leads to counterintuitive diffusion behavior, which we identify by comparing a single-layer two-dimensional electron gas (2DEG) and graphene. When subject to a temperature gradient $\nabla T$, thermodiffusion of massless Dirac electrons in graphene exhibits an anomalous behavior with electrons moving along $\nabla T$ and accumulating in hot regions, in contrast to normal electron diffusion in a 2DEG with parabolic dispersion, where net motion against $\nabla T$ is observed, accompanied by electron depletion in hot regions. These findings have fundamentally importance for the understanding of the spatial electron dynamics in emerging material, establishing close relations with other branches of physics dealing with electron systems under nonuniform temperature conditions.

cond-mat.mes-hall↗

Is friction essential for dilatancy and shear jamming in granular matter?

Granular packings display the remarkable phenomenon of dilatancy [1], wherein their volume increases upon shear deformation. Conventional wisdom and previous results suggest that dilatancy, as also the related phenomenon of shear-induced jamming, requires frictional interactions [2, 3]. Here, we investigate the occurrence of dilatancy and shear jamming in frictionless packings. We show that the existence of isotropic jamming densities ϕj above the minimal density, the J-point density ϕJ [4, 5], leads both to the emergence of shear-induced jamming and dilatancy. Packings at ϕJ form a significant threshold state into which systems evolve in the limit of vanishing pressure under constant pressure shear, irrespective of the initial jamming density ϕj. While packings for different ϕj display equivalent scaling properties under compression [6], they exhibit striking differences in rheological behaviour under shear. The yield stress under constant volume shear increases discontinuously with density when ϕj > ϕJ, contrary to the continuous behavior in generic packings that jam at ϕJ [4, 7].

cond-mat.stat-mech↗

On the Learning Property of Logistic and Softmax Losses for Deep Neural Networks

Deep convolutional neural networks (CNNs) trained with logistic and softmax losses have made significant advancement in visual recognition tasks in computer vision. When training data exhibit class imbalances, the class-wise reweighted version of logistic and softmax losses are often used to boost performance of the unweighted version. In this paper, motivated to explain the reweighting mechanism, we explicate the learning property of those two loss functions by analyzing the necessary condition (e.g., gradient equals to zero) after training CNNs to converge to a local minimum. The analysis immediately provides us explanations for understanding (1) quantitative effects of the class-wise reweighting mechanism: deterministic effectiveness for binary classification using logistic loss yet indeterministic for multi-class classification using softmax loss; (2) disadvantage of logistic loss for single-label multi-class classification via one-vs.-all approach, which is due to the averaging effect on predicted probabilities for the negative class (e.g., non-target classes) in the learning process. With the disadvantage and advantage of logistic loss disentangled, we thereafter propose a novel reweighted logistic loss for multi-class classification. Our simple yet effective formulation improves ordinary logistic loss by focusing on learning hard non-target classes (target vs. non-target class in one-vs.-all) and turned out to be competitive with softmax loss. We evaluate our method on several benchmark datasets to demonstrate its effectiveness.

cs.LG↗