Searcharxiv⌕ Search

arXiv subjects

Da Xu

Publications and source records attributed to Da Xu.

At least 55 records · Page 3Linked to original sources

An Efficient Group-based Search Engine Marketing System for E-Commerce

With the increasing scale of search engine marketing, designing an efficient bidding system is becoming paramount for the success of e-commerce companies. The critical challenges faced by a modern industrial-level bidding system include: 1. the catalog is enormous, and the relevant bidding features are of high sparsity; 2. the large volume of bidding requests induces significant computation burden to both the offline and online serving. Leveraging extraneous user-item information proves essential to mitigate the sparsity issue, for which we exploit the natural language signals from the users' query and the contextual knowledge from the products. In particular, we extract the vector representations of ads via the Transformer model and leverage their geometric relation to building collaborative bidding predictions via clustering. The two-step procedure also significantly reduces the computation stress of bid evaluation and optimization. In this paper, we introduce the end-to-end structure of the bidding system for search engine marketing for Walmart e-commerce, which successfully handles tens of millions of bids each day. We analyze the online and offline performances of our approach and discuss how we find it as a production-efficient solution.

cs.CL↗

A universal deep learning strategy for designing high-quality-factor photonic resonances

Resonance is instrumental in modern optics and photonics for novel phenomena such as cavity quantum electrodynamics and electric-field-induced transparency. While one can use numerical simulations to sweep geometric and material parameters of optical structures, these simulations usually require considerably long calculation time (spanning from several hours to several weeks) and substantial computational resources. Such requirements significantly limit their applicability in understanding and inverse designing structures with desired resonance performances. Recently, the introduction of artificial intelligence allows for faster predictions of resonance with less demanding computational requirements. However, current end-to-end deep learning approaches generally fail to predict resonances with high quality-factors (Q-factor). Here, we introduce a universal deep learning strategy that can predict ultra-high Q-factor resonances by decomposing spectra with an adaptive data acquisition (ADA) method while incorporating resonance information. We exploit bound states in the continuum (BICs) with an infinite Q-factor to testify this resonance-informed deep learning (RIDL) strategy. The trained RIDL strategy achieves high-accuracy prediction of reflection spectra and photonic band structures while using a considerably small training dataset. We further develop an inverse design algorithm based on the RIDL strategy for a symmetry-protected BIC on a suspended silicon nitride photonic crystal (PhC) slab. The predicted and measured angle-resolved band structures show minimum differences. We expect the RIDL strategy to apply to many other physical phenomena which exhibit Gaussian, Lorentzian, and Fano resonances.

physics.optics↗

Exceptional Point and Cross-Relaxation Effect in a Hybrid Quantum System

Exceptional points (EPs) are exotic degeneracies of non-Hermitian systems, where the eigenvalues and the corresponding eigenvectors simultaneously coalesce in parameter space, and these degeneracies are sensitive to tiny perturbations on the system. Here we report an experimental observation of the EP in a hybrid quantum system consisting of dense nitrogen (P1) centers in diamond coupled to a coplanar-waveguide resonator. These P1 centers can be divided into three subensembles of spins, and cross relaxation occurs among them. As a new method to demonstrate this EP, we pump a given spin subensemble with a drive field to tune the magnon-photon coupling in a wide range. We observe the EP in the middle spin subensemble coupled to the resonator mode, irrespective of which spin subensemble is actually driven. This robustness of the EP against pumping reveals the key role of the cross relaxation in P1 centers. It offers a novel way to convincingly prove the existence of the cross-relaxation effect via the EP.

quant-ph↗

Understanding the role of importance weighting for deep learning

The recent paper by Byrd & Lipton (2019), based on empirical observations, raises a major concern on the impact of importance weighting for the over-parameterized deep learning models. They observe that as long as the model can separate the training data, the impact of importance weighting diminishes as the training proceeds. Nevertheless, there lacks a rigorous characterization of this phenomenon. In this paper, we provide formal characterizations and theoretical justifications on the role of importance weighting with respect to the implicit bias of gradient descent and margin-based learning theory. We reveal both the optimization dynamics and generalization performance under deep learning models. Our work not only explains the various novel phenomenons observed for importance weighting in deep learning, but also extends to the studies where the weights are being optimized as part of the model, which applies to a number of topics under active research.

cs.LG↗

A Temporal Kernel Approach for Deep Learning with Continuous-time Information

Sequential deep learning models such as RNN, causal CNN and attention mechanism do not readily consume continuous-time information. Discretizing the temporal data, as we show, causes inconsistency even for simple continuous-time processes. Current approaches often handle time in a heuristic manner to be consistent with the existing deep learning architectures and implementations. In this paper, we provide a principled way to characterize continuous-time systems using deep learning tools. Notably, the proposed approach applies to all the major deep learning architectures and requires little modifications to the implementation. The critical insight is to represent the continuous-time system by composing neural networks with a temporal kernel, where we gain our intuition from the recent advancements in understanding deep learning with Gaussian process and neural tangent kernel. To represent the temporal kernel, we introduce the random feature approach and convert the kernel learning problem to spectral density estimation under reparameterization. We further prove the convergence and consistency results even when the temporal kernel is non-stationary, and the spectral density is misspecified. The simulations and real-data experiments demonstrate the empirical effectiveness of our temporal kernel approach in a broad range of settings.

cs.LG↗

6 nm super-resolution optical transmission and scattering spectroscopic imaging of carbon nanotubes using a nanometer-scale white light source

Optical hyperspectral imaging based on absorption and scattering of photons at the visible and adjacent frequencies denotes one of the most informative and inclusive characterization methods in material research. Unfortunately, restricted by the diffraction limit of light, it is unable to resolve the nanoscale inhomogeneity in light-matter interactions, which is diagnostic of the local modulation in material structure and properties. Moreover, many nanomaterials have highly anisotropic optical properties that are outstandingly appealing yet hard to characterize through conventional optical methods. Therefore, there has been a pressing demand in the diverse fields including electronics, photonics, physics, and materials science to extend the optical hyperspectral imaging into the nanometer length scale. In this work, we report a super-resolution hyperspectral imaging technique that simultaneously measures optical absorption and scattering spectra with the illumination from a tungsten-halogen lamp. We demonstrated sub-5 nm spatial resolution in both visible and near-infrared wavelengths (415 to 980 nm) for the hyperspectral imaging of strained single-walled carbon nanotubes (SWNT) and reconstructed true-color images to reveal the longitudinal and transverse optical transition-induced light absorption and scattering in the SWNTs. This is the first time transverse optical absorption in SWNTs were clearly observed experimentally. The new technique provides rich near-field spectroscopic information that had made it possible to analyze the spatial modulation of band-structure along a single SWNT induced through strain engineering.

physics.optics↗

Theoretical Understandings of Product Embedding for E-commerce Machine Learning

Product embeddings have been heavily investigated in the past few years, serving as the cornerstone for a broad range of machine learning applications in e-commerce. Despite the empirical success of product embeddings, little is known on how and why they work from the theoretical standpoint. Analogous results from the natural language processing (NLP) often rely on domain-specific properties that are not transferable to the e-commerce setting, and the downstream tasks often focus on different aspects of the embeddings. We take an e-commerce-oriented view of the product embeddings and reveal a complete theoretical view from both the representation learning and the learning theory perspective. We prove that product embeddings trained by the widely-adopted skip-gram negative sampling algorithm and its variants are sufficient dimension reduction regarding a critical product relatedness measure. The generalization performance in the downstream machine learning task is controlled by the alignment between the embeddings and the product relatedness measure. Following the theoretical discoveries, we conduct exploratory experiments that supports our theoretical insights for the product embeddings.

cs.LG↗

Adversarial Counterfactual Learning and Evaluation for Recommender System

The feedback data of recommender systems are often subject to what was exposed to the users; however, most learning and evaluation methods do not account for the underlying exposure mechanism. We first show in theory that applying supervised learning to detect user preferences may end up with inconsistent results in the absence of exposure information. The counterfactual propensity-weighting approach from causal inference can account for the exposure mechanism; nevertheless, the partial-observation nature of the feedback data can cause identifiability issues. We propose a principled solution by introducing a minimax empirical risk formulation. We show that the relaxation of the dual problem can be converted to an adversarial game between two recommendation models, where the opponent of the candidate model characterizes the underlying exposure mechanism. We provide learning bounds and conduct extensive simulation studies to illustrate and justify the proposed approach over a broad range of recommendation settings, which shed insights on the various benefits of the proposed approach.

cs.IR↗

Sparse Symmetric Tensor Regression for Functional Connectivity Analysis

Tensor regression models, such as CP regression and Tucker regression, have many successful applications in neuroimaging analysis where the covariates are of ultrahigh dimensionality and possess complex spatial structures. The high-dimensional covariate arrays, also known as tensors, can be approximated by low-rank structures and fit into the generalized linear models. The resulting tensor regression achieves a significant reduction in dimensionality while remaining efficient in estimation and prediction. Brain functional connectivity is an essential measure of brain activity and has shown significant association with neurological disorders such as Alzheimer's disease. The symmetry nature of functional connectivity is a property that has not been explored in previous tensor regression models. In this work, we propose a sparse symmetric tensor regression that further reduces the number of free parameters and achieves superior performance over symmetrized and ordinary CP regression, under a variety of simulation settings. We apply the proposed method to a study of Alzheimer's disease (AD) and normal ageing from the Berkeley Aging Cohort Study (BACS) and detect two regions of interest that have been identified important to AD.

cs.LG↗

Inductive Representation Learning on Temporal Graphs

Inductive representation learning on temporal graphs is an important step toward salable machine learning on real-world dynamic networks. The evolving nature of temporal dynamic graphs requires handling new nodes as well as capturing temporal patterns. The node embeddings, which are now functions of time, should represent both the static node features and the evolving topological structures. Moreover, node and topological features can be temporal as well, whose patterns the node embeddings should also capture. We propose the temporal graph attention (TGAT) layer to efficiently aggregate temporal-topological neighborhood features as well as to learn the time-feature interactions. For TGAT, we use the self-attention mechanism as building block and develop a novel functional time encoding technique based on the classical Bochner's theorem from harmonic analysis. By stacking TGAT layers, the network recognizes the node embeddings as functions of time and is able to inductively infer embeddings for both new and observed nodes as the graph evolves. The proposed approach handles both node classification and link prediction task, and can be naturally extended to include the temporal edge features. We evaluate our method with transductive and inductive tasks under temporal settings with two benchmark and one industrial dataset. Our TGAT model compares favorably to state-of-the-art baselines as well as the previous temporal graph embedding approaches.

cs.LG↗

Self-attention with Functional Time Representation Learning

Sequential modelling with self-attention has achieved cutting edge performances in natural language processing. With advantages in model flexibility, computation complexity and interpretability, self-attention is gradually becoming a key component in event sequence models. However, like most other sequence models, self-attention does not account for the time span between events and thus captures sequential signals rather than temporal patterns. Without relying on recurrent network structures, self-attention recognizes event orderings via positional encoding. To bridge the gap between modelling time-independent and time-dependent event sequence, we introduce a functional feature map that embeds time span into high-dimensional spaces. By constructing the associated translation-invariant time kernel function, we reveal the functional forms of the feature map under classic functional function analysis results, namely Bochner's Theorem and Mercer's Theorem. We propose several models to learn the functional time representation and the interactions with event representation. These methods are evaluated on real-world datasets under various continuous-time event sequence prediction tasks. The experiments reveal that the proposed methods compare favorably to baseline models while also capturing useful time-event interactions.

cs.LG↗

Product Knowledge Graph Embedding for E-commerce

In this paper, we propose a new product knowledge graph (PKG) embedding approach for learning the intrinsic product relations as product knowledge for e-commerce. We define the key entities and summarize the pivotal product relations that are critical for general e-commerce applications including marketing, advertisement, search ranking and recommendation. We first provide a comprehensive comparison between PKG and ordinary knowledge graph (KG) and then illustrate why KG embedding methods are not suitable for PKG learning. We construct a self-attention-enhanced distributed representation learning model for learning PKG embeddings from raw customer activity data in an end-to-end fashion. We design an effective multi-task learning schema to fully leverage the multi-modal e-commerce data. The Poincare embedding is also employed to handle complex entity structures. We use a real-world dataset from grocery.walmart.com to evaluate the performances on knowledge completion, search ranking and recommendation. The proposed approach compares favourably to baselines in knowledge completion and downstream tasks.

cs.LG↗

Knowledge-aware Complementary Product Representation Learning

Learning product representations that reflect complementary relationship plays a central role in e-commerce recommender system. In the absence of the product relationships graph, which existing methods rely on, there is a need to detect the complementary relationships directly from noisy and sparse customer purchase activities. Furthermore, unlike simple relationships such as similarity, complementariness is asymmetric and non-transitive. Standard usage of representation learning emphasizes on only one set of embedding, which is problematic for modelling such properties of complementariness. We propose using knowledge-aware learning with dual product embedding to solve the above challenges. We encode contextual knowledge into product representation by multi-task learning, to alleviate the sparsity issue. By explicitly modelling with user bias terms, we separate the noise of customer-specific preferences from the complementariness. Furthermore, we adopt the dual embedding framework to capture the intrinsic properties of complementariness and provide geometric interpretation motivated by the classic separating hyperplane theory. Finally, we propose a Bayesian network structure that unifies all the components, which also concludes several popular models as special cases. The proposed method compares favourably to state-of-art methods, in downstream classification and recommendation tasks. We also develop an implementation that scales efficiently to a dataset with millions of items and customers.

cs.IR↗

Synchronization and temporal nonreciprocity of optical microresonators via spontaneous symmetry breaking

Synchronization is of importance in both fundamental and applied physics, but their demonstration at the micro/nanoscale is mainly limited to low-frequency oscillations like mechanical resonators. Here, we report the synchronization of two coupled optical microresonators, in which the high-frequency resonances in optical domain are aligned with reduced noise. It is found that two types of synchronization emerge with either the first- or second-order transition, both presenting a process of spontaneous symmetry breaking. In the second-order regime, the synchronization happens with an invariant topological character number and a larger detuning than that of the first-order case. Furthermore, an unconventional hysteresis behavior is revealed for a time-dependent coupling strength, breaking the static limitation and the temporal reciprocity. The synchronization of optical microresonators offers great potential in reconfigurable simulations of many-body physics and scalable photonic devices on a chip.

physics.optics↗

Quantum Simulation of the Fermion-Boson Composite Quasi-Particles with a Driven Qubit-Magnon Hybrid Quantum System

We experimentally demonstrate strong coupling between the ferromagnetic magnons in a small yttrium-iron-garnet (YIG) sphere and the drive-field-induced dressed states of a superconducting qubit, which gives rise to the double dressing of the superconducting qubit. The YIG sphere and the superconducting qubit are embedded in a microwave cavity and the effective coupling between them is mediated by the virtual cavity photons. The theoretical results fit the experimental observations well in a wide region of the drive-field power resonantly applied to the superconducting qubit and reveal that the driven qubit-magnon hybrid quantum system can be harnessed to emulate a particle-hole-symmetric pair coupled to a bosonic mode. This hybrid quantum system offers a novel platform for quantum simulation of the composite quasi-particles consisting of fermions and bosons.

quant-ph↗

Generative Graph Convolutional Network for Growing Graphs

Modeling generative process of growing graphs has wide applications in social networks and recommendation systems, where cold start problem leads to new nodes isolated from existing graph. Despite the emerging literature in learning graph representation and graph generation, most of them can not handle isolated new nodes without nontrivial modifications. The challenge arises due to the fact that learning to generate representations for nodes in observed graph relies heavily on topological features, whereas for new nodes only node attributes are available. Here we propose a unified generative graph convolutional network that learns node representations for all nodes adaptively in a generative model framework, by sampling graph generation sequences constructed from observed graph data. We optimize over a variational lower bound that consists of a graph reconstruction term and an adaptive Kullback-Leibler divergence regularization term. We demonstrate the superior performance of our approach on several benchmark citation network datasets.

cs.LG↗

Spontaneous $\mathcal{T}$-symmetry breaking and exceptional points in cavity quantum electrodynamics systems

Spontaneous symmetry breaking has revolutionized the understanding in numerous fields of modern physics. Here, we theoretically demonstrate the spontaneous time-reversal symmetry breaking in a cavity quantum electrodynamics system in which an atomic ensemble interacts coherently with a single resonant cavity mode. The interacting system can be effectively described by two coupled oscillators with positive and negative mass when the two-level atoms are prepared in their excited states. The occurrence of symmetry breaking is controlled by the atomic detuning and the coupling to the cavity mode, which naturally divides the parameter space into the symmetry broken and symmetry unbroken phases. The two phases are separated by a spectral singularity, a so-called exceptional point, where the eigenstates of the Hamiltonian coalesce. When encircling the singularity in the parameter space, the quasi-adiabatic dynamics shows chiral mode switching which enables topological manipulation of quantum states.

quant-ph↗

Synthesis of antisymmetric spin exchange interaction and entanglement generation with chiral spin states in a superconducting circuit

We have synthesized the anti-symmetric spin exchange interaction (ASI), which is also called the Dzyaloshinskii-Moriya interaction, in a superconducting circuit containing five superconducting qubits connected to a bus resonator, by periodically modulating the transition frequencies of the qubits with different modulation phases. This allows us to show the chiral spin dynamics in three-, four- and five-spin clusters. We also demonstrate a three-spin chiral logic gate and entangle up to five qubits in Greenberger-Horne-Zeilinger states. Our results pave the way for quantum simulation of magnetism with ASI and quantum computation with chiral spin states.

quant-ph↗