Searcharxiv⌕ Search

arXiv subjects

Sen Li

Publications and source records attributed to Sen Li.

At least 73 records · Page 4Linked to original sources

Rethinking the Role of Pre-ranking in Large-scale E-Commerce Searching System

E-commerce search systems such as Taobao Search, the largest e-commerce searching system in China, aim at providing users with the most preferred items (e.g., products). Due to the massive data and limited time for response, a typical industrial ranking system consists of three or more modules, including matching, pre-ranking, and ranking. The pre-ranking is widely considered a mini-ranking module, as it needs to rank hundreds of times more items than the ranking under limited latency. Existing researches focus on building a lighter model that imitates the ranking model. As such, the metric of a pre-ranking model follows the ranking model using Area Under ROC (AUC) for offline evaluation. However, such a metric is inconsistent with online A/B tests in practice, so engineers have to perform costly online tests to reach a convincing conclusion. In our work, we rethink the role of the pre-ranking. We argue that the primary goal of the pre-ranking stage is to return an optimal unordered set rather than an ordered list of items because it is the ranking that determines the final exposures. Since AUC measures the quality of an ordered item list, it is not suitable for evaluating the quality of the output unordered set. This paper proposes a new evaluation metric called All-Scenario Hitrate (ASH) for pre-ranking. ASH is proven effective in the offline evaluation and consistent with online A/B tests based on numerous experiments in Taobao Search. We also introduce an all-scenario-based multi-objective learning framework (ASMOL), which improves the ASH significantly. Surprisingly, the new pre-ranking model can outperforms the ranking model when outputting thousands of items. The phenomenon validates that the pre-ranking stage should not imitate the ranking blindly. With the improvements in ASH consistently translating to online improvement, it makes a 1.2% GMV improvement on Taobao Search.

cs.IR↗

DABS: Data-Agnostic Backdoor attack at the Server in Federated Learning

Federated learning (FL) attempts to train a global model by aggregating local models from distributed devices under the coordination of a central server. However, the existence of a large number of heterogeneous devices makes FL vulnerable to various attacks, especially the stealthy backdoor attack. Backdoor attack aims to trick a neural network to misclassify data to a target label by injecting specific triggers while keeping correct predictions on original training data. Existing works focus on client-side attacks which try to poison the global model by modifying the local datasets. In this work, we propose a new attack model for FL, namely Data-Agnostic Backdoor attack at the Server (DABS), where the server directly modifies the global model to backdoor an FL system. Extensive simulation results show that this attack scheme achieves a higher attack success rate compared with baseline methods while maintaining normal accuracy on the clean data.

cs.CR↗

Memory-augmented Contrastive Learning for Talking Head Generation

Given one reference facial image and a piece of speech as input, talking head generation aims to synthesize a realistic-looking talking head video. However, generating a lip-synchronized video with natural head movements is challenging. The same speech clip can generate multiple possible lip and head movements, that is, there is no one-to-one mapping relationship between them. To overcome this problem, we propose a Speech Feature Extractor (SFE) based on memory-augmented self-supervised contrastive learning, which introduces the memory module to store multiple different speech mapping results. In addition, we introduce the Mixed Density Networks (MDN) into the landmark regression task to generate multiple predicted facial landmarks. Extensive qualitative and quantitative experiments show that the quality of our facial animation is significantly superior to that of the state-of-the-art (SOTA). The code has been released at https://github.com/Yaxinzhao97/MACL.git.

cs.MM↗

MAKE: Vision-Language Pre-training based Product Retrieval in Taobao Search

Taobao Search consists of two phases: the retrieval phase and the ranking phase. Given a user query, the retrieval phase returns a subset of candidate products for the following ranking phase. Recently, the paradigm of pre-training and fine-tuning has shown its potential in incorporating visual clues into retrieval tasks. In this paper, we focus on solving the problem of text-to-multimodal retrieval in Taobao Search. We consider that users' attention on titles or images varies on products. Hence, we propose a novel Modal Adaptation module for cross-modal fusion, which helps assigns appropriate weights on texts and images across products. Furthermore, in e-commerce search, user queries tend to be brief and thus lead to significant semantic imbalance between user queries and product titles. Therefore, we design a separate text encoder and a Keyword Enhancement mechanism to enrich the query representations and improve text-to-multimodal matching. To this end, we present a novel vision-language (V+L) pre-training methods to exploit the multimodal information of (user query, product title, product image). Extensive experiments demonstrate that our retrieval-specific pre-training model (referred to as MAKE) outperforms existing V+L pre-training methods on the text-to-multimodal retrieval task. MAKE has been deployed online and brings major improvements on the retrieval system of Taobao Search.

cs.IR↗

Multi-Objective Personalized Product Retrieval in Taobao Search

In large-scale e-commerce platforms like Taobao, it is a big challenge to retrieve products that satisfy users from billions of candidates. This has been a common concern of academia and industry. Recently, plenty of works in this domain have achieved significant improvements by enhancing embedding-based retrieval (EBR) methods, including the Multi-Grained Deep Semantic Product Retrieval (MGDSPR) model [16] in Taobao search engine. However, we find that MGDSPR still has problems of poor relevance and weak personalization compared to other retrieval methods in our online system, such as lexical matching and collaborative filtering. These problems promote us to further strengthen the capabilities of our EBR model in both relevance estimation and personalized retrieval. In this paper, we propose a novel Multi-Objective Personalized Product Retrieval (MOPPR) model with four hierarchical optimization objectives: relevance, exposure, click and purchase. We construct entire-space multi-positive samples to train MOPPR, rather than the single-positive samples for existing EBR models.We adopt a modified softmax loss for optimizing multiple objectives. Results of extensive offline and online experiments show that MOPPR outperforms the baseline MGDSPR on evaluation metrics of relevance estimation and personalized retrieval. MOPPR achieves 0.96% transaction and 1.29% GMV improvements in a 28-day online A/B test. Since the Double-11 shopping festival of 2021, MOPPR has been fully deployed in mobile Taobao search, replacing the previous MGDSPR. Finally, we discuss several advanced topics of our deeper explorations on multi-objective retrieval and ranking to contribute to the community.

cs.IR↗

TJ4DRadSet: A 4D Radar Dataset for Autonomous Driving

The next-generation high-resolution automotive radar (4D radar) can provide additional elevation measurement and denser point clouds, which has great potential for 3D sensing in autonomous driving. In this paper, we introduce a dataset named TJ4DRadSet with 4D radar points for autonomous driving research. The dataset was collected in various driving scenarios, with a total of 7757 synchronized frames in 44 consecutive sequences, which are well annotated with 3D bounding boxes and track ids. We provide a 4D radar-based 3D object detection baseline for our dataset to demonstrate the effectiveness of deep learning methods for 4D radar point clouds. The dataset can be accessed via the following link: https://github.com/TJRadarLab/TJ4DRadSet.

cs.CV↗

A deep domain decomposition method based on Fourier features

In this paper we present a Fourier feature based deep domain decomposition method (F-D3M) for partial differential equations (PDEs). Currently, deep neural network based methods are actively developed for solving PDEs, but their efficiency can degenerate for problems with high frequency modes. In this new F-D3M strategy, overlapping domain decomposition is conducted for the spatial domain, such that high frequency modes can be reduced to relatively low frequency ones. In each local subdomain, multi Fourier feature networks (MFFNets) are constructed, where efficient boundary and interface treatments are applied for the corresponding loss functions. We present a general mathematical framework of F-D3M, validate its accuracy and demonstrate its efficiency with numerical experiments.

math.NA↗

Modeling User Behavior with Graph Convolution for Personalized Product Search

User preference modeling is a vital yet challenging problem in personalized product search. In recent years, latent space based methods have achieved state-of-the-art performance by jointly learning semantic representations of products, users, and text tokens. However, existing methods are limited in their ability to model user preferences. They typically represent users by the products they visited in a short span of time using attentive models and lack the ability to exploit relational information such as user-product interactions or item co-occurrence relations. In this work, we propose to address the limitations of prior arts by exploring local and global user behavior patterns on a user successive behavior graph, which is constructed by utilizing short-term actions of all users. To capture implicit user preference signals and collaborative patterns, we use an efficient jumping graph convolution to explore high-order relations to enrich product representations for user preference modeling. Our approach can be seamlessly integrated with existing latent space based methods and be potentially applied in any product retrieval method that uses purchase history to model user preferences. Extensive experiments on eight Amazon benchmarks demonstrate the effectiveness and potential of our approach. The source code is available at \url{https://github.com/floatSDSDS/SBG}.

cs.IR↗

Research on Event Accumulator Settings for Event-Based SLAM

Event cameras are a new type of sensors that are different from traditional cameras. Each pixel is triggered asynchronously by event. The trigger event is the change of the brightness irradiated on the pixel. If the increment or decrement of brightness is higher than a certain threshold, an event is output. Compared with traditional cameras, event cameras have the advantages of high dynamic range and no motion blur. Accumulating events to frames and using traditional SLAM algorithm is a direct and efficient way for event-based SLAM. Different event accumulator settings, such as slice method of event stream, processing method for no motion, using polarity or not, decay function and event contribution, can cause quite different accumulating results. We conducted the research on how to accumulate event frames to achieve a better event-based SLAM performance. For experiment verification, accumulated event frames are fed to the traditional SLAM system to construct an event-based SLAM system. Our strategy of setting event accumulator has been evaluated on the public dataset. The experiment results show that our method can achieve better performance in most sequences compared with the state-of-the-art event frame based SLAM algorithm. In addition, the proposed approach has been tested on a quadrotor UAV to show the potential of applications in real scenario. Code and results are open sourced to benefit the research community of event cameras.

cs.RO↗

Adversarial Learning for Incentive Optimization in Mobile Payment Marketing

Many payment platforms hold large-scale marketing campaigns, which allocate incentives to encourage users to pay through their applications. To maximize the return on investment, incentive allocations are commonly solved in a two-stage procedure. After training a response estimation model to estimate the users' mobile payment probabilities (MPP), a linear programming process is applied to obtain the optimal incentive allocation. However, the large amount of biased data in the training set, generated by the previous biased allocation policy, causes a biased estimation. This bias deteriorates the performance of the response model and misleads the linear programming process, dramatically degrading the performance of the resulting allocation policy. To overcome this obstacle, we propose a bias correction adversarial network. Our method leverages the small set of unbiased data obtained under a full-randomized allocation policy to train an unbiased model and then uses it to reduce the bias with adversarial learning. Offline and online experimental results demonstrate that our method outperforms state-of-the-art approaches and significantly improves the performance of the resulting allocation policy in a real-world marketing campaign.

cs.LG↗

On-Demand Valet Charging for Electric Vehicles: Economic Equilibrium, Infrastructure Planning and Regulatory Incentives

Many city residents cannot install their private electric vehicle (EV) chargers due to the lack of dedicated parking spaces or insufficient grid capacity. This presents a significant barrier towards large-scale EV adoption. To address this concern, this paper considers a novel business model, on-demand valet charging, that unlocks the potential of under-utilized public charging infrastructure to promise higher EV penetration. In the proposed model, a platform recruits a fleet of couriers that shuttle between customers and public charging stations to provide on-demand valet charging services to EV owners at an affordable price. Couriers are dispatched to pick up low-battery EVs from customers, deliver the EVs to charging stations, plug them in, and then return the fully-charged EVs to customers. To depict the proposed business model, we develop a queuing network to represent the stochastic matching dynamics, and further formulate an economic equilibrium model to capture the incentives of couriers, customers as well as the platform. These models are used to examine how charging infrastructure planning and regulatory intervention will affect the market outcome. First, we find that the optimal charging station densities for distinct stakeholders are different: couriers prefer a lower density; the platform prefers a higher density; while the density in-between leads to the highest EV penetration as it balances the time traveling to and queuing at charging stations. Second, we evaluate a regulatory policy that imposes a tax on the platform and invests the tax revenue in public charging infrastructure. Numerical results suggest that this regulation can suppress the platform's market power associated with monopoly pricing, increase social welfare, and facilitate the market expansion.

math.OC↗

Spatial Pricing in Ride-Sourcing Markets under a Congestion Charge

This paper studies the optimal spatial pricing for a ride-sourcing platform subject to a congestion charge. The platform determines the ride prices over the transportation network to maximize its profit, while the regulatory agency imposes the congestion charge to reduce traffic congestion in the urban core. A network economic equilibrium model is proposed to capture the intimate interactions among passenger demand, driver supply, passenger and driver waiting times, platform pricing, vehicle repositioning and flow balance over the transportation network. The overall optimal pricing problem is cast as a non-convex program. An algorithm is proposed to approximately compute its optimal solution, and a tight upper bound is established to evaluate its performance loss with respect to the globally optimal solution. Using the proposed model, we compare the impacts of three forms of congestion charge: (a) a one-directional cordon charge on ride-sourcing vehicles that enter the congestion area; (b) a bi-directional cordon charge on ride-sourcing vehicles that enter or exit the congestion area; (c) a trip-based congestion charge on all ride-sourcing trips. We show that the one-directional congestion charge not only reduces the ride-sourcing traffic in the congestion area, but also reduces the travel cost outside the congestion zone and benefits passengers in these underserved areas. We further show that compared to other congestion charges, the one-directional cordon charge is more effective in congestion mitigation: to achieve the same congestion-mitigation target, it imposes a smaller cost on passengers, drivers, and the platform. On the other hand, compared with the other charges, the trip-based congestion charge is more effective in revenue-raising: to raise the same tax revenue, it leads to a smaller loss to passengers, drivers, and the platform.

math.OC↗

Embedding-based Product Retrieval in Taobao Search

Nowadays, the product search service of e-commerce platforms has become a vital shopping channel in people's life. The retrieval phase of products determines the search system's quality and gradually attracts researchers' attention. Retrieving the most relevant products from a large-scale corpus while preserving personalized user characteristics remains an open question. Recent approaches in this domain have mainly focused on embedding-based retrieval (EBR) systems. However, after a long period of practice on Taobao, we find that the performance of the EBR system is dramatically degraded due to its: (1) low relevance with a given query and (2) discrepancy between the training and inference phases. Therefore, we propose a novel and practical embedding-based product retrieval model, named Multi-Grained Deep Semantic Product Retrieval (MGDSPR). Specifically, we first identify the inconsistency between the training and inference stages, and then use the softmax cross-entropy loss as the training objective, which achieves better performance and faster convergence. Two efficient methods are further proposed to improve retrieval relevance, including smoothing noisy training data and generating relevance-improving hard negative samples without requiring extra knowledge and training procedures. We evaluate MGDSPR on Taobao Product Search with significant metrics gains observed in offline experiments and online A/B tests. MGDSPR has been successfully deployed to the existing multi-channel retrieval system in Taobao Search. We also introduce the online deployment scheme and share practical lessons of our retrieval system to contribute to the community.

cs.IR↗

Interpreting and Boosting Dropout from a Game-Theoretic View

This paper aims to understand and improve the utility of the dropout operation from the perspective of game-theoretic interactions. We prove that dropout can suppress the strength of interactions between input variables of deep neural networks (DNNs). The theoretic proof is also verified by various experiments. Furthermore, we find that such interactions were strongly related to the over-fitting problem in deep learning. Thus, the utility of dropout can be regarded as decreasing interactions to alleviate the significance of over-fitting. Based on this understanding, we propose an interaction loss to further improve the utility of dropout. Experimental results have shown that the interaction loss can effectively improve the utility of dropout and boost the performance of DNNs.

cs.LG↗

Impact of Congestion Charge and Minimum Wage on TNCs: A Case Study for San Francisco

This paper describes the impact on transportation network companies (TNCs) of the imposition of a congestion charge and a driver minimum wage. The impact is assessed using a market equilibrium model to calculate the changes in the number of passenger trips and trip fare, number of drivers employed, the TNC platform profit, the number of TNC vehicles, and city revenue. Two charges are considered: (a) a charge per TNC trip similar to an excise tax, and (b) a charge per vehicle operating hour (whether or not it has a passenger) similar to a road tax. Both charges reduce the number of TNC trips, but this reduction is limited by the wage floor, and the number of TNC vehicles reduced is not significant. The time-based charge is preferable to the trip-based charge since, by penalizing idle vehicle time, the former increases vehicle occupancy. In a case study for San Francisco, the time-based charge is found to be Pareto superior to the trip-based charge as it yields higher passenger surplus, higher platform profits, and higher tax revenue for the city.

econ.EM↗

Selling Demand Response Using Options

Wholesale electricity markets in many jurisdictions use a two-settlement structure: a day-ahead market for bulk power transactions and a real-time market for fine-grain supply-demand balancing. This paper explores trading demand response assets within this two-settlement market structure. We consider two approaches for trading demand response assets: (a) an intermediate spot market with contingent pricing, and (b) an over-the-counter options contract. In the first case, we characterize the competitive equilibrium of the spot market, and show that it is socially optimal. Economic orthodoxy advocates spot markets, but these require expensive infrastructure and regulatory blessing. In the second case, we characterize competitive equilibria and compare its efficiency with the idealized spot market. Options contract are private bilateral over-the-counter transactions and do not require regulatory approval. We show that the optimal social welfare is, in general, not supported. We then design optimal option prices that minimize the social welfare gap. This optimal design serves to approximate the ideal spot market for demand response using options with modest loss of efficiency. Our results are validated through numerical simulations.

eess.SY↗

Off-Street Parking for TNC Vehicles to Reduce Cruising Traffic

This paper considers off-street parking for the cruising vehicles of transportation network companies (TNCs) to reduce the traffic congestion. We propose a novel business that integrates the shared parking service into the TNC platform. In the proposed model, the platform (a) provides interfaces that connect passengers, drivers and garage operators (commercial or private garages); (b) determines the ride fare, driver payment, and parking rates; (c) matches passengers to TNC vehicles for ride-hailing services; and (d) matches vacant TNC vehicles to unoccupied parking garages to reduce the cruising cost. A queuing-theoretic model is proposed to capture the matching process of passengers, drivers, and parking garages. A market-equilibrium model is developed to capture the incentives of the passengers, drivers, and garage operators. An optimization-based model is formulated to capture the optimal pricing of the TNC platform. Through a realistic case study, we show that the proposed business model will offer a Pareto improvement that benefits all stakeholders, which leads to higher passenger surplus, higher drivers surplus, higher garage operator surplus, higher platform profit, and reduced traffic congestion.

math.OC↗

Spin orbit field in a physically defined p type MOS silicon double quantum dot

We experimentally and theoretically investigate the spin orbit (SO) field in a physically defined, p type metal oxide semiconductor double quantum dot in silicon. We measure the magnetic field dependence of the leakage current through the double dot in the Pauli spin blockade. A finite magnetic field lifts the blockade, with the lifting least effective when the external and SO fields are parallel. In this way, we find that the spin flip of a tunneling hole is due to a SO field pointing perpendicular to the double dot axis and almost fully out of the quantum well plane. We augment the measurements by a derivation of SO terms using group symmetric representations theory. It predicts that without in plane electric fields (a quantum well case), the SO field would be mostly within the plane, dominated by a sum of a Rashba and a Dresselhaus like term. We, therefore, interpret the observed SO field as originated in the electric fields with substantial in plane components.

cond-mat.mes-hall↗