Searcharxiv⌕ Search

arXiv subjects

Sen Li

Publications and source records attributed to Sen Li.

At least 55 records · Page 3Linked to original sources

Enhancing Cyber-Resilience in Integrated Energy System Scheduling with Demand Response Using Deep Reinforcement Learning

Optimally scheduling multi-energy flow is an effective method to utilize renewable energy sources (RES) and improve the stability and economy of integrated energy systems (IES). However, the stable demand-supply of IES faces challenges from uncertainties that arise from RES and loads, as well as the increasing impact of cyber-attacks with advanced information and communication technologies adoption. To address these challenges, this paper proposes an innovative model-free resilience scheduling method based on state-adversarial deep reinforcement learning (DRL) for integrated demand response (IDR)-enabled IES. The proposed method designs an IDR program to explore the interaction ability of electricity-gas-heat flexible loads. Additionally, the state-adversarial Markov decision process (SA-MDP) model characterizes the energy scheduling problem of IES under cyber-attack, incorporating cyber-attacks as adversaries directly into the scheduling process. The state-adversarial soft actor-critic (SA-SAC) algorithm is proposed to mitigate the impact of cyber-attacks on the scheduling strategy, integrating adversarial training into the learning process to against cyber-attacks. Simulation results demonstrate that our method is capable of adequately addressing the uncertainties resulting from RES and loads, mitigating the impact of cyber-attacks on the scheduling strategy, and ensuring a stable demand supply for various energy sources. Moreover, the proposed method demonstrates resilience against cyber-attacks. Compared to the original soft actor-critic (SAC) algorithm, it achieves a 10% improvement in economic performance under cyber-attack scenarios.

eess.SY↗

Physical Informed-Inspired Deep Reinforcement Learning Based Bi-Level Programming for Microgrid Scheduling

To coordinate the interests of operator and users in a microgrid under complex and changeable operating conditions, this paper proposes a microgrid scheduling model considering the thermal flexibility of thermostatically controlled loads and demand response by leveraging physical informed-inspired deep reinforcement learning (DRL) based bi-level programming. To overcome the non-convex limitations of karush-kuhn-tucker (KKT)-based methods, a novel optimization solution method based on DRL theory is proposed to handle the bi-level programming through alternate iterations between levels. Specifically, by combining a DRL algorithm named asynchronous advantage actor-critic (A3C) and automated machine learning-prioritized experience replay (AutoML-PER) strategy to improve the generalization performance of A3C to address the above problems, an improved A3C algorithm, called AutoML-PER-A3C, is designed to solve the upper-level problem; while the DOCPLEX optimizer is adopted to address the lower-level problem. In this solution process, AutoML is used to automatically optimize hyperparameters and PER improves learning efficiency and quality by extracting the most valuable samples. The test results demonstrate that the presented approach manages to reconcile the interests between multiple stakeholders in MG by fully exploiting various flexibility resources. Furthermore, in terms of economic viability and computational efficiency, the proposal vastly exceeds other advanced reinforcement learning methods.

eess.SY↗

A Watermark-Conditioned Diffusion Model for IP Protection

The ethical need to protect AI-generated content has been a significant concern in recent years. While existing watermarking strategies have demonstrated success in detecting synthetic content (detection), there has been limited exploration in identifying the users responsible for generating these outputs from a single model (owner identification). In this paper, we focus on both practical scenarios and propose a unified watermarking framework for content copyright protection within the context of diffusion models. Specifically, we consider two parties: the model provider, who grants public access to a diffusion model via an API, and the users, who can solely query the model API and generate images in a black-box manner. Our task is to embed hidden information into the generated contents, which facilitates further detection and owner identification. To tackle this challenge, we propose a Watermark-conditioned Diffusion model called WaDiff, which manipulates the watermark as a conditioned input and incorporates fingerprinting into the generation process. All the generative outputs from our WaDiff carry user-specific information, which can be recovered by an image extractor and further facilitate forensic identification. Extensive experiments are conducted on two popular diffusion models, and we demonstrate that our method is effective and robust in both the detection and owner identification tasks. Meanwhile, our watermarking framework only exerts a negligible impact on the original generation and is more stealthy and efficient in comparison to existing watermarking strategies.

cs.CR↗

A Two-Stage Online Algorithm for EV Charging Station Energy Management and Carbon Trading

The increasing electric vehicle (EV) adoption challenges the energy management of charging stations (CSs) due to the large number of EVs and the underlying uncertainties. Moreover, the carbon footprint of CSs is growing significantly due to the rising charging power demand. This makes it important for CSs to properly manage their energy usage and ensure their carbon footprint stay within their carbon emission quotas. This paper proposes a two-stage online algorithm for this purpose, considering the different time scales of energy management and carbon trading. In the first stage, the CS characterizes the real-time aggregate EV power flexibility, in terms of upper and lower bounds on the total charging power, by a Lyapunov optimization-based online algorithm. In the second stage, the CS co-optimizes energy management and carbon trading, with EV charging power chosen within the aggregate flexibility region provided by the first stage. A generalized battery model is proposed to capture the dynamic carbon footprint changes and carbon trading. A virtual carbon queue is designed to develop an online algorithm for the second stage, which can ensure the carbon footprint of CS be within its carbon emission quota and its total operation cost is nearly offline optimal. Case studies validate the effectiveness and advantages of the proposed algorithm.

math.OC↗

Invisible Backdoor Attacks on Diffusion Models

In recent years, diffusion models have achieved remarkable success in the realm of high-quality image generation, garnering increased attention. This surge in interest is paralleled by a growing concern over the security threats associated with diffusion models, largely attributed to their susceptibility to malicious exploitation. Notably, recent research has brought to light the vulnerability of diffusion models to backdoor attacks, enabling the generation of specific target images through corresponding triggers. However, prevailing backdoor attack methods rely on manually crafted trigger generation functions, often manifesting as discernible patterns incorporated into input noise, thus rendering them susceptible to human detection. In this paper, we present an innovative and versatile optimization framework designed to acquire invisible triggers, enhancing the stealthiness and resilience of inserted backdoors. Our proposed framework is applicable to both unconditional and conditional diffusion models, and notably, we are the pioneers in demonstrating the backdooring of diffusion models within the context of text-guided image editing and inpainting pipelines. Moreover, we also show that the backdoors in the conditional generation can be directly applied to model watermarking for model ownership verification, which further boosts the significance of the proposed framework. Extensive experiments on various commonly used samplers and datasets verify the efficacy and stealthiness of the proposed framework. Our code is publicly available at https://github.com/invisibleTriggerDiffusion/invisible_triggers_for_diffusion.

cs.LG↗

MuLan: Multimodal-LLM Agent for Progressive and Interactive Multi-Object Diffusion

Existing text-to-image models still struggle to generate images of multiple objects, especially in handling their spatial positions, relative sizes, overlapping, and attribute bindings. To efficiently address these challenges, we develop a training-free Multimodal-LLM agent (MuLan), as a human painter, that can progressively generate multi-object with intricate planning and feedback control. MuLan harnesses a large language model (LLM) to decompose a prompt to a sequence of sub-tasks, each generating only one object by stable diffusion, conditioned on previously generated objects. Unlike existing LLM-grounded methods, MuLan only produces a high-level plan at the beginning while the exact size and location of each object are determined upon each sub-task by an LLM and attention guidance. Moreover, MuLan adopts a vision-language model (VLM) to provide feedback to the image generated in each sub-task and control the diffusion model to re-generate the image if it violates the original prompt. Hence, each model in every step of MuLan only needs to address an easy sub-task it is specialized for. The multi-step process also allows human users to monitor the generation process and make preferred changes at any intermediate step via text prompts, thereby improving the human-AI collaboration experience. We collect 200 prompts containing multi-objects with spatial relationships and attribute bindings from different benchmarks to evaluate MuLan. The results demonstrate the superiority of MuLan in generating multiple objects over baselines and its creativity when collaborating with human users. The code is available at https://github.com/measure-infinity/mulan-code.

cs.CV↗

Towards a Multimodal Charging Network: Joint Planning of Charging Stations and Battery Swapping Stations for Electrified Ride-Hailing Fleets

This paper considers a multimodal charging network in which charging stations and battery swapping stations are jointly built to support an electric ride-hailing fleet synergistically. Our argument is based on the observation that charging an EV is a time-consuming burden, and battery swapping faces scaling issues due to its deployment costs. However, charging stations are cost-effective, making them ideal for scaling up EV fleets, while battery swapping stations offer quick turnaround and can be deployed in tandem with charging stations to improve fleet utilization and reduce operational costs. To fulfill this vision, we consider a ride-hailing platform that jointly builds charging and battery swapping stations to support an EV fleet. An optimization model is proposed to capture the platform's planning and operational decisions. In particular, the model incorporates essential components such as elastic passenger demand, spatial charging equilibrium, charging and swapping congestion, etc. The overall problem is formulated as a nonconcave program. Instead of pursuing the globally optimal solution, we establish a tight upper bound through relaxation and decomposition, allowing us to evaluate the solution optimality even in the absence of concavity. Through case studies for Manhattan, New York City, we find that joint planning of charging and battery swapping stations outperforms deploying only one of them, yielding a total profit that is 11.7% higher than swapping-only deployment under a limited budget, and 17.5% higher than charging-only deployment under a sufficient budget. These results underscore the complementary benefit between charging and battery swapping facilities.

math.OC↗

Task-Space Riccati Feedback based Whole Body Control for Underactuated Legged Locomotion

This manuscript primarily aims to enhance the performance of whole-body controllers(WBC) for underactuated legged locomotion. We introduce a systematic parameter design mechanism for the floating-base feedback control within the WBC. The proposed approach involves utilizing the linearized model of unactuated dynamics to formulate a Linear Quadratic Regulator(LQR) and solving a Riccati gain while accounting for potential physical constraints through a second-order approximation of the log-barrier function. And then the user-tuned feedback gain for the floating base task is replaced by a new one constructed from the solved Riccati gain. Extensive simulations conducted in MuJoCo with a point bipedal robot, as well as real-world experiments performed on a quadruped robot, demonstrate the effectiveness of the proposed method. In the different bipedal locomotion tasks, compared with the user-tuned method, the proposed approach is at least 12% better and up to 50% better at linear velocity tracking, and at least 7% better and up to 47% better at angular velocity tracking. In the quadruped experiment, linear velocity tracking is improved by at least 3% and angular velocity tracking is improved by at least 23% using the proposed method.

cs.RO↗

SeisT: A foundational deep learning model for earthquake monitoring tasks

Seismograms, the fundamental seismic records, have revolutionized earthquake research and monitoring. Recent advancements in deep learning have further enhanced seismic signal processing, leading to even more precise and effective earthquake monitoring capabilities. This paper introduces a foundational deep learning model, the Seismogram Transformer (SeisT), designed for a variety of earthquake monitoring tasks. SeisT combines multiple modules tailored to different tasks and exhibits impressive out-of-distribution generalization performance, outperforming or matching state-of-the-art models in tasks like earthquake detection, seismic phase picking, first-motion polarity classification, magnitude estimation, back-azimuth estimation, and epicentral distance estimation. The performance scores on the tasks are 0.96, 0.96, 0.68, 0.95, 0.86, 0.55, and 0.81, respectively. The most significant improvements, in comparison to existing models, are observed in phase-P picking, phase-S picking, and magnitude estimation, with gains of 1.7%, 9.5%, and 8.0%, respectively. Our study, through rigorous experiments and evaluations, suggests that SeisT has the potential to contribute to the advancement of seismic signal processing and earthquake research.

physics.geo-ph↗

Regulating Transportation Network Companies with a Mixture of Autonomous Vehicles and For-Hire Human Drivers

This paper investigates the equity impacts of autonomous vehicles (AV) on for-hire human drivers and passengers in a ride-hailing market, and examines regulation policies that protect human drivers and improve transport equity for ride-hailing passengers. We consider a transportation network companies (TNC) that employs a mixture of AVs and human drivers to provide ride-hailing services. The TNC platform determines the spatial prices, fleet size, human driver payments, and vehicle relocation strategies to maximize its profit, while individual passengers choose between different transport modes to minimize their travel costs. A market equilibrium model is proposed to capture the interactions among passengers, human drivers, AVs, and TNC over the transportation network. The overall problem is formulated as a non-concave program, and an algorithm is developed to derive its approximate solution with a theoretical performance guarantee. Our study shows that TNC prioritizes AV deployment in higher-demand areas to make a higher profit. As AVs flood into these higher-demand areas, they compete with human drivers in the urban core and push them to relocate to suburbs. This leads to reduced earning opportunities for human drivers and increased spatial inequity for passengers. To mitigate these concerns, we consider: (a) a minimum wage for human drivers; and (b) a restrictive pickup policy that prohibits AVs from picking up passengers in higher-demand areas. In the former case, we show that a minimum wage for human drivers will protect them from the negative impact of AVs with negligible impacts on passengers. However, there exists a threshold beyond which the minimum wage will trigger the platform to replace the majority of human drivers with AVs.

math.OC↗

Spatiotemporal Pricing and Fleet Management of Autonomous Mobility-on-Demand Networks: A Decomposition and Dynamic Programming Approach with Bounded Optimality Gap

This paper studies spatiotemporal pricing and fleet management for autonomous mobility-on-demand (AMoD) systems while taking elastic demand into account. We consider a platform that offers ride-hailing services using a fleet of autonomous vehicles and makes pricing, rebalancing, and fleet sizing decisions in response to demand fluctuations. A network flow model is developed to characterize the evolution of system states over space and time, which captures the vehicle-passenger matching process and demand elasticity with respect to price and waiting time. The platform's objective of maximizing profit is formulated as a constrained optimal control problem, which is highly nonconvex due to the nonlinear demand model and complex supply-demand interdependence. To address this challenge, an integrated decomposition and dynamic programming approach is proposed, where we first relax the problem through a change of variable, then separate the relaxed problem into a few small-scale subproblems via dual decomposition, and finally solve each subproblem using dynamic programming. Despite the nonconvexity, our approach establishes a theoretical upper bound to evaluate the solution optimality. The proposed model and methodology are validated in numerical studies for Manhattan. We find that compared to the benchmark case, the proposed upper bound is significantly tighter. We also find that compared to pricing alone, joint pricing and fleet rebalancing can only offer a minor profit improvement when demand can be accurately predicted. However, during unanticipated demand surges, joint pricing and rebalancing can lead to substantially improved profits, and the impacts of demand shocks, despite being more widespread, can dissipate faster.

math.OC↗

Active Surface with Passive Omni-Directional Adaptation of Soft Polyhedral Fingers for In-Hand Manipulation

Track systems effectively distribute loads, augmenting traction and maneuverability on unstable terrains, leveraging their expansive contact areas. This tracked locomotion capability also aids in hand manipulation of not only regular objects but also irregular objects. In this study, we present the design of a soft robotic finger with an active surface on an omni-adaptive network structure, which can be easily installed on existing grippers and achieve stability and dexterity for in-hand manipulation. The system's active surfaces initially transfer the object from the fingertip segment with less compliance to the middle segment of the finger with superior adaptability. Despite the omni-directional deformation of the finger, in-hand manipulation can still be executed with controlled active surfaces. We characterized the soft finger's stiffness distribution and simplified models to assess the feasibility of repositioning and reorienting a grasped object. A set of experiments on in-hand manipulation was performed with the proposed fingers, demonstrating the dexterity and robustness of the strategy.

cs.RO↗

Charging Autonomous Electric Vehicle Fleet for Mobility-on-Demand Services: Plug in or Swap out?

This paper compares two prevalent charging strategies for electric vehicles, plug-in charging and battery swapping, to investigate which charging strategy is superior for electric autonomous mobility-on-demand (AMoD) systems. To this end, we use a queueing-theoretic model to characterize the vehicle waiting time at charging stations and battery swapping stations, respectively. The model is integrated into an economic analysis of the electric AMoD system operated by a transportation network company (TNC), where the incentives of passengers, the charging/operating shift of TNC vehicles, the operational decisions of the platform, and the planning decisions of the government are captured. Overall, a bi-level optimization framework is proposed for charging infrastructure planning of the electric AMoD system. Based on the proposed framework, we compare the socio-economic performance of plug-in charging and battery swapping, and investigate how this comparison depends on the evolving charging technologies (such as charging speed, battery capacity, and infrastructure cost). At the planning level, we find that when choosing plug-in charging, increased charging speed leads to a transformation of infrastructure from sparsely distributed large stations to densely distributed small stations, while enlarged battery capacity transforms the infrastructure from densely distributed small stations to sparsely distributed large stations. On the other hand, when choosing battery swapping, both increased charging speed and enlarged battery capacity will lead to a smaller number of battery swapping stations. At the operational level, we find that improved charging speed leads to increased TNC profit when choosing plug-in charging, whereas improved charging speed may lead to smaller TNC profit under battery swapping. The above insights are validated through realistic numerical studies.

math.OC↗

Regulating For-Hire Autonomous Vehicles for An Equitable Multimodal Transportation Network

This paper assesses the equity impacts of for-hire autonomous vehicles (AVs) and investigates regulatory policies that promote spatial and social equity in future autonomous mobility ecosystems. To this end, we consider a multimodal transportation network, where a ride-hailing platform operates a fleet of AVs to offer mobility-on-demand services in competition with a public transit agency that offers transit services on a transportation network. A game-theoretic model is developed to characterize the intimate interactions between the ride-hailing platform, the transit agency, and multiclass passengers with distinct income levels. An algorithm is proposed to compute the Nash equilibrium of the game and conduct an ex-post evaluation of the performance of the obtained solution. Based on the proposed framework, we evaluate the spatial and social equity in transport accessibility using the Theil index, and find that although the proliferation of for-hire AVs in the ride-hailing network improves overall accessibility, the benefits are not fairly distributed among distinct locations or population groups, implying that the deployment of AVs will enlarge the existing spatial and social inequity gaps in the transportation network if no regulatory intervention is in place. To address this concern, we investigate two regulatory policies that can improve transport equity: (a) a minimum service-level requirement on ride-hailing services, which improves the spatial equity in the transport network; (b) a subsidy on transit services by taxing ride-hailing services, which promotes the use of public transit and improves the spatial and social equity of the transport network. We show that the minimum service-level requirement entails a trade-off: as a higher minimum service level is imposed, the spatial inequity reduces, but the social inequity will be exacerbated. On the other hand ...

math.OC↗

Piggyback on Idle Ride-Sourcing Drivers for Integrated On-Demand and Flexible Intracity Parcel Delivery Services

This paper investigates the spatial pricing and fleet management strategies for an integrated platform that provides both ride-sourcing services and intracity parcel delivery services over a transportation network utilizing the idle time of ride-sourcing drivers. Specifically, the integrated platform simultaneously offers on-demand ride-sourcing services for passengers and multiple modes of parcel delivery services for customers, including: (1) on-demand delivery, where drivers immediately pick up and deliver parcels upon receiving a delivery request; and (2) flexible delivery, where drivers can pick up (or drop off) parcels only when they are idle and waiting for the next ride-sourcing request. A continuous-time Markov Chain (CTMC) model is proposed to characterize the status change of drivers under joint movement of passengers and parcels over the transportation network with limited vehicle capacity, where the service quality of ride-sourcing services, on-demand delivery services, and flexible delivery services are rigorously quantified. Building on the CTMC model, incentives for ride-sourcing passengers, delivery customers, drivers, and the platform are captured through an economic equilibrium model, and the optimal spatial pricing decisions of the platform are derived by solving a non-convex profit-maximizing problem. We prove the well-posedness of the model and develop a tailored algorithm to compute the optimal decisions of the platform. Furthermore, we validate the proposed model in a comprehensive case study for San Francisco, demonstrating that joint management of ride-sourcing services and intracity package delivery services can lead to a Pareto improvement that benefits all stakeholders in the integrated ride-sourcing and parcel delivery market under realistic parcel and passenger demand patterns.

eess.SY↗

Graph Contrastive Learning with Multi-Objective for Personalized Product Retrieval in Taobao Search

In e-commerce search, personalized retrieval is a crucial technique for improving user shopping experience. Recent works in this domain have achieved significant improvements by the representation learning paradigm, e.g., embedding-based retrieval (EBR) and collaborative filtering (CF). EBR methods do not sufficiently exploit the useful collaborative signal and are difficult to learn the representations of long-tail item well. Graph-based CF methods improve personalization by modeling collaborative signal within the user click graph. However, existing Graph-based methods ignore user's multiple behaviours, such as click/purchase and the relevance constraint between user behaviours and items.In this paper, we propose a Graph Contrastive Learning with Multi-Objective (GCL-MO) collaborative filtering model, which solves the problems of weak relevance and incomplete personalization in e-commerce search. Specifically, GCL-MO builds a homogeneous graph of items and then optimizes a multi-objective function of personalization and relevance. Moreover, we propose a modified contrastive loss for multi-objectives graph learning, which avoids the mutual suppression among positive samples and thus improves the generalization and robustness of long-tail item representations. These learned item embeddings are then used for personalized retrieval by constructing an efficient offline-to-online inverted table. GCL-MO outperforms the online collaborative filtering baseline in both offline/online experimental metrics and shows a significant improvement in the online A/B testing of Taobao search.

cs.IR↗

Emotional Talking Head Generation based on Memory-Sharing and Attention-Augmented Networks

Given an audio clip and a reference face image, the goal of the talking head generation is to generate a high-fidelity talking head video. Although some audio-driven methods of generating talking head videos have made some achievements in the past, most of them only focused on lip and audio synchronization and lack the ability to reproduce the facial expressions of the target person. To this end, we propose a talking head generation model consisting of a Memory-Sharing Emotion Feature extractor (MSEF) and an Attention-Augmented Translator based on U-net (AATU). Firstly, MSEF can extract implicit emotional auxiliary features from audio to estimate more accurate emotional face landmarks.~Secondly, AATU acts as a translator between the estimated landmarks and the photo-realistic video frames. Extensive qualitative and quantitative experiments have shown the superiority of the proposed method to the previous works. Codes will be made publicly available.

cs.CV↗

MAVD: The First Open Large-Scale Mandarin Audio-Visual Dataset with Depth Information

Audio-visual speech recognition (AVSR) gains increasing attention from researchers as an important part of human-computer interaction. However, the existing available Mandarin audio-visual datasets are limited and lack the depth information. To address this issue, this work establishes the MAVD, a new large-scale Mandarin multimodal corpus comprising 12,484 utterances spoken by 64 native Chinese speakers. To ensure the dataset covers diverse real-world scenarios, a pipeline for cleaning and filtering the raw text material has been developed to create a well-balanced reading material. In particular, the latest data acquisition device of Microsoft, Azure Kinect is used to capture depth information in addition to the traditional audio signals and RGB images during data acquisition. We also provide a baseline experiment, which could be used to evaluate the effectiveness of the dataset. The dataset and code will be released at https://github.com/SpringHuo/MAVD.

cs.SD↗