SearcharxivSearch

arXiv subjects

Lele Zhang

Publications and source records attributed to Lele Zhang.

At least 19 recordsLinked to original sources

A decentralised forward-backward-type algorithm with network-independent heterogenous agent step sizes

Consider the problem of finding a zero of a finite sum of maximally monotone operators, where some operators are Lipschitz continuous and the rest are potentially set-valued. We propose a forward-backward-type algorithm for this problem suitable for decentralised implementation. In each iteration, agents evaluate a Lipschitz continuous operator and the resolvent of a potentially set-valued operator, and then communicate with neighbouring agents. Agents choose their step sizes independently using only local information, and the step size upper bound has no dependence on the communication graph. We demonstrate the potential advantages of the proposed algorithm with numerical results for min-max problems and aggregative games.

math.OC

An LLM-powered MILP modelling engine for workforce scheduling guided by expert knowledge

Formulating mathematical models from real-world decision problems is a core task in Operational Research, yet it typically requires considerable human expertise and effort, limiting practical application. Recent advances in large language models (LLMs) have sparked interest in automating this process from natural language descriptions. However, challenges including limited modelling expertise, dependence on large-scale training data, and hallucination affect the reliable application of LLMs in optimisation modelling. To address these challenges, we propose SMILO, an expert-knowledge-driven framework that integrates optimisation modelling expertise with LLMs to generate mixed-integer linear programming models. SMILO uses a three-stage architecture built on reusable modelling graphs and associated resources: identifying relevant modelling components, extracting instance-specific information using LLMs, and constructing models through expert-defined templates. This modular architecture separates information extraction from formula generation, enhancing modelling accuracy, transparency, and reproducibility. We demonstrate the implementation of our modelling framework using workforce scheduling problems spanning manufacturing, logistic, and service operations as illustrative cases. Experiments show that SMILO consistently generates correct models in 90% of test instances across five trials, outperforming one-step LLM baselines by at least 35%. This work offers a generalisable paradigm for integrating LLMs with expert knowledge across diverse decision-making contexts, advancing automation in optimisation modelling.

math.OC

Advancing Antiferromagnetic Nitrides via Metal Alloy Nitridation

Nitride materials, valued for their structural stability and exceptional physical properties, have garnered significant interest in both fundamental research and technological applications. The fabrication of high-quality nitride thin films is essential for advancing their use in microelectronics and spintronics. Yet, achieving single-crystal nitride thin films with excellent structural integrity remains a challenge. Here, we introduce a straightforward yet innovative metallic alloy nitridation technique for the synthesis of stable single-crystal nitride thin films. By subjecting metal alloy thin films to a controlled nitridation process, nitrogen atoms integrate into the lattice, driving structural transformations while preserving high epitaxial quality. Combining nanoscale magnetic imaging with a diamond nitrogen-vacancy (NV) probe, X-ray magnetic linear dichroism, and comprehensive transport measurements, we confirm that the nitridated films exhibit a robust antiferromagnetic character with a zero net magnetic moment. This work not only provides a refined and reproducible strategy for the fabrication of nitride thin films but also lays a robust foundation for exploring their burgeoning device applications.

cond-mat.mtrl-sci

Scalable Hierarchical Reinforcement Learning for Hyper Scale Multi-Robot Task Planning

To improve the efficiency of warehousing system and meet huge customer orders, we aim to solve the challenges of dimension disaster and dynamic properties in hyper scale multi-robot task planning (MRTP) for robotic mobile fulfillment system (RMFS). Existing research indicates that hierarchical reinforcement learning (HRL) is an effective method to reduce these challenges. Based on that, we construct an efficient multi-stage HRL-based multi-robot task planner for hyper scale MRTP in RMFS, and the planning process is represented with a special temporal graph topology. To ensure optimality, the planner is designed with a centralized architecture, but it also brings the challenges of scaling up and generalization that require policies to maintain performance for various unlearned scales and maps. To tackle these difficulties, we first construct a hierarchical temporal attention network (HTAN) to ensure basic ability of handling inputs with unfixed lengths, and then design multi-stage curricula for hierarchical policy learning to further improve the scaling up and generalization ability while avoiding catastrophic forgetting. Additionally, we notice that policies with hierarchical structure suffer from unfair credit assignment that is similar to that in multi-agent reinforcement learning, inspired of which, we propose a hierarchical reinforcement learning algorithm with counterfactual rollout baseline to improve learning performance. Experimental results demonstrate that our planner outperform other state-of-the-art methods on various MRTP instances in both simulated and real-world RMFS. Also, our planner can successfully scale up to hyper scale MRTP instances in RMFS with up to 200 robots and 1000 retrieval racks on unlearned maps while keeping superior performance over other methods.

cs.RO

Pedestrian Volume Prediction Using a Diffusion Convolutional Gated Recurrent Unit Model

Effective models for analysing and predicting pedestrian flow are important to ensure the safety of both pedestrians and other road users. These tools also play a key role in optimising infrastructure design and geometry and supporting the economic utility of interconnected communities. The implementation of city-wide automatic pedestrian counting systems provides researchers with invaluable data, enabling the development and training of deep learning applications that offer better insights into traffic and crowd flows. Benefiting from real-world data provided by the City of Melbourne pedestrian counting system, this study presents a pedestrian flow prediction model, as an extension of Diffusion Convolutional Grated Recurrent Unit (DCGRU) with dynamic time warping, named DCGRU-DTW. This model captures the spatial dependencies of pedestrian flow through the diffusion process and the temporal dependency captured by Gated Recurrent Unit (GRU). Through extensive numerical experiments, we demonstrate that the proposed model outperforms the classic vector autoregressive model and the original DCGRU across multiple model accuracy metrics.

cs.LG

Evolution and Efficiency in Neural Architecture Search: Bridging the Gap Between Expert Design and Automated Optimization

The paper provides a comprehensive overview of Neural Architecture Search (NAS), emphasizing its evolution from manual design to automated, computationally-driven approaches. It covers the inception and growth of NAS, highlighting its application across various domains, including medical imaging and natural language processing. The document details the shift from expert-driven design to algorithm-driven processes, exploring initial methodologies like reinforcement learning and evolutionary algorithms. It also discusses the challenges of computational demands and the emergence of efficient NAS methodologies, such as Differentiable Architecture Search and hardware-aware NAS. The paper further elaborates on NAS's application in computer vision, NLP, and beyond, demonstrating its versatility and potential for optimizing neural network architectures across different tasks. Future directions and challenges, including computational efficiency and the integration with emerging AI domains, are addressed, showcasing NAS's dynamic nature and its continued evolution towards more sophisticated and efficient architecture search methods.

cs.NE

FedEmb: A Vertical and Hybrid Federated Learning Algorithm using Network And Feature Embedding Aggregation

Federated learning (FL) is an emerging paradigm for decentralized training of machine learning models on distributed clients, without revealing the data to the central server. The learning scheme may be horizontal, vertical or hybrid (both vertical and horizontal). Most existing research work with deep neural network (DNN) modelling is focused on horizontal data distributions, while vertical and hybrid schemes are much less studied. In this paper, we propose a generalized algorithm FedEmb, for modelling vertical and hybrid DNN-based learning. The idea of our algorithm is characterised by higher inference accuracy, stronger privacy-preserving properties, and lower client-server communication bandwidth demands as compared with existing work. The experimental results show that FedEmb is an effective method to tackle both split feature & subject space decentralized problems, shows 0.3% to 4.2% inference accuracy improvement with limited privacy revealing for datasets stored in local clients, and reduces 88.9 % time complexity over vertical baseline method.

cs.LG

Sample-based Dynamic Hierarchical Transformer with Layer and Head Flexibility via Contextual Bandit

Transformer requires a fixed number of layers and heads which makes them inflexible to the complexity of individual samples and expensive in training and inference. To address this, we propose a sample-based Dynamic Hierarchical Transformer (DHT) model whose layers and heads can be dynamically configured with single data samples via solving contextual bandit problems. To determine the number of layers and heads, we use the Uniform Confidence Bound while we deploy combinatorial Thompson Sampling in order to select specific head combinations given their number. Different from previous work that focuses on compressing trained networks for inference only, DHT is not only advantageous for adaptively optimizing the underlying network architecture during training but also has a flexible network for efficient inference. To the best of our knowledge, this is the first comprehensive data-driven dynamic transformer without any additional auxiliary neural networks that implement the dynamic system. According to the experiment results, we achieve up to 74% computational savings for both training and inference with a minimal loss of accuracy.

cs.LG

Joint Detection Algorithm for Multiple Cognitive Users in Spectrum Sensing

Spectrum sensing technology is a crucial aspect of modern communication technology, serving as one of the essential techniques for efficiently utilizing scarce information resources in tight frequency bands. This paper first introduces three common logical circuit decision criteria in hard decisions and analyzes their decision rigor. Building upon hard decisions, the paper further introduces a method for multi-user spectrum sensing based on soft decisions. Then the paper simulates the false alarm probability and detection probability curves corresponding to the three criteria. The simulated results of multi-user collaborative sensing indicate that the simulation process significantly reduces false alarm probability and enhances detection probability. This approach effectively detects spectrum resources unoccupied during idle periods, leveraging the concept of time-division multiplexing and rationalizing the redistribution of information resources. The entire computation process relies on the calculation principles of power spectral density in communication theory, involving threshold decision detection for noise power and the sum of noise and signal power. It provides a secondary decision detection, reflecting the perceptual decision performance of logical detection methods with relative accuracy.

eess.SP

Synthesizing mixed-integer linear programming models from natural language descriptions

Numerous real-world decision-making problems can be formulated and solved using Mixed-Integer Linear Programming (MILP) models. However, the transformation of these problems into MILP models heavily relies on expertise in operations research and mathematical optimization, which restricts non-experts' accessibility to MILP. To address this challenge, we propose a framework for automatically formulating MILP models from unstructured natural language descriptions of decision problems, which integrates Large Language Models (LLMs) and mathematical modeling techniques. This framework consists of three phases: i) identification of decision variables, ii) classification of objective and constraints, and iii) finally, generation of MILP models. In this study, we present a constraint classification scheme and a set of constraint templates that can guide the LLMs in synthesizing a complete MILP model. After fine-tuning LLMs, our approach can identify and synthesize logic constraints in addition to classic demand and resource constraints. The logic constraints have not been studied in existing work. To evaluate the performance of the proposed framework, we extend the NL4Opt dataset with more problem descriptions and constraint types, and with the new dataset, we compare our framework with one-step model generation methods offered by LLMs. The experimental results reveal that with respect to the accuracies of generating the correct model, objective, and constraints, our method which integrates constraint classification and templates with LLMs significantly outperforms the others. The prototype system that we developed has a great potential to capture more constraints for more complex MILPs. It opens up opportunities for developing training tools for operations research practitioners and has the potential to be a powerful tool for automatic decision problem modeling and solving in practice.

math.OC

On the retrieval of forward-scattered waveforms from acoustic reflection and transmission data with the Marchenko equation

A Green's function in an acoustic medium can be retrieved from reflection data by solving a multidimensional Marchenko equation. This procedure requires a-priori knowledge of the initial focusing function, which can be interpreted as the inverse of a transmitted wavefield as it would propagate through the medium, excluding (multiply) reflected waveforms. In practice, the initial focusing function is often replaced by a time-reversed direct wave, which is computed with help of a macro velocity model. Green's functions that are retrieved under this (direct-wave) approximation typically lack forward-scattered waveforms and their associated multiple reflections. We examine whether this problem can be mitigated by incorporating transmission data. Based on these transmission data, we derive an auxiliary equation for the forward-scattered components of the initial focusing function. We demonstrate that this equation can be solved in an acoustic medium with mass density contrast and constant propagation velocity. By solving the auxiliary and Marchenko equation successively, we can include forward-scattered waveforms in our Green's function estimates, as we demonstrate with a numerical example.

physics.app-ph

A Restless Bandit Model for Dynamic Ride Matching with Reneging Travelers

This paper studies a large-scale ride-matching problem with a large number of travelers who are either drivers with vehicles or riders looking for sharing vehicles. Drivers can match riders that have similar itineraries and share the same vehicle; and reneging travelers, who become impatient and leave the service system after waiting a long time for shared rides, are considered in our model. The aim is to maximize the long-run average revenue of the ride service vendor, which is defined as the difference between the long-run average reward earned by providing ride services and the long-run average penalty incurred by reneging travelers. The problem is complicated by its scale, the heterogeneity of travelers (in terms of origins, destinations, and travel preferences), and the reneging behaviors. To this end, we formulate the ride-matching problem as a specific Markov decision process and propose a scalable ride-matching policy, referred to as Bivariate Index (BI) policy. The BI policy prioritizes travelers according to a ranking of their bivariate indices, which we prove, in a special case, leads to an optimal policy to the relaxed version of the ride-matching problem. For the general case, through extensive numerical simulations for systems with real-world travel demands, it is demonstrated that the BI policy significantly outperforms baseline policies.

math.OC

An overview of Marchenko methods

Since the introduction of the Marchenko method in geophysics, many variants have been developed. Using a compact unified notation, we review redatuming by multidimensional deconvolution and by double focusing, virtual seismology, double dereverberation and transmission-compensated Marchenko multiple elimination, and discuss the underlying assumptions, merits and limitations of these methods.

physics.geo-ph

A new role for adaptive filters in Marchenko equation-based methods for the attenuation of internal multiples

We have seen many developments in Marchenko equation-based methods for internal multiple attenuation in the past years. Starting from a wave-equation based method that required a smooth velocity model, there are now Marchenko equation-based methods that do not require any model information or user-input. In principle, these methods accurately predict internal multiples. Therefore, the role of the adaptive filter has changed for these methods. Rather than needing an aggressive adaptive filter to compensate for inaccurate internal multiple predictions, only a conservative adaptive filter is needed to compensate for minor amplitude and/or phase errors in the internal multiple predictions caused by imperfect acquisition and preprocessing of the input data. We demonstate that a conservative adaptive filter can be used to improve the attenuation of internal multiples when applying a Marchenko multiple elimination (MME) method to a 2D line of streamer data. In addition, we suggest that an adaptive filter can be used as a feedback mechanism to improve the preprocessing of the input data.

physics.geo-ph

Data-driven retrieval of primary plane-wave responses

Seismic images provided by reverse time migration can be contaminated by artefacts associated with the migration of multiples. Multiples can corrupt seismic images, producing both false positives, i.e. by focusing energy at unphysical interfaces, and false negatives, i.e. by destructively interfering with primaries. Multiple prediction / primary synthesis methods are usually designed to operate on point source gathers, and can therefore be computationally demanding when large problems are considered. A computationally attractive scheme that operates on plane-wave datasets is derived by adapting a data-driven point source gathers method, based on convolutions and cross-correlations of the reflection response with itself, to include plane-wave concepts. As a result, the presented algorithm allows fully data-driven synthesis of primary reflections associated with plane-wave source responses. Once primary plane-wave responses are estimated, they are used for multiple-free imaging via plane-wave reverse time migration. Numerical tests of increasing complexity demonstrate the potential of the proposed algorithm to produce multiple-free images from only a small number of plane-wave datasets.

physics.geo-ph

Behaviour of traffic on a link with traffic light boundaries

This paper considers a single link with traffic light boundary conditions at both ends, and investigates the traffic evolution over time with various signal and system configurations. A hydrodynamic model and a modified stochastic domain wall theory are proposed to describe the local density variation. The Nagel-Schreckenberg model (NaSch), an agent based stochastic model, is used as a benchmark. The hydrodynamic model provides good approximations over short time scales. The domain wall model is found to reproduce the time evolution of local densities, in good agreement with the NaSch simulations for both short and long time scales. A systematic investigation of the impact of network parameters, including system sizes, cycle lengths, phase splits and signal offsets, on traffic flows suggests that the stationary flow is dominated by the boundary with the smaller split. Nevertheless, the signal offset plays an important role in determining the flow. Analytical expressions of the flow in relation to those parameters are obtained for the deterministic domain wall model and match the deterministic NaSch simulations. The analytic results agree qualitatively with the general stochastic models. When the cycle is sufficiently short, the stationary state is governed by effective inflow and outflow rates, and the density profile is approximately linear and independent of time.

physics.soc-ph

Traffic disruption and recovery in road networks

We study the impact of disruptions on road networks, and the recovery process after the disruption is removed from the system. Such disruptions could be caused by vehicle breakdown or illegal parking. We analyze the transient behavior using domain wall theory, and compare these predictions with simulations of a stochastic cellular automaton model. We find that the domain wall model can reproduce the time evolution of flow and density during the disruption and the recovery processes, for both one-dimensional systems and two-dimensional networks.

nlin.CG

A Comparison of Tram Priority at Signalized Intersections

We study tram priority at signalized intersections using a stochastic cellular automaton model for multimodal traffic flow. We simulate realistic traffic signal systems, which include signal linking and adaptive cycle lengths and split plans, with different levels of tram priority. We find that tram priority can improve service performance in terms of both average travel time and travel time variability. We consider two main types of tram priority, which we refer to as full and partial priority. Full tram priority is able to guarantee service quality even when traffic is saturated, however, it results in significant costs to other road users. Partial tram priority significantly reduces tram delays while having limited impact on other traffic, and therefore achieves a better result in terms of the overall network performance. We also study variations in which the tram priority is only enforced when trams are running behind schedule, and we find that those variations retain almost all of the benefit for tram operations but with reduced negative impact on the network.

nlin.CG