SearcharxivSearch

arXiv subjects

Vikrant Vaze

Publications and source records attributed to Vikrant Vaze.

7 recordsLinked to original sources

Optimistic Policy Regularization

Deep reinforcement learning agents frequently suffer from premature convergence, where early entropy collapse causes the policy to discard exploratory behaviors before discovering globally optimal strategies. We introduce Optimistic Policy Regularization (OPR), a lightweight mechanism designed to preserve and reinforce historically successful trajectories during policy optimization. OPR maintains a dynamic buffer of high-performing episodes and biases learning toward these behaviors through directional log-ratio reward shaping and an auxiliary behavioral cloning objective. When instantiated on Proximal Policy Optimization (PPO), OPR substantially improves sample efficiency on the Arcade Learning Environment. Across 49 Atari games evaluated at the 10-million step benchmark, OPR achieves the highest score in 22 environments despite baseline methods being reported at the standard 50-million step horizon. Beyond arcade benchmarks, OPR also generalizes to the CAGE Challenge 2 cyber-defense environment, surpassing the competition-winning Cardiff agent while using the same PPO architecture. These results demonstrate that anchoring policy updates to empirically successful trajectories can improve both sample efficiency and final performance.

cs.LG

Strategic Cyber Defense via Reinforcement Learning-Guided Combinatorial Auctions

Cyber defense operations increasingly require long-term strategic planning under uncertainty and resource constraints. We propose a new use of combinatorial auctions for allocating defensive action bundles in a realistic cyber environment, using host-specific valuations derived from reinforcement learning (RL) Q-values. These Q-values encode long-term expected utility, allowing upstream planning. We train CAFormer, a differentiable Transformer-based auction mechanism, to produce allocations that are approximately incentive-compatible under misreporting. Rather than benchmarking against existing agents, we explore the qualitative and strategic properties of the learned mechanisms. Compared to oracle and heuristic allocations, our method achieves competitive revenue while offering robustness to misreporting. In addition, we find that allocation patterns correlate with adversarial and defensive activity, suggesting implicit alignment with operational priorities. Our results demonstrate the viability of auction-based planning in cyber defense and highlight the interpretability benefits of RL-derived value structures.

cs.GT

Physics-aware Truck and Drone Delivery Planning Using Optimization & Machine Learning

Combining an energy-efficient drone with a high-capacity truck for last-mile package delivery can benefit operators and customers by reducing delivery times and environmental impact. However, directly integrating drone flight dynamics into the combinatorially hard truck route planning problem is challenging. Simplified models that ignore drone flight physics can lead to suboptimal delivery plans. We propose an integrated formulation for the joint problem of truck route and drone trajectory planning and a new end-to-end solution approach that combines optimization and machine learning to generate high-quality solutions in practical online runtimes. Our solution method trains neural network predictors based on offline solutions to the drone trajectory optimization problem instances to approximate drone flight times, and uses these approximations to optimize the overall truck-and-drone delivery plan by augmenting an existing order-first-split-second heuristic. Our method explicitly incorporates key kinematics and energy equations in drone trajectory optimization, and thereby outperforms state-of-the-art benchmarks that ignore drone flight physics. Extensive experimentation using synthetic datasets and real-world case studies shows that the integration of drone trajectories into package delivery planning substantially improves system performance in terms of tour duration and drone energy consumption. Our modeling and computational framework can help delivery planners achieve annual savings worth millions of dollars while also benefiting the environment.

math.OC

Rural School Bus Routing and Scheduling

Long school bus rides adversely affect student performance and well-being. Rural school bus rides are particularly long, incentivizing parents to drive their children to school rather than to opt for the school bus. This in turn exacerbates the traffic congestion around schools, further compounding the problem of long bus rides, creating a vicious cycle. It also results in underutilized school buses and higher bus operating costs per rider. To address these challenges, this paper focuses on the design of rural school bus routes and schedules, a particularly challenging problem due to its unique operational complexities, including mixed loading and irregular road networks. We formalize a rural school bus routing and scheduling model that tackles these complexities while minimizing the total commute time of students. We develop an original road network-aware cluster-then-route heuristic that leverages our problem formulation to produce high-quality solutions. For real-world case studies, our approach outperforms status quo solutions by reducing the commute times of students by 22-25 %. Our solutions also make the school bus more attractive, helping address both the underutilization of school buses and the prevalence of car commutes. Our routing and scheduling approach can improve school bus use by 14-15 % and reduce car trips that induce congestion near schools by 10-13 %. Many rural school districts share the operational characteristics modeled in this study, including long bus rides, high operational expenditures, mixed loading, and a high proportion of car-based school commutes, suggesting the broad applicability of our approach. Ultimately, by reducing student travel times, increasing school bus utilization, and alleviating congestion near schools, our approach enables rural school district planners to address transportation-related barriers to student performance and well-being.

math.OC

Advancing Differentiable Economics: A Neural Network Framework for Revenue-Maximizing Combinatorial Auction Mechanisms

Differentiable economics, which uses neural networks as function approximators and gradient-based optimization in automated mechanism design (AMD), marked a significant breakthrough with the introduction of RegretNet \citep{regretnet_paper}. It combines the flexibility of deep learning with a regret-based approach to relax incentive compatibility, allowing for approximations of revenue-maximizing auctions. However, applying these techniques to combinatorial auctions (CAs) - where bidders value bundles rather than individual items, capturing item interdependencies - remains a challenge, primarily due to the lack of methodologies that can effectively deal with combinatorial constraints. To tackle this, we propose two architectures: CANet, a fully connected neural network, and CAFormer, a transformer-based model designed to learn optimal randomized mechanisms. Unlike existing methods in traditional AMD, our approach is more scalable and free of assumptions about the structures of allowable bundles or bidder valuations. We demonstrate that our models match current methods in non-combinatorial settings and set new benchmarks for CAs. Specifically, our models consistently outperform benchmark mechanisms derived from heuristic approaches and provide empirical solutions where analytical results are unavailable. This work bridges the gap in applying differentiable economics to combinatorial auctions, offering a scalable and flexible framework for designing revenue-maximizing mechanisms.

cs.GT

Operational Research: Methods and Applications

Throughout its history, Operational Research has evolved to include a variety of methods, models and algorithms that have been applied to a diverse and wide range of contexts. This encyclopedic article consists of two main sections: methods and applications. The first aims to summarise the up-to-date knowledge and provide an overview of the state-of-the-art methods and key developments in the various subdomains of the field. The second offers a wide-ranging list of areas where Operational Research has been applied. The article is meant to be read in a nonlinear fashion. It should be used as a point of reference or first-port-of-call for a diverse pool of readers: academics, researchers, students, and practitioners. The entries within the methods and applications sections are presented in alphabetical order. The authors dedicate this paper to the 2023 Turkey/Syria earthquake victims. We sincerely hope that advances in OR will play a role towards minimising the pain and suffering caused by this and future catastrophes.

math.OC

Multimodal Transportation Pricing Alliance Design: Large-Scale Optimization for Rapid Gains

Transit agencies have the opportunity to outsource certain services to established Mobility-on-Demand (MOD) providers. Such alliances can improve service quality, coverage, and ridership; reduce public sector costs and vehicular emissions; and integrate the passenger experience. To amplify the effectiveness of such alliances, we develop a fare-setting model that jointly optimizes fares and discounts across a multimodal network. We capture commuters' travel decisions with a discrete choice model, resulting in a large-scale, mixed-integer, non-convex optimization problem. To solve this challenging problem, we develop a two-stage decomposition with the pricing decisions in the first stage and a mixed-integer linear optimization of fare discounts and passengers' travel decisions in the second stage. To solve the decomposition, we develop a new solution approach combining tailored coordinate descent, parsimonious second-stage evaluations, and interpolations using special ordered sets. This approach, enhanced by acceleration techniques based on slanted traversal, randomization and warm-start, significantly outperforms algorithmic benchmarks. Different alliance priorities result in qualitatively different fare designs: flat fares decrease the total vehicle-miles traveled, while geographically-informed discounts improve passenger happiness. The model responds appropriately to equity-oriented and passenger-centric priorities, improving system utilization and lowering prices for low-income and long-distance commuters. Our profit allocation mechanism improves outcomes for both types of operators, thus incentivizing profit-oriented MOD operators to adopt transit priorities.

math.OC