SearcharxivSearch

arXiv subjects

Maximilian Schiffer

Publications and source records attributed to Maximilian Schiffer.

At least 19 recordsLinked to original sources

Linear Decision Tree Policies for Integer Linear Programs

We study optimal decision policies, represented as linear decision trees, for integer linear programs with a fixed feasible set and varying cost vectors. Once synthesized for a given feasible set, they return an optimal solution for any queried cost vector through a sequence of linear tests. We show that there exists a policy performing this operation in a polynomial number of arithmetic operations in the worst case. In contrast, deciding whether there exists an exact policy with a prescribed maximum number of leaves is $Σ_2^p$-complete. Alongside these theoretical results, we develop a practical construction framework to synthesize policies within a specific subclass of linear decision trees. Our computational experiments show that, although policy synthesis can be time-intensive, it allows one to retrieve optimal solutions orders of magnitude faster than classical and specialized solution methods on repeated queries. Overall, this paradigm provides a different perspective on the solution of integer linear programs and offers a principled offline-online approach for repeated optimization.

math.OC

Dynamic capacity allocation of hybrid transportation units for cargo-hitching in urban public transportation systems

To improve the utilization of public transportation systems (PTSs) during off-peak hours, we present an algorithmic framework that designs PTSs with hybrid transportation units (HTUs), which can transport passengers or freight by leveraging a flexible interior. Against this background, we study a capacitated network design problem to enable cargo-hitching in existing PTSs. Specifically, we study a setting with fixed vehicle routes and timetables in which vehicles can be equipped with HTUs to enable cargo-hitching. We optimize the network design from a total cost perspective to account for normalized network design costs tied to the investment in HTUs and freight routing costs. We present an algorithmic framework that encodes some of the problem's constraints in a spatially and temporally expanded, layered graph, and solves the resulting network design problem with a price-and-branch algorithm. We apply this framework to a case study based on the subway network in the city of Munich. Our algorithm outscales commercial solvers by a factor of six and yields integer feasible solutions with a median integrality gap of less than 1.56% for all instances. We show that cargo-hitching with HTUs increases the utilization of PTSs, especially during off-peak hours, without cannibalizing passenger service level and quality. We quantify the value of hybrid transportation units (HTUs) at up to 3.2% of the total cost. Moreover, we present a sensitivity analysis that indicates that cargo-hitching is worthwhile if truck-based transport occurs at an externality cost of more than 1.5 EUR per vehicle and kilometer and loading and unloading costs of less than 2.0 EUR per passenger equivalent.

math.OC

Disentangling generalization and memorization in large language models using chess

Large Language Models (LLMs) exhibit remarkable capabilities, yet it remains unclear to what extent these reflect sophisticated recall or genuine reasoning ability. We introduce chess as a controlled testbed aimed at disentangling these faculties. Leveraging the game's structure and scalable engine evaluations, we construct a taxonomy of positions varying in density of relevant priors - ranging from common states solvable by memorization to completely novel ones requiring generalization. Crucially, our approach achieves this distinction without requiring explicit knowledge of the models' training data. Applying this taxonomy, we combine a longitudinal analysis of the GPT lineage with a rigorous evaluation of contemporary models, including Claude Opus and Gemini. Our analysis reveals a steep gradient: performance consistently degrades as the density of relevant priors decreases. Notably, for tasks with few relevant priors, base model performance regresses to the random-play baseline. While newer models improve, progress slows significantly for tasks with sparse priors. Furthermore, while reasoning-augmented inference improves performance, its relative marginal benefit per token decreases in the absence of relevant priors. These results suggest limitations in systematic generalization, highlighting the need for mechanisms beyond scale to achieve robust performance when deprived of relevant priors.

cs.CL

Target-Aligned Reinforcement Learning

Many value-based deep reinforcement learning algorithms rely on target networks - lagged copies of the online network - to stabilize training. While effective, this mechanism introduces a fundamental stability-recency tradeoff: slower target updates improve stability but reduce the recency of learning signals, hindering convergence speed. We propose Target-Aligned Reinforcement Learning (TARL), a simple drop-in refinement for existing algorithms that emphasizes transitions for which the target and online network estimates are highly aligned. By focusing updates on well-aligned targets, TARL mitigates the adverse effects of stale target estimates while retaining the stabilizing benefits of target networks. We empirically demonstrate consistent improvements within discrete and continuous control algorithms across various benchmark environments without any hyperparameter tuning, including a 38.18% peak score gain on Atari-10, while incurring less than a 4% increase in wall-clock time.

cs.LG

Neural Cluster First, Route Second: One-Shot Capacitated Vehicle Routing via Differentiable Optimal Transport

The Capacitated Vehicle Routing Problem (CVRP) underpins modern last-mile logistics. Current Neural Combinatorial Optimization (NCO) methods construct CVRP solutions autoregressively, inheriting sequential decoding bottlenecks, sensitivity to spatial symmetries, and brittle out-of-distribution behavior. We revisit the classical Cluster-First-Route-Second (CFRS) paradigm -- long known to be asymptotically optimal but largely overlooked by NCO -- and argue that it is structurally aligned with the core strengths of deep learning: similarity and assignment over global context, rather than the construction of long sequential tours. We introduce Neural CFRS, the first purely non-autoregressive one-shot neural CFRS framework for the CVRP. It enforces global fleet-capacity constraints end-to-end via a differentiable entropic Optimal Transport layer, producing a continuous transport plan to sparsify an exact capacitated assignment solver. We provide formal theoretical guarantees that our architecture intrinsically abstracts away $E(2)$ spatial, inter-route permutation, and intra-route traversal symmetries. By equipping the framework with a pre-trained spatial vocabulary, we unlock extreme parameter efficiency and zero-shot scaling. Designed primarily for real-world spatial distributions under a constant capacity setting, Neural CFRS scales robustly to out-of-distribution $N=1000$ instances with a < 4% gap -- retaining an approximate 5% gap at this scale even as an ultra-lightweight, single-layer architecture. Furthermore, when deployed out-of-the-box on standard benchmarks, we achieve a highly competitive 2.73% optimality gap on size-100 problems.

cs.LG

Breaking the Grid: Distance-Guided Reinforcement Learning in Large Discrete Action Spaces

Reinforcement Learning (RL) is increasingly applied to large-scale decision-making problems like logistics, scheduling, and recommender systems, but existing algorithms struggle with the curse of dimensionality in such large discrete action spaces. We propose Distance-Guided Reinforcement Learning (DGRL), combining Sampled Dynamic Neighborhoods and Distance-Based Updates to enable efficient RL in problems with up to $10^{20}$ actions. Unlike prior methods, DGRL performs stochastic volumetric exploration and transforms policy optimization into a stable regression task, decoupling gradient variance from action space cardinality. On structured tasks, DGRL provably guarantees local value improvement. DGRL naturally generalizes to hybrid continuous-discrete action spaces. We demonstrate performance improvements of up to 66% against state-of-the-art benchmarks across regularly and irregularly structured environments, while simultaneously improving convergence speed and computational complexity.

cs.LG

Optimizing a Worldwide-Scale Shipper Transportation Planning in a Carmaker Inbound Supply Chain

We study the shipper-side design of large-scale inbound transportation networks, motivated by the global supply chain of the carmaker Renault. We formalize the Shipper Transportation Planning Problem (STPP), which integrates discrete flow consolidation via explicit bin-packing, time-expanded routing, and operational regularity constraints. To solve this high-complexity combinatorial problem at an industrial scale, we propose a tailored Iterated Local Search (ILS) metaheuristic. The algorithm combines large-neighborhood search with MILP-based perturbations and leverages bundle-specific decompositions to obtain scalable lower bounds and effective search guidance. Computational experiments on real industrial data involving more than 700,000 commodities and 1.2 million arcs demonstrate that the ILS achieves an average gap of 7.9% to the best available lower bound. The results reveal a 23.2% cost-reduction potential compared to legacy planning benchmarks. Most significantly, the proposed framework is currently deployed in production at Renault, where it supports weekly strategic decisions and generates realized cost savings estimated at approximately 20 million euros per year. Our analysis yields key managerial insights: we demonstrate that explicit 1D bin-packing is a critical step forward for realistic consolidation modeling, that transport regularity offers a robust balance between cost and stability, and that high-volume global networks benefit significantly from in-house strategic planning over third-party outsourcing. To the best of our knowledge, this is the first work to successfully solve a shipper-side transportation design problem at this magnitude.

math.OC

Synthetic Monitoring Environments for Reinforcement Learning

Reinforcement Learning (RL) lacks benchmarks that enable precise, white-box diagnostics of agent behavior. Current environments often entangle complexity factors and lack ground-truth optimality metrics, making it difficult to isolate why algorithms fail. We introduce Synthetic Monitoring Environments (SMEs), an infinite suite of continuous control tasks. SMEs provide fully configurable task characteristics and known optimal policies. As such, SMEs allow for the exact calculation of instantaneous regret. Their rigorous geometric state space bounds allow for systematic within-distribution (WD) and out-of-distribution (OOD) evaluation. We demonstrate the framework's benefit through multidimensional ablations of PPO, TD3, and SAC, revealing how specific environmental properties - such as action or state space size, reward sparsity and complexity of the optimal policy - impact WD and OOD performance. We thereby show that SMEs offer a standardized, transparent testbed for transitioning RL evaluation from empirical benchmarking toward rigorous scientific analysis.

cs.LG

Combinatorial Optimization Augmented Machine Learning

Combinatorial optimization augmented machine learning (COAML) has recently emerged as a powerful paradigm for integrating predictive models with combinatorial decision-making. By embedding combinatorial optimization oracles into learning pipelines, COAML enables the construction of policies that are both data-driven and feasibility-preserving, bridging the traditions of machine learning, operations research, and stochastic optimization. This paper provides a comprehensive overview of the state of the art in COAML. We introduce a unifying framework for COAML pipelines, describe their methodological building blocks, and formalize their connection to empirical cost minimization. We then develop a taxonomy of problem settings based on the form of uncertainty and decision structure. Using this taxonomy, we review algorithmic approaches for static and dynamic problems, survey applications across domains such as scheduling, vehicle routing, stochastic programming, and reinforcement learning, and synthesize methodological contributions in terms of empirical cost minimization, imitation learning, and reinforcement learning. Finally, we identify key research frontiers. This survey aims to serve both as a tutorial introduction to the field and as a roadmap for future research at the interface of combinatorial optimization and machine learning.

cs.LG

A Note on Piecewise Affine Decision Rules for Robust, Stochastic, and Data-Driven Optimization

Multi-stage decision-making under uncertainty, where decisions are taken under sequentially revealing uncertain problem parameters, is often essential to faithfully model managerial problems. Given the significant computational challenges involved, these problems are typically solved approximately. This short note introduces an algorithmic framework that revisits a popular approximation scheme for multi-stage stochastic programs by Georghiou et al. (2015) and improves upon it to deliver superior policies in the stochastic setting, as well as extend its applicability to robust optimization and a contemporary Wasserstein-based data-driven setting. We demonstrate how the policies of our framework can be computed efficiently, and we present numerical experiments that highlight the benefits of our method.

math.OC

Reproducibility in the Control of Autonomous Mobility-on-Demand Systems

Autonomous Mobility-on-Demand (AMoD) systems, powered by advances in robotics, control, and Machine Learning (ML), offer a promising paradigm for future urban transportation. AMoD offers fast and personalized travel services by leveraging centralized control of autonomous vehicle fleets to optimize operations and enhance service performance. However, the rapid growth of this field has outpaced the development of standardized practices for evaluating and reporting results, leading to significant challenges in reproducibility. As AMoD control algorithms become increasingly complex and data-driven, a lack of transparency in modeling assumptions, experimental setups, and algorithmic implementation hinders scientific progress and undermines confidence in the results. This paper presents a systematic study of reproducibility in AMoD research. We identify key components across the research pipeline, spanning system modeling, control problems, simulation design, algorithm specification, and evaluation, and analyze common sources of irreproducibility. We survey prevalent practices in the literature, highlight gaps, and propose a structured framework to assess and improve reproducibility. Specifically, concrete guidelines are offered, along with a "reproducibility checklist", to support future work in achieving replicable, comparable, and extensible results. While focused on AMoD, the principles and practices we advocate generalize to a broader class of cyber-physical systems that rely on networked autonomy and data-driven control. This work aims to lay the foundation for a more transparent and reproducible research culture in the design and deployment of intelligent mobility systems.

cs.RO

Reliability-Adjusted Prioritized Experience Replay

Experience replay enables data-efficient learning from past experiences in online reinforcement learning agents. Traditionally, experiences were sampled uniformly from a replay buffer, regardless of differences in experience-specific learning potential. In an effort to sample more efficiently, researchers introduced Prioritized Experience Replay (PER). In this paper, we propose an extension to PER by introducing a novel measure of temporal difference error reliability. We theoretically show that the resulting transition selection algorithm, Reliability-adjusted Prioritized Experience Replay (ReaPER), enables more efficient learning than PER. We further present empirical results showing that ReaPER outperforms PER across various environment types, including the Atari-10 benchmark.

cs.LG

Public Transport Under Epidemic Conditions: Nonlinear Trade-Offs Between Risk and Accessibility

Epidemics expose critical tensions between protecting public health and maintaining essential urban mobility. Public transport systems face this dilemma most acutely: they enable access to jobs, education, and services, yet also facilitate close contact among travelers. We develop an integrated modeling framework that couples agent-based epidemic simulation (EpiSim) with an optimization-based public transport flow model under capacity constraints. Using Munich as a case study, we analyze how combinations of facility closures and transport restrictions shape epidemic outcomes and accessibility. The results reveal three key insights. First, epidemic interventions redistribute rather than simply reduce infection risks, shifting transmission to households. Second, epidemic and transport policies interact nonlinearly - moderate demand suppression can offset large capacity cuts. Third, epidemic pressures amplify temporal and spatial inequalities, disproportionately affecting peripheral and peak-hour travelers. These findings highlight that blanket restrictions are both inefficient and inequitable, calling for targeted, time- and space-differentiated measures to build epidemic-resilient and socially fair transport systems.

physics.soc-ph

Structured Reinforcement Learning for Combinatorial Decision-Making

Reinforcement learning (RL) is increasingly applied to real-world problems involving complex and structured decisions, such as routing, scheduling, and assortment planning. These settings challenge standard RL algorithms, which struggle to scale, generalize, and exploit structure in the presence of combinatorial action spaces. We propose Structured Reinforcement Learning (SRL), a novel actor-critic paradigm that embeds combinatorial optimization-layers into the actor neural network. We enable end-to-end learning of the actor via Fenchel-Young losses and provide a geometric interpretation of SRL as a primal-dual algorithm in the dual of the moment polytope. Across six environments with exogenous and endogenous uncertainty, SRL matches or surpasses the performance of unstructured RL and imitation learning on static tasks and improves over these baselines by up to 92% on dynamic problems, with improved stability and convergence speed.

cs.LG

Versatile Top-Down Patterning Technique for Perovskite On-Chip Integration

Metal-halide perovskites (MHPs) have exciting optoelectronic properties and are under investigation for various applications, such as photovoltaics, light-emitting diodes, and lasers. An essential step towards exploiting the full potential of this class of materials is their large-scale, on-chip integration with high-resolution, top-down patterning. The development of such patterning methods for perovskite films is challenging because of their intrinsic ionic nature and adverse reactions with the solvents used in standard lithography processes. Here, we introduce a versatile and precise method comprising photolithography and reactive ion etching (RIE) processes that can be tuned to accommodate different perovskite compositions and morphologies. Our method utilizes conventional photoresists at reduced temperatures to create micron-sized features down to 1 $μ$m, providing high reproducibility from chip to chip. The patterning technique is validated through atomic force microscopy (AFM), X-ray diffraction (XRD), optical spectroscopy, and scanning electron microscopy (SEM). It enables the scalable and high-throughput on-chip monolithic integration of MHPs.

cond-mat.mtrl-sci

Direct probe of magnetic field effects on phonons by ultrasound propagation in a quasi-two-dimensional honeycomb magnet Na$_2$Co$_2$TeO$_6$

We study the phonon behavior of a Co-based honeycomb frustrated magnet Na$_2$Co$_2$TeO$_6$ under magnetic field applied perpendicular to the honeycomb plane. The temperature and field dependence of the sound velocity and sound attenuation unveil prominent spin-lattice coupling in this material, promoting ultrasound as a sensitive probe for magnetic properties. An out-of-plane ferrimagnetic order is determined below the Néel temperature $T_N=27$~K. A comprehensive analysis of our data further supports a triple-Q ground state of Na$_2$Co$_2$TeO$_6$. Furthermore, the ultrasound data were systematically compared to the thermal transport results from literature, to unveil the importance of phononic contribution to the observed transport behaviors.

cond-mat.str-el

Integrated Balanced and Staggered Routing in Autonomous Mobility-on-Demand Systems

Autonomous mobility-on-demand (AMoD) systems, centrally coordinated fleets of self-driving vehicles, offer a promising alternative to traditional ride-hailing by improving traffic flow and reducing operating costs. Centralized control in AMoD systems enables two complementary routing strategies: balanced routing, which distributes traffic across alternative routes to ease congestion, and staggered routing, which delays departures to smooth peak demand over time. In this work, we introduce a unified framework that jointly optimizes both route choices and departure times to minimize system travel times. We formulate the problem as an optimization model and show that our congestion model yields an unbiased estimate of travel times derived from a discretized version of Vickrey's bottleneck model. To solve large-scale instances, we develop a custom metaheuristic based on a large neighborhood search framework. We assess our method through a case study on the Manhattan street network using real-world taxi data. In a setting with exclusively centrally controlled AMoD vehicles, our approach reduces total traffic delay by up to 25 percent and mitigates network congestion by up to 35 percent compared to selfish routing. We also consider mixed-traffic settings with both AMoD and conventional vehicles, comparing a welfare-oriented operator that minimizes total system travel time with a profit-oriented one that optimizes only the fleet's travel time. Independent of the operator's objective, the analysis reveals a win-win outcome: across all control levels, both autonomous and non-autonomous traffic benefit from the implementation of balancing and staggering strategies.

math.OC

Multi-Agent Soft Actor-Critic with Coordinated Loss for Autonomous Mobility-on-Demand Fleet Control

We study a sequential decision-making problem for a profit-maximizing operator of an autonomous mobility-on-demand system. Optimizing a central operator's vehicle-to-request dispatching policy requires efficient and effective fleet control strategies. To this end, we employ a multi-agent Soft Actor-Critic algorithm combined with weighted bipartite matching. We propose a novel vehicle-based algorithm architecture and adapt the critic's loss function to appropriately consider coordinated actions. Furthermore, we extend our algorithm to incorporate rebalancing capabilities. Through numerical experiments, we show that our approach outperforms state-of-the-art benchmarks by up to 12.9% for dispatching and up to 38.9% with integrated rebalancing.

eess.SY