Searcharxiv⌕ Search

arXiv subjects

Mehrdad Mohammadi

Publications and source records attributed to Mehrdad Mohammadi.

4 recordsLinked to original sources

Vector-Valued Distributional Reinforcement Learning Policy Evaluation: A Hilbert Space Embedding Approach

We propose Kernel Embedding Distributional Reinforcement Learning (KE-DRL) an offline method to estimate kernel mean embeddings of conditional multivariate return distributions from observed offline trajectories with continuous state-action inputs and vector-valued rewards. KE-DRL approximates the conditional embedding and estimates its a finite-dictionary representation coefficients through a maximum mean discrepancy Bellman criterion. For a regular class of return distributions, we show that a Matérn embedding and a Bellman-invariant Sobolev-moment class identifies the unique distributional Bellman fixed point and and derive a Hölder-type bound that controls Wasserstein error by the population embedding Bellman residual. For directly observed responses, we provide finite-sample and uniform error bounds for a regularized conditional mean embedding estimator. Simulation experiments evaluate pointwise embedding recovery against Monte Carlo benchmarks across multiple behavior-target policy pairs. An application to Expedia hotel-search data illustrates conditional multivariate return evaluation under two policies.

cs.LG↗

Toscani-Fourier Distance on Probability Measures: Wasserstein Control, Topological Equivalence on Model Classes, and Duality

Comparing probability measures in machine learning trades transport geometry against computational cost: Wasserstein distances encode the geometry of $\mathbb R^d$ but require solving a transport problem, while kernel discrepancies are cheap to evaluate yet depend delicately on their test class. We study the Toscani--Fourier family $\mathrm T_{s,p}$, the weighted $L^p$ norm of the difference of two characteristic functions, as a continuous Fourier-side discrepancy on $\mathbb R^d$. For $1\le p<\infty$ we show that $d/p<s<1+d/p$ is exactly the window in which $\mathrm T_{s,p}$ is finite on $\mathcal P_p(\mathbb R^d)$, both endpoints already failing for a pair of Dirac measures, and we establish the metric, embedding, and compactness structure of the resulting space, which we prove to be complete. Duality identifies $\mathrm T_{s,p}$ as an integral probability metric over a homogeneous Fourier--Lebesgue ball, with an explicit extremizer when $1<p<\infty$. We prove the global bound $\mathrm T_{s,p}\lesssim W_p^{\,s-d/p}$, whose exponent is sharp, show that no global converse of any form can hold, and recover topological equivalence with $W_p$ on bounded-support and uniform-tail classes, together with explicit reverse moduli on bounded-support classes that improve the imported energy-kernel exponent at $p=2$. There, $\mathrm T_{s,2}$ is a constant multiple of the classical energy distance, which yields an exact finite-sample identity for the mean of the empirical discrepancy; the numerical experiments are otherwise diagnostic.

stat.ML↗

Integrated trucks assignment and scheduling problem with mixed service mode docks: A Q-learning based adaptive large neighborhood search algorithm

Mixed service mode docks enhance efficiency by flexibly handling both loading and unloading trucks in warehouses. However, existing research often predetermines the number and location of these docks prior to planning truck assignment and sequencing. This paper proposes a new model integrating dock mode decision, truck assignment, and scheduling, thus enabling adaptive dock mode arrangements. Specifically, we introduce a Q-learning-based adaptive large neighborhood search (Q-ALNS) algorithm to address the integrated problem. The algorithm adjusts dock modes via perturbation operators, while truck assignment and scheduling are solved using destroy and repair local search operators. Q-learning adaptively selects these operators based on their performance history and future gains, employing the epsilon-greedy strategy. Extensive experimental results and statistical analysis indicate that the Q-ALNS benefits from efficient operator combinations and its adaptive mechanism, consistently outperforming benchmark algorithms in terms of optimality gap and Pareto front discovery. In comparison to the predetermined service mode, our adaptive strategy results in lower average tardiness and makespan, highlighting its superior adaptability to varying demands.

cs.LG↗

Dynamic operator management in meta-heuristics using reinforcement learning: an application to permutation flowshop scheduling problems

This study develops a framework based on reinforcement learning to dynamically manage a large portfolio of search operators within meta-heuristics. Using the idea of tabu search, the framework allows for continuous adaptation by temporarily excluding less efficient operators and updating the portfolio composition during the search. A Q-learning-based adaptive operator selection mechanism is used to select the most suitable operator from the dynamically updated portfolio at each stage. Unlike traditional approaches, the proposed framework requires no input from the experts regarding the search operators, allowing domain-specific non-experts to effectively use the framework. The performance of the proposed framework is analyzed through an application to the permutation flowshop scheduling problem. The results demonstrate the superior performance of the proposed framework against state-of-the-art algorithms in terms of optimality gap and convergence speed.

cs.LG↗