Searcharxiv⌕ Search

arXiv subjects

Berkay Anahtarci

Publications and source records attributed to Berkay Anahtarci.

7 recordsLinked to original sources

NL-PAC: Specification Ambiguity and Certified Minimax Risk Floors in LLM-Mediated Supervision

Large language models increasingly provide labels, evaluations, and feedback for tasks specified in natural language. When a specification admits multiple readings but the supervision channel does not reveal which is operative, additional labels reduce sampling error without resolving the resulting identification problem. We introduce Natural Language PAC (NL-PAC), a framework that uses a fixed model's thresholded decoding law to define admissible labels and candidate targets. The probability that multiple labels are admissible equals the diameter of the pointwise-admissible target class, and under target-blind supervision every learner incurs worst-case risk of at least half this diameter, at every sample size; the exact randomized minimax risk over this class is attained by a data-independent strategy. Finite-sample confidence bounds make these quantities certifiable from held-out unlabeled inputs. In a frozen Qwen~2.5--3B audit, one prespecified prompt yields a positive model-relative certificate, whereas a paraphrase and exact-rule controls yield zero. A held-out bridge audit finds that supplied candidate reading clauses fail the admissibility condition needed to transfer the certificate to coherent readings. The guarantee is specific to the audited model, prompt, threshold, and input distribution; extending it to human interpretations requires external validation.

cs.LG↗

Kernel Based Maximum Entropy Inverse Reinforcement Learning for Mean-Field Games

We consider the maximum causal entropy inverse reinforcement learning (IRL) problem for infinite-horizon stationary mean-field games (MFG), in which we model the unknown reward function within a reproducing kernel Hilbert space (RKHS). This allows the inference of rich and potentially nonlinear reward structures directly from expert demonstrations, in contrast to most existing approaches for MFGs that typically restrict the reward to a linear combination of a fixed finite set of basis functions and rely on finite-horizon formulations. We introduce a Lagrangian relaxation that enables us to reformulate the problem as an unconstrained log-likelihood maximization and obtain a solution via a gradient ascent algorithm. To establish the theoretical consistency of the algorithm, we prove the smoothness of the log-likelihood objective through the Fréchet differentiability of the related soft Bellman operators with respect to the parameters in the RKHS. To illustrate the practical advantages of the RKHS formulation, we validate our framework on a mean-field traffic routing game exhibiting state-dependent preference reversal, where the kernel-based method reduces policy recovery error by over an order of magnitude compared to a linear reward baseline with a comparable parameter count. Furthermore, we extend the framework to the finite-horizon non-stationary setting. We demonstrate that the log-likelihood reformulation is structurally unavailable in this regime and instead develop an alternative gradient descent algorithm on the convex dual via Danskin's theorem, establishing smoothness and convergence guarantees.

cs.LG↗

Maximum Causal Entropy IRL in Mean-Field Games and GNEP Framework for Forward RL

This paper explores the use of Maximum Causal Entropy Inverse Reinforcement Learning (IRL) within the context of discrete-time stationary Mean-Field Games (MFGs) characterized by finite state spaces and an infinite-horizon, discounted-reward setting. Although the resulting optimization problem is non-convex with respect to policies, we reformulate it as a convex optimization problem in terms of state-action occupation measures by leveraging the linear programming framework of Markov Decision Processes. Based on this convex reformulation, we introduce a gradient descent algorithm with a guaranteed convergence rate to efficiently compute the optimal solution. Moreover, we develop a new method that conceptualizes the MFG problem as a Generalized Nash Equilibrium Problem (GNEP), enabling effective computation of the mean-field equilibrium for forward reinforcement learning (RL) problems and marking an advancement in MFG solution techniques. We further illustrate the practical applicability of our GNEP approach by employing this algorithm to generate data for numerical MFG examples.

eess.SY↗

Q-Learning in Regularized Mean-field Games

In this paper, we introduce a regularized mean-field game and study learning of this game under an infinite-horizon discounted reward function. Regularization is introduced by adding a strongly concave regularization function to the one-stage reward function in the classical mean-field game model. We establish a value iteration based learning algorithm to this regularized mean-field game using fitted Q-learning. The regularization term in general makes reinforcement learning algorithm more robust to the system components. Moreover, it enables us to establish error analysis of the learning algorithm without imposing restrictive convexity assumptions on the system components, which are needed in the absence of a regularization term.

math.OC↗

Value Iteration Algorithm for Mean-field Games

In the literature, existence of mean-field equilibria has been established for discrete-time mean field games under both the discounted cost and the average cost optimality criteria. In this paper, we provide a value iteration algorithm to compute mean-field equilibrium for both the discounted cost and the average cost criteria, whose existence proved previously. We establish that the value iteration algorithm converges to the fixed point of a mean-field equilibrium operator. Then, using this fixed point, we construct a mean-field equilibrium. In our value iteration algorithm, we use $Q$-functions instead of value functions.

eess.SY↗

Asymptotics of spectral gaps of 1D Dirac operator with two exponential terms potential

The one-dimensional Dirac operator \begin{equation*} L = i \begin{pmatrix} 1 & 0 \\ 0 & -1 \end{pmatrix} \frac{d}{dx} +\begin{pmatrix} 0 & P(x) \\ Q(x) & 0 \end{pmatrix}, \quad P,Q \in L^2 ([0,π]), \end{equation*} considered on $[0,π]$ with periodic and antiperiodic boundary conditions, has discrete spectra. For large enough $|n|,\, n \in \mathbb{Z}, $ there are two (counted with multiplicity) eigenvalues $λ_n^-,λ_n^+ $ (periodic if $n$ is even, or antiperiodic if $n$ is odd) such that $|λ_n^\pm - n |<1/2.$ We study the asymptotics of spectral gaps $γ_n =λ_n^+ - λ_n^-$ in the case $$P(x)=a e^{-2ix} + A e^{2ix}, \quad Q(x)=b e^{-2ix} + B e^{2ix},$$ where $a, A, b, B$ are nonzero complex numbers. We show, for large enough $m,$ that $γ_{\pm 2m}=0 $ and \begin{align*} γ_{2m+1} = \pm 2 \frac{\sqrt{(Ab)^m (aB)^{m+1}}}{4^{2m} (m!)^2 } \left[ 1 + O \left( \frac{\log^2 m}{m^2}\right) \right], \end{align*} \begin{align*} γ_{-(2m+1)} = \pm 2\frac{\sqrt{(Ab)^{m+1} (aB)^m}}{4^{2m} (m!)^2} \left[ 1 + O \left( \frac{\log^2 m}{m^2}\right) \right]. \end{align*}

math.SP↗

Improved asymptotics of the spectral gap for the Mathieu operator

The Mathieu operator {equation*} L(y)=-y"+2a \cos{(2x)}y, \quad a\in \mathbb{C},\;a\neq 0, {equation*} considered with periodic or anti-periodic boundary conditions has, close to $n^2$ for large enough $n$, two periodic (if $n$ is even) or anti-periodic (if $n$ is odd) eigenvalues $λ_n^-$, $λ_n^+$. For fixed $a$, we show that {equation*} λ_n^+ - λ_n^-= \pm \frac{8(a/4)^n}{[(n-1)!]^2} [1 - \frac{a^2}{4n^3}+ O (\frac{1}{n^4})], \quad n\rightarrow\infty. {equation*} This result extends the asymptotic formula of Harrell-Avron-Simon, by providing more asymptotic terms.

math.SP↗