SearcharxivSearch

arXiv subjects

Florian Lorkowski

Publications and source records attributed to Florian Lorkowski.

3 recordsLinked to original sources

Infra-Bayesian Reinforcement Learning Agents Outperform Classical RL For Worst-Case Robustness

Classical reinforcement learning assumes the agent interacts with a fixed environment whose behavior does not depend on the agent's policy. This assumption breaks down in non-realizable settings where other actors might anticipate the agent's behavior, including environments crucial to AI safety, where the agent interacts with predictors, humans, other AI agents, and institutions. In such settings, the agent's model class fails to capture the world in which it operates. Under such misspecification, classical Bayesian methods can produce confidently wrong posteriors, unreliable decisions, and unbounded regret, as realizability fails to obtain. Infra-Bayesianism is a decision-theoretic framework that addresses these failures by distinguishing ordinary probabilistic uncertainty, where priors can be reasonably chosen, from Knightian uncertainty, where no grounds exist for the construction of such a prior. It does so by evaluating actions on their worst-case outcomes, rather than from posterior expectations or weighted averaging. We present the first proof-of-concept implementation of an infra-Bayesian reinforcement learning architecture for finite-outcome stateless decision problems. Our agent maintains a set of imprecise hypotheses, updates them using infra-Bayesian conditioning, and selects actions by maximizing worst-case expected value. We apply this implementation of the infra-Bayesian maximin decision process to an environment with Knightian uncertainty, and demonstrate a lower worst-case regret as compared to classical reinforcement learning agents. We also investigate Newcomb's problem and show that the infra-Bayesian agent picks the optimal strategy, outperforming classical decision theory agents. Our results provide a step towards reinforcement learning agents that remain robust under model misspecification and policy-dependent uncertainty.

cs.LG

Elliptic Multiple Polylogarithms with Arbitrary Arguments in \textsc{GiNaC}

We present an algorithm for the numerical evaluation of elliptic multiple polylogarithms for arbitrary arguments and to arbitrary precision. The cornerstone of our approach is a procedure to obtain a convergent $q$-series representation of elliptic multiple polylogarithms. Its coefficients are expressed in terms of ordinary multiple polylogarithms, which can be evaluated efficiently using existing libraries. In a series of preparation steps the elliptic polylogarithms are mapped into a region where the $q$-series converges rapidly. We also present an implementation of our algorithm into the \texttt{GiNaC} framework. This release constitutes the first public package capable of evaluating elliptic multiple polylogarithms to high precision and for arbitrary values of the arguments.

hep-ph

Treatment of QED corrections in jet production in deep inelastic scattering at ZEUS

A new measurement of inclusive jet production in deep inelastic scattering was recently published by the ZEUS Collaboration. This contribution presents a detailed discussion of the treatment of higher-order QED effects in this measurement. A comprehensive treatment of these effects is crucial for a more direct comparison between ever more precise measurements and theoretical calculations. The present analysis is the only measurement of jet production in deep inelastic scattering that can be compared to full NNLO QCD + NLO electroweak predictions.

hep-ex