SearcharxivSearch

arXiv subjects

Florian Fuchs

Publications and source records attributed to Florian Fuchs.

14 recordsLinked to original sources

Reward-Adaptive Iterative Discovery: A Case Study on Automated Game Testing for NHL26

Testing is a major effort for the gaming industry, requiring a significant part of development budget and people power. We present a case study on a development version of the ice hockey game EA SPORTS NHL 26, for which human playtesters test the goalie AI for behavioral exploits. To reduce the effort of re-testing the goalie AI after every game or behavior modification in the development phase, we propose Reward-Adaptive Iterative Discovery (RAID), a novel approach to automatically find exploits using an iterative Reinforcement Learning (RL) approach that trains a population of goal scoring agents. While previous approaches can already successfully find exploits, RL algorithms tend to overfit to a single solution. We introduce a simple extension on top of existing RL algorithms, such that they find multiple diverse high-quality solutions. For our first deployment of this approach, within a single experiment we were able to find six hockey scoring exploit strategies that were qualitatively similar to those that playtesters had found in hours-long manual testing sessions.

cs.LG

Coachable agents for interactive gameplay

Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from game playing to robotics to foundation models. Through trial-and-error, these AI systems typically learn one, near-optimal behavior to solve their tasks. However, there are many use cases in which one would like to assert some level of control, preferably in real time, over how the task is solved. We refer to these modifications of a core task as styles. We combine universal value function approximators (UVFAs) with carefully selected training scenarios, learning algorithms, and data augmentation to create a framework for coaching agents that exhibit styles in complex domains. We demonstrate the framework's application in the AAA video games Horizon Forbidden West and Gran Turismo, and in an open-source humanoid test domain. Despite the different nature of the domains -- car racing, stylized game combat, and humanoid walking -- each agent shows strong coherence to the style requests while still satisfying the main task in its domain. Importantly, the techniques outlined in this paper allow an end user to choose the final behavior at run time, giving them flexible control over the final executed performance.

cs.AI

Hierarchical Control in Multi-Agent Games: LLM-based Planning and RL Execution

Reinforcement learning (RL) has achieved strong performance in sequential decision-making, yet scaling to complex multi-agent environments remains challenging due to sparse rewards, large state-action spaces, and the difficulty of learning coordinated strategies. We propose a hierarchical architecture where a pretrained large language model (LLM) acts as a centralized strategic controller that selects among specialized RL skill policies for a team of agents, while RL policies handle reactive low-level execution. We evaluate this hybrid system in a competitive 2v2 King of the Hill environment against behavior tree (BT) and \emph{``Flat''} RL (end-to-end training without skill decomposition) baselines. The LLM+RL system achieves task performance statistically equivalent to hand-crafted BT (46.4\% vs 51.5\% win rate, $p=0.103$) while both significantly outperform Flat RL trained without skill decomposition. A user study ($n=15$) reveals that 60\% of participants perceive LLM+RL agents as the most human-like ($p=0.027$), citing behavioral adaptability and tactical variability. These results demonstrate that pretrained LLM reasoning can effectively orchestrate pretrained RL skills, achieving competitive multi-agent coordination and superior perceived believability without manual rule engineering.

cs.LG

Augmenting Game AI with Deep Reinforcement Learning

Immersion in video games depends not only on graphics, audio, and game mechanics, but also on the quality of in-game characters. Producing believable characters, or game AI, remains a significant challenge as behavioral complexity is hard to capture with hand-coded systems. Game AI is a source of immersion and engagement; however, the limitations stemming from the challenges of creating game AI often lead to frustration and the breaking of the illusion of realism within the game. The introduction of machine learning models opens the door to creating more believable, authentic, and relatable characters in games. The promise is that they either learn from interacting with the game, or from player data, to develop true human-like behavior. In this paper, we envision more applications of reinforcement learning for game AI in the future. For this to materialize, current research limitations are prohibitive to broad deployment across game genres. Therefore, we propose a framework for training reinforcement learning models with a set of requirements in mind that are suited towards game AI and game development. We present examples of games with reinforcement learning-augmented game AI and describe the practicalities of deploying player-facing machine learning agents in modern games. Furthermore, we identify bottlenecks and hard problems in these areas, which we believe offer promising research directions to accelerate the adoption of machine learning in game AI for the video game industry.

cs.AI

Quantum confinement in semiconductor random alloys: a case study on Si/SiGe/Si

Local composition fluctuations in random alloys become crucial when one or more dimensions are reduced to the nanoscale. Using extended H\"uckel theory, we study the semiconductor random alloy SiGe sandwiched between Si due to its relevance for transistor devices. We evaluate the effects of the alloy composition, layer thickness, and local fluctuations of the Ge concentration on the band alignment and the band gap. The results are compared with the finite quantum well model. That model captures the essential physics and can act as a computationally faster alternative.

cond-mat.mes-hall

Human-Like Goalkeeping in a Realistic Football Simulation: a Sample-Efficient Reinforcement Learning Approach

While several high profile video games have served as testbeds for Deep Reinforcement Learning (DRL), this technique has rarely been employed by the game industry for crafting authentic AI behaviors. Previous research focuses on training super-human agents with large models, which is impractical for game studios with limited resources aiming for human-like agents. This paper proposes a sample-efficient DRL method tailored for training and fine-tuning agents in industrial settings such as the video game industry. Our method improves sample efficiency of value-based DRL by leveraging pre-collected data and increasing network plasticity. We evaluate our method training a goalkeeper agent in EA SPORTS FC 25, one of the best-selling football simulations today. Our agent outperforms the game's built-in AI by 10% in ball saving rate. Ablation studies show that our method trains agents 50% faster compared to standard DRL methods. Finally, qualitative evaluation from domain experts indicates that our approach creates more human-like gameplay compared to hand-crafted agents. As a testament to the impact of the approach, the method has been adopted for use in the most recent release of the series.

cs.AI

Solving Integrated Periodic Railway Timetabling with Satisfiability Modulo Theories: A Scalable Approach to Routing and Vehicle Circulation

This paper introduces a novel approach for jointly solving the periodic Train Timetabling Problem (TTP), train routing, and Vehicle Circulation Problem (VCP) through a unified optimization model. While these planning stages are traditionally addressed sequentially, their interdependencies often lead to suboptimal vehicle usage. We propose the VCR-PESP, an integrated formulation that minimizes fleet size while ensuring feasible and infrastructure-compliant periodic timetables. We present the first Satisfiability Modulo Theories (SMT)-based method for the VCR-PESP to solve the resulting large-scale instances. Unlike the Boolean Satisfiability Problem (SAT), which requires time discretisation, SMT supports continuous time via difference constraints, eliminating the trade-off between temporal precision and encoding size. Our approach avoids rounding artifacts and scales effectively, outperforming both SAT and Mixed Integer Program (MIP) models across non-trivial instances. Using real-world data from the Swiss narrow-gauge operator RhB, we conduct extensive experiments to assess the impact of time discretisation, vehicle circulation strategies, route flexibility, and planning integration. We show that discrete models inflate vehicle requirements and that fully integrated solutions substantially reduce fleet needs compared to sequential approaches. Our framework consistently delivers high-resolution solutions with tractable runtimes, even in large and complex networks. By combining modeling accuracy with scalable solver technology, this work establishes SMT as a powerful tool for integrated railway planning. It demonstrates how relaxing discretisation and solving across planning layers enables more efficient and implementable timetables.

math.OC

First-principles analysis of the effect of magnetic states on the oxygen vacancy formation energy in doped La$_{0.5}$Sr$_{0.5}$CoO$_3$ perovskite

Oxygen vacancies are critical for determining the electrochemical performance of fast oxygen ion conductors. The perovskite La$_{0.5}$Sr$_{0.5}$CoO$_3$, known for its excellent mixed ionic-electronic conduction, has attracted significant attention due to its favorable vacancy characteristics. In this study, we employ first-principles calculations to systematically investigate the impact of 3$d$ transition-metal doping on the oxygen vacancy formation energies in the perovskite. Two magnetic states, namely the ferromagnetic and paramagnetic states, are considered in our models to capture the influence of magnetic effects on oxygen vacancy energetics. Our results reveal that the oxygen vacancy formation energies are strongly dependent on both the dopant species and the magnetic state. Notably, the magnetic states alter the vacancy formation energy in a dopant-specific manner due to double exchange interactions, indicating that relying solely on the ferromagnetic ground state may result in misleading trends in doping behavior. These findings emphasise the importance of accounting for magnetic effects when investigating oxygen vacancy properties in perovskite oxides.

cond-mat.mtrl-sci

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games

Reinforcement Learning (RL) in games has gained significant momentum in recent years, enabling the creation of different agent behaviors that can transform a player's gaming experience. However, deploying RL agents in production environments presents two key challenges: (1) designing an effective reward function typically requires an RL expert, and (2) when a game's content or mechanics are modified, previously tuned reward weights may no longer be optimal. Towards the latter challenge, we propose an automated approach for iteratively fine-tuning an RL agent's reward function weights, based on a user-defined language based behavioral goal. A Language Model (LM) proposes updated weights at each iteration based on this target behavior and a summary of performance statistics from prior training rounds. This closed-loop process allows the LM to self-correct and refine its output over time, producing increasingly aligned behavior without the need for manual reward engineering. We evaluate our approach in a racing task and show that it consistently improves agent performance across iterations. The LM-guided agents show a significant increase in performance from $9\%$ to $74\%$ success rate in just one iteration. We compare our LM-guided tuning against a human expert's manual weight design in the racing task: by the final iteration, the LM-tuned agent achieved an $80\%$ success rate, and completed laps in an average of $855$ time steps, a competitive performance against the expert-tuned agent's peak $94\%$ success, and $850$ time steps.

cs.AI

Local laser-induced solid-phase recrystallization of phosphorus-implanted Si/SiGe heterostructures for contacts below 4.2 K

Si/SiGe heterostructures are of high interest for high mobility transistor and qubit applications, specifically for operations below 4.2 K. In order to optimize parameters such as charge mobility, built-in strain, electrostatic disorder, charge noise and valley splitting, these heterostructures require Ge concentration profiles close to mono-layer precision. Ohmic contacts to undoped heterostructures are usually facilitated by a global annealing step activating implanted dopants, but compromising the carefully engineered layer stack due to atom diffusion and strain relaxation in the active device region. We demonstrate a local laser-based annealing process for recrystallization of ion-implanted contacts in SiGe, greatly reducing the thermal load on the active device area. To quickly adapt this process to the constantly evolving heterostructures, we deploy a calibration procedure based exclusively on optical inspection at room-temperature. We measure the electron mobility and contact resistance of laser annealed Hall bars at temperatures below 4.2 K and obtain values similar or superior than that of a globally annealed reference samples. This highlights the usefulness of laser-based annealing to take full advantage of high-performance Si/SiGe heterostructures.

physics.app-ph

Super-Human Performance in Gran Turismo Sport Using Deep Reinforcement Learning

Autonomous car racing is a major challenge in robotics. It raises fundamental problems for classical approaches such as planning minimum-time trajectories under uncertain dynamics and controlling the car at the limits of its handling. Besides, the requirement of minimizing the lap time, which is a sparse objective, and the difficulty of collecting training data from human experts have also hindered researchers from directly applying learning-based approaches to solve the problem. In the present work, we propose a learning-based system for autonomous car racing by leveraging a high-fidelity physical car simulation, a course-progress proxy reward, and deep reinforcement learning. We deploy our system in Gran Turismo Sport, a world-leading car simulator known for its realistic physics simulation of different race cars and tracks, which is even used to recruit human race car drivers. Our trained policy achieves autonomous racing performance that goes beyond what had been achieved so far by the built-in AI, and, at the same time, outperforms the fastest driver in a dataset of over 50,000 human players.

cs.AI

Noise recovery for Lévy-driven CARMA processes and high-frequency behaviour of approximating Riemann sums

We consider high-frequency sampled continuous-time autoregressive moving average (CARMA) models driven by finite-variance zero-mean Lévy processes. An L^2-consistent estimator for the increments of the driving Lévy process without order selection in advance is proposed if the CARMA model is invertible. In the second part we analyse the high-frequency behaviour of approximating Riemann sum processes, which represent a natural way to simulate continuous-time moving average processes on a discrete grid. We shall compare their autocovariance structure with the one of sampled CARMA processes, where the rule of integration plays a crucial role. Moreover, new insight into the kernel estimation procedure of Brockwell et al. (2012a) is given.

math.PR

Stationarity and Geometric Ergodicity of BEKK Multivariate GARCH Models

Conditions for the existence of strictly stationary multivariate GARCH processes in the so-called BEKK parametrisation, which is the most general form of multivariate GARCH processes typically used in applications, and for their geometric ergodicity are obtained. The conditions are that the driving noise is absolutely continuous with respect to the Lebesgue measure and zero is in the interior of its support and that a certain matrix built from the GARCH coefficients has spectral radius smaller than one. To establish the results semi-polynomial Markov chains are defined and analysed using algebraic geometry.

math.PR

Spectral Representation of Multivariate Regularly Varying Lévy and CARMA processes

A spectral representation for regularly varying Lévy processes with index between one and two is established and the properties of the resulting random noise are discussed in detail giving also new insight in the $L^2$-case where the noise is a random orthogonal measure. This allows a spectral definition of multivariate regularly varying Lévy-driven continuous time autoregressive moving average (CARMA) processes. It is shown that they extend the well-studied case with finite second moments and coincide with definitions previously used in the infinite variance case when they apply.

math.PR