SearcharxivSearch

arXiv subjects

Felix Schmitt

Publications and source records attributed to Felix Schmitt.

14 recordsLinked to original sources

Reward (Mis)design for Autonomous Driving

This article considers the problem of diagnosing certain common errors in reward design. Its insights are also applicable to the design of cost functions and performance metrics more generally. To diagnose common errors, we develop 8 simple sanity checks for identifying flaws in reward functions. These sanity checks are applied to reward functions from past work on reinforcement learning (RL) for autonomous driving (AD), revealing near-universal flaws in reward design for AD that might also exist pervasively across reward design for other tasks. Lastly, we explore promising directions that may aid the design of reward functions for AD in subsequent research, following a process of inquiry that can be adapted to other domains.

cs.LG

Hierarchies of Planning and Reinforcement Learning for Robot Navigation

Solving robotic navigation tasks via reinforcement learning (RL) is challenging due to their sparse reward and long decision horizon nature. However, in many navigation tasks, high-level (HL) task representations, like a rough floor plan, are available. Previous work has demonstrated efficient learning by hierarchal approaches consisting of path planning in the HL representation and using sub-goals derived from the plan to guide the RL policy in the source task. However, these approaches usually neglect the complex dynamics and sub-optimal sub-goal-reaching capabilities of the robot during planning. This work overcomes these limitations by proposing a novel hierarchical framework that utilizes a trainable planning policy for the HL representation. Thereby robot capabilities and environment conditions can be learned utilizing collected rollout data. We specifically introduce a planning policy based on value iteration with a learned transition model (VI-RL). In simulated robotic navigation tasks, VI-RL results in consistent strong improvement over vanilla RL, is on par with vanilla hierarchal RL on single layouts but more broadly applicable to multiple layouts, and is on par with trainable HL path planning baselines except for a parking task with difficult non-holonomic dynamics where it shows marked improvements.

cs.RO

On the Scalability of Data Reduction Techniques in Current and Upcoming HPC Systems from an Application Perspective

We implement and benchmark parallel I/O methods for the fully-manycore driven particle-in-cell code PIConGPU. Identifying throughput and overall I/O size as a major challenge for applications on today's and future HPC systems, we present a scaling law characterizing performance bottlenecks in state-of-the-art approaches for data reduction. Consequently, we propose, implement and verify multi-threaded data-transformations for the I/O library ADIOS as a feasible way to trade underutilized host-side compute potential on heterogeneous systems for reduced I/O latency.

cs.PF

Exact Maximum Entropy Inverse Optimal Control for Modelling Human Attention Switching and Control

Maximum Causal Entropy (MCE) Inverse Optimal Control (IOC) has become an effective tool for modelling human behaviour in many control tasks. Its advantage over classic techniques for estimating human policies is the transferability of the inferred objectives: Behaviour can be predicted in variations of the control task by policy computation using a relaxed optimality criterion. However, exact policy inference is often computationally intractable in control problems with imperfect state observation. In this work, we present a model class that allows modelling human control of two tasks of which only one be perfectly observed at a time requiring attention switching. We show how efficient and exact objective and policy inference via MCE can be conducted for these control problems. Both MCE-IOC and Maximum Causal Likelihood (MCL)-IOC, a variant of the original MCE approach, as well as Direct Policy Estimation (DPE) are evaluated using simulated and real behavioural data. Prediction error and generalization over changes in the control process are both considered in the evaluation. The results show a clear advantage of both IOC methods over DPE, especially in the transfer over variation of the control process. MCE and MCL performed similar when training on a large set of simulated data, but differed significantly on small sets and real data.

eess.SY

Inverse Reinforcement Learning with Simultaneous Estimation of Rewards and Dynamics

Inverse Reinforcement Learning (IRL) describes the problem of learning an unknown reward function of a Markov Decision Process (MDP) from observed behavior of an agent. Since the agent's behavior originates in its policy and MDP policies depend on both the stochastic system dynamics as well as the reward function, the solution of the inverse problem is significantly influenced by both. Current IRL approaches assume that if the transition model is unknown, additional samples from the system's dynamics are accessible, or the observed behavior provides enough samples of the system's dynamics to solve the inverse problem accurately. These assumptions are often not satisfied. To overcome this, we present a gradient-based IRL approach that simultaneously estimates the system's dynamics. By solving the combined optimization problem, our approach takes into account the bias of the demonstrations, which stems from the generating policy. The evaluation on a synthetic MDP and a transfer learning task shows improvements regarding the sample efficiency as well as the accuracy of the estimated reward functions and transition models.

cs.AI

Predicting Lane Keeping Behavior of Visually Distracted Drivers Using Inverse Suboptimal Control

Driver distraction strongly contributes to crash-risk. Therefore, assistance systems that warn the driver if her distraction poses a hazard to road safety, promise a great safety benefit. Current approaches either seek to detect critical situations using environmental sensors or estimate a driver's attention state solely from her behavior. However, this neglects that driving situation, driver deficiencies and compensation strategies altogether determine the risk of an accident. This work proposes to use inverse suboptimal control to predict these aspects in visually distracted lane keeping. In contrast to other approaches, this allows a situation-dependent assessment of the risk posed by distraction. Real traffic data of seven drivers are used for evaluation of the predictive power of our approach. For comparison, a baseline was built using established behavior models. In the evaluation our method achieves a consistently lower prediction error over speed and track-topology variations. Additionally, our approach generalizes better to driving speeds unseen in training phase.

eess.SY

Observation of universal strong orbital-dependent correlation effects in iron chalcogenides

Establishing the appropriate theoretical framework for unconventional superconductivity in the iron-based materials requires correct understanding of both the electron correlation strength and the role of Fermi surfaces. This fundamental issue becomes especially relevant with the discovery of the iron chalcogenide (FeCh) superconductors, the only iron-based family in proximity to an insulating phase. Here, we use angle-resolved photoemission spectroscopy (ARPES) to measure three representative FeCh superconductors, FeTe0.56Se0.44, K0.76Fe1.72Se2, and monolayer FeSe film grown on SrTiO3. We show that, these FeChs are all in a strongly correlated regime at low temperatures, with an orbital-selective strong renormalization in the dxy bands despite having drastically different Fermi-surface topologies. Furthermore, raising temperature brings all three compounds from a metallic superconducting state to a phase where the dxy orbital loses all spectral weight while other orbitals remain itinerant. These observations establish that FeChs display universal orbital-selective strong correlation behaviors that are insensitive to the Fermi surface topology, and are close to an orbital-selective Mott phase (OSMP), hence placing strong constraints for theoretical understanding of iron-based superconductors.

cond-mat.supr-con

Visualizing the Radiation of the Kelvin-Helmholtz Instability

Emerging new technologies in plasma simulations allow tracking billions of particles while computing their radiative spectra. We present a visualization of the relativistic Kelvin-Helmholtz Instability from a simulation performed with the fully relativistic particle-in-cell code PIConGPU powered by 18,000 GPUs on the USA's fastest supercomputer Titan [1].

physics.plasm-ph

Direct observation of the transition from indirect to direct bandgap in atomically thin epitaxial MoSe2

Quantum systems in confined geometries are host to novel physical phenomena. Examples include quantum Hall systems in semiconductors and Dirac electrons in graphene. Interest in such systems has also been intensified by the recent discovery of a large enhancement in photoluminescence quantum efficiency and a potential route to valleytronics in atomically thin layers of transition metal dichalcogenides, MX2 (M = Mo, W; X = S, Se, Te), which are closely related to the indirect to direct bandgap transition in monolayers. Here, we report the first direct observation of the transition from indirect to direct bandgap in monolayer samples by using angle resolved photoemission spectroscopy on high-quality thin films of MoSe2 with variable thickness, grown by molecular beam epitaxy. The band structure measured experimentally indicates a stronger tendency of monolayer MoSe2 towards a direct bandgap, as well as a larger gap size, than theoretically predicted. Moreover, our finding of a significant spin-splitting of 180 meV at the valence band maximum of a monolayer MoSe2 film could expand its possible application to spintronic devices.

cond-mat.mtrl-sci

Route-Based Detection of Conflicting ATC Clearances on Airports

Runway incursions are among the most serious safety concerns in air traffic control. Traditional A-SMGCS level 2 safety systems detect runway incursions with the help of surveillance information only. In the context of SESAR, complementary safety systems are emerging that also use other information in addition to surveillance, and that aim at warning about potential runway incursions at earlier points in time. One such system is "conflicting ATC clearances", which processes the clearances entered by the air traffic controller into an electronic flight strips system and cross-checks them for potentially dangerous inconsistencies. The cross-checking logic may be implemented directly based on the clearances and on surveillance data, but this is cumbersome. We present an approach that instead uses ground routes as an intermediate layer, thereby simplifying the core safety logic.

eess.SY

Software Design Principles of a DFS Tower A-CWP Prototype

SESAR is supposed to boost the development of new operational procedures together with the supporting systems in order to modernize the pan-European air traffic management (ATM). One consequence of this development is that more and more information is presented to - and has to be processed by - air traffic control officers (ATCOs). Thus, there is a strong need for a software design concept that fosters the development of an advanced (tower) controller working position (A-CWP) that comprehensively integrates the still counting amount of information while reducing the data management workload of ATCOs. We report on our first hands-on experiences obtained during the development of an A-CWP prototype that was used in two SESAR validation sessions.

cs.SE

Ab-initio phase diagram of ultracold 87-Rb in an one-dimensional two-color superlattice

We investigate the ab-initio phase diagram of ultracold 87-Rb atoms in an one-dimensional two-color superlattice. Using single-particle band structure calculations we map the experimental setup onto the parameters of the Bose-Hubbard model. This ab-initio ansatz allows us to express the phase diagrams in terms of the experimental control parameters, i.e., the intensities of the lasers that form the optical superlattice. In order to solve the many-body problem for experimental system sizes we adopt the density-matrix renormalization-group algorithm. A detailed study of convergence and finite-size effects for all observables is presented. Our results show that all relevant quantum phases, i.e., superfluid, Mott-insulator, and quasi Bose-glass, can be accessed through intensity variation of the lasers alone. However, it turns out that the phase diagram is strongly affected by the longitudinal trapping potential.

cond-mat.quant-gas

Phase Diagram of Bosons in Two-Color Superlattices from Experimental Parameters

We study the zero-temperature phase diagram of a gas of bosonic 87-Rb atoms in two-color superlattice potentials starting directly from the experimental parameters, such as wavelengths and intensities of the two lasers generating the superlattice. In a first step, we map the experimental setup to a Bose-Hubbard Hamiltonian with site-dependent parameters through explicit band-structure calculations. In the second step, we solve the many-body problem using the density-matrix renormalization group (DMRG) approach and compute observables such as energy gap, condensate fraction, maximum number fluctuations and visibility of interference fringes. We study the phase diagram as function of the laser intensities s_2 and s_1 as control parameters and show that all relevant quantum phases, i.e. superfluid, Mott-insulator, and quasi Bose-glass phase, and the transitions between them can be investigated through a variation of these intensities alone.

cond-mat.quant-gas

Ultracold Bose gases in time-dependent 1D superlattices: response and quasimomentum structure

The response of ultracold atomic Bose gases in time-dependent optical lattices is discussed based on direct simulations of the time-evolution of the many-body state in the framework of the Bose-Hubbard model. We focus on small-amplitude modulations of the lattice potential as implemented in several recent experiment and study different observables in the region of the first resonance in the Mott-insulator phase. In addition to the energy transfer we investigate the quasimomentum structure of the system which is accessible via the matter-wave interference pattern after a prompt release. We identify characteristic correlations between the excitation frequency and the quasimomentum distribution and study their structure in the presence of a superlattice potential.

cond-mat.stat-mech