SearcharxivSearch

arXiv subjects

Michael Fink

Publications and source records attributed to Michael Fink.

At least 19 recordsLinked to original sources

The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality

We introduce The FACTS Leaderboard, an online leaderboard suite and associated set of benchmarks that comprehensively evaluates the ability of language models to generate factually accurate text across diverse scenarios. The suite provides a holistic measure of factuality by aggregating the performance of models on four distinct sub-leaderboards: (1) FACTS Multimodal, which measures the factuality of responses to image-based questions; (2) FACTS Parametric, which assesses models' world knowledge by answering closed-book factoid questions from internal parameters; (3) FACTS Search, which evaluates factuality in information-seeking scenarios, where the model must use a search API; and (4) FACTS Grounding (v2), which evaluates whether long-form responses are grounded in provided documents, featuring significantly improved judge models. Each sub-leaderboard employs automated judge models to score model responses, and the final suite score is an average of the four components, designed to provide a robust and balanced assessment of a model's overall factuality. The FACTS Leaderboard Suite will be actively maintained, containing both public and private splits to allow for external participation while guarding its integrity. It can be found at https://www.kaggle.com/benchmarks/google/facts .

cs.CL

Towards an AI-Augmented Textbook

Textbooks are a cornerstone of education, but they have a fundamental limitation: they are a one-size-fits-all medium. Any new material or alternative representation requires arduous human effort, so that textbooks cannot be adapted in a scalable manner. We present an approach for transforming and augmenting textbooks using generative AI, adding layers of multiple representations and personalization while maintaining content integrity and quality. We refer to the system built with this approach as Learn Your Way. We report pedagogical evaluations of the different transformations and augmentations, and present the results of a a randomized control trial, highlighting the advantages of learning with Learn Your Way over regular textbook usage.

cs.CY

Safe and Non-Conservative Trajectory Planning for Autonomous Driving Handling Unanticipated Behaviors of Traffic Participants

Trajectory planning for autonomous driving is challenging because the unknown future motion of traffic participants must be accounted for, yielding large uncertainty. Stochastic Model Predictive Control (SMPC)-based planners provide non-conservative planning, but do not rule out a (small) probability of collision. We propose a control scheme that yields an efficient trajectory based on SMPC when the traffic scenario allows, still avoiding that the vehicle causes collisions with traffic participants if the latter move according to the prediction assumptions. If some traffic participant does not behave as anticipated, no safety guarantee can be given. Then, our approach yields a trajectory which minimizes the probability of collision, using Constraint Violation Probability Minimization techniques. Our algorithm can also be adapted to minimize the anticipated harm caused by a collision. We provide a thorough discussion of the benefits of our novel control scheme and compare it to a previous approach through numerical simulations from the CommonRoad database.

eess.SY

Minimal Constraint Violation Probability in Model Predictive Control for Linear Systems

Handling uncertainty in model predictive control comes with various challenges, especially when considering state constraints under uncertainty. Most methods focus on either the conservative approach of robustly accounting for uncertainty or allowing a small probability of constraint violation. In this work, we propose a linear model predictive control approach that minimizes the probability that linear state constraints are violated in the presence of additive uncertainty. This is achieved by first determining a set of inputs that minimize the probability of constraint violation. Then, this resulting set is used to define admissible inputs for the optimal control problem. Recursive feasibility is guaranteed and input-to-state stability is proved under assumptions. Numerical results illustrate the benefits of the proposed model predictive control approach.

eess.SY

Optimal Control for Indoor Vertical Farms Based on Crop Growth

Vertical farming allows for year-round cultivation of a variety of crops, overcoming environmental limitations and ensuring food security. This closed and highly controlled system allows the plants to grow in optimal conditions, so that they reach maturity faster and yield more than on a conventional outdoor farm. However, one of the challenges of vertical farming is the high energy consumption. In this work, we optimize wheat growth using an optimal control approach with two objectives: first, we optimize inputs such as water, radiation, and temperature for each day of the growth cycle, and second, we optimize the duration of the plant's growth period to achieve the highest possible yield over a whole year. For this, we use a nonlinear, discrete-time hybrid model based on a simple universal crop model that we adapt to make the optimization more efficient. Using our approach, we find an optimal trade-off between used resources, net profit of the yield, and duration of a cropping period, thus increasing the annual yield of crops significantly while keeping input costs as low as possible. This work demonstrates the high potential of control theory in the discipline of vertical farming.

eess.SY

Comparison of Dynamic Tomato Growth Models for Optimal Control in Greenhouses

As global demand for efficiency in agriculture rises, there is a growing interest in high-precision farming practices. Particularly greenhouses play a critical role in ensuring a year-round supply of fresh produce. In order to maximize efficiency and productivity while minimizing resource use, mathematical techniques such as optimal control have been employed. However, selecting appropriate models for optimal control requires domain expertise. This study aims to compare three established tomato models for their suitability in an optimal control framework. Results show that all three models have similar yield predictions and accuracy, but only two models are currently applicable for optimal control due to implementation limitations. The two remaining models each have advantages in terms of economic yield and computation times, but the differences in optimal control strategies suggest that they require more accurate parameter identification and calibration tailored to greenhouses.

math.OC

Dynamic Planning in Open-Ended Dialogue using Reinforcement Learning

Despite recent advances in natural language understanding and generation, and decades of research on the development of conversational bots, building automated agents that can carry on rich open-ended conversations with humans "in the wild" remains a formidable challenge. In this work we develop a real-time, open-ended dialogue system that uses reinforcement learning (RL) to power a bot's conversational skill at scale. Our work pairs the succinct embedding of the conversation state generated using SOTA (supervised) language models with RL techniques that are particularly suited to a dynamic action space that changes as the conversation progresses. Trained using crowd-sourced data, our novel system is able to substantially exceeds the (strong) baseline supervised model with respect to several metrics of interest in a live experiment with real users of the Google Assistant.

cs.CL

Implementation of Linear Model Predictive Control -- Tutorial

This tutorial shows an overview of Model Predictive Control with a linear discrete-time system and constrained states and inputs. The focus is on the implementation of the method under consideration of stability and recursive feasibility. The MATLAB code for the examples and plots is available online.

eess.SY

Baseline Detection in Historical Documents using Convolutional U-Nets

Baseline detection is still a challenging task for heterogeneous collections of historical documents. We present a novel approach to baseline extraction in such settings, turning out the winning entry to the ICDAR 2017 Competition on Baseline detection (cBAD). It utilizes deep convolutional nets (CNNs) for both, the actual extraction of baselines, as well as for a simple form of layout analysis in a pre-processing step. To the best of our knowledge it is the first CNN-based system for baseline extraction applying a U-net architecture and sliding window detection, profiting from a high local accuracy of the candidate lines extracted. Final baseline post-processing complements our approach, compensating for inaccuracies mainly due to missing context information during sliding window detection. We experimentally evaluate the components of our system individually on the cBAD dataset. Moreover, we investigate how it generalizes to different data by means of the dataset used for the baseline extraction task of the ICDAR 2017 Competition on Layout Analysis for Challenging Medieval Manuscripts (HisDoc). A comparison with the results reported for HisDoc shows that it also outperforms the contestants of the latter.

cs.CV

Three-dimensional simulations of gravitationally confined detonations compared to observations of SN 1991T

The gravitationally confined detonation (GCD) model has been proposed as a possible explosion mechanism for Type Ia supernovae in the single-degenerate evolution channel. Driven by buoyancy, a deflagration flame rises in a narrow cone towards the surface. For the most part, the flow of the expanding ashes remains radial, but upon reaching the outer, low-pressure layers of the white dwarf, an additional lateral component develops. This makes the deflagration ashes converge again at the opposite side, where the compression heats fuel and a detonation may be launched. To test the GCD explosion model, we perform a 3D simulation for a model with an ignition spot offset near the upper limit of what is still justifiable, 200 km. This simulation meets our deliberately optimistic detonation criteria and we initiate a detonation. The detonation burns through the white dwarf and leads to its complete disruption. We determine nucleosynthetic yields by post-processing 10^6 tracer particles with a 384 nuclide reaction network and we present multi-band light curves and time-dependent optical spectra. We find that our synthetic observables show a prominent viewing-angle sensitivity in UV and blue bands, which is in tension with observed SNe Ia. The strong dependence on viewing-angle is caused by the asymmetric distribution of the deflagration ashes in the outer ejecta layers. Finally, we perform a comparison of our model to SN 1991T. The overall flux-level of the model is slightly too low and the model predicts pre-maximum light spectral features due to Ca, S, and Si that are too strong. Furthermore, the model chemical abundance stratification qualitatively disagrees with recent abundance tomography results in two key areas: our model lacks low velocity stable Fe and instead has copious amounts of high-velocity 56Ni and stable Fe. We therefore do not find good agreement of the model with SN 1991T.

astro-ph.SR

A model building framework for Answer Set Programming with external computations

As software systems are getting increasingly connected, there is a need for equipping nonmonotonic logic programs with access to external sources that are possibly remote and may contain information in heterogeneous formats. To cater for this need, HEX programs were designed as a generalization of answer set programs with an API style interface that allows to access arbitrary external sources, providing great flexibility. Efficient evaluation of such programs however is challenging, and it requires to interleave external computation and model building; to decide when to switch between these tasks is difficult, and existing approaches have limited scalability in many real-world application scenarios. We present a new approach for the evaluation of logic programs with external source access, which is based on a configurable framework for dividing the non-ground program into possibly overlapping smaller parts called evaluation units. The latter will be processed by interleaving external evaluation and model building using an evaluation graph and a model graph, respectively, and by combining intermediate results. Experiments with our prototype implementation show a significant improvement compared to previous approaches. While designed for HEX-programs, the new evaluation approach may be deployed to related rule-based formalisms as well.

cs.AI

Towards Ideal Semantics for Analyzing Stream Reasoning

The rise of smart applications has drawn interest to logical reasoning over data streams. Recently, different query languages and stream processing/reasoning engines were proposed in different communities. However, due to a lack of theoretical foundations, the expressivity and semantics of these diverse approaches are given only informally. Towards clear specifications and means for analytic study, a formal framework is needed to define their semantics in precise terms. To this end, we present a first step towards an ideal semantics that allows for exact descriptions and comparisons of stream reasoning systems.

cs.AI

Workshop Notes of the 6th International Workshop on Acquisition, Representation and Reasoning about Context with Logic (ARCOE-Logic 2014)

ARCOE-Logic 2014, the 6th International Workshop on Acquisition, Representation and Reasoning about Context with Logic, was held in co-location with the 19th International Conference on Knowledge Engineering and Knowledge Management (EKAW 2014) on November 25, 2014 in Linköping, Sweden. These notes contain the five papers which were accepted and presented at the workshop.

cs.AI

Causal Graph Justifications of Logic Programs

In this work we propose a multi-valued extension of logic programs under the stable models semantics where each true atom in a model is associated with a set of justifications. These justifications are expressed in terms of causal graphs formed by rule labels and edges that represent their application ordering. For positive programs, we show that the causal justifications obtained for a given atom have a direct correspon- dence to (relevant) syntactic proofs of that atom using the program rules involved in the graphs. The most interesting contribution is that this causal information is obtained in a purely semantic way, by algebraic op- erations (product, sum and application) on a lattice of causal values whose ordering relation expresses when a justification is stronger than another. Finally, for programs with negation, we define the concept of causal stable model by introducing an analogous transformation to Gelfond and Lifschitz's program reduct. As a result, default negation behaves as "absence of proof" and no justification is derived from negative liter

cs.AI

The white dwarf's carbon fraction as a secondary parameter of Type Ia supernovae

Binary stellar evolution calculations predict that Chandrasekhar-mass carbon/oxygen white dwarfs (WDs) show a radially varying profile for the composition with a carbon depleted core. Many recent multi-dimensional simulations of Type Ia supernovae (SNe Ia), however, assume the progenitor WD has a homogeneous chemical composition. In this work, we explore the impact of different initial carbon profiles of the progenitor WD on the explosion phase and on synthetic observables in the Chandrasekhar-mass delayed detonation model. Spectra and light curves are compared to observations to judge the validity of the model. The explosion phase is simulated using the finite volume supernova code LEAFS, which is extended to treat different compositions of the progenitor WD. The synthetic observables are computed with the Monte Carlo radiative transfer code ARTIS. Differences in binding energies of carbon and oxygen lead to a lower nuclear energy release for carbon depleted material; thus, the burning fronts that develop are weaker and the total nuclear energy release is smaller. For otherwise identical conditions, carbon depleted models produce less Ni-56. Comparing different models with similar Ni-56 yields shows lower kinetic energies in the ejecta for carbon depleted models, but only small differences in velocity distributions and line velocities in spectra. The light curve width-luminosity relation (WLR) obtained for models with differing carbon depletion is roughly perpendicular to the observed WLR, hence the carbon mass fraction is probably only a secondary parameter in the family of SNe Ia.

astro-ph.SR

Proceedings of Answer Set Programming and Other Computing Paradigms (ASPOCP 2013), 6th International Workshop, August 25, 2013, Istanbul, Turkey

This volume contains the papers presented at the sixth workshop on Answer Set Programming and Other Computing Paradigms (ASPOCP 2013) held on August 25th, 2013 in Istanbul, co-located with the 29th International Conference on Logic Programming (ICLP 2013). It thus continues a series of previous events co-located with ICLP, aiming at facilitating the discussion about crossing the boundaries of current ASP techniques in theory, solving, and applications, in combination with or inspired by other computing paradigms.

cs.AI

Spectral modelling of the "Super-Chandra" Type Ia SN 2009dc - testing a 2 M_sun white dwarf explosion model and alternatives

Extremely luminous, super-Chandrasekhar (SC) Type Ia Supernovae (SNe Ia) are as yet an unexplained phenomenon. We analyse a well-observed SN of this class, SN 2009dc, by modelling its photospheric spectra with a spectral synthesis code, using the technique of 'Abundance Tomography'. We present spectral models based on different density profiles, corresponding to different explosion scenarios, and discuss their consistency. First, we use a density structure of a simulated explosion of a 2 M_sun rotating C-O white dwarf (WD), which is often proposed as a possibility to explain SC SNe Ia. Then, we test a density profile empirically inferred from the evolution of line velocities (blueshifts). This model may be interpreted as a core-collapse SN with an ejecta mass ~ 3 M_sun. Finally, we calculate spectra assuming an interaction scenario. In such a scenario, SN 2009dc would be a standard WD explosion with a normal intrinsic luminosity, and this luminosity would be augmented by interaction of the ejecta with a H-/He-poor circumstellar medium. We find that no model tested easily explains SN 2009dc. With the 2 M_sun WD model, our abundance analysis predicts small amounts of burning products in the intermediate-/high-velocity ejecta (v > 9000 km/s). However, in the original explosion simulations, where the nuclear energy release per unit mass is large, burned material is present at high v. This contradiction can only be resolved if asymmetries strongly affect the radiative transfer or if C-O WDs with masses significantly above 2 M_sun exist. In a core-collapse scenario, low velocities of Fe-group elements are expected, but the abundance stratification in SN 2009dc seems 'SN Ia-like'. The interaction-based model looks promising, and we have some speculations on possible progenitor configurations. However, radiation-hydro simulations will be needed to judge whether this scenario is realistic at all.

astro-ph.SR

Proceedings of Answer Set Programming and Other Computing Paradigms (ASPOCP 2012), 5th International Workshop, September 4, 2012, Budapest, Hungary

This volume contains the papers presented at the fifth workshop on Answer Set Programming and Other Computing Paradigms (ASPOCP 2012) held on September 4th, 2012 in Budapest, co-located with the 28th International Conference on Logic Programming (ICLP 2012). It thus continues a series of previous events co-located with ICLP, aiming at facilitating the discussion about crossing the boundaries of current ASP techniques in theory, solving, and applications, in combination with or inspired by other computing paradigms.

cs.AI