SearcharxivSearch

arXiv subjects

Peter Mühlbacher

Publications and source records attributed to Peter Mühlbacher.

8 recordsLinked to original sources

Automating Forecasting Question Generation and Resolution for AI Evaluation

Forecasting future events is highly valuable in decision-making and is a robust measure of general intelligence. As forecasting is probabilistic, developing and evaluating AI forecasters requires generating large numbers of diverse and difficult questions, and accurately resolving them. Previous efforts to automate this laborious work relied on recurring data sources (e.g., weather, stocks), limiting diversity and utility. In this work, we present a system for generating and resolving high-quality forecasting questions automatically and at scale using LLM-powered web research agents. We use this system to generate 1499 diverse, real-world forecasting questions, and to resolve them several months later. We estimate that our system produces verifiable, unambiguous questions approximately 96% of the time, exceeding the rate of Metaculus, a leading human-curated forecasting platform. We also find that our system resolves questions at approximately 95% accuracy. We verify that forecasting agents powered by more intelligent LLMs perform better on these questions (Brier score of 0.134 for Gemini 3 Pro, 0.149 for GPT-5, and 0.179 for Gemini 2.5 Flash). Finally, we demonstrate how our system can be leveraged to directly improve forecasting, by evaluating a question decomposition strategy on a generated question set, yielding a significant improvement in Brier scores (0.132 vs. 0.141).

cs.LG

Bench to the Future: A Pastcasting Benchmark for Forecasting Agents

Forecasting is a challenging task that offers a clearly measurable way to study AI systems. Forecasting requires a large amount of research on the internet, and evaluations require time for events to happen, making the development of forecasting benchmarks challenging. To date, no forecasting benchmark provides a realistic, hermetic, and repeatable environment for LLM forecasters. We introduce Bench To the Future (BTF), a "pastcasting" benchmark with hundreds of high-quality questions for which the resolution is already known. Each question is accompanied by a large offline corpus of tens of thousands of relevant web pages, enabling a way to elicit realistic "forecasts" on past events from LLMs. Results suggest that our pastcasting environment can produce results comparable to those based on forecasts using the internet on at-the-time unresolved questions. We show results benchmarking agent and chain-of-thought forecasting approaches using several LLMs, including the recently-released Claude 4 models, and demonstrate BTF's ability to track steady forecasting capability progress over time. We intend this to be a living benchmark, with new questions added continually to account for increasing training data cutoff dates. We invite researchers to contact us at hello@futuresearch.ai to utilize our benchmark or tooling for their own research.

cs.CL

Deep Research Bench: Evaluating AI Web Research Agents

Amongst the most common use cases of modern AI is LLM chat with web search enabled. However, no direct evaluations of the quality of web research agents exist that control for the continually-changing web. We introduce Deep Research Bench, consisting of 89 multi-step web research task instances of varying difficulty across 8 diverse task categories, with the answers carefully worked out by skilled humans. We provide a "RetroSearch" environment with a large frozen set of scraped web pages, and demonstrate that offline "RetroSearch" agents perform comparably to "live web" agents, enabling reliable evaluations of models over time. We provide robust agent tooling and scaffolding to benchmark major LLMs as they are released, including "thinking" models like o3 and Gemini 2.5 Pro. We include automated evaluations of the lengthy agent traces to report progress over time in hallucinations, tool use, and forgetting. Finally, we evaluate the major web research products branded as "Deep Research", "Deep Search", "Search", or "Research." Results are available on a public leaderboard at https://drb.futuresearch.ai/.

cs.AI

Towards a Realistic Long-Term Benchmark for Open-Web Research Agents

We present initial results of a forthcoming benchmark for evaluating LLM agents on white-collar tasks of economic value. We evaluate agents on real-world "messy" open-web research tasks of the type that are routine in finance and consulting. In doing so, we lay the groundwork for an LLM agent evaluation suite where good performance directly corresponds to a large economic and societal impact. We built and tested several agent architectures with o1-preview, GPT-4o, Claude-3.5 Sonnet, Llama 3.1 (405b), and GPT-4o-mini. On average, LLM agents powered by Claude-3.5 Sonnet and o1-preview substantially outperformed agents using GPT-4o, with agents based on Llama 3.1 (405b) and GPT-4o-mini lagging noticeably behind. Across LLMs, a ReAct architecture with the ability to delegate subtasks to subagents performed best. In addition to quantitative evaluations, we qualitatively assessed the performance of the LLM agents by inspecting their traces and reflecting on their observations. Our evaluation represents the first in-depth assessment of agents' abilities to conduct challenging, economically valuable analyst-style research on the real open web.

cs.CL

Poisson-Dirichlet distributions and weakly first-order spin-nematic phase transitions

We provide a quantitative characterization of generic weakly first-order thermal phase transitions out of planar spin-nematic states in three-dimensional spin-one quantum magnets, based on calculations using Poisson-Dirichlet distributions (PD) within a universal loop model formulation, combined with large-scale quantum Monte Carlo calculations. In contrast to earlier claims, the thermal melting of the nematic state is not continuous, instead a weakly first-order transition is identified from both thermal properties and the distribution of the nematic order parameter. Furthermore, based on PD calculations, we obtain exact results for the order parameter distribution and Binder cumulants at the discontinuous melting transition. Our findings establish the thermal melting of planar spin-nematic states as a generic platform for quantitative approaches to weakly first-order phase transitions in quantum systems with a continuous SU(2) internal symmetry.

cond-mat.str-el

Dimerization in quantum spin chains with $O(n)$ symmetry

We consider quantum spins with $S\geq1$, and two-body interactions with $O(2S+1)$ symmetry. We discuss the ground state phase diagram of the one-dimensional system. We give a rigorous proof of dimerization for an open region of the phase diagram, for $S$ sufficiently large. We also prove the existence of a gap for excitations.

math-ph

Critical Parameters for Loop and Bernoulli Percolation

We consider a class of random loop models (including the random interchange process) that are parametrised by a time parameter $β\geq 0$. Intuitively, larger $β$ means more randomness. In particular, at $β=0$ we start with loops of length 1 and as $β$ crosses a critical value $β_c$, infinite loops start to occur almost surely. Our random loop models admit a natural comparison to bond percolation with $p=1-e^{-β}$ on the same graph to obtain a lower bound on $β_c$. For those graphs of diverging vertex degree where $β_c$ and the critical parameter for percolation have been calculated explicitly, that inequality has been found to be an equality. In contrast, we show in this paper that for graphs of bounded degree the inequality is strict, i.e. we show existence of an interval of values of $β$ where there are no infinite loops, but infinite percolation clusters almost surely.

math.PR

Bounds on the norm of Wigner-type random matrices

We consider a Wigner-type ensemble, i.e. large hermitian $N\times N$ random matrices $H=H^*$ with centered independent entries and with a general matrix of variances $S_{xy}=\mathbb E|H_{xy}|^2$. The norm of $H$ is asymptotically given by the maximum of the support of the self-consistent density of states. We establish a bound on this maximum in terms of norms of powers of $S$ that substantially improves the earlier bound $2\| S\|^{1/2}_\infty$ given in [arXiv:1506.05098]. The key element of the proof is an effective Markov chain approximation for the contributions of the weighted Dyck paths appearing in the iterative solution of the corresponding Dyson equation.

math.PR