SearcharxivSearch

arXiv subjects

Andrew Schwartz

Publications and source records attributed to Andrew Schwartz.

8 recordsLinked to original sources

Can AI agents conduct open-ended AI research? Early evidence from two case studies

Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.

cs.AI

Open-World Evaluations for Measuring Frontier AI Capabilities

Benchmark-based evaluation remains important for tracking frontier AI progress. But it can both overstate and understate deployed capability because it privileges tasks that can be precisely specified, automatically graded, easy to optimize for, and run with low budgets and short time horizons. We advocate for a complementary class of evaluations, which we term open-world evaluations: long-horizon, messy, real-world tasks assessed through small-sample qualitative analysis rather than benchmark-scale automation. In this paper we survey recent open-world evaluations, identify their strengths and limitations, and introduce CRUX (Collaborative Research for Updating AI eXpectations), a project for conducting such evaluations regularly. As a first instance, we task an AI agent with developing and publishing a simple iOS application to the Apple App Store. The agent completed the task with only a single avoidable manual intervention, suggesting that open-world evaluations can provide early warning of capabilities that may soon become widespread. We conclude with recommendations for designing and reporting open-world evals.

cs.AI

FareShare: A Tool for Labor Organizers to Estimate Lost Wages and Contest Arbitrary AI and Algorithmic Deactivations

What happens when a rideshare driver is suddenly locked out of the platform connecting them to riders, wages, and daily work? Deactivation-the abrupt removal of gig workers' platform access-typically occurs through arbitrary AI and algorithmic decisions with little explanation or recourse. This represents one of the most severe forms of algorithmic control and often devastates workers' financial stability. Recent U.S. state policies now mandate appeals processes and recovering compensation during the period of wrongful deactivation based on past earnings. Yet, labor organizers still lack effective tools to support these complex, error-prone workflows. We designed FareShare, a computational tool automating lost wage estimation for deactivated drivers, through a 6 month partnership with the State of Washington's largest rideshare labor union. Over the following 3 months, our field deployment of FareShare registered 178 account signups. We observed that the tool could reduce lost wage calculation time by over 95%, eliminate manual data entry errors, and enable legal teams to generate arbitration-ready reports more efficiently. Beyond these gains, the deployment also surfaced important socio-technical challenges around trust, consent, and tool adoption in high-stakes labor contexts.

cs.CY

FairFare: A Tool for Crowdsourcing Rideshare Data to Empower Labor Organizers

Rideshare workers experience unpredictable working conditions due to gig work platforms' reliance on opaque AI and algorithmic systems. In response to these challenges, we found that labor organizers want data to help them advocate for legislation to increase the transparency and accountability of these platforms. To address this need, we collaborated with a Colorado-based rideshare union to develop FairFare, a tool that crowdsources and analyzes workers' data to estimate the take rate -- the percentage of the rider price retained by the rideshare platform. We deployed FairFare with our partner organization that collaborated with us in collecting data on 76,000+ trips from 45 drivers over 18 months. During evaluation interviews, organizers reported that FairFare helped influence the bill language and passage of Colorado Senate Bill 24-75, calling for greater transparency and data disclosure of platform operations, and create a national narrative. Finally, we reflect on complexities of translating quantitative data into policy outcomes, nature of community based audits, and design implications for future transparency tools.

cs.HC

The Zero Forcing Numbers of Peony Graphs and Web Graphs

The concept of zero forcing involves a dynamic coloring process by which blue vertices cause white vertices to become blue, with the goal of forcing the entire graph blue while choosing as few as possible vertices to be initially blue. Past research in this area has focused on structural arguments, with approaches varying from graph substructures to the interplay between local and global graph structures. This paper explores the use of these structural concepts when determining the zero forcing number of complex classes of graphs, specifically two infinite classes of graphs each defined on multiple parameters.

math.CO

The zero forcing numbers and propagation times of gear graphs and helm graphs

Zero forcing is a dynamic coloring process on graphs. Initially, each vertex of a graph is assigned a color of either blue or white, and then a process begins by which blue vertices force white vertices to become blue. The zero forcing number is the cardinality of the smallest set of initially blue vertices which can force the entire graph to become blue, and the propagation time is the minimum number of steps in such a zero forcing process. In this paper we will determine the zero forcing numbers and propagation times of two infinite classes of graphs called gear graphs and helm graphs.

math.CO

The frequency, temperature, and magnetic field dependence of ferromagnetic resonance and anti-resonance in La$_{0.8}$Sr$_{0.2}$MnO$_3$

Employing a broadband microwave reflection configuration, we have measured the complex surface impedance, $Z_S(ω,T,H)$, of single crystal La$_{0.8}$Sr$_{0.2}$MnO$_3$, as a function of frequency (0.045-45 GHz), temperature (250-325 K), and magnetic field (0-1.9 kOe). The microwave surface impedance depends not only on the resistivity of the material, but also on the magnetic permeability, $\hatμ(ω,T,H)$, which gives rise to ferromagnetic resonance (FMR) and ferromagnetic anti-resonance (FMAR). The broadband nature of this experiment allows us to follow the FMR to low frequency and to deduce the behavior of both the local internal fields and the local magnetization in the sample.

cond-mat.str-el

Determination of the magnetization scaling exponent for single crystal La$_{0.8}$Sr$_{0.2}$MnO$_3$ by broadband microwave surface impedance measurements

Employing a broadband microwave reflection configuration, we have measured the complex surface impedance, $Z_S(ω,T)$, of single crystal La$_{0.8}$Sr$_{0.2}$MnO$_3$, as a function of frequency (0.045-45 GHz) and temperature (250-325 K). Through the dependence of the microwave surface impedance on the magnetic permeability, $\hatμ(ω,T)$, we have studied the local magnetic behavior of this material, and have extracted the spontaneous magnetization, $M_0(T)$, in {\em zero applied field}. The broadband nature of these measurements and the fact that no external field is applied to the material provide a unique opportunity to analyze the critical behavior of the spontaneous magnetization at temperatures very close to the ferromagnetic phase transition. We find a Curie temperature $T_C=305.5\pm 0.5$ K and scaling exponent $β=0.45\pm 0.05$, in agreement with the prediction of mean-field theory. We also discuss other recent determinations of the magnetization critical exponent in this and similar materials and show why our results are more definitive.

cond-mat.str-el