SearcharxivSearch

arXiv subjects

Jason Poulos

Publications and source records attributed to Jason Poulos.

12 recordsLinked to original sources

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not sufficiently difficult to meaningfully measure frontier models. To this end, we present Terminal-Bench 2.0: a carefully curated hard benchmark composed of 89 tasks in computer terminal environments inspired by problems from real workflows. Each task features a unique environment, human-written solution, and comprehensive tests for verification. We show that frontier models and agents score less than 65\% on the benchmark and conduct an error analysis to identify areas for model and agent improvement. We publish the dataset and evaluation harness to assist developers and researchers in future work at https://www.tbench.ai/ .

cs.SE

Targeted learning in observational studies with multi-valued treatments: An evaluation of antipsychotic drug treatment safety

We investigate estimation of causal effects of multiple competing (multi-valued) treatments in the absence of randomization. Our work is motivated by an intention-to-treat study of the relative cardiometabolic risk of assignment to one of six commonly prescribed antipsychotic drugs in a cohort of nearly 39,000 adults with serious mental illnesses. Doubly-robust estimators, such as targeted minimum loss-based estimation (TMLE), require correct specification of either the treatment model or outcome model to ensure consistent estimation; however, common TMLE implementations estimate treatment probabilities using multiple binomial regressions rather than multinomial regression. We implement a TMLE estimator that uses multinomial treatment assignment and ensemble machine learning to estimate average treatment effects. Our multinomial implementation improves coverage, but does not necessarily reduce bias, relative to the binomial implementation in simulation experiments with varying treatment propensity overlap and event rates. Evaluating the causal effects of the antipsychotics on three-year diabetes risk or death, we find a safety benefit of moving from a second-generation drug considered among the safest of the second-generation drugs to an infrequently prescribed first-generation drug known for having low cardiometabolic risk.

stat.AP

Gender gaps in frontier entrepreneurship? Evidence from 1901 Oklahoma land lottery winners

The paper investigates gender differences in entrepreneurship by exploiting a large-scale land lottery in Oklahoma at the turn of the 20$^{\text{th}}$ century. Lottery winners claimed land in the order in which their names were drawn, so the draw number is an approximate rank ordering of lottery wealth. This mechanism allows for the estimation of a dose-response function, which relates each draw number to the expected outcome under each draw. I estimate dose-response functions on a linked dataset of lottery winners and land patent records, and find the probability of purchasing land from the government to be decreasing as a function of lottery wealth, which is evidence for the presence of liquidity constraints. I find female winners were more effective in leveraging lottery wealth to purchase additional land, as evidenced by significantly higher median dose-responses compared to those of male winners. For a sample of winners linked to the 1910 Census, I find that male winners have higher median dose-responses compared to female winners in terms of farm or home ownership. These results suggest that liquidity constraints may have been more binding for female entrepreneurs in the market economy.

econ.GN

Retrospective causal inference via matrix completion, with an evaluation of the effect of European integration on cross-border employment

We propose a method of retrospective counterfactual imputation in panel data settings with later-treated and always-treated units, but no never-treated units. We use the observed outcomes to impute the counterfactual outcomes of the later-treated using a matrix completion estimator. We propose a novel propensity-score and elapsed-time weighting of the estimator's objective function to correct for differences in the observed covariate and unobserved fixed effects distributions, and elapsed time since treatment between groups. Our methodology is motivated by studying the effect of two milestones of European integration -- the Free Movement of persons and the Schengen Agreement -- on the share of cross-border workers in sending border regions. We apply the proposed method to the European Labour Force Survey (ELFS) data and provide evidence that opening the border almost doubled the probability of working beyond the border in Eastern European regions.

stat.ME

Amnesty Policy and Elite Persistence in the Postbellum South: Evidence from a Regression Discontinuity Design

This paper investigates the impact of Reconstruction-era amnesty policy on the officeholding and wealth of elites in the postbellum South. Amnesty policy restricted the political and economic rights of Southern elites for nearly three years during Reconstruction. I estimate the effect of being excluded from amnesty on elites' future wealth and political power using a regression discontinuity design that compares individuals just above and below a wealth threshold that determined exclusion from amnesty. Results on a sample of Reconstruction convention delegates show that exclusion from amnesty significantly decreased the likelihood of ex-post officeholding. I find no evidence that exclusion impacted later census wealth for Reconstruction delegates or for a larger sample of known slaveholders who lived in the South in 1860. These findings are in line with previous studies evidencing both changes to the identity of the political elite, and the continuity of economic mobility among the planter elite across the Civil War and Reconstruction.

econ.GN

Are deep learning models superior for missing data imputation in large surveys? Evidence from an empirical comparison

Multiple imputation (MI) is a popular approach for dealing with missing data arising from non-response in sample surveys. Multiple imputation by chained equations (MICE) is one of the most widely used MI algorithms for multivariate data, but it lacks theoretical foundation and is computationally intensive. Recently, missing data imputation methods based on deep learning models have been developed with encouraging results in small studies. However, there has been limited research on evaluating their performance in realistic settings compared to MICE, particularly in big surveys. We conduct extensive simulation studies based on a subsample of the American Community Survey to compare the repeated sampling properties of four machine learning based MI methods: MICE with classification trees, MICE with random forests, generative adversarial imputation networks, and multiple imputation using denoising autoencoders. We find the deep learning imputation methods are superior to MICE in terms of computational time. However, with the default choice of hyperparameters in the common software packages, MICE with classification trees consistently outperforms, often by a large margin, the deep learning imputation methods in terms of bias, mean squared error, and coverage under a range of realistic settings.

cs.LG

Adversarial Machine Learning: Bayesian Perspectives

Adversarial Machine Learning (AML) is emerging as a major field aimed at protecting machine learning (ML) systems against security threats: in certain scenarios there may be adversaries that actively manipulate input data to fool learning systems. This creates a new class of security vulnerabilities that ML systems may face, and a new desirable property called adversarial robustness essential to trust operations based on ML outputs. Most work in AML is built upon a game-theoretic modelling of the conflict between a learning system and an adversary, ready to manipulate input data. This assumes that each agent knows their opponent's interests and uncertainty judgments, facilitating inferences based on Nash equilibria. However, such common knowledge assumption is not realistic in the security scenarios typical of AML. After reviewing such game-theoretic approaches, we discuss the benefits that Bayesian perspectives provide when defending ML-based systems. We demonstrate how the Bayesian approach allows us to explicitly model our uncertainty about the opponent's beliefs and interests, relaxing unrealistic assumptions, and providing more robust inferences. We illustrate this approach in supervised learning settings, and identify relevant future research problems.

cs.AI

State-Building through Public Land Disposal? An Application of Matrix Completion for Counterfactual Prediction

This paper examines how homestead policies, which opened vast frontier lands for settlement, influenced the development of American frontier states. It uses a treatment propensity-weighted matrix completion model to estimate the counterfactual size of these states without homesteading. In simulation studies, the method shows lower bias and variance than other estimators, particularly in higher complexity scenarios. The empirical analysis reveals that homestead policies significantly and persistently reduced state government expenditure and revenue. These findings align with continuous difference-in-differences estimates using 1.46 million land patent records. This study's extension of the matrix completion method to include propensity score weighting for causal effect estimation in panel data, especially in staggered treatment contexts, enhances policy evaluation by improving the precision of long-term policy impact assessments.

econ.GN

Estimating population average treatment effects from experiments with noncompliance

Randomized control trials (RCTs) are the gold standard for estimating causal effects, but often use samples that are non-representative of the actual population of interest. We propose a reweighting method for estimating population average treatment effects in settings with noncompliance. Simulations show the proposed compliance-adjusted population estimator outperforms its unadjusted counterpart when compliance is relatively low and can be predicted by observed covariates. We apply the method to evaluate the effect of Medicaid coverage on health care use for a target population of adults who may benefit from expansions to the Medicaid program. We draw RCT data from the Oregon Health Insurance Experiment, where less than one-third of those randomly selected to receive Medicaid benefits actually enrolled.

stat.ME

Character-Based Handwritten Text Transcription with Attention Networks

The paper approaches the task of handwritten text recognition (HTR) with attentional encoder-decoder networks trained on sequences of characters, rather than words. We experiment on lines of text from popular handwriting datasets and compare different activation functions for the attention mechanism used for aligning image pixels and target characters. We find that softmax attention focuses heavily on individual characters, while sigmoid attention focuses on multiple characters at each step of the decoding. When the sequence alignment is one-to-one, softmax attention is able to learn a more precise alignment at each step of the decoding, whereas the alignment generated by sigmoid attention is much less precise. When a linear function is used to obtain attention weights, the model predicts a character by looking at the entire sequence of characters and performs poorly because it lacks a precise alignment between the source and target. Future research may explore HTR in natural scene images, since the model is capable of transcribing handwritten text without the need for producing segmentations or bounding boxes of text in images.

cs.CV

RNN-based counterfactual prediction, with an application to homestead policy and public schooling

This paper proposes a method for estimating the effect of a policy intervention on an outcome over time. We train recurrent neural networks (RNNs) on the history of control unit outcomes to learn a useful representation for predicting future outcomes. The learned representation of control units is then applied to the treated units for predicting counterfactual outcomes. RNNs are specifically structured to exploit temporal dependencies in panel data, and are able to learn negative and nonlinear interactions between control unit outcomes. We apply the method to the problem of estimating the long-run impact of U.S. homestead policy on public school spending.

stat.ML

Missing Data Imputation for Supervised Learning

Missing data imputation can help improve the performance of prediction models in situations where missing data hide useful information. This paper compares methods for imputing missing categorical data for supervised classification tasks. We experiment on two machine learning benchmark datasets with missing categorical data, comparing classifiers trained on non-imputed (i.e., one-hot encoded) or imputed data with different levels of additional missing-data perturbation. We show imputation methods can increase predictive accuracy in the presence of missing-data perturbation, which can actually improve prediction accuracy by regularizing the classifier. We achieve the state-of-the-art on the Adult dataset with missing-data perturbation and k-nearest-neighbors (k-NN) imputation.

stat.ML