PopPy: Opportunistically Exploiting Parallelism in Python Agentic Workflows
Agentic workflows, which compose calls to ML models using a general-purpose programming language like Python, are widely used for a variety of user-facing tasks, from software engineering to enterprise automation, making their end-to-end latency a critical bottleneck. To improve their latency, developers can either manually parallelize them or use restricted agentic frameworks that cannot express or exploit all forms of parallelization. We address this by developing PopPy, a system that parallelizes external calls, e.g., AI models, in agentic workflows written in a very expressive subset of Python. PopPy combines an ahead-of-time compiler with a runtime, addressing three key challenges in extracting parallelism from Python applications: language complexity, dynamic dispatch, and variable mutation. On a set of real-world agentic workflows, PopPy achieves up to $6.7\times$ speedups in end-to-end execution time, compared to standard Python execution and several agentic frameworks, while preserving the sequential program semantics.