SearcharxivSearch

arXiv subjects

Pranav Mehta

Publications and source records attributed to Pranav Mehta.

3 recordsLinked to original sources

Macro-Prudential AI Governance: A Two-Layer Early Warning and Response System for Frontier AI

Frontier-AI governance today faces a problem structurally analogous to the one banking regulation faced pre-2008, and which post-2008 reforms (Basel III, Dodd-Frank) have since addressed. Two gaps recur: discovering a risk is not tantamount to acting on it, and individual-model review is unlike managing correlated build-up across the sector. Drawing on the Basel III framework and the U.S. financial-stability architecture, I propose a macro-prudential early warning and response system ("MEWRS") for internal frontier AI. These are systems deployed for labs' own internal research, testing, and production workflows, as distinct from externally released products. Layer A adapts the finder-coordinator-defender early-warning model to route structured reports on dual-use capabilities, autonomy indicators, and security compromises through a government clearinghouse to domain-specific defender working groups. Layer B calibrates operational controls via three quantitative buffer metrics, namely Effective Compute-at-Risk (ECAR), Cumulative Red-Team Hours (CRTH), and an Alignment Robustness Score (ARS), so that faster capability scaling automatically triggers stronger safeguards, analogously to how risk-weighted assets drive capital ratios under Basel III. I outline the reporting schema, map six Basel III mechanisms onto AI-governance analogues, identify seven failure modes with concrete mitigations, and sketch an exercise-based validation plan. MEWRS is designed to detect correlated risk build-ups across the frontier-AI sector and create pre-committed off-ramps before a cascade unfolds.

cs.CY

APEX-SWE

We introduce the AI Productivity Index for Software Engineering (APEX-SWE), a benchmark for assessing whether frontier AI models can execute economically valuable software engineering work. Unlike existing evaluations that focus on narrow, well-defined tasks, APEX-SWE assesses two novel task types that reflect real-world software engineering: (1) Integration tasks (n=100), which require constructing end-to-end systems across heterogeneous cloud primitives, business applications, and infrastructure-as-code services, and (2) Observability tasks (n=100), which require debugging production failures using telemetry signals such as logs and dashboards, as well as unstructured context. We evaluated eleven frontier models for the APEX-SWE leaderboard. Claude Opus 4.6 leads the APEX-SWE leaderboard with 40.5% Pass@1, followed by Claude Opus 4.5 at 38.7%. Our analysis shows that strong performance is primarily driven by epistemic discipline, defined as the capacity to distinguish between assumptions and verified facts. It is often combined with systematic verification prior to acting. We open-source the APEX-SWE evaluation harness and a dev set (n=50).

cs.SE

Cell spheroid viscoelasticity is deformation-dependent

Tissue surface tension influences cell sorting and tissue fusion. Earlier mechanical studies suggest that multicellular spheroids actively reinforce their surface tension with applied force. Here we study this open question through high-throughput microfluidic micropipette aspiration measurements on cell spheroids to identify the role of force duration and cell contractility. We find that larger spheroid deformations lead to faster cellular retraction once the pressure is released, regardless of the applied force and cellular contractility. These new insights demonstrate that spheroid viscoelasticity is deformation-dependent and challenge whether surface tension truly reinforces.

cond-mat.soft