Searcharxiv⌕ Search

arXiv subjects

Maksym Nechepurenko

Publications and source records attributed to Maksym Nechepurenko.

At least 19 recordsLinked to original sources

Two Models of Event Finality: Functional Alignment, Contestability, and Empirical Comparability on Polymarket and Kalshi

Event contracts reach economic finality through different institutional paths. Polymarket distinguishes oracle adjudication, adapter consumption, Conditional Tokens payout recording, technical redeemability, and optional holder redemption. Kalshi distinguishes venue determination, public finalization, lifecycle messages, and exact REST settlement fields. Resolved, settled, and finalized do not themselves define comparable endpoints. We define a mechanism-aware comparison certificate using functional roles rather than common labels. A scalar cross-venue duration is admissible only when payoff semantics, endpoint functions, calendar support, censoring, observation grades, and dependence units are jointly aligned. We separate exact paired, interval-qualified paired, and standardized unpaired estimands. Polymarket evidence contains 108,638 exact-linked conditions, 99,283 protocol payout records, 92,158 observed redemptions of any amount, and 91,817 observed positive-payout redemptions. Kalshi evidence contains 152,694 ordinary markets reconstructed as at risk at the enrollment boundary and 7,611,594 exact MVE market objects. Exact public endpoints cover 71,657 ordinary markets and 7,357,576 MVE objects. Exact determination-to-endpoint pairs total 70,979 for ordinary markets and 126,806 for MVE. These are Kalshi-native finality results, not Polymarket-Kalshi estimates. REST endpoint completeness can be high when exact lifecycle paths are sparse, and transport gaps can preserve two exact clocks while preventing claims about intermediate revisions. Current MVE timer fields do not identify historical version-consistent timer rules, and public finalization does not identify member cash. Cross-venue numerical comparison remains blocked until Paper 7.3 supplies blind first/stable decidability clocks and a semantic/calendar registry establishes genuinely comparable events and endpoints.

q-fin.TR↗

Price Discovery at the Boundary of Contractual Decidability: Terminal-Value Gaps, Trading Availability, and Venue Finality on Kalshi

This paper studies price discovery around contractual decidability rather than an arbitrary venue label. Its upstream lifecycle and decidability clocks are specified in Papers 7.1 and 7.3. Historical venue endpoints remain useful background: 152,694 ordinary markets form the retrospective feasibility denominator, 71,657 have an exact public endpoint, and 70,979 have an exact determination-to-endpoint pair. Those fields do not supply a contractual-decidability clock. The completed historical recovery produced no historically admissible contractual-decidability cohort. The prospective infrastructure shakedown has passed and evidence enrollment is active, but production price extraction has not started. The primary binary cohort will be drawn from the prospectively enrolled and blind-adjudicated contractual-decidability frame, with an accepted exact or interval first-decidability clock, exact terminal payoff, trading-availability classification, and admissible bounded non-block trade or real-candle coverage. Markets tradable after decidability enter a reaction cohort; markets closed before decidability enter a stale-terminal cohort. Closure is a competing event for convergence, not ordinary missingness. A 20-market price pilot may validate acquisition, block-trade treatment, synthetic-candle rejection, staleness, and clock alignment only after blind packet lock and only from accepted prospective clock candidates. It does not create price estimates. The paper specifies a prospective primary cohort and a targeted pilot protocol. Price estimation awaits independently reconstructed first/stable-decidability clocks and source-release clusters. No broad exchange-wide trade crawl or post hoc clock substitution is permitted.

q-fin.TR↗

Outcome Determination and Settlement Finality on Kalshi: Public State Paths, Prospective Measurement, and Empirical Identification

Event-contract settlement is a state path rather than a universal timestamp. Using a registered seven-day enrollment and seven-day administrative follow-up on Kalshi, this paper distinguishes market-object membership, lifecycle messages, endpoint states, and observation gaps. The MVE layer contains 7,611,594 tickers created inside the registered half-open interval; the reconstructed ordinary stock contains 152,694 markets at risk at the enrollment boundary, 3,477,227 already terminal records, and 11,410 boundary-uncertain records. The follow-up identifies exact public endpoints for 7,357,576 MVE market objects: 7,237,993 through REST settlement time alone, 119,212 agreeing across WebSocket and REST, and 371 WebSocket-only; 254,018 lack an admissible exact endpoint. In the ordinary boundary cohort, 71,657 markets have an exact WebSocket endpoint and 81,037 do not. These are venue-level public-finality outcomes, not member-cash observations. Exact determination-to-endpoint pairs are available for 126,806 MVE and 70,979 ordinary tickers, with respective mean durations of 9.676 and 322.50 seconds and observed maxima of 15 and 7,133 seconds. The contrast is descriptive, not causal. Transport gaps preserve two exact endpoint clocks but prevent the same rows from proving complete intermediate paths or absence of revisions; current settlement timers cannot be bound to historical rule versions. Public endpoint completion is therefore extensive, especially on REST, while path completeness remains narrower and family-specific. The exact-key historical crosswalk records earlier interim observations without changing the final cohort or endpoint facts.

q-fin.TR↗

From Public Evidence to Contractual Outcome: First and Stable Decidability on Kalshi

Public evidence can become sufficient to settle a prediction-market contract before the venue records its first determination, but the relevant boundary depends on the applicable rule version, exact release object, source hierarchy, correction history, and unfinished contract conditions. This paper defines two Kalshi clocks: first decidability, the earliest contemporaneous singleton in the rule-evidence mapping, and stable decidability, the retrospective earliest time after which the same singleton remains unchanged through finalization. A completed retrospective identification test establishes a narrow feasibility result. In a frozen blind pilot, all 25 identities and blinding checks passed and current rule text was recovered for all 25; no exact or bounded historical rule version and no exact or bounded official source-release object was recovered. The full historical recovery covered 152,694 ordinary tickers, 11,530 exact event identities, and 6,540 read-only official requests, with zero historically eligible events and zero historically eligible tickers. This is an observability result, not a claim that no market was decidable or that public evidence never existed. A completed prospective infrastructure shakedown established observation capability for three source programmes across 25 markets, with integrity revalidation of 781,266 lifecycle frames, 22 closed lower-bounded reconnect receipts, no unresolved reconnect gap, no due-but-missed official release, and an inactive price layer. Production evidence enrollment is active. The prospective sample is constructed only at enrollment close from prospectively frozen identities and pre-outcome fields; its frozen target size is selected mechanically under the registered full, reduced, exploratory, or no-go support disposition. No contractual-decidability clock, human-adjudication, price, or cross-venue result is reported here.

q-fin.TR↗

Resolution Is Not Settlement, Part I: Oracle Adjudication and Semantic Governance on Polymarket

Prediction-market resolution is often reduced to a terminal outcome and one timestamp. That representation is inadequate for leveraged event claims because rule versioning, request creation, proposal, dispute, reset, Oracle finality, and adapter terminality are distinct states with different observation precision and balance-sheet consequences. We reconstruct those states for Polymarket using Oracle request generations as the unit of adjudication. The population is frozen at Polygon block 79,721,080 and contains 185,550 initialized adapter-question instances and 350,703 decoded adapter logs. Exact requester-filtered extraction yields 504,332 decoded Oracle lifecycle events: 184,148 request creations, 159,447 proposals, 1,604 disputes, and 159,133 settlements. The accounting closes as 182,671 questions with at least one request plus 1,477 successor generations. Immutable chain identity and deployed request semantics provide exact linkage; unfinished histories remain right-censored. Request age is not semantic resolution age. Median request-to-first-proposal time is 182 seconds on the legacy route but 176,388-744,151 seconds on modern routes; post-reset successor proposals arrive within 300-2,909 seconds at the median. This descriptive contrast does not establish causal efficiency. Exact stable-ID metadata linkage recovers 104,032 of 185,550 questions (56.07%), leaves 81,518 unmatched, and produces no ambiguous exact match. External-source publication and contractual-decidability clocks remain unmeasured population-wide and are not replaced by mechanism timestamps. The results establish an event-sourced account of Oracle adjudication, semantic governance, and adapter terminality. Companion Part II reconstructs Conditional Tokens payout recording and observed redemption, preserving the boundary between adjudication, protocol settlement, and holder realization.

q-fin.TR↗

Resolution Is Not Settlement, Part II: Protocol Finality and Observed Redemption on Polymarket

An Oracle result is not yet a protocol payout, a redeemable position is not yet collateral in a holder's account, and a redemption event is not a complete measure of economic entitlement. This companion paper develops an event-sourced framework for Polymarket conditions from preparation through protocol finality and observed holder realization. The empirical design uses three Conditional Tokens Framework event families derived from a pinned contract application binary interface (ABI): ConditionPreparation, ConditionResolution, and PayoutRedemption. It separates the contract-wide acquisition universe from the frozen Polymarket adapter-question cohort. The exact bridge contains 108,638 linked conditions. Of these, 99,283 have an observed protocol-resolution event by the fixed snapshot at Polygon block 90,114,204 (2026-07-12T17:11:41Z). Among resolved exact-linked conditions, 92,158 have an observed redemption of any amount and 91,817 have an observed positive-payout redemption; condition-specific Kaplan-Meier medians from first protocol resolution are 182 and 200 seconds respectively. The exact-linked payout taxonomy contains 53,847 canonical (0,1) vectors, 45,024 canonical (1,0) vectors, 410 fifty-fifty vectors, two other valid vectors, and 9,355 conditions with no observed resolution. Cross-contract Oracle-adapter-protocol ordering is reported conservatively: 823 conditions have an interval-qualified terminal generation, 48 have multiple candidate generations, 91,638 have no compatible terminal generation in the frozen evidence, and 16,129 are right-censored or otherwise unevaluable. The formal results show that Oracle finality does not identify protocol finality, protocol finality does not identify holder realization, and redemption events alone do not identify the fraction of entitlement redeemed without an independent balance-consistent entitlement denominator.

q-fin.TR↗

On-Demand Combinatorial Event Markets on Kalshi: Instantiation, Concentration, and Effective Market Breadth

Kalshi's multivariate-event architecture produces market objects on demand from exact selected legs. Across a registered seven-day interval, 190 independently validated temporal shards yield 7,611,594 unique REST MVE market tickers after excluding 5,777 boundary-overlap observations; the population was created at an average rate of 1.087 million objects per day, with strong hourly burstiness. The hierarchy is sharply compressed relative to the market-object count: 5,262,526 exact event keys, three collection keys, 83,701 selected-leg primitives, and 66,344,938 selected-leg occurrences. The largest collection accounts for 7,217,085 objects (94.82 percent), with an effective collection count of 1.11; the effective primitive count is approximately 720. REST and WebSocket are distinct observation surfaces: 1,487,330 created-notification tickers intersect the REST population in only 276,177 exact tickers. All three observed collection keys have current endpoint confirmation without establishing historical collection or rule versions. The exact structural signature yields 7,611,594 unique signatures with zero collision groups. At retrieval, 2,692,787 objects (35.38 percent) have positive cumulative volume or open interest; all observed 24-hour-volume values are zero and trade-count fields are unavailable. This is current-snapshot activity eligibility, not historical trading. The central result is that a very large on-demand market-object population is generated by a concentrated collection layer and reused primitive vocabulary; economic breadth must be measured at several hierarchical levels rather than by ticker count alone. Exact public endpoint evidence is available for 7,357,576 MVE objects, whereas only 126,806 have an exact determination-to-endpoint pair: terminal-state coverage and lifecycle-path coverage are distinct statistical objects.

physics.soc-ph↗

Axient: Debt-Free Finality for Leveraged Binary Event Markets

Leveraged event positions combine a repayable loan with an outcome claim that may become non-tradable before oracle payout is final. This paper specifies Axient, a physically backed margin layer for binary event markets that separates leverage maturity from claim maturity and makes the hard-flat decision under explicit execution uncertainty. The model distinguishes quoted book proceeds, matched proceeds, settled proceeds, and redemption. At decision time, the protocol selects the smallest sale whose lower settled-proceeds envelope covers an upper bound on debt at the settlement horizon plus a buffer. We prove robust ex-ante debt clearing, pathwise debt-extinguishment and debt-free-finality invariants, maximal residual spot exposure, payout-vector and dispute-duration invariance of lender principal after debt extinction, and an impossibility boundary when execution, signer control, settlement, or market closure leave the registered operating set. We also derive a book-dependent leverage envelope, aggregate hard-flat capacity without double-counting shared liquidity, and scenario-conditional reserve bounds. A deterministic verifier covers step books, partial fills, settlement delay, adversarial book transformations, shared-book liquidation, reserve allocation, zero liquidity, and multiple payout vectors. The operating and stress sets are author-specified; empirical calibration is separate. The contribution is a conditional mechanism-design result and reference-implementation boundary, not a production-safety claim.

q-fin.TR↗

Axient: On-Chain Credit and Loss Allocation for Leveraged Event Markets: A Venue-Agnostic Protocol for Traders, Credit Providers, Market Makers, and Liquidation Backstops

A physically backed leveraged event position requires real credit: if collateral C receives leverage L, the protocol supplies (L-1)C and uses the combined amount to acquire recognized event exposure. This paper develops a venue-agnostic on-chain credit architecture for that capital layer and an endogenous model of its capital market. It separates traders, Senior Credit LPs, market makers, liquidators, and Liquidation Backstop Providers; formalizes pool and debt shares, utilization- and risk-sensitive interest, collateral-locked position accounts, venue capabilities, market-maker commitments, withdrawal queues, isolated pools, non-redeemable reserves, and a deterministic loss waterfall; and models endogenous provider participation, leverage demand, liquidity withdrawal, liquidator entry, reserve replenishment, runs, and common-factor contagion. Formal results establish balanced real- and integer-unit accounting, settlement-confirmed debt priority, trader residual ownership, idempotent partial settlement, non-dilutive share issuance, junior-before-Senior impairment, loss-participating withdrawal queues, utilization-equilibrium conditions, and loss-allocation and contagion bounds. The release preserves 28 exact fixtures and 31,082 deterministic checks and adds fixed-seed agent-based experiments with 70,207,488 scalar invariant evaluations and zero failures. The experiments show that layered protection reduces but does not eliminate Senior loss, that 5x leverage materially increases capital pressure, and that market-maker capacity can raise aggregate shortfall if admission expands too quickly. Results are synthetic mechanism comparisons under author-specified behavior, not forecasts of APY, defaults, venue liquidity, or production safety.

q-fin.TR↗

Fill-Side Behavioral Concentration on Polymarket: Identification Limits under Record-Level Attribution

This paper studies behavioral concentration in Polymarket's public executed-fill record and formalizes what that record can and cannot identify. A pre-publication reconciliation corrects the empirical scope: the archived extraction covers the legacy CTF Exchange over Polygon blocks 86,008,447-86,107,178, approximately 25 April 2026 17:09 UTC through 28 April 2026 00:00 UTC, rather than the full 21-27 April week stated previously. It contains 13,356,931 OrderFilled records, 77,204 addresses with at least five attributed records, and 43,116 token identifiers; negative-risk markets are absent. The archived feature construction credits both maker and taker addresses on each record. This convention is not invariant to match fragmentation, and mint/burn executions do not admit a universal buyer/seller interpretation. The reported one-cluster result is therefore retained only as a null under the original record-level representation, not as evidence that the participant population is intrinsically unimodal. The concentration table is likewise attribution-weighted arithmetic under that convention. Two methodological results remain durable: public fills do not identify the quote lifecycle required to infer market making, spoofing, or strategic withdrawal; and a one-cluster result rejects density separation in an observed feature space, not latent economic heterogeneity. A match-normalized, mint-aware, multi-window replication is required before treating the cluster null or exact cohort shares as stable venue properties.

q-fin.TR↗

Resolution-Aware Perpetual Futures on Binary Prediction Markets: Failure Modes and Mechanical Stress Tests Using Polymarket Data

We study whether crypto-style perpetual-futures mechanics can be applied to a binary event claim that ultimately pays 0 or 1. A synthetic long entered at price p0 with collateral xp0/L has terminal equity -xp0(1-1/L) when the claim pays 0, if the position survives to resolution without a top-up, close, or conversion. Thus any L > 1 creates an adverse-outcome account shortfall under these conditions, independent of the pre-resolution mark. We also derive a funding trilemma: basis-only funding loses relative force near a boundary, while uniform relative-basis correction requires unbounded transfers and conflicts with payer solvency or participation. We then apply a mechanical stress test to observed Polymarket paths from 21-27 April 2026. Of 61,087 enriched candidate markets, 13,298 pass the stated adequacy gates. Two structural diagnostics pass, but three of five pre-specified materiality tests fail. Dynamic margin and leverage compression pre-empt more observed paths than the static baseline, pooled drawdown falls by only 5.1 percent, and a staged halt reduces final-hour liquidations mechanically while leaving terminal shortfall incidence slightly worse. The contribution is a corrected non-portability analysis, a reusable observed-path replay design, and negative design lessons; the results do not establish equilibrium performance or deployment safety.

q-fin.TR↗

A Taxonomy of Event-Linked Perpetual Futures: Design Axes, Failure Modes, and Empirical Evaluability

The label event-linked perpetual often conflates mathematically different contracts. We replace a flat product list with a four-axis taxonomy: underlying geometry, temporal structure, settlement structure, and venue-oracle composition. The taxonomy covers a single binary probability, conditional ratios, event spreads, baskets, path functionals, liquidity indices, rolling sequences, and flow-only swaps. We derive a corrected inheritance map from the single-event case and prove six narrow results. Terminal shortfall depends on exact settlement support and collateral, not on a generic bounded-event label. Conditional ratios become locally ill-conditioned as the conditioning value approaches zero. Once one spread leg is final, the residual position is an affine single-leg exposure. Basket risk depends on feasible joint outcomes and weights, not leg count. Binary entropy settles deterministically to zero, whereas realized variation requires an explicit sampling and terminal-jump convention. Fixed-path replay does not identify deployment behavior when a liquidity contract changes its own reference process. No new empirical estimates are reported. We assign evidence grades and minimum data requirements to each design. The central conclusion is that no universal event-perpetual engine exists: each contract must separately define support, clocks, settlement, source hierarchy, and the controls implied by those choices.

q-fin.TR↗

Manipulation, Informed Trading, and Regulation in Leveraged Event-Linked Markets

Leverage does not create manipulation or informed trading in event markets, but it changes their economics. We separate four conduct channels: market-price manipulation, real-world outcome manipulation, resolution-process manipulation, and informed trading that exploits non-public information without changing the event or resolution rule. A capital-constrained amplification model shows that gross directional gains scale with financed exposure while information, influence, concealment, execution, and enforcement costs need not scale proportionally. Leverage therefore weakly enlarges profitable influence opportunities under fixed-cost conditions, but endogenous price impact, convex detection costs, or enforceable position limits can reverse the result. We connect this model to the revised non-portability, taxonomy, and fill-side evidence of the companion papers. Historical-path replay identifies mechanical rule channels, not equilibrium conduct; fill-side address data support concentration analysis but not address-level spoofing or quote-withdrawal claims. The regulatory contribution is a functions-first framework covering product authorization, market-integrity surveillance, event-influence controls, resolution governance, and loss-bearing capital. A July 2026 source snapshot indicates an evolving United States framework and no harmonized cross-jurisdictional category for leveraged event contracts. Legal status remains product- and jurisdiction-specific.

q-fin.TR↗

ForesightFlow: An Information Leakage Score Framework for Prediction Markets

ForesightFlow is an Information Leakage Score (ILS) framework for detecting informed trading on decentralized prediction markets. For an event-resolved binary market, the score quantifies the fraction of the terminal information move priced in before the public news event. Three operational scope conditions (edge effect, non-trivial total move, anchor sensitivity) are stated as preconditions for interpretation. The score admits a Murphy-decomposition reading that connects label generation to the proper-scoring-rule literature. A pilot empirical evaluation surfaces three findings. First, a resolution-anchored proxy for the public-event timestamp does not separate event-resolved markets from a matched control population (Mann-Whitney p = 1e-6, separation reversed), demonstrating that proxy quality is itself a binding constraint. Second, the article-derived timestamp on a single high-stakes case shifts the score by 0.444 in magnitude relative to the proxy and lies on the opposite side of zero. Third, an audit of the publicly documented Polymarket insider record reveals that documented cases are systematically deadline-resolved, falling outside the original ILS scope (0 of 24 FFIC inventory markets satisfied original scope conditions). This last finding motivates a deadline-ILS extension introduced in Section 7, anchored at the public-event timestamp rather than the news timestamp, and equipped with a per-category exponential hazard baseline for the time-to-event distribution. The extension closes the gap between the methodology and the population in which insider trading has been empirically documented. An end-to-end evaluation of the extension on the 2026 U.S.-Iran conflict cluster is reported in a companion paper. We release the FFIC inventory, the resolution-typology classification of the 911,237-market corpus, and all code at github.com/ForesightFlow.

q-fin.TR↗

Empirical Evaluation of Deadline-Resolved Information Leakage on Documented Polymarket Insider Cases

This paper reports an end-to-end empirical evaluation of the deadline-Information Leakage Score (ILS-dl) extension introduced in the companion methodology paper. The deadline-ILS extends the original ILS to deadline-resolved prediction-market contracts, the dominant structural form of publicly documented insider trading on Polymarket. We anchor the evaluation in the 2026 U.S.-Iran conflict cluster of the ForesightFlow Insider Cases (FFIC) inventory, the largest documented deadline cluster. The evaluation has four parts: per-category exponential-hazard estimation, a single-case ILS-dl computation, cross-market wallet analysis, and methodological refinements. Hazard-rate estimation produces an adequate exponential fit for military-geopolitics markets (KS p = 0.426, half-life 2.9 days, n = 18) and a preliminary fit for corporate-disclosure markets (n = 5). The regulatory-decision category is rejected as bimodal (p = 0.023). On the largest applicable FFIC contract ("US forces enter Iran by April 30," $269M volume), the article-derived public-event timestamp yields ILS-dl = +0.113 versus a resolution-anchored proxy value of -0.331: a 0.444 shift in magnitude on opposite sides of zero, demonstrating that the extension distinguishes signal from proxy artefact. Pre-event drift is mild, and short-window variants (30-min, 2-hour) are exactly zero. Cross-market wallet analysis identifies 332 wallets active in both major Iran-cluster markets, but the available trade history covers only the resolution-settlement window. v2 (May 2026) corrects the hazard fit to the full Tier-3 population; the v1 estimate lies inside the v2 95% CI.

q-fin.TR↗

Per-Market Information Leakage and Order-Flow Skill: Two Methodological Lenses on Informed Trading in Decentralized Prediction Markets

April 2026 saw notable methodological convergence in the academic study of informed trading on decentralized prediction markets. Three approaches surfaced almost simultaneously: Mitts and Ofir (2026) apply a composite screen to over 210,000 wallet-market pairs; Gomez-Cram et al. (2026) apply an event-level sign-randomization test to Polymarket's complete transaction history, classifying 3.14% of accounts as "skilled winners" and separately flagging 1,950 accounts as "insiders" via a lifecycle heuristic; Nechepurenko (2026) develops the Information Leakage Score (ILS) framework, which quantifies per-market information front-loading at an article-derived public-event timestamp. This paper provides a methodological comparison. The central claim is that these are three distinct layers of detection, not competing methods on a single layer. Sign-randomization is best understood as an account-level test of persistent directional skill conditional on opportunity selection -- not a direct test of insider trading, and not a per-market measure. The heuristic insider flag is separate from the skill classifier, applies to a population the classifier excludes by design, and has unknown precision. The Polymarket sample pools politics, sports, crypto, and other categories with different information technologies, so a platform-wide "skilled winner" classification is mechanism-ambiguous. The January 2026 U.S.-Venezuela operation cluster, where the DOJ indictment of Master Sergeant Gannon Van Dyke provides a rare external enforcement benchmark, illustrates how the layers stack: lifecycle heuristics identify suspicious accounts; legal investigation addresses non-public-information possession; per-market scoring would quantify how much information was leaked into each contract. A combined pipeline gains in precision because each layer filters a different dimension.

q-fin.TR↗

Coordination as an Architectural Layer for LLM-Based Multi-Agent Systems

Multi-agent LLM systems fail in production at rates between 41% and 87%, mostly due to coordination defects rather than base-model capability. Existing responses split between cataloguing failure modes empirically and shipping declarative orchestration frameworks as engineering tools; neither delivers a principled mapping from coordination configuration to predictable failure-mode signature. We argue that coordination should be treated as a configurable architectural layer, separable from agent logic and from information access, enabling architectural reasoning rather than only engineering productivity. We instantiate this with an information-controlled design on prediction markets: a single LLM, fixed tools, fixed per-call output cap, and fixed prompt template across five reference coordination configurations, with total compute per question treated as an endogenous architectural output. The Murphy decomposition of the Brier score separates calibration from discriminative power, so configurations leave distinguishable signatures even when aggregate scores coincide. On 100 Polymarket binary markets resolved after the model's training cutoff (claude-opus-4-6) we report Murphy signatures, a cost-quality Pareto frontier, category-conditioned analysis, and a bootstrap power-projection. Three of five pre-specified predictions are upheld in direction; two configurations dominate the Pareto frontier within this regime; exploratory bootstrap intervals separate consensus alignment from others, though pairwise tests do not survive Bonferroni correction at n=100. We also deploy the same configurations as live agents on Foresight Arena under web-search-enabled conditions, as an on-chain replication channel accumulating in parallel. Harness, trace dataset, and production agents are released. We position this as a methodology-validating first instantiation, not a general cross-model claim.

cs.MA↗

Foresight Arena: An On-Chain Benchmark for Evaluating AI Forecasting Agents

Evaluating the true forecasting ability of AI agents requires environments that are resistant to environments resistant to overfitting, free from centralized trust, and grounded in incentive-compatible scoring. Existing benchmarks either rely on static datasets vulnerable to training-data contamination, or measure trading PnL -- a metric conflating predictive accuracy with timing, sizing, and risk appetite. We introduce Foresight Arena, the first permissionless, on-chain benchmark for evaluating AI forecasting agents on real-world prediction markets. Agents submit probabilistic forecasts on binary Polymarket markets via a commit-reveal protocol enforced by Solidity smart contracts on Polygon PoS; outcomes are resolved trustlessly through the Gnosis Conditional Token Framework. Performance is measured by the Brier Score and a novel Alpha Score -- proper scoring rules that incentivize honest probability reporting and isolate predictive edge over market consensus. We provide a formal analysis: closed-form variance for per-market Alpha, the connection to Murphy's classical Brier decomposition, and a power analysis characterizing the number of rounds required to reliably distinguish agents of different skill levels. We show that detecting a true edge of $α^* = 0.02$ at 80% power requires approximately 350 resolved binary predictions (50 rounds of 7 markets), while $α^* = 0.01$ requires four times more. We complement these analytical results with a deterministic, seed-controlled simulation study calibrated to literature-reported Brier-score ranges, illustrating how Murphy decomposition distinguishes well-calibrated agents from market-tracking agents that fail through reduced resolution. Live results from the deployed benchmark will be reported in a future revision. All smart contracts and evaluation infrastructure are open-source.

cs.MA↗