SearcharxivSearch

arXiv subjects

David Tan

Publications and source records attributed to David Tan.

5 recordsLinked to original sources

Two Kinds of Nothing: What Insignificant Results in Finance Actually Show

Claims of the form "we find no evidence that X affects Y" appear throughout the applied finance literature, yet whether such a claim contains evidence of absence or absence of evidence depends entirely on its confidence interval. The term "statistically insignificant" is routinely read to mean zero economic effect. However, a more honest description is that zero could not be rejected along with a range of other coefficient effect sizes. The crucial question is whether effect sizes in that range are consequential. This note distinguishes two kinds of insignificant results that are indistinguishable in a standard regression table: bounded null claims where the intervals reject effect sizes of consequence and thus represent a genuine finding, and vacuous null claims where even consequential effects remain unrejected and therefore establish nothing. I propose a minimal reporting standard for regression results in applied finance, where the smallest consequential effect size is stated (in the units of the decision, per a named increment of the regressor) alongside the descriptive statistics and compared with the relevant edges of the confidence intervals of null claims. Using only the reported coefficient and standard error, authors can distinguish bounded (informative) null claims that reject consequential effect sizes from vacuous null claims that establish no information, perhaps due to deficiencies in data or the identification strategy. The symmetric phrase "no effect" conceals, in particular, the frequent split verdict: bounded in one direction, vacuous in the other. In ongoing work, I apply this framework to published null claims in leading finance journals, beginning with my own.

q-fin.GN

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence

Apollo did not reach the Moon merely because its engineers could solve difficult equations. It succeeded by turning a distant ambition into a mission architecture of explicit objectives, simulation, verification, and repeated correction. AI now faces a similar transition: frontier models can solve difficult tasks once the problem, tools, and success criteria are specified, yet consequential real-world challenges rarely arrive in an executable or verifiable form. We introduce Apodex Discovery, a framework for building and evaluating discoverative AI through the heavy-duty solver, a system comprising a foundation model, harness, tools, and control policies that pursues extended, stateful, verifiable investigations. It has three core components. First, a problem-scouting process surveyed 561 industries across 16 sectors, assembled 423 high-value real-world problems, and selected 20 for the initial release. Second, a common environment-task-episode abstraction provides data, tools, constraints, feedback, trajectory recording, and verification of intermediate artifacts and final submissions. Third, HDS6 evaluates Tools, Repair, Alternatives, Coherence, Evidence, and Scope independently of final-task success. In AAV capsid design, Apodex surpassed the published state of the art by 7% across viability, tropism, structure prediction, and generative design. In drug repurposing and reformulation, a task-specific biomedical environment improved the mean normalized prediction score of GPT-5.5 and GPT-5.6-sol by 2.5 and 7.6 points over the same closed-book backbone. Controlled ablations show that the fixed TRACES episode interface enables attribution of performance differences to specific solver components. Apodex Discovery moves AI evaluation beyond predefined benchmarks toward verifiable investigations aimed at genuine discovery.

cs.AI

Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural considerations. We therefore advocate for source-contrastive evaluation and instantiate it with Cultivar, a localised subset of FLORES, which enables locale-specific translation evaluation. When paired with unlocalised counterparts, performance discrepancy allows the probing of data contamination and localisation robustness. We benchmark 32 open-weight models and find that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US content better than that of other locales, regardless of language.

cs.CL

When Flores Bloomz Wrong: Cross-Direction Contamination in Machine Translation Evaluation

Large language models (LLMs) can be benchmark-contaminated, resulting in inflated scores that mask memorization as generalization, and in multilingual settings, this memorization can even transfer to "uncontaminated" languages. Using the FLORES-200 translation benchmark as a diagnostic, we study two 7-8B instruction-tuned multilingual LLMs: Bloomz, which was trained on FLORES, and Llama as an uncontaminated control. We confirm Bloomz's FLORES contamination and demonstrate that machine translation contamination can be cross-directional, artificially boosting performance in unseen translation directions due to target-side memorization. Further analysis shows that recall of memorized references often persists despite various source-side perturbation efforts like paraphrasing and named entity replacement. However, replacing named entities leads to a consistent decrease in BLEU, suggesting an effective probing method for memorization in contaminated models.

cs.CL

Strategic Preemption Under Shared Catastrophic Risk: The Suicide Region and the Race to Artificial General Intelligence

We analyze a continuous-time preemption game with shared catastrophic externalities. When the cost of catastrophe is embedded in both players' payoffs, the risk term cancels out in the equilibrium indifference condition. This creates a "suicide region" where competitive pressures force rational agents to deploy despite negative risk-adjusted net present values. We apply this framework to the race for artificial general intelligence (AGI). We show that this suicide region widens as the cost of systemic ruin grows: higher catastrophic risk does not deter the race but instead enlarges the set of conditions under which rational actors deploy despite negative social value. We characterize the resulting welfare distortion against a social planner's benchmark and demonstrate how two complementary mechanisms - private liability and prize-sharing - can close the suicide region. Private liability raises the cost of unsafe deployment while prize-sharing reduces the strategic imperative to deploy first. "Warning shots" (sub-existential disasters) will fail to deter AGI acceleration, as the winner-takes-all nature of the race remains intact.

q-fin.RM