Making AI Scientists Auditable from Evidence to Claim
AI research systems can report improved results without establishing that the tested implementation matches the method described in their conclusions. We introduce Xcientist, a research harness that connects literature review, component-based idea generation, experiment validation and reporting. An evidence graph retains source passages, methods, baselines and evaluation conditions; component records track design revisions; and validation contracts specify the implementation, comparisons and artifacts required at each stage. We define claim drift as an unresolved mismatch between an accepted claim and its attributed sources, code or experiments. A retrospective audit across agent memory, traffic forecasting and physics-informed learning found audit claim drift rates of 3.6--16.7\% for Xcientist, lower than those of each of three comparison systems in all nine task-by-relation comparisons. This measure includes both confirmed mismatches and unverifiable claims; most of the difference arose from fewer unverifiable records. The case studies also document performance gains, unsuccessful revisions and limits on component attribution. This technical report describes the system, its implementation and the records needed to check how a proposed idea became a tested method and a final claim.