Searcharxiv⌕ Search

arXiv subjects

Van Khue Nguyen

Publications and source records attributed to Van Khue Nguyen.

2 recordsLinked to original sources

Evaluating Agents Across Runtime Contracts: When Mismatch Costs Efficiency or Quality

In CodeAct, language-model agents write Python that calls tools and use execution feedback to choose actions. Persistent runtimes preserve Python variables between actions; stateless runtimes clear them without resetting task progress. Training traces demonstrate task-solving strategies and runtime-specific ways to store and recover intermediate results. We study runtime transfer: whether agents trained under one contract remain effective under the other, a dependence that fixed-runtime evaluations can conceal. Across three tasks requiring a working record built from tool feedback, we fine-tune separate Qwen3-8B agents per task and runtime on instance-paired persistent and stateless traces and evaluate all four training-deployment combinations. We vary the per-turn tool-call cap: how many calls one Python action may execute. In our primary task, agents inspect hidden item attributes and select a high-value subset under a weight limit. At 80 calls per turn, a persistent-trained agent deployed statelessly scores 0.61 of optimal value against 0.77 under matched persistent deployment, but uses 4.4 times as many tokens rebuilding its record. Tightening the cap to 25 reduces mismatched quality to 0.07 while matched quality remains 0.66. Trace replay shows that the cap interrupts inspection, the reset erases the partial record, and rebuilding consumes calls needed for progress. The widening gap replicates across training seeds, a second rollout, and two additional base models. Across the settings studied, cap tightening produces a mismatch-specific collapse only where interruptions are frequent and matched execution resumes while mismatched execution restarts. Evaluations of trace-fine-tuned agents should specify episode limits, test runtime transfer across caps, and report cost alongside quality; otherwise, the same mismatch may appear as redundant computation or near-total quality loss.

cs.AI↗

MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources

We present MixtureVitae, an open-access pretraining corpus built to minimize legal risk while providing strong downstream performance. MixtureVitae follows a permissive-first, risk-mitigated sourcing strategy that combines public-domain and permissively licensed text (e.g., CC-BY/Apache) with carefully justified low-risk additions (e.g., government works and EU TDM-eligible sources). MixtureVitae adopts a simple, single-stage pretraining recipe that integrates a large proportion of permissive synthetic instruction and reasoning data-signals typically introduced during post-training and generally scarce in permissive web corpora. We categorize all sources into a three-tier scheme that reflects varying risk levels and provide shard-level provenance metadata to enable risk-aware usage. In controlled experiments using the open-sci-ref training protocol (fixed architectures and hyperparameters; 50B and 300B token budgets across 130M-1.7B parameters), models trained on MixtureVitae consistently outperform other permissive datasets across a suite of standard benchmarks, and at the 1.7B-parameters/300B-tokens setting, they surpass FineWeb-Edu and approach DCLM late in training. Performance is particularly strong on MMLU and on math and code benchmarks: a 1.7B model pretrained on 300B MixtureVitae tokens matches or exceeds a strong 1.7B instruction-tuned baseline on GSM8K, HumanEval, and MBPP, despite using over 36 times fewer tokens (300B vs. ~11T). Supported by a thorough decontamination analysis, these results show that permissive-first data with high instruction and reasoning density, tiered by licensing and provenance-related risk, can provide a practical and risk-mitigated foundation for training capable LLMs, reducing reliance on broad web scrapes without sacrificing competitiveness. Code: https://github.com/ontocord/mixturevitae

cs.CL↗