Searcharxiv⌕ Search

arXiv · 2610.07762

ES-Trace: Auditing Ethical-Sourcing Disclosure of Code Generation Models Beyond Model Cards

Abstract

Code generation models have been increasingly used in software development, but their development raises ethical-sourcing concerns involving intellectual property, privacy, fairness, labour practices, and environmental impact. Although prior work has defined ethical-sourcing criteria for code generation, it remains unclear how much evidence existing models disclose and where that evidence can be found. We introduce ES-Trace, a framework for ethical-sourcing disclosure audits that traces disclosed evidence beyond model cards using the Model Documentation Traceability Graph (MDTG), which represents relationships among models, versions, and documentation artifacts. We apply ES-Trace to 26 models from 10 publishers across 77 documents and 20 ES-CodeGen aspects. Model-card-only auditing yields a mean score of 1.77/5, while resolving the declared references increases it to 2.82/5, with most of the increase arising from documents that the publisher declares in structured metadata. The key findings of our study include: (1) social and labour-related aspects remain poorly documented, even when expanding the audit to the full documentation scope, (2) resolving documentation references substantially increases observed disclosure, raising the mean score from 1.77/5 to 2.82/5, (3) model-card-only audits can mischaracterize release-level disclosure changes, and (4) documentation mismatches can associate evidence with the wrong model or version, highlighting the need for explicit model--version binding and consistency across documentation artifacts. Our study calls for reference-aware ethical-sourcing disclosure audits, explicit model--version binding and consistency across documentation artifacts, and stronger documentation of currently underreported social and labour-related aspects.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zhuolin Xu, Haibo Wang, Shin Hwei Tan. 2026-10-06. ES-Trace: Auditing Ethical-Sourcing Disclosure of Code Generation Models Beyond Model Cards. https://arxiv.org/abs/2610.07762

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

From Verification Failures to Reusable Guidance for Coding Agents

Coding agents need to establish that a program satisfies a specification and that the specification captures the requested behavior. We study how expert diagnosis of verification failures can become reusable guidance for this work. Our approach combines executable language definitions in the K framework with a kit of procedures for constructing specifications, repairing proofs, and auditing their adequacy. A human-guided development campaign on HumanEval, a benchmark of 164 Python programming tasks, achieves a 164/164 success rate with the semantics and the kit, measured by final AI audit Pass verdicts after two targeted repairs. To examine whether auditing detects problems that successful proofs leave unresolved, we construct 12 author-reviewed pairs of clean and defective packages. Every package passes its K proofs, and completed audits identify all defects and accept all clean packages. We then use KleverBench to test specification and proof construction for 31 programs with changed operator meanings. Comparisons with complete acceptance rules and equally long generic advice yield mixed results across two model and budget settings, motivating further work on selecting useful guidance within resource limits. Human-reviewed Optimism proofs establish expected pause reverts for six operations within declared input bounds under London semantics with unbounded gas. We report progress, difficulties, and lessons toward agents that deliver programs with checkable correctness arguments.

cs.SE↗

PreMaQ: Predicting Maintainability-Related Quality of LLM-Generated Code Before Generation

As large language models (LLMs) become increasingly capable of code generation, adopting generated code in software development requires assessing not only its functional correctness but also its maintainability-related quality. If such quality could be estimated before generation, developers could avoid the cost of generating, reviewing, and discarding low-quality code. Although prior work has shown that the functional correctness of the LLM-generated code can be predicted in advance, it remains unclear whether maintainability-related quality is similarly predictable. We introduce Pre-Generation Maintainability-Related Quality Prediction (PreMaQ), which predicts the Code Smell Score (CSS) and Maintainability Index (MI) of generated code from the internal representations of LLMs before generation. Our evaluation covers four open-weight LLMs and four Python code generation benchmarks, comprising 2,695 tasks in total. Our results show that predicted CSS and MI consistently correlate with their observed values across all 16 model-benchmark combinations, achieving mean Spearman rank correlations of 0.57 and 0.65, respectively. When used for model selection, PreMaQ achieves 59.5% of the maintainability-related quality improvement attainable by an ideal maintainability-based selector over random selection on tasks for which multiple models generate functionally correct code. Combining predictions from PreMaQ and prompt embeddings increases this proportion to 60.7%, indicating that the two signals are complementary. These findings suggest that PreMaQ can be used to predict and improve the maintainability-related quality of LLM-generated code.

cs.SE↗

RAPO-Sol: Retrieval-Augmented Preference Optimization for Repository-Level Solidity Code Generation

Smart contracts written in Solidity manage assets, permissions, and irreversible state changes, making code generation both useful and security-critical. Repository-level Solidity generation is challenging because models must synthesize complete contracts or libraries while preserving consistency across state variables, modifiers, events, inheritance, external calls, and access-control logic. We present RAPO-Sol, a two-stage training framework for repository-level Solidity code generation. First, Retrieval-Augmented Fine-Tuning (RAFT) augments each training input with similar Solidity examples, helping the model learn recurring contract-level patterns while remaining retrieval-free at inference time. Second, Direct Preference Optimization (DPO) trains the model to prefer reference contracts over close but semantically flawed alternatives. We construct rejected samples using Solidity Semantic-Anchor Perturbation (SAP), which perturbs validation statements, visibility modifiers, data-location keywords, context variables, payment operations, and low-level calls. Experiments on SolidityBench with CodeLlama-7B-Instruct, DeepSeek-Coder-6.7B-Instruct, and Qwen2.5-Coder-7B-Instruct show that RAFT consistently improves over supervised fine-tuning, while SAP-based DPO provides further gains in BLEU and SolidityScore. The full RAFT+DPO pipeline achieves the best performance across all three models, demonstrating complementary benefits from retrieval during training and Solidity-aware preference optimization without adding retrieval cost at inference.

cs.SE↗