Searcharxiv⌕ Search

arXiv · 2609.34983

SmartMemory: Detecting On-chain-off-chain Communication Inconsistency for Smart Contract via Memory-based Agent

Abstract

Smart contracts underpin decentralized finance, where growing demand for on-chain/off-chain communication(OFC) has driven diverse applications such as cross-chain bridges, real-world asset tokenization, and fiat-backed stablecoins. TheOFC-related security incidents in these applications are increasingly frequent, but prior studies address separate vulnerability categories within OFC applications rather than providing a unified view, causing vulnerabilities outside known patterns to be missed.In this paper, we identify OFC inconsistency (OFCI) as a root cause of OFC vulnerabilities, which arises from business-logic flaw and ultimately breaks the equivalence between the on-chain and off-chain asset representations to induce inconsistency.Automatically detecting OFCIs faces two challenges including (1)locating heterogeneous business logic, and (2) transferring existing vulnerability knowledge to identify unseen OFCI instances. To this end, we propose SmartMemory, the first framework to leverage a memory-based agent for OFCI detection. To address heterogeneity, SmartMemory maps diverse implementations ofOFC contracts into a canonical business-semantic representation to locate the business logic for OFCI inspection. For knowledge reuse, SmartMemory integrates a memory-based agent to distill vulnerability knowledge from features into patterns and detection rules, enabling knowledge transfer across cases to identify unseenOFCIs. Lastly, SmartMemory performs taint analysis to verify the reachability, type, and impact of each candidate OFCI. We construct the first real-world OFCI dataset comprising 48 DApps with 81 OFCIs for evaluation, on which SmartMemory achieves80.68% precision and 87.65% recall. In addition, through an analysis of 325 real-world OFC applications, SmartMemory detects 36 previously unknown OFCIs, all of which have been confirmed and fixed by corresponding parties.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zeqin Liao, Yuhong Nan, Henglong Liang, Zixu Gao, Lianyu Hu, Yuqiang Sun, Zhijie Zhong, Xiaoyu Ma, Zibin Zheng, Yang Liu. 2026-09-28. SmartMemory: Detecting On-chain-off-chain Communication Inconsistency for Smart Contract via Memory-based Agent. https://arxiv.org/abs/2609.34983

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Is Agent Code Less Maintainable Than Human Code?

Maintainability is a core dimension of software engineering, shaping how code is written, reviewed, and developed over time. While coding agents have demonstrated strong performance on single-issue tasks, it remains unclear how maintainable their code is when future agents build on top of it, potentially leading to compounding downstream effects. We investigate how agent code compares to human code in these maintenance settings, presenting CodeThread, a framework to construct controlled experiments from repository-level coding benchmarks. Applying CodeThread to four frontier coding agents and four benchmarks, we find that agents are less effective at resolving tasks when building on agent code compared to human code, with task resolve rate drops of up to 13.1%. Regression analysis reveals that many traditional software engineering maintainability metrics do not explain this difference. Instead, the clearest signals are subtler behavioral differences in agent code, such as changes to input validation and error handling, along with differences in downstream code size and task difficulty. These findings highlight the need to evaluate these systems not only by immediate task resolution but also by code maintainability, and point to potential sources of downstream errors introduced by agent code.

cs.SE↗

CURATE: Leveraging LLM Agents to Compose, Catalog, and Deploy Reproducible Workflows

Agentic code generation has the potential to accelerate the development of computational workflows while also reducing barriers to entry. However, a key gap remains: existing coding agents focus on code generation and do not address the entire workflow lifecycle, including deployment and sharing. As a result, users develop and stitch modules independently while managing deployment on their own. To address this gap, we propose CURATE (Composition, User-in-the-loop, Reuse, and Automated Task Execution), a novel human-in-the-loop multi-agent system that uses LLM agents to manage and develop composable workflows across their entire lifecycle. A key feature of the system is a catalog that allows for the storage and reuse of modules across workflows. Module catalogs provide a foundation that can be expanded to support FAIR principles by facilitating the sharing and reuse of curated modules and subgraphs. We demonstrate the feasibility of our system with an initial prototype and 6 experiments.

cs.SE↗

MCP Error Messages Written for Developers Hurt the Most Capable Agents Most

Many Model Context Protocol (MCP) servers wrap web APIs built for human developers, and their error messages tell the reader to run a command, edit a configuration, open a web page or wait. Many agents that read them can only call the server's tools. In 150 widely used MCP servers, 949 of 3,001 error messages tell the caller what to do next, and half of these steps depend on something the server cannot see about the caller. On credential errors, 62 of 67 steps ask for a terminal command, a configuration change or a web page; on rate limits, 20 of 30 say to wait and retry without naming the call to repeat. We tested five OpenAI models that act only through the tools of Berkeley Function Calling Leaderboard tasks, and the agents did what the step said. On expired credentials, a terminal command in the step left 45% of tasks recovered, and the loss it caused grew from 18 points for GPT-5.5 to 69 for GPT-6 Astra. On a rate limit, GitHub's "Wait before retrying." left 6%. We tested two remedies. For MCP developers, naming a server tool in the step raised recovery on expired credentials to 84%, with the login tool in place of the command, and on a rate limit to 88%, with the call to repeat in place of the bare wait. For agent developers, deleting the step with a one-sentence prompt before the model reads it raised recovery on expired credentials to 82%.

cs.SE↗