Searcharxiv⌕ Search

arXiv · 2610.05999

No More Translation at Runtime: LLM-Empowered Static Binary Translation

Abstract

While AArch64 CPUs are becoming strong market contenders, their software ecosystem lags behind the mature x86-64 environment, hindering the adoption of the new architectures and impacting user experience. Binary translation bridges this divide by converting binary code from one architecture (e.g., x86-64) to run on another (e.g., AArch64), allowing legacy software to benefit from modern hardware's performance and energy efficiency advantages. Current translation methods are typically either dynamic, which adds significant runtime overhead, or static, which struggles with reliability due to the inherent complexities of binary analysis. This paper introduces a new static, assembly-to-assembly translation paradigm that transforms binary code ahead of execution, generating portable, efficient native-like binaries that run on AArch64 devices without runtime frameworks. Benefiting from recent breakthroughs in large language models (LLMs), we provide a practical and automated translation engine that produces high-quality code with minimal human intervention. To ensure correctness, we introduce a crucial verification step, where we split the assembly code into simplified snippets, enabling efficient and scalable semantic verification. Our evaluation shows that this approach significantly outperforms existing open-source solutions with a large margin, producing binaries with near-native performance. Furthermore, it shows substantial improvements over the leading industrial translator, ExaGear, illuminating a promising new direction for cross-architecture binary translation research.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zhibo Liu, Huaijin Wang, Wai Kin Wong, Daoyuan Wu, Shuai Wang. 2026-10-05. No More Translation at Runtime: LLM-Empowered Static Binary Translation. https://doi.org/10.1145/3767295.3803600

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

AKTS: Sub-Microsecond Kernel Policy Switching for Language-Model Agents

GPU-backed LLM servers multiplex interactive requests with background batch work on the same CPUs; a fixed kernel policy serves one objective and loses the other, so agentic OS control needs to switch scheduler behavior as the workload changes. The hard part is applying the switch safely and fast enough for the kernel: scheduler events occur every 1-10 $μ$s, and any code running there must satisfy the eBPF verifier. Scalar knobs are fast but limited, while generating eBPF policy code puts compilation, verification and possible rejection on the runtime path. We present AKTS, which verifies a policy library once, at load time, and reduces the agent's runtime action to writing an integer index into an in-kernel array of preverified policies, resolved by a tail call. Because the agent emits an index rather than code, verifier failure is not a runtime outcome. On Linux 6.14, AKTS applies a policy switch in 920 ns (p50); makes an invalid index inert across 60,217 invocations on an attached scheduler; and switches policies in a vLLM workload to capture 97% of a throughput policy's batch work while matching a latency policy's burst response.

cs.OS↗

Behind a Simple Read: Understanding Work and Waiting in the Linux I/O Stack

A simple read interface provides uniform functional semantics, but not equally simple or predictable performance behavior. We use synchronous large-buffer reads in Linux as an observation window and decompose the buffered-read path end to end, distinguishing work volume, processing time, and critical-path exposure. We find that Linux reduces metadata work through large folios and overlaps most cold-read copying with device waiting, but these mechanisms depend on access advice, folio granularity, cache state, and backend execution. Using an experimental kernel prototype, we validate additional opportunities from out-of-order early copying and opportunistic parallelism, while showing that faster request submission does not improve end-to-end performance when SSD supply is already sufficient. To explain device waiting, we abstract first-completion wait $F$ and subsequent completion capacity $B_{CQ}$ from finite request-batch completion timelines. Measurements on a real SSD characterize how they vary with request size and batch size, while MQSim experiments connect them to internal device mechanisms and distinguish finite-batch completion from sustained throughput. We then construct a cross-layer wait chain along the actual folio--bio--request mapping. Independently calibrated device parameters predict waiting at the block layer and at read_pages with errors no greater than approximately 6.9% and 3.2%, respectively. Finally, four use cases apply the model and wait chain to parallel submission, Linux readahead, dependent-read layout, and polling versus sleeping. Their gains, no gains, and gain reversals form a closed loop from observation and modeling to explanation and control, providing a measurable basis for performance decisions in layered I/O stacks.

cs.OS↗

FDP: The Data Placement Promise of Modern NVMe SSDs

NVMe SSDs are now widely deployed as the storage tier in data centers. As SSDs have evolved over the past decade, the commu- nity has continued to debate the interfaces they expose and how operating systems and storage systems should exploit them. The NVMe Flexible Data Placement (FDP) proposal is the latest point in this design space. FDP introduces an interface based on Reclaim Units that enables explicit data placement to reduce device write amplification without the software engineering costs of sequential- write constraints and host garbage collection. FDP-enabled SSDs are emerging in commercial products and early data center de- ployments. Their compatibility with conventional block I/O allows existing applications to run unchanged, allowing a frictionless adop- tion in industry. This paper presents an experimental evaluation of FDP SSDs to characterize their data placement guarantees over the raw device interface. We then revisit two widely deployed and distinct open- source storage systems, MySQL and RocksDB, and examine whether lifetime-based data separation and distinct write patterns built into their architectures can be mapped onto FDP SSDs without inva- sive changes. Our evaluation shows end-to-end WAF reductions at higher device utilization, along with QoS and throughput improve- ments under synthetic and real-world workloads. These results demonstrate that FDP provides a practical and deployable cross- layer mechanism for data placement with open-source ecosystem support on Linux. They also highlight why FDP SSDs are gaining traction in industry.

cs.OS↗