Searcharxiv⌕ Search

arXiv · 2610.08156

Vibe Building

Abstract

Automated building design must comply with seismic and wind codes and satisfy structural mechanics constraints, yet most existing agents produce visually plausible models without verification grounded in mechanical analysis and code compliance. We introduce the Vibe Building task and propose PE-Loop (Physics-Engine-in-the-Loop), an agent in which a deterministic physics engine is the sole source of evaluation signals, mapping code constraints to a physics process reward, while the language model is confined to proposing discrete revisions (section menu, topology, and lateral system). Designs are verified by held-out seismic and wind time-history checks and a constructability gate. On VB-Bench, 3,577 physics-adjudicated building instances across six code families, PE-Loop achieves the highest verified success rate under three of four backbone LLMs, the highest held-out seismic pass rate under all four, and the highest held-out wind pass rate under three. Replacing the physics verdict with a language-model judge, all else fixed, leaves 58.43% of accepted designs noncompliant. These results suggest that reliable structural design rests less on a stronger LLM proposer than on an adjudicator the proposer cannot influence, a division of labor for agents whose outputs must hold up in the physical world.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yongqing Jiang, Haoran Luo, Jianze Wang, Xin Zhou, Kaoshan Dai, Zhiqi Shen. 2026-10-06. Vibe Building. https://arxiv.org/abs/2610.08156

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Hours-of-service-aware siting of charging and battery-swapping stations for long-haul electric trucks under adoption uncertainty

Planning en-route charging and battery-swapping infrastructure for long-haul battery-electric trucks (BETs) requires models that reflect how trucks actually operate. This paper develops a mixed-integer programming framework that jointly sites charging or swapping stations and schedules each truck's charging, swapping and mandatory driver rest, so that charging time overlaps with regulated rest instead of being added to it. Energy use is derived segment by segment from road terrain with a tractive-force model, and the truck battery is modelled as a set of independently swappable packs. Staged investment under uncertain BET adoption is formulated as a multistage stochastic program with Markovian demand and solved by stochastic dual dynamic integer programming (SDDiP) with Lagrangian cuts; we show that the Lagrangian multipliers can be bounded by each station's annualised cost without weakening the cuts. Applied to twelve freight corridors on Australia's East Coast, the algorithm jointly optimises charging and battery-swapping events as well as mandatory break events. A +/- 30\% change in adoption alters the final charge-only network by only -12\% to +13\% of stations, with over 95\% of stations built by the second stage. Hedging against uncertainty mainly changes which sites are chosen (71\% overlap with a deterministic rolling-horizon model), not when they are built.

cs.CE↗

Tool-calling retrieval versus vector RAG for a small Greek--English knowledge base: accuracy and robustness to how users type Greek

Assistants grounded in a small, frequently edited knowledge base can retrieve through tool calls to a live data interface or through vector retrieval-augmented generation (RAG). We compare the two on KyGround, a benchmark of 198 questions drawn from the published records of a Greek--English agricultural platform on Kythera, Greece, with answers verified automatically against the records and each question posed in up to nine forms, including Greek without accents, in capitals and in three Latin-script (Greeklish) schemes. With Claude Haiku 4.5 as router and answer model, a reconstruction of the platform's tool agent answered 71.6\% of canonical Greek questions correctly and vector RAG 95.3\% (difference $-23.6$ percentage points, 95\% CI $-33.1$ to $-15.1$). Letting the router write the vector query changed nothing, and placing the whole knowledge base of about 26,000 tokens in the prompt reached 99.3\%. The tool agent's losses arose in retrieval. Its literal searches returned nothing when the router's arguments did not occur verbatim in a record, for example when it transliterated Greek into Latin script or combined words that occur in a record but not as one phrase, and the agent then abstained. Unaccented and capitalised questions cost the tool agent about 20 points and vector RAG at most 2; accent-insensitive search removed this loss, and matching stemmed tokens raised the tool agent to 83.8\% on canonical Greek. Greeklish cost both designs about 21 to 32 points. Tool interfaces for community knowledge bases need search that tolerates how users type.

cs.CE↗

Closing the realism gap in physics-based gait simulations with a learned state prior

Predictive simulation of human movement is a promising tool for studying ``what-if'' scenarios in human movement and its underlying motor control, yet its realism is often limited. To address this gap, we incorporate a learned state prior that is trained on a large-scale dataset of human gait kinematics and external forces into predictive simulations. Resulting gait simulations yield kinematics and kinetics across diverse walking and running speeds that better match experimental data than current physics-based simulations, achieving accuracy comparable to data-driven models that reproduce learned data. Furthermore, our method enables robust hypothesis testing by demonstrating how varying optimality assumptions, muscle weakness, and footwear choices influence predicted gait. We also show that this prior generalizes well beyond its training data, successfully reconstructing full-body kinematics for curved running and cutting maneuvers from sparse marker sets. Ultimately, these results suggest that state priors should be broadly integrated into predictive simulations.

cs.CE↗