LMBuild: Evaluating LLM Agents for GeneratingBuildable and Functional Structures
Siebel School of Computing and Data ScienceUniversity of Illinois Urbana-Champaign
{jiateng5, hengji}@illinois.edu
Paper arXiv Leaderboard Data Code Citation
Can an LLM agent build an object that does more than look right: one that is functional as intended and buildable in the real world?
LMBuild evaluates LLM agents on building structures from a text prompt and an image: agents retrieve, create and place parts, then declare the joints, materials and assembly sequence that make a design buildable and functional.
The taskFrom a photo and one sentence to an assembled structure
Each task gives the agent an input image and a one-line instruction. Each frame adds the parts of one placement step. Six real product-level builds:
FindingsKey takeaways
- 1
Looking right is no longer the bottleneck; working right is. Frontier models score in the high 80s on soundness and around 70 on design, but at most 42.7 on kinematics (A.3) and 22.1 on simulated operability (R.3).
- 2
Strong models create, weaker models retrieve. Frontier agents gain from authoring their own parts. Weaker models lose ground once creation is allowed: Qwen3-VL-32B drops from 61.7 with retrieval only to 34.1 with creation only.
- 3
Specifying beats revising. Two extra revision rounds add about 2 points. Spelling out functional requirements adds +10.3 on functional parts and +12.0 on operability.
- 4
Open models have a part-creation gap. Closed models succeed on 95.7% of part-creation calls, open models on 51.7%. Most failures break CAD-kernel conventions or the program structure, not the geometry. Qwen3.5-27B (89.4%) is the exception.
- 5
Open models lose track of the scene. They refer to parts that were never placed 7.2 times per 100 calls, against 0.3 for closed models.
ResultsLeaderboard
All 30 systems on LMBuild-Core under the baseline protocol (retrieve + create). GPT-6 Astra leads the LLM agents, closely followed by Claude Fable 5.1 and Claude Opus 5; Qwen3.5-27B is the strongest open model.
How the scores are computed
- ScaleEvery metric is on a common 0–100 scale. Cell colour runs from blue (low) through grey to orange (high), as in Table 2 of the paper.
- MeanThe unweighted mean of the 12 metrics, shown only for systems that produce all 12. It is a ranking aid on this page; the paper reports the 12 metrics separately.
- N/AThe system cannot produce the required output, for example a generator with no joints, materials or assembly order.
- † (P)The generator's parts are passed through Particulate to obtain joints.
ExploreThe benchmark
Click any task to watch every working build, step by step
LMBuild-Core: 200 reference structures
AnalysisDiscussion
Functional affordance and physical operability are the largest gaps
Most systems build connected, collision-free structures, and frontier models already produce sound decompositions and aligned designs. Functional geometry (A.1) and parts (A.2) remain imperfect, and joints (A.3) are hard for all.
The gap carries into realization: agents name plausible assembly orders (R.1) and materials (R.2), but simulated operability (R.3) stays low for every system.
Strong models create, weaker models retrieve
Weaker models degrade sharply when they must author every part, and even offering creation alongside retrieval can hurt them: more freedom means more decisions.
Frontier models choose to create when both tools are available, and stay as strong or stronger without retrieval. For them, the catalog is the constraint.
Figure 3. Scores under retrieve-only (A), retrieve + create (B) and create-only (C), on 25 sampled objects.
Tell the agent what the object must do
Two extra rounds of inspecting error reports and revising add only about 2 points. Functional descriptions grounded in Wikipedia add +10.3 on functional parts and +12.0 on operability (p < 0.001).
Agents turn stated functional constraints into structure far more reliably than they infer them from an object name alone.
Figure 4. Average score change from two revision rounds, attribute descriptions and functional descriptions (95% bootstrap CIs).
FAQ
What is LMBuild?
LMBuild tests whether LLM agents can generate 3D structures that can be built and that work. Each output is an assembled structure: parts, joints, materials and an assembly sequence.
What does the agent get, and what must it produce?
An image and a one-line instruction, e.g. “According to image W1, build a desk.” The agent retrieves or creates parts, places them, then declares joints, materials and an assembly order.
What are the interaction protocols?
- Retrieve + create (baseline): catalog parts and newly created parts.
- Retrieve only: catalog parts only.
- Create only: every part is authored by the agent.
How is a build scored?
Twelve metrics in four levels: soundness (connectivity, collision, stability), affordance (geometry, parts, kinematics), design (decomposition, aesthetics, alignment) and realization (sequence, material, operability).
Which systems are evaluated?
Six frontier closed-source APIs, thirteen open-source LLMs and eleven domain-specific generators: BrickGPT, LegoACE, PartCrafter, PartPacker, Cube3D + CubePart, their Particulate variants and PhysX-Anything.
Where do the tasks come from?
LMBuild-Core (200 tasks) and LMBuild-Full (2,549) repurpose BrickNet, BrickComposer, PartNeXt, Fusion 360 Gallery, Artiverse and open-source hardware CAD, grounded in Wikipedia and Wikidata. See the Dataset tab.
Citation
@misc{liu2026lmbuildevaluatingllmagents,
title={LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures},
author={Jiateng Liu and Rushi Wang and Cheng Qian and Xuejun Zhang and Sun Li and Jiayu Liu and Yifan Shen and Xu Cao and Jiarui Yao and Bingxuan Li and Ruhi Sarikaya and Heng Ji},
year={2026},
eprint={2610.04292},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2610.04292},
}