LMBuild: Evaluating LLM Agents for GeneratingBuildable and Functional Structures

Jiateng LiuRushi WangCheng QianXuejun Zhang Sun LiJiayu LiuYifan ShenXu Cao Jiarui YaoBingxuan LiRuhi SarikayaHeng Ji

Illinois Block ISiebel School of Computing and Data ScienceUniversity of Illinois Urbana-Champaign

{jiateng5, hengji}@illinois.edu

Can an LLM agent build an object that does more than look right: one that is functional as intended and buildable in the real world?

LMBuild evaluates LLM agents on building structures from a text prompt and an image: agents retrieve, create and place parts, then declare the joints, materials and assembly sequence that make a design buildable and functional.

30 systems evaluated 200 core tasks 2,549 full-set tasks 12 metrics, 4 levels

The taskFrom a photo and one sentence to an assembled structure

Each task gives the agent an input image and a one-line instruction. Each frame adds the parts of one placement step. Six real product-level builds:

LMBuild framework: a builder agent inspects, selects, creates and places parts, then declares joints, materials and an assembly sequence

FindingsKey takeaways

ResultsLeaderboard

All 30 systems on LMBuild-Core under the baseline protocol (retrieve + create). GPT-6 Astra leads the LLM agents, closely followed by Claude Fable 5.1 and Claude Opus 5; Qwen3.5-27B is the strongest open model.

How the scores are computed
  • ScaleEvery metric is on a common 0–100 scale. Cell colour runs from blue (low) through grey to orange (high), as in Table 2 of the paper.
  • MeanThe unweighted mean of the 12 metrics, shown only for systems that produce all 12. It is a ranking aid on this page; the paper reports the 12 metrics separately.
  • N/AThe system cannot produce the required output, for example a generator with no joints, materials or assembly order.
  • † (P)The generator's parts are passed through Particulate to obtain joints.
Generated structures and the hierarchical evaluation

ExploreThe benchmark

Click any task to watch every working build, step by step

AnalysisDiscussion

Functional affordance and physical operability are the largest gaps

Most systems build connected, collision-free structures, and frontier models already produce sound decompositions and aligned designs. Functional geometry (A.1) and parts (A.2) remain imperfect, and joints (A.3) are hard for all.

The gap carries into realization: agents name plausible assembly orders (R.1) and materials (R.2), but simulated operability (R.3) stays low for every system.

Strong models create, weaker models retrieve

Weaker models degrade sharply when they must author every part, and even offering creation alongside retrieval can hurt them: more freedom means more decisions.

Frontier models choose to create when both tools are available, and stay as strong or stronger without retrieval. For them, the catalog is the constraint.

Performance under three part-access settings

Figure 3. Scores under retrieve-only (A), retrieve + create (B) and create-only (C), on 25 sampled objects.

Tell the agent what the object must do

Two extra rounds of inspecting error reports and revising add only about 2 points. Functional descriptions grounded in Wikipedia add +10.3 on functional parts and +12.0 on operability (p < 0.001).

Agents turn stated functional constraints into structure far more reliably than they infer them from an object name alone.

Score change from revision rounds and richer prompts

Figure 4. Average score change from two revision rounds, attribute descriptions and functional descriptions (95% bootstrap CIs).

FAQ

What is LMBuild?

LMBuild tests whether LLM agents can generate 3D structures that can be built and that work. Each output is an assembled structure: parts, joints, materials and an assembly sequence.

What does the agent get, and what must it produce?

An image and a one-line instruction, e.g. “According to image W1, build a desk.” The agent retrieves or creates parts, places them, then declares joints, materials and an assembly order.

What are the interaction protocols?
  • Retrieve + create (baseline): catalog parts and newly created parts.
  • Retrieve only: catalog parts only.
  • Create only: every part is authored by the agent.
How is a build scored?

Twelve metrics in four levels: soundness (connectivity, collision, stability), affordance (geometry, parts, kinematics), design (decomposition, aesthetics, alignment) and realization (sequence, material, operability).

Which systems are evaluated?

Six frontier closed-source APIs, thirteen open-source LLMs and eleven domain-specific generators: BrickGPT, LegoACE, PartCrafter, PartPacker, Cube3D + CubePart, their Particulate variants and PhysX-Anything.

Where do the tasks come from?

LMBuild-Core (200 tasks) and LMBuild-Full (2,549) repurpose BrickNet, BrickComposer, PartNeXt, Fusion 360 Gallery, Artiverse and open-source hardware CAD, grounded in Wikipedia and Wikidata. See the Dataset tab.

Citation

BibTeX
@misc{liu2026lmbuildevaluatingllmagents,
      title={LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures},
      author={Jiateng Liu and Rushi Wang and Cheng Qian and Xuejun Zhang and Sun Li and Jiayu Liu and Yifan Shen and Xu Cao and Jiarui Yao and Bingxuan Li and Ruhi Sarikaya and Heng Ji},
      year={2026},
      eprint={2610.04292},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2610.04292},
}