LEGO-Anything研究:编码智能体可将单张照片转成可执行Blender 3D场景程序但几何自评接近随机
AI agents build 3D scenes from photos but have no idea if they got it right
LEGO-Anything让编码智能体从单张图像迭代编写可执行的Blender场景程序,并推出LEGO-Bench基准(208张图像、104个场景、443个注册资产),用仿真渲染提供精确3D真值。

The approach is called "Image-to-Code." A coding agent receives a single image and writes code for Blender, the widely used 3D software. Rather than generating the scene in one pass, the agent works iteratively: it writes code, runs it, looks at the result, and revises until the scene matches the original.
Because the output is a program, it captures objects, geometry, layout, and camera position explicitly. You can run, check, and modify the scene like any other piece of code.
Simulator scenes provide the exact ground truth
To measure how well agents perform, the team introduces LEGO-Bench. It contains 208 images from 104 indoor and outdoor scenes and uses 443 registered assets.
Real photos don't provide a precise 3D ground truth to compare against, according to the researchers. Simple synthetic scenes look unrealistic. So LEGO-Bench splits the difference by rendering its images from professionally built simulator scenes. The inputs look natural, while the exact geometry, depth, and object assignments stay hidden and serve as the answer key for automated scoring. Scene complexity can also be ramped up without changing lighting or camera settings.

The benchmark scores each scene on three axes: validity checks whether a usable scene artifact was delivered at all. Reconstruction measures how accurate the visible geometry is. Appearance captures how closely the look matches the original by re-rendering the submitted scene and comparing it pixel by pixel against the reference image.
Agents deliver usable artifacts but struggle with geometry
All six tested GPT configurations delivered a working scene almost every time. Accuracy varied wildly, though. GPT-6 Astra, the best tested model, hit 53.4 percent on indoor scenes and 39.6 percent on outdoor scenes. Weaker configurations scored around 15 percent.

The more complex a scene, the more accuracy drops, and outdoor scenes are harder than interiors. When the researchers increased the models' reasoning budget, the GPT-6 variants improved a lot. Astra's score on an office test subset jumped from 32.3 to 61.8 percent.
To understand why, the researchers analyzed the agents' work steps. They found poor initial attempts, revisions that undid earlier progress, and unreliable self-assessment as the most common issues.

That last point is the most telling. When models had to pick which of two versions better matched the original, their geometric judgments landed near or below chance level. An agent basically can't tell whether its own scene has gotten better. The researchers conclude that refinement should rely on concrete measurements, not the agent's own judgment.

That insight led the authors to build LEGO-Plugin, an extension that needs no extra training. It anchors the starting scene in the reference image, swaps the unreliable self-judgment for concrete measurements, and shields correct progress from regressive edits. The plugin improved all six models. Weaker agents saw the biggest gains, with boosts up to 62.7 percent. The already strong top model gained only about two percentage points.

Reconstructed scenes aren't accurate enough yet
Finally, the team tested whether reconstructed scenes could serve as a basis for standard vision tasks. Because each scene is an executable program, object detection, segmentation, and depth estimation can be pulled directly from it.
Without any extra training, the scenes produced usable but unremarkable results across all three tasks. Object detection fared best, with the reconstructed scenes reaching roughly half the performance of the specialized model DINO. For segmentation and depth estimation, the gap to specialized models like SAM 3 and Depth Anything 3 was larger. The authors say executable scene programs from current coding agents show promise but aren't accurate enough. A big gap remains between a working result and a faithful reconstruction.
GPT-6 Astra's commanding lead in LEGO-Bench lines up with other observations. AI researcher Yoav Artzi sees the model as a major leap in spatial understanding and suspects it was trained on large amounts of 3D data like Blender scenes. 3D software makers are already gearing up for these kinds of agents. Unity has released official plugins for Claude Code and Codex. Other approaches skip code entirely and reconstruct scenes directly inside the model, like the Atlas world model from World Labs. Google Deepmind takes yet another route with GenCeption, using a video model for depth estimation and segmentation that matches the performance of specialized models.
来源:The Decoder:AI News · the-decoder.com