跳到正文
The Decoder:AI News· Jonathan Kemper·· 3 小时前AI 评分54

LEGO-Anything研究:编码智能体可将单张照片转成可执行Blender 3D场景程序但几何自评接近随机

AI agents build 3D scenes from photos but have no idea if they got it right

AI 导读

LEGO-Anything让编码智能体从单张图像迭代编写可执行的Blender场景程序,并推出LEGO-Bench基准(208张图像、104个场景、443个注册资产),用仿真渲染提供精确3D真值。

正文
LEGO-Anything overview showing how a coding agent converts a single photo into an executable Blender program for editable 3D scenes
LEGO-Anything has a coding agent write an executable Blender program from a single photo. The resulting scene can be edited and queried for image analysis tasks. | Image: Li et al.

The approach is called "Image-to-Code." A coding agent receives a single image and writes code for Blender, the widely used 3D software. Rather than generating the scene in one pass, the agent works iteratively: it writes code, runs it, looks at the result, and revises until the scene matches the original.

Because the output is a program, it captures objects, geometry, layout, and camera position explicitly. You can run, check, and modify the scene like any other piece of code.

Simulator scenes provide the exact ground truth

To measure how well agents perform, the team introduces LEGO-Bench. It contains 208 images from 104 indoor and outdoor scenes and uses 443 registered assets.

Real photos don't provide a precise 3D ground truth to compare against, according to the researchers. Simple synthetic scenes look unrealistic. So LEGO-Bench splits the difference by rendering its images from professionally built simulator scenes. The inputs look natural, while the exact geometry, depth, and object assignments stay hidden and serve as the answer key for automated scoring. Scene complexity can also be ramped up without changing lighting or camera settings.

LEGO-Bench evaluation metrics for validity, geometric accuracy, and visual similarity
LEGO-Bench scores scenes on validity, geometric accuracy, and visual similarity to the original. | Image: Li et al.

The benchmark scores each scene on three axes: validity checks whether a usable scene artifact was delivered at all. Reconstruction measures how accurate the visible geometry is. Appearance captures how closely the look matches the original by re-rendering the submitted scene and comparing it pixel by pixel against the reference image.

Agents deliver usable artifacts but struggle with geometry

All six tested GPT configurations delivered a working scene almost every time. Accuracy varied wildly, though. GPT-6 Astra, the best tested model, hit 53.4 percent on indoor scenes and 39.6 percent on outdoor scenes. Weaker configurations scored around 15 percent.

Reconstruction examples comparing GPT-6 Astra output to reference images and older models
GPT-6 Astra comes closest to the reference images. Older models frequently miss camera angles, lighting, or entire objects. | Image: Li et al.

The more complex a scene, the more accuracy drops, and outdoor scenes are harder than interiors. When the researchers increased the models' reasoning budget, the GPT-6 variants improved a lot. Astra's score on an office test subset jumped from 32.3 to 61.8 percent.

To understand why, the researchers analyzed the agents' work steps. They found poor initial attempts, revisions that undid earlier progress, and unreliable self-assessment as the most common issues.

Regression example where GPT-6 Astra degrades its own scene from 33.9 to 4.4 percent late in the process
Even GPT-6 Astra wrecks its own scene late in the process, dropping from 33.9 to 4.4 percent. | Image: Li et al.

That last point is the most telling. When models had to pick which of two versions better matched the original, their geometric judgments landed near or below chance level. An agent basically can't tell whether its own scene has gotten better. The researchers conclude that refinement should rely on concrete measurements, not the agent's own judgment.

Heatmap showing model judgment accuracy on geometry near chance level
When judging geometry, the models perform near chance level, even when evaluating their own scenes. | Image: Li et al.

That insight led the authors to build LEGO-Plugin, an extension that needs no extra training. It anchors the starting scene in the reference image, swaps the unreliable self-judgment for concrete measurements, and shields correct progress from regressive edits. The plugin improved all six models. Weaker agents saw the biggest gains, with boosts up to 62.7 percent. The already strong top model gained only about two percentage points.

LEGO-Plugin benchmark results showing improvements across all tested models
The plugin improves all tested models. Weaker agents benefit the most. | Image: Li et al.

Reconstructed scenes aren't accurate enough yet

Finally, the team tested whether reconstructed scenes could serve as a basis for standard vision tasks. Because each scene is an executable program, object detection, segmentation, and depth estimation can be pulled directly from it.

Without any extra training, the scenes produced usable but unremarkable results across all three tasks. Object detection fared best, with the reconstructed scenes reaching roughly half the performance of the specialized model DINO. For segmentation and depth estimation, the gap to specialized models like SAM 3 and Depth Anything 3 was larger. The authors say executable scene programs from current coding agents show promise but aren't accurate enough. A big gap remains between a working result and a faithful reconstruction.

GPT-6 Astra's commanding lead in LEGO-Bench lines up with other observations. AI researcher Yoav Artzi sees the model as a major leap in spatial understanding and suspects it was trained on large amounts of 3D data like Blender scenes. 3D software makers are already gearing up for these kinds of agents. Unity has released official plugins for Claude Code and Codex. Other approaches skip code entirely and reconstruct scenes directly inside the model, like the Atlas world model from World Labs. Google Deepmind takes yet another route with GenCeption, using a video model for depth estimation and segmentation that matches the performance of specialized models.

来源:The Decoder:AI News · the-decoder.com