The aim: have Claude, an AI agent, do real 3D work, with architectural visualization as the test. Continuity is the known weakness of AI image and video models: ask for the same room twice and it changes. So the workflow is entirely AI but grounded in 3D. Claude builds the house in Blender, a person directs it in plain language without touching the software, and image and video models repaint the renders. One 3D model, so the house is identical in every shot.
Generative image and video models make a new scene on every request. Ask for the same room twice and the furniture moves and the light changes. That breaks continuity, and environments need continuity for storytelling: a room has to be the same room every time the story returns to it, and audiences notice when it is not.
What started this project was inconsistency between rooms. The kitchen seen through the living room doorway should be the kitchen from the previous shot, and even with reference photographs and detailed instructions the models could not hold it: references suggest, they do not constrain. Only traditional anchoring keeps things accurate, the method film studios have always used. If the world exists as real geometry, nothing has to be remembered. Build the set.
Claude builds it, but cannot reliably recall Blender technique from memory, so tutorials, the supervisor's rules and the lessons of every build are compiled into a skill library and indexed into a database it queries before building anything.
With that library Claude designs and builds the whole house in real 3D: plan, walls, furniture, lighting, cameras. One person directs with words, pictures and taste and never opens the software. The set does not need to look finished, it needs to be correct. The image and video models, Seedream 5.0 and Seedance 2.5, then produce the final pictures on top of its renders.
Read the diagram top down. The supervisor sits above everything: brief and references in, rulings out. Below that is the skill library, consulted before anything is built. Production runs left to right, from design to a human ruling; every Blender render is reviewed and every AI output judged. The blue line along the bottom is the anchor: the approved Blender render, which the AI may repaint but not rearrange. A problem found late goes back to the step that caused it and the work reruns from there.
Drag to pan · click it, then scroll or pinch to zoom · the square button fits the whole pipeline.
Claude works from a plan and a strict procedure. Before any 3D existed it drew the house in two dimensions: full floor plans, every room at real dimensions in metres, down to the sink, stove and fridge. Photo references came from Nano Banana Pro, Google's image model, and from the supervisor, who ruled on the drawings; the plan went through sixteen drawn versions before the first wall, and the layout kept changing during the build. Below is the plan as built, cut from the finished 3D model 1.4 metres above each floor, so every room and dimension can be checked against the renders.
Each floor starts in full view. Drag to pan · click it, then scroll or pinch to zoom · the square button fits the floor again.
The Claude agents work inside a running Blender session through two software links: one carries build commands, measurements and checks; the other stays open for renders that run for minutes and would time out on the first. The build order is fixed: slab, walls and openings, the second floor and roof, the exterior finish, the environment, the interiors object by object, then props, set dressing, shading and lighting. Every object starts with a query to the skill library, not the model's memory. Below are the main stages, not every step: the first image is the plan as built, then Blender viewport, flat shadeds of the scene as it grew, and a final render.








Every render is judged before it goes anywhere. Gemini 3.1 Pro compares the frame with real photographs and names the defects: faceted curves, gaps between parts that should touch, missing everyday items, wrong proportions. Each goes to where it gets fixed, in the 3D scene or in the image. Then the supervisor signs off on the frame pair. AI finishing never starts from an unapproved render.
Every shot, in order:
Drag the handle to compare Blender's render with the AI's still, then use the player: both clips run in lock step. Wipe across the moving frame, set them side by side, or watch one alone. The ticks at the ends of the timeline are the two anchor stills. The AI changes color, materials and light and fills in the world past the property line; every wall, window and cabinet stays where Blender put it. The set holds across shots: the room seen through the kitchen doorway is the family room shot.
Blender renderAI finish
Blender renderAI finish
Blender renderAI finish
Blender renderAI finish
Blender renderAI finish
Two mechanisms, both forced by failures, make each build need less supervision than the last.
A tree of skills, 164 pages in nine branches from architecture and lighting down to door hardware. Claude Opus agents condense tutorials, design theory and reference images into it; Claude Fable builds from it. The tree is embedded into a vector database of 3,017 chunks, queried per object before that object is built: the chair before the chair. Nothing is built from memory. Results are weighted by authority, so a supervisor's rule outranks the tutorial it corrects. Every edit is re-indexed.
A defect goes back to the step that caused it: a wrong window grille to the 3D scene, a wrong color cast to the image stage, and the flow reruns from there. Every failure becomes a lesson in the database and every supervisor ruling a law that outranks what comes after. Hard problems are kept as idea trees: the symptom at the root, each attempt a branch, the verdict at the leaf, so the next run walks the tree instead of repeating the dead ends.
None of this happened in one pass; it took a great deal of human input and feedback. The biggest gap is judging quality. The model finds measurable geometry problems on its own: surfaces lying on each other (coplanar faces), inverted surfaces, holes, intersecting parts. Whether an image looks right is still human judgment.
The speed limit is the agent, not the image models and not the reviewer. The time goes into the agent's first mistakes and the corrections that follow. As models improve and the library grows, the agent repeats fewer mistakes and the same supervision yields more finished work. That is why the goal is one supervisor rather than none.
The method transfers: recipes written for one chair build chairs that appeared in no tutorial, because the library stores the method, not the chair. In January 2026 this ran on Opus 4.5, an earlier Claude, with heavy correction at every step: uneven quality room to room, rules read but not followed, the library consulted then ignored. Fable, the current model, runs the same method with far less of that. Improvements in the models compound on a library that never forgets. That is a forecast, not a promise.