OpenAI

How to Turn an Image Into a 3D World with Sora 2

Use Sora 2 as the downstream camera-move engine for an Image-to-3D-World workflow on Astorie — the captured stills from the navigable world feed directly into Sora 2 video nodes for matched-angle motion shots. The world node's output is a canvas-internal navigable scene preview, not a portable .obj, .fbx, .glb, or USD mesh. Sora 2 takes the captured stills as starting frames and produces video clips that all share the same locked location, with cinematographic camera moves that respect the spatial structure of the source world.

Step-by-Step Guide

1

Source the navigable world from Nano Banana 2 or FLUX.2

Sora 2 is the camera-move engine, not the world generator. Upstream: drop a Nano Banana 2 or FLUX.2 image node, generate the source scene at 4K, then wire into a World Labs or Image-to-3D-World node. The navigable scene preview lands inside the canvas (~5 minute generation) — orbit and capture stills before moving to Sora 2.

2

Capture four matched angles from the world

Inside the navigable preview, capture stills from the four-angle pattern: front view, three-quarter left, three-quarter right, back/over-shoulder. Each capture lands as an image node. These are your Sora 2 starting frames — capture more than you need. Re-running the world node produces a different scene, so screenshot first, iterate later.

3

Wire each captured still into its own Sora 2 video node

Sora 2 has deep understanding of 3D space, motion, and scene continuity — captured stills from a navigable world are the cleanest input format. Wire each captured angle into its own Sora 2 image-to-video node. The video model inherits the spatial structure from the still and produces motion that respects parallax, occlusion, and depth.

4

Write cinematographic motion prompts per shot

Use cinematographic verbs in the Sora 2 prompt: "slow camera push forward," "gentle orbit clockwise," "static camera, character moves out of frame right," "dolly forward with parallax." Sora 2 maps these directly to its training distribution. Avoid generic verbs ("move closer," "spin") — they leave the model guessing and produce inconsistent shot-to-shot behavior.

5

Stitch shots with last-frame to first-frame chaining

For sequences longer than one Sora 2 clip, route the last frame of clip N into a frame-extraction tool node, then feed it as the starting frame of clip N+1. Combined with the locked 3D world reference, this gives both spatial AND temporal continuity. The world locks the location; frame chaining locks the motion thread.

6

Export the multi-shot sequence to NLE

Drop the Sora 2 outputs into Astorie's sequence builder in story order. Each clip is 5-10s. Cut markers preserved. Layer audio (ElevenLabs Eleven v3 + Minimax Music). Export as native sequence to Premiere, DaVinci Resolve, or Final Cut. The locked 3D world is what made the multi-shot read as one place; the NLE export is the final delivery.

Prompt Examples

Establishing shot via slow push. The captured still locks the location; Sora 2 adds the camera move.

[Captured still from world: front view of empty mid-century living room] + Sora 2 prompt: slow camera push forward through the living room toward the fireplace, soft afternoon light from the windows on camera left, no character, 8 seconds, 16:9.

Medium shot via orbit. Same world; new angle. The orbit instruction maps to Sora 2's 3D-aware training.

[Captured still: three-quarter left angle of same living room] + Sora 2 prompt: gentle orbit clockwise around the center of the room, lighting unchanged, depth of field shallow on the armchair, 6 seconds, 16:9.

Detail close-up via static zoom. The static instruction tells Sora 2 not to add unwanted parallax.

[Captured still: tight angle on the fireplace] + Sora 2 prompt: static camera, slow zoom in toward the fireplace mantelpiece, no other motion in frame, atmospheric, 5 seconds, 16:9.

Reverse shot with parallax dolly. Sora 2's strongest move type — the depth structure of the still drives the parallax effect.

[Captured still: reverse over-shoulder angle, looking back toward the windows] + Sora 2 prompt: dolly forward with parallax, soft afternoon light revealing dust motes in the air, 7 seconds, 16:9.

Parameter Tips

Sora 2 is the camera-move engine, not the world generator. The world comes from a Nano Banana 2 / FLUX.2 source + World Labs / Image-to-3D-World node upstream.

Capture stills from the world BEFORE running Sora 2. Re-running the world produces a different scene; capture once, fan out to many Sora 2 nodes.

Use cinematographic verbs (dolly, orbit, push, pull, static, parallax) — they map to Sora 2's training distribution. Generic verbs produce inconsistent results.

For sequences, use last-frame chaining: Sora 2 clip N's last frame = clip N+1's starting frame. Combined with the locked world, both spatial and temporal continuity are preserved.

Sora 2 image-to-video clips are 5-10s. For longer takes, chain shorter clips rather than asking for a single impossibly-long clip.

The world node output is canvas-internal — Sora 2 uses captured stills, not the navigable world directly. Export from Astorie = NLE-ready video sequence, not a 3D file.

What to Expect

Sora 2 returns 5-10s 1080p video clips per node, with strong 3D spatial reasoning that respects the depth structure of captured stills from a navigable world. Generation time 60-120s per clip. Cinematographic camera moves (dolly, orbit, push, parallax) are Sora 2's strongest territory. Output drops onto the canvas; chain via sequence builder for multi-shot delivery, NLE export for native Premiere/DaVinci sequences. The world remains canvas-internal; Sora 2 outputs are exportable video deliverables.

Use Sora 2 on Astorie

Connect Sora 2 with other AI models on Astorie's infinite canvas. No GPU required — start free.

Get Started Free

Related features

Docs

Related reading

Try Other Models for This Task

How to Turn an Image Into a 3D World