AI depth estimation for video

AI Video Depth Map Generator: turn any clip into a depth pass

Drop in a video and get a depth map back, one for every frame: brighter is nearer the camera, darker is further away. The motion, camera move, framing and timing all survive; colour and texture do not. That is the point — it gives a video model, a compositor or a 3D tool the structure of real footage without dragging its appearance along.

Output

Greyscale is the one every downstream tool reads. The colour ramps are for looking at.

8 credits per second of video

Up to 100 MB per clip

40 credits

The depth video opens in a session you can come back to, ready to download or carry into another tool.

How it works

Make a depth map in three steps

  1. Step 1

    Upload your video

    Drop in an MP4, MOV or WebM up to 100 MB. Resolution and frame rate follow the source, so nothing is resized or resampled on the way in.

  2. Step 2

    Depth is estimated per frame

    The model reads the whole clip rather than each frame alone, and writes out how far every pixel sits from the camera. Because it looks at neighbouring frames, a wall that is not moving keeps the same depth instead of shifting from frame to frame.

  3. Step 3

    Generate and download

    The depth video opens in a session you can come back to, so you can download it, run the next clip, or carry the result into the workflow editor as the control input for a video model.

Before and after

The same shot, written as distance

  • Dozens of people at dozens of distances, each one separated from the next. The whole tonal range is in use, from the near-white figures at the bottom edge to the black at the vanishing point.
  • A few subjects at very different distances, and the fine structure survives: the plant's separate leaves are resolved, and the room through the doorway correctly reads as further away than the wall around it.

What it is for

When a model or a compositor needs structure, not pixels

Depth-guided video generation

Feed the depth pass to a video model as a control input and it will follow your footage's layout, camera move and timing while inventing an entirely different look. Depth carries the geometry without carrying the faces, colours or textures of the original.

Compositing and grading

A depth pass drives depth-of-field, atmospheric haze, fog and relighting in After Effects, Nuke, Fusion or Resolve. It is the matte you would otherwise roto by hand, and it updates itself as the shot moves.

3D, spatial video and parallax

Displace a plane by the depth map and a flat clip gains real parallax — the 2.5D camera move, a spatial-video conversion, a wiggle-stereo pass. Depth is the one channel those effects all need and no camera records.

Separating near from far

Because depth is continuous rather than a hard cutout, you can threshold it anywhere: isolate the foreground subject, drop only the background, or grade the middle distance on its own without keying on colour.

How the estimation works

Built for video, which is what keeps it steady

It reads the frames around each frame

Running an image depth model over a clip frame by frame is what makes depth boil: every frame gets its own independent guess, and the guesses disagree slightly, so flat surfaces crawl and edges buzz. A video depth model reads neighbouring frames together, so a surface that is not moving holds the same value and the pass stays usable in an edit.

Relative depth, not metres

The output is relative: it tells you what is nearer and what is further within a shot, not how far anything is in real units, and the scale is re-fitted per clip. That is what ControlNet, displacement and depth-of-field all want. It is not what a measuring tool wants, and no single-camera model can give you that.

Greyscale to work with, colour to look at

Greyscale is the interchange format: every tool that consumes a depth map reads brightness, so that is the default. The four colour ramps — turbo, inferno, magma and viridis — spread the same values across a hue range, which makes the depth ordering far easier to read on screen but has to be converted back before anything downstream can use it.

Pay by the second, not by the month

Every plan includes credits you can spend on this page or anywhere else on YouArt, and a new account starts with free trial credits.

Basic

For hobbyists and explorers

$9.99/mo
Select Plan
  • 1000 credits
  • Up to ~200 images/month
  • At least ~1000s video/month
  • Intelligent creative agent
  • Video editor
  • Latest image models, including GPT Image 2 and Nano Banana Pro
  • Latest video models, including Seedance 2
  • Realistic face uploads
  • Voice generation with ElevenLabs
  • No watermark
  • Unlimited template access

Pro

Most Popular

For creators and pro users

$29.99/mo
Select Plan
  • 3300 credits
  • Up to ~1000 images/month
  • At least ~3300s video/month
  • Intelligent creative agent
  • Video editor
  • Latest image models, including GPT Image 2 and Nano Banana Pro
  • Latest video models, including Seedance 2
  • Realistic face uploads
  • Voice generation with ElevenLabs
  • No watermark
  • Unlimited template access

Max

For power users and teams

$149.99/mo
Select Plan
  • 18000 credits
  • Up to ~6000 images/month
  • At least ~18000s video/month
  • Intelligent creative agent
  • Video editor
  • Latest image models, including GPT Image 2 and Nano Banana Pro
  • Latest video models, including Seedance 2
  • Realistic face uploads
  • Voice generation with ElevenLabs
  • No watermark
  • Unlimited template access

Team

For teams and studios

$329.99/mo
Select Plan
  • 36300 credits
  • Up to ~12000 images/month
  • At least ~36300s video/month
  • Realistic face uploads
  • Share canvas, workflows, and assets with your team
  • Up to 10 members per team
  • Up to 5 teams
  • Role-based management
  • Team credit management and spending caps
  • Per-member usage tracking

Questions

Video depth map FAQ

What is a video depth map?
It is a second video the same length as yours in which brightness means distance: the brighter a pixel, the nearer that point is to the camera. It keeps the original's motion, camera move, framing and timing and discards colour and texture, which is what makes it useful as an input to something else rather than as something to watch.
Why not just run an image depth model on every frame?
You can, and the result usually flickers. Each frame is estimated independently, so a wall that never moves gets a slightly different depth in every frame and appears to crawl or boil. This model reads the frames around each frame, so static geometry holds still. That temporal stability is the whole reason a video-specific depth model exists.
Is this metric depth? Can I measure distances with it?
No. It is relative depth — it ranks what is nearer and further within a shot, and the scale is fitted per clip rather than expressed in metres. That is exactly what depth-guided generation, displacement, fog and depth-of-field need. If you need true measurements you need a depth sensor or a calibrated multi-camera rig, which no single-video model can substitute for.
Can I use the output to control an AI video model?
Yes, and it is the most common reason to make one. A depth pass hands a video model your footage's geometry, camera move and timing while leaving the appearance entirely open, so you can restyle a shot completely and still keep its blocking. Use the greyscale output for this — that is what control inputs expect.
Which colour output should I choose?
Greyscale unless you are only looking at it. Every downstream consumer — a control input, a compositor's depth channel, a displacement modifier — reads brightness, so a turbo or viridis file has to be converted back before it will work. The colour ramps are genuinely better for judging depth by eye, which makes them useful for review and presentation.
What does it cost, and can I try it for free?
It is billed by the length of the clip, at 8 credits per second: a ten-second video costs 80 credits and a minute costs 480. A new account starts with free trial credits, so you can run a short clip and see the output before deciding.
Does the depth video keep my audio?
No. The output is silent — a depth map is picture data and any soundtrack on the source is dropped. If your final cut needs the original audio, bring it back in the video editor alongside the depth pass.
Which video files can I upload?
MP4, MOV and WebM, up to 100 MB per file. Resolution and frame rate follow the source rather than being forced to a preset, so what comes back lines up frame for frame with what went in.

More tools

Other things to do with the same file

Each one does a single job on a file you already have — no prompt to write and no editor to learn.

Browse all 9 AI tools

Beyond one pass

Want to use the depth, not just download it?

This page runs one fixed step. Open the workflow editor to wire the depth pass straight into a video model as a control input, run a batch of clips through the same graph, or chain it after an upscale.