Depth Anything V2

byTikTok / ByteDance & HKU

Extract high-contrast relative depth maps with razor-sharp silhouette boundaries and clean structural separation

Depth Anything V2

How Depth Anything V2 works

Convert any 2D image into clean, high-resolution relative depth conditioning in three simple steps.

Supply source scene

Supply source scene

Provide any clear 2D photograph, render, or illustration with defined spatial depth and foreground contrast.

Extract relative geometry

Extract relative geometry

The vision transformer analyses spatial contours and converts scene distance into a high-contrast grayscale field.

Export depth conditioning

Export depth conditioning

Feed the clean depth map directly into ControlNet conditioning pipelines, 3D parallax plates, or post-production blur shaders.

What Depth Anything V2 is good at

Engineered with a distilled vision transformer backbone that delivers clean silhouette isolation without haloing or diffusion latency.

Sub-pixel silhouette boundaries

Sub-pixel silhouette boundaries

Isolates complex contours like delicate foliage, fine hair strands, and slender furniture legs without edge bleeding, haloing, or artifact smearing.

Continuous spatial gradients

Continuous spatial gradients

Maps continuous distances from pure foreground white to deep background black, preserving subtle slope transitions across complex floorplans and landscapes.

Distilled transformer speed

Distilled transformer speed

Runs over ten times faster than diffusion-based estimators by using an efficient teacher-student architecture trained on 62 million synthetic depth labels.

ControlNet-ready luminance scale

ControlNet-ready luminance scale

Generates standardized grayscale depth maps calibrated for direct spatial guidance in modern diffusion workflows without manual histogram rebalancing.

Made with Depth Anything V2

A gallery of architectural, product, and cinematic compositions highlighting structural fidelity and edge precision in depth estimation.

Documentary archival plate with distinct geological separation

Documentary archival plate with distinct geological separation

Commercial product study with dynamic fluid detail

Commercial product study with dynamic fluid detail

Graphic album sleeve with sharp fragmented silhouettes

Graphic album sleeve with sharp fragmented silhouettes

E-commerce product still highlighting tactile materials

E-commerce product still highlighting tactile materials

Monolithic interior layout demonstrating spatial gradience

Monolithic interior layout demonstrating spatial gradience

What people build with Depth Anything V2

From generative ControlNet conditioning to post-production optical blur, discover how spatial engineers and visual artists leverage relative depth.

ControlNet generative conditioning

01

Guide image synthesis in diffusion models by locking scene perspective, camera angles, and structural anatomy with clean grayscale guidance.

Depth-of-field and bokeh synthesis

02

Generate physically accurate optical depth-of-field maps for compositing, enabling custom focal plane selection and realistic lens blur.

3D scene layout and spatial extraction

03

Reconstruct spatial geometry and room dimensions from single 2D interior renders and architectural photographs for virtual staging.

Product isolation and relighting

04

Separate products cleanly from their backdrops using crisp luminance boundaries, accelerating catalog relighting and background replacements.

Parallax plate extraction

05

Slice still photographs into multi-plane 2.5D visual effects plates for camera projection, UI animations, and immersive video transitions.

Frequently asked questions

Depth Anything V2 is designed for extracting high-contrast relative depth maps from 2D images. It is ideal for ControlNet spatial conditioning in diffusion pipelines, post-production depth-of-field blur, foreground subject isolation, and 2.5D parallax plate generation.

Depth Anything V2 replaces noisy real-world depth annotations with synthetic data distillation, eliminating the edge haloing, boundary bleeding, and smeared silhouettes common in older models. It produces significantly cleaner structural outlines and runs over ten times faster than diffusion-based alternatives.

No, Depth Anything V2 generates relative depth maps rather than metric physical measurements. The output maps proximity as a continuous grayscale gradient from near (white) to far (black), rather than true real-world units like meters or feet.

The model struggles with mirrors, highly reflective surfaces, and transparent elements like glass or smoke. Because monocular depth estimation relies on visual cues, reflections and refractions can cause depth inversion or false holes in the resulting depth map.

Depth Anything V2 was created by researchers at TikTok/ByteDance and the University of Hong Kong (HKU), published in June 2024 by Lihe Yang and collaborators.

Try Depth Anything V2 on Fuser

Extract high-contrast relative depth maps with razor-sharp silhouette boundaries and clean structural separation