DA3 SMALL

Providerdepth-anything
Categorydepth-estimation
Licenseapache-2.0
Downloads255
Stars0

Overview

DA3 Small is a compact, high-efficiency depth estimation model designed for real-time spatial analysis. Unlike heavy vision transformers, this model focuses on providing precise relative depth maps from single RGB images with minimal computational overhead. It is particularly suited for developers building robotics pipelines, AR applications, or edge-computing vision systems where latency is critical. With an Apache-2.0 license, it offers flexible integration into commercial stacks. Compared to larger depth models, DA3 Small prioritizes inference speed and memory efficiency while maintaining robust boundary definition, making it an ideal choice for deployment on mobile devices or embedded hardware.

Highlights

  • Real-time relative depth estimation from single RGB images
  • Optimized for low-latency edge and mobile deployment
  • Permissive Apache-2.0 license for commercial integration
  • Efficient memory footprint for embedded vision pipelines

Usage

Install
# Install Hugging Face transformers
pip install transformers torch
SDK Usage
# Load model with transformers
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("depth-anything/DA3-SMALL")
tokenizer = AutoTokenizer.from_pretrained("depth-anything/DA3-SMALL")

Hugging Face Download

We recommend downloading the model via the Hugging Face CLI or Hub SDK.

Guidance:Before downloading, install huggingface_hub with:

Guidance
pip install -U huggingface_hub

CLI Download

Download the full repository

Download the full repository
huggingface-cli download depth-anything/DA3-SMALL

Download a single file to a local folder (e.g. config.json into ./dir)

Download a single file to a local folder (e.g. config.json into ./dir)
huggingface-cli download depth-anything/DA3-SMALL config.json --local-dir ./dir

See the official docs for more CLI options

SDK Download

SDK Download
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('depth-anything/DA3-SMALL')

Git Download

Make sure git-lfs is installed first

Git Download
git lfs install
git clone https://huggingface.co/depth-anything/DA3-SMALL

To skip LFS large-file downloads, use:

Skip LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/depth-anything/DA3-SMALL

Model files are hosted on the Hugging Face Hub — download directly via HF CLI / SDK / Git, not through this site.

PyTorch / Transformers Usage

Install Transformers

Install Transformers
pip install -U transformers torch

Load the model and run inference

Load the model and run inference
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained('depth-anything/DA3-SMALL')
tokenizer = AutoTokenizer.from_pretrained('depth-anything/DA3-SMALL')

Model Download

We recommend downloading the model via the ModelScope CLI or SDK.

Guidance:Before downloading, install ModelScope with:

Guidance
pip install modelscope

CLI Download

Download the full repository

Download the full repository
modelscope download --model depth-anything/DA3-SMALL

Download a single file to a local folder (e.g. README.md into ./dir)

Download a single file to a local folder (e.g. README.md into ./dir)
modelscope download --model depth-anything/DA3-SMALL README.md --local_dir ./dir

See the docs for more CLI options

SDK Download

SDK Download
# 模型下载
from modelscope import snapshot_download
model_dir = snapshot_download('depth-anything/DA3-SMALL')

Git Download

Make sure git-lfs is installed first

Git Download
git lfs install
git clone https://www.modelscope.cn/depth-anything/DA3-SMALL.git

To skip LFS large-file downloads, use:

Skip LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://www.modelscope.cn/depth-anything/DA3-SMALL.git

ModelScope 模型页直接下载模型文件;无需将模型文件放在本站服务器。

Notebook Quickstart

Install the ModelScope library

Install the ModelScope library
pip install "modelscope[audio,cv,nlp,multi-modal,science]" -f https://modelscope.oss-cn-beijing.aliyuncs.com/releases/repo.html

Load the model and run inference

Load the model and run inference
from modelscope.pipelines import pipeline
from modelscope.utils.constant import Tasks

p = pipeline('text-generation', 'depth-anything/DA3-SMALL')

Full Documentation

来源: HuggingFace

---
license: apache-2.0
tags:

  • depth-estimation

  • computer-vision

  • monocular-depth

  • multi-view-geometry

  • pose-estimation

library_name: depth-anything-3
pipeline_tag: depth-estimation
---

Depth Anything 3: DA3-SMALL

<div align="center">

![Project Page](https://depth-anything-3.github.io)
![Paper](https://arxiv.org/abs/)
![Demo](https://huggingface.co/spaces/depth-anything/Depth-Anything-3) # noqa: E501
<!-- Benchmark badge removed as per request -->

</div>

Model Description

DA3 Small model for multi-view depth estimation and camera pose estimation. Efficient foundation model with unified depth-ray representation.

| Property | Value |
|----------|-------|
| Model Series | Any-view Model |
| Parameters | 0.08B |
| License | Apache 2.0 |

Capabilities

  • ✅ Relative Depth
  • ✅ Pose Estimation
  • ✅ Pose Conditioning

Quick Start

Installation

bash
git clone https://github.com/ByteDance-Seed/depth-anything-3
cd depth-anything-3
pip install -e .

Basic Example

python
import torch
from depth_anything_3.api import DepthAnything3

Load model from Hugging Face Hub

device = torch.device("cuda" if torch.cuda.is_available() else "cpu") model = DepthAnything3.from_pretrained("depth-anything/da3-small") model = model.to(device=device)

Run inference on images

images = ["image1.jpg", "image2.jpg"] # List of image paths, PIL Images, or numpy arrays prediction = model.inference( images, export_dir="output", export_format="glb" # Options: glb, npz, ply, mini_npz, gs_ply, gs_video )

Access results

print(prediction.depth.shape) # Depth maps: [N, H, W] float32 print(prediction.conf.shape) # Confidence maps: [N, H, W] float32 print(prediction.extrinsics.shape) # Camera poses (w2c): [N, 3, 4] float32 print(prediction.intrinsics.shape) # Camera intrinsics: [N, 3, 3] float32

Command Line Interface

bash
# Process images with auto mode
da3 auto path/to/images \
    --export-format glb \
    --export-dir output \
    --model-dir depth-anything/da3-small

Use backend for faster repeated inference

da3 backend --model-dir depth-anything/da3-small da3 auto path/to/images --export-format glb --use-backend

Model Details

  • Developed by: ByteDance Seed Team
  • Model Type: Vision Transformer for Visual Geometry
  • Architecture: Plain transformer with unified depth-ray representation
  • Training Data: Public academic datasets only

Key Insights

💎 A single plain transformer (e.g., vanilla DINO encoder) is sufficient as a backbone without architectural specialization. # noqa: E501

✨ A singular depth-ray representation obviates the need for complex multi-task learning.

Performance

🏆 Depth Anything 3 significantly outperforms:

  • Depth Anything 2 for monocular depth estimation

  • VGGT for multi-view depth estimation and pose estimation

For detailed benchmarks, please refer to our paper. # noqa: E501

Limitations

  • The model is trained on academic datasets and may have limitations on certain domain-specific images # noqa: E501
  • Performance may vary depending on image quality, lighting conditions, and scene complexity

Citation

If you find Depth Anything 3 useful in your research or projects, please cite:

bibtex
@article{depthanything3,
  title={Depth Anything 3: Recovering the visual space from any views},
  author={Haotong Lin and Sili Chen and Jun Hao Liew and Donny Y. Chen and Zhenyu Li and Guang Shi and Jiashi Feng and Bingyi Kang},  # noqa: E501
  journal={arXiv preprint arXiv:XXXX.XXXXX},
  year={2025}
}

Links

Authors

Haotong Lin · Sili Chen · Junhao Liew · Donny Y. Chen · Zhenyu Li · Guang Shi · Jiashi Feng · Bingyi Kang # noqa: E501

Join our Telegram