depth anything base hf

提供商LiheYoung
分类depth-estimation
许可证apache-2.0
下载量48.2K
星标0

简介

Depth Anything Base 是一个强大的单目深度估计模型,旨在将 2D 图像转化为高精度的深度图。与传统深度模型依赖特定数据集不同,它通过大规模无监督学习获得了极强的泛化能力,能够处理各种复杂场景而无需针对性微调。对于开发者而言,该模型上手门槛低,非常适合集成到 3D 视觉重建、背景虚化、自动驾驶感知或 AR 增强现实等项目中。它可以作为视觉管线的前置模块,为后续的 3D 空间分析提供可靠的深度基准。

核心亮点

  • 极强的泛化能力,无需微调即可适配多种场景
  • 高效生成高分辨率深度图,边缘细节还原精准
  • 单目输入即可实现空间深度感知,部署门槛低
  • Apache-2.0 协议,适合商业化集成与二次开发

使用方法

安装依赖
# 安装 Hugging Face transformers
pip install transformers torch
SDK 使用
# 使用 transformers 加载模型
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("LiheYoung/depth-anything-base-hf")
tokenizer = AutoTokenizer.from_pretrained("LiheYoung/depth-anything-base-hf")

Hugging Face 下载

我们推荐使用命令行或者 Hugging Face Hub SDK 来进行模型的下载。

操作指引:在下载前,请先通过如下命令安装 huggingface_hub:

操作指引
pip install -U huggingface_hub

命令行下载

下载完整模型库

下载完整模型库
huggingface-cli download LiheYoung/depth-anything-base-hf

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)
huggingface-cli download LiheYoung/depth-anything-base-hf config.json --local-dir ./dir

更多命令行下载选项,可参见官方文档

SDK 下载

SDK 下载
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('LiheYoung/depth-anything-base-hf')

Git 下载

请确保 lfs 已经被正确安装

Git 下载
git lfs install
git clone https://huggingface.co/LiheYoung/depth-anything-base-hf

如果您希望跳过 lfs 大文件下载,可以使用如下命令

跳过 LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/LiheYoung/depth-anything-base-hf

模型文件托管在 Hugging Face Hub,使用 HF CLI / SDK / Git 直接下载,不经过本站。

PyTorch / Transformers 使用

安装 Transformers

安装 Transformers
pip install -U transformers torch

模型加载和推理

模型加载和推理
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained('LiheYoung/depth-anything-base-hf')
tokenizer = AutoTokenizer.from_pretrained('LiheYoung/depth-anything-base-hf')

完整文档

来源: HuggingFace

---
license: apache-2.0
tags:

  • vision

pipeline_tag: depth-estimation
widget:
  • inference: false

---

Depth Anything (base-sized model, Transformers version)

Depth Anything model. It was introduced in the paper Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data by Lihe Yang et al. and first released in this repository.

Online demo is also provided.

Disclaimer: The team releasing Depth Anything did not write a model card for this model so this model card has been written by the Hugging Face team.

Model description

Depth Anything leverages the DPT architecture with a DINOv2 backbone.

The model is trained on ~62 million images, obtaining state-of-the-art results for both relative and absolute depth estimation.

<img src="https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/depth_anything_overview.jpg"
alt="drawing" width="600"/>

<small> Depth Anything overview. Taken from the <a href="https://arxiv.org/abs/2401.10891">original paper</a>.</small>

Intended uses & limitations

You can use the raw model for tasks like zero-shot depth estimation. See the model hub to look for
other versions on a task that interests you.

How to use

Here is how to use this model to perform zero-shot depth estimation:

python
from transformers import pipeline
from PIL import Image
import requests

load pipe

pipe = pipeline(task="depth-estimation", model="LiheYoung/depth-anything-base-hf")

load image

url = 'http://images.cocodataset.org/val2017/000000039769.jpg' image = Image.open(requests.get(url, stream=True).raw)

inference

depth = pipe(image)["depth"]

Alternatively, one can use the classes themselves:

python
from transformers import AutoImageProcessor, AutoModelForDepthEstimation
import torch
import numpy as np
from PIL import Image
import requests

url = "http://images.cocodataset.org/val2017/000000039769.jpg"
image = Image.open(requests.get(url, stream=True).raw)

image_processor = AutoImageProcessor.from_pretrained("LiheYoung/depth-anything-base-hf")
model = AutoModelForDepthEstimation.from_pretrained("LiheYoung/depth-anything-base-hf")

prepare image for the model

inputs = image_processor(images=image, return_tensors="pt")

with torch.no_grad():
outputs = model(**inputs)
predicted_depth = outputs.predicted_depth

interpolate to original size

prediction = torch.nn.functional.interpolate( predicted_depth.unsqueeze(1), size=image.size[::-1], mode="bicubic", align_corners=False, )
For more code examples, we refer to the documentation.

BibTeX entry and citation info

bibtex
@misc{yang2024depth,
      title={Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data}, 
      author={Lihe Yang and Bingyi Kang and Zilong Huang and Xiaogang Xu and Jiashi Feng and Hengshuang Zhao},
      year={2024},
      eprint={2401.10891},
      archivePrefix={arXiv},
      primaryClass={cs.CV}
}