Azure AI 视觉图像分析 Python SDK

azure-ai-vision-imageanalysis-py
分类数据
作者Agentic Awesome Skills 社区
许可MIT
评分4.70/5
使用10.5K

Azure AI Vision 图像分析 Python SDK

Azure AI Vision 4.0 图像分析的客户端库,包含标题生成、标签、对象检测、OCR 等功能。

安装

bash
pip install azure-ai-vision-imageanalysis

环境变量

bash
VISION_ENDPOINT=https://<resource>.cognitiveservices.azure.com
VISION_KEY=<your-api-key>  # 如果使用 API 密钥

身份验证

API 密钥

python
import os
from azure.ai.vision.imageanalysis import ImageAnalysisClient
from azure.core.credentials import AzureKeyCredential

endpoint = os.environ["VISION_ENDPOINT"]
key = os.environ["VISION_KEY"]

client = ImageAnalysisClient(
endpoint=endpoint,
credential=AzureKeyCredential(key)
)

Entra ID (推荐)

python
from azure.ai.vision.imageanalysis import ImageAnalysisClient
from azure.identity import DefaultAzureCredential

client = ImageAnalysisClient(
endpoint=os.environ["VISION_ENDPOINT"],
credential=DefaultAzureCredential()
)

通过 URL 分析图像

python
from azure.ai.vision.imageanalysis.models import VisualFeatures

image_url = "https://example.com/image.jpg"

result = client.analyze_from_url(
image_url=image_url,
visual_features=[
VisualFeatures.CAPTION,
VisualFeatures.TAGS,
VisualFeatures.OBJECTS,
VisualFeatures.READ,
VisualFeatures.PEOPLE,
VisualFeatures.SMART_CROPS,
VisualFeatures.DENSE_CAPTIONS
],
gender_neutral_caption=True,
language="en"
)

通过文件分析图像

python
with open("image.jpg", "rb") as f:
    image_data = f.read()

result = client.analyze(
image_data=image_data,
visual_features=[VisualFeatures.CAPTION, VisualFeatures.TAGS]
)

图像标题 (Image Caption)

python
result = client.analyze_from_url(
    image_url=image_url,
    visual_features=[VisualFeatures.CAPTION],
    gender_neutral_caption=True
)

if result.caption:
print(f"Caption: {result.caption.text}")
print(f"Confidence: {result.caption.confidence:.2f}")

密集标题 (Dense Captions - 多个区域)

python
result = client.analyze_from_url(
    image_url=image_url,
    visual_features=[VisualFeatures.DENSE_CAPTIONS]
)

if result.dense_captions:
for caption in result.dense_captions.list:
print(f"Caption: {caption.text}")
print(f" Confidence: {caption.confidence:.2f}")
print(f" Bounding box: {caption.bounding_box}")

标签 (Tags)

python
result = client.analyze_from_url(
    image_url=image_url,
    visual_features=[VisualFeatures.TAGS]
)

if result.tags:
for tag in result.tags.list:
print(f"Tag: {tag.name} (confidence: {tag.confidence:.2f})")

对象检测 (Object Detection)

python
result = client.analyze_from_url(
    image_url=image_url,
    visual_features=[VisualFeatures.OBJECTS]
)

if result.objects:
for obj in result.objects.list:
print(f"Object: {obj.tags[0].name}")
print(f" Confidence: {obj.tags[0].confidence:.2f}")
box = obj.bounding_box
print(f" Bounding box: x={box.x}, y={box.y}, w={box.width}, h={box.height}")

OCR (文本提取)

python
result = client.analyze_from_url(
    image_url=image_url,
    visual_features=[VisualFeature
s.READ] )

if result.read:
for block in result.read.blocks:
for line in block.lines:
print(f"Line: {line.text}")
print(f" Bounding polygon: {line.bounding_polygon}")

# Word-level details
for word in line.words:
print(f" Word: {word.text} (confidence: {word.confidence:.2f})")

code
## 人员检测
python
result = client.analyze_from_url(
image_url=image_url,
visual_features=[VisualFeatures.PEOPLE]
)

if result.people:
for person in result.people.list:
print(f"Person detected:")
print(f" Confidence: {person.confidence:.2f}")
box = person.bounding_box
print(f" Bounding box: x={box.x}, y={box.y}, w={box.width}, h={box.height}")

code
## 智能裁剪
python
result = client.analyze_from_url(
image_url=image_url,
visual_features=[VisualFeatures.SMART_CROPS],
smart_crops_aspect_ratios=[0.9, 1.33, 1.78] # 纵向, 4:3, 16:9
)

if result.smart_crops:
for crop in result.smart_crops.list:
print(f"Aspect ratio: {crop.aspect_ratio}")
box = crop.bounding_box
print(f" Crop region: x={box.x}, y={box.y}, w={box.width}, h={box.height}")

code
## 异步客户端
python
from azure.ai.vision.imageanalysis.aio import ImageAnalysisClient
from azure.identity.aio import DefaultAzureCredential

async def analyze_image():
async with ImageAnalysisClient(
endpoint=endpoint,
credential=DefaultAzureCredential()
) as client:
result = await client.analyze_from_url(
image_url=image_url,
visual_features=[VisualFeatures.CAPTION]
)
print(result.caption.text)

code
## 视觉特征

| 特征 | 描述 |
|---------|-------------|
| CAPTION | 描述图像的单句标题 |
| DENSE_CAPTIONS | 针对多个区域的详细标题 |
| TAGS | 内容标签(对象、场景、动作) |
| OBJECTS | 带有边界框的对象检测 |
| READ | OCR 文本提取 |
| PEOPLE | 带有边界框的人员检测 |
| SMART_CROPS | 为缩略图建议的裁剪区域 |

错误处理

python from azure.core.exceptions import HttpResponseError

try:
result = client.analyze_from_url(
image_url=image_url,
visual_features=[VisualFeatures.CAPTION]
)
except HttpResponseError as e:
print(f"Status code: {e.status_code}")
print(f"Reason: {e.reason}")
print(f"Message: {e.error.message}")
```

图像要求

  • 格式:JPEG, PNG, GIF, BMP, WEBP, ICO, TIFF, MPO
  • 最大尺寸:20 MB
  • 分辨率:50x50 至 16000x16000 像素

最佳实践

1. 仅选择必要的特征以优化延迟和成本
2. 在高吞吐量场景下使用异步客户端
3. 处理 HttpResponseError 以应对无效图像或认证问题
4. 启用 gender_neutral_caption 以获得更具包容性的描述
5. 指定语言以获取本地化标题
6. 使用与缩略图需求匹配的 smart_crops_aspect_ratios
7. 多次分析同一图像时缓存结果

适用场景

此技能适用于执行概览中描述的工作流或操作。

局限性

  • 仅在任务明确符合上述范围时使用此技能。
  • 不要将输出视为环境特定验证、测试或专家评审的替代方案。
  • 如果需要输入、权限、安全边界或成功标准不明确,请停止并请求澄清。
ria 丢失了。