CLIP ViT-B/32 is a versatile vision-language model designed for zero-shot image and text understanding. Unlike tradition...
openai
image-text-retrieval
21.0M
0
CLIP ViT-L/14 is a powerful vision-language model designed for zero-shot image and text understanding. Unlike traditiona...
openai
image-text-retrieval
6.6M
0
CLIP ViT-L/14@336 is a high-resolution vision-language model designed for precise image-text alignment. By utilizing a V...
openai
image-text-retrieval
3.7M
0
CLIP ViT-B/32 (trained on LAION-2B) is a robust vision-language model designed for high-performance image-text alignment...
laion
image-text-retrieval
3.6M
0
Fashion CLIP is a domain-specific adaptation of the CLIP architecture, fine-tuned specifically for the fashion industry ...
patrickjohncyh
image-text-retrieval
2.6M
0
This model is a high-performance vision-language encoder based on the ConvNeXt architecture, trained on the massive LAIO...
laion
image-text-retrieval
2.0M
0
CLIP ViT-B/16 is a versatile vision-language model designed to map images and text into a shared embedding space. Unlike...
openai
image-text-retrieval
2.0M
0
Tiny CLIP Text 2 is a lightweight image-text retrieval model designed for developers who need efficient embedding genera...
peft-internal-testing
image-text-retrieval
1.9M
0
The CLIP ViT-L/14 (laion2B s32B b82K) is a high-performance vision-language model optimized for cross-modal retrieval an...
laion
image-text-retrieval
597.3K
0
ImageTextRetrieval is a specialized model designed for cross-modal alignment, allowing developers to perform efficient s...
ohgnues
image-text-retrieval
5
0