Categories

AI Models Marketplace

Explore the latest and most popular AI models — LLM, vision, audio and more — to pick and integrate the right one.

MiniCPM V

Apache-2.0
MiniCPM V is a compact yet powerful vision-language model designed for efficient multimodal processing. Unlike monolithi...
openbmb visual-question-answering
89.4K 52

vilt b32 finetuned vqa

apache-2.0
The ViLT b32 finetuned VQA model is a streamlined vision-and-language transformer designed for Visual Question Answering...
dandelin visual-question-answering
74.8K 0

MiniCPM V 2

Apache-2.0
MiniCPM-V 2 is a compact yet powerful vision-language model designed to bridge the gap between edge-device efficiency an...
openbmb visual-question-answering
22.0K 48

llava med v1.5 mistral 7b hf

apache-2.0
LLaVA-Med v1.5 Mistral 7B HF is a domain-specific multimodal model designed for biomedical visual question answering. By...
chaoyinshe visual-question-answering
3.4K 0

MiniCPM V 4 5 GGUF

Apache-2.0
MiniCPM-V 2.6 (GGUF) is a compact yet powerful vision-language model optimized for edge deployment and local inference. ...
gaianet visual-question-answering
2.5K 0

MiniCPM V 4 5 GGUF

Apache-2.0
MiniCPM-V 2.6 (provided here in GGUF format) is a high-efficiency multimodal model designed for edge deployment and reso...
second-state visual-question-answering
2.2K 0

uAI NEXUS MedVLM 1.0a 7B RL

apache-2.0
uAI NEXUS MedVLM 1.0a 7B RL is a specialized vision-language model optimized for medical image analysis and clinical que...
UII-AI visual-question-answering
2.0K 0

OpenCaption 4B VL SFT v1.0 i1 GGUF

apache-2.0
OpenCaption 4B VL SFT v1.0 is a vision-language model optimized for high-accuracy image captioning and visual question a...
mradermacher visual-question-answering
1.5K 0

VideoLLaMA2.1 7B AV

apache-2.0
VideoLLaMA2.1 7B AV is a multimodal LLM optimized for high-fidelity video understanding and visual question answering. U...
DAMO-NLP-SG visual-question-answering
1.1K 1

VideoLLaMA2.1 7B 16F Base

apache-2.0
VideoLLaMA2.1 7B 16F Base is a specialized multimodal model designed for high-fidelity video understanding and visual qu...
DAMO-NLP-SG visual-question-answering
990 0

deplot

apache-2.0
DePlot is a specialized visual-question-answering (VQA) model designed to bridge the gap between plotted data and struct...
google visual-question-answering
572 0

pix2struct ai2d base

apache-2.0
Pix2Struct ai2d-base is a vision-encoder-decoder model designed specifically for parsing visual information into structu...
google visual-question-answering
422 0

Qwen3.5 2B MedVL

apache-2.0
Qwen3.5 2B MedVL is a compact vision-language model specifically tuned for the medical domain. Designed for efficiency, ...
OpenMed visual-question-answering
388 3

blip vqa base

bsd-3-clause
BLIP-VQA Base is a vision-language model designed for visual question answering, bridging the gap between image encoding...
Salesforce visual-question-answering
308 0

blip vqa capfilt large

bsd-3-clause
The BLIP VQA Capfilt Large is a specialized vision-language model designed for precise visual question answering. Unlike...
Salesforce visual-question-answering
230 0

VideoScore2

apache-2.0
VideoScore2 is a specialized visual-question-answering (VQA) model designed to quantify video quality and semantic align...
TIGER-Lab visual-question-answering
223 0
Join our Telegram