MiniCPM V is a compact yet powerful vision-language model designed for efficient multimodal processing. Unlike monolithi...
openbmb
visual-question-answering
89.4K
52
The ViLT b32 finetuned VQA model is a streamlined vision-and-language transformer designed for Visual Question Answering...
dandelin
visual-question-answering
74.8K
0
MiniCPM-V 2 is a compact yet powerful vision-language model designed to bridge the gap between edge-device efficiency an...
openbmb
visual-question-answering
22.0K
48
LLaVA-Med v1.5 Mistral 7B HF is a domain-specific multimodal model designed for biomedical visual question answering. By...
chaoyinshe
visual-question-answering
3.4K
0
MiniCPM-V 2.6 (GGUF) is a compact yet powerful vision-language model optimized for edge deployment and local inference. ...
gaianet
visual-question-answering
2.5K
0
MiniCPM-V 2.6 (provided here in GGUF format) is a high-efficiency multimodal model designed for edge deployment and reso...
second-state
visual-question-answering
2.2K
0
uAI NEXUS MedVLM 1.0a 7B RL is a specialized vision-language model optimized for medical image analysis and clinical que...
UII-AI
visual-question-answering
2.0K
0
OpenCaption 4B VL SFT v1.0 is a vision-language model optimized for high-accuracy image captioning and visual question a...
mradermacher
visual-question-answering
1.5K
0
VideoLLaMA2.1 7B AV is a multimodal LLM optimized for high-fidelity video understanding and visual question answering. U...
DAMO-NLP-SG
visual-question-answering
1.1K
1
VideoLLaMA2.1 7B 16F Base is a specialized multimodal model designed for high-fidelity video understanding and visual qu...
DAMO-NLP-SG
visual-question-answering
990
0
DePlot is a specialized visual-question-answering (VQA) model designed to bridge the gap between plotted data and struct...
google
visual-question-answering
572
0
Pix2Struct ai2d-base is a vision-encoder-decoder model designed specifically for parsing visual information into structu...
google
visual-question-answering
422
0
Qwen3.5 2B MedVL is a compact vision-language model specifically tuned for the medical domain. Designed for efficiency, ...
OpenMed
visual-question-answering
388
3
BLIP-VQA Base is a vision-language model designed for visual question answering, bridging the gap between image encoding...
Salesforce
visual-question-answering
308
0
The BLIP VQA Capfilt Large is a specialized vision-language model designed for precise visual question answering. Unlike...
Salesforce
visual-question-answering
230
0
VideoScore2 is a specialized visual-question-answering (VQA) model designed to quantify video quality and semantic align...
TIGER-Lab
visual-question-answering
223
0