Server detail
The ai-vision-mcp server bridges the gap between LLMs and visual data by integrating Google Gemini and Vertex AI directly into the Model Context Protocol. Unlike basic image-to-text tools, this server provides structured visual analysis, making it particularly useful for developers handling UI/UX audits or automated visual regression testing. It allows an AI agent to 'see' and interpret interfaces across different operating systems, enabling tasks like identifying layout shifts, verifying element placement, or analyzing video frames for behavioral bugs. By exposing these multimodal capabilities as MCP tools, it removes the need to manually upload screenshots to a chat interface, allowing the AI to trigger visual inspections programmatically during the development workflow.
An MCP server that exposes tan-yong-sheng/ai-vision-mcp capabilities to MCP-compatible AI clients.
Collections featuring this MCP
Tool testing
tan-yong-sheng-ai-vision-mcp
Call the MCP capabilities provided by tan-yong-sheng/ai-vision-mcp and return a structured result.
inputPass arguments according to the server tool schema.Connection modes
{
"mcpServers": {
"tan-yong-sheng/ai-vision-mcp": {
"url": "Generated by the provider after deployment"
}
}
}The Remote endpoint is generated by the provider after deployment; this page does not fabricate an unusable endpoint.
No npm package is recorded. Open the source repository to complete command and args.If no npm package is registered, follow the installation method in the source repository.
How to use
- 01Step 1
Review server capabilities and permission scope.
- 02Step 2
Copy the install command or JSON configuration.
- 03Step 3
Run a small connection test in your client.
- 04Step 4
Adopt it long term only after reviewing access and maintenance.
Discussions
Use this space to keep checking source information, usage experience and maintenance status.
Open source page