Skip to content
  • Models
  • Rankings
  • Ori
Sign Up
Sign Up
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Tools
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status
  • AI Site Map

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Collections/Vision Models

AI Models with Vision: Multimodal LLMs for Image Understanding

Model rankings updated September 2026 based on real usage data.

Models are ranked by total prompt and completion tokens processed through the OpenRouter API over the trailing 7 days. Showing the top 10 models. Rankings measure usage on OpenRouter and reflect adoption, not model quality or benchmark performance.

Vision models are multimodal LLMs that analyze images, read documents, interpret charts, and answer questions about visual content alongside text. This collection ranks vision-capable models by their usage on OpenRouter over the past week. The current top models are DeepSeek V4.1 Flash, GLM 5.3 Flash, and GPT-5.6 Luna. Access models from Anthropic, Google, OpenAI, and other providers through a single API, and compare context length, pricing, and capabilities.

Browse All ModelsCompare Models

Top Vision Models on OpenRouter

Favicon for deepseek

DeepSeek: DeepSeek V4.1 Flash

19.5T tokens
Academia (#2)
Finance (#1)
Health (#2)
Legal (#5)
Marketing (#1)

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on output from a 552B-parameter backbone, an asymmetric split that keeps per-token compute low relative to the model's total size. Image understanding is native to the architecture, with visual and text embeddings trained jointly from the start of pre-training rather than added afterward as in the earlier experimental V4 Flash Vision Exp.

It is suited for coding, terminal, and computer-use agents, along with long-horizon tasks that must run to completion across many steps and long-context analysis. Compressed KV caching cuts cache memory to roughly a quarter of the previous Flash generation, significantly reducing costs on agentic workloads. DeepSeek positions it as the cost-efficient tier of the V4.1 family and reports that it exceeds V4 Pro on performance, speed, and task completion time.

by deepseek1.05M context$0.04/M input tokens$0.49/M output tokens
Favicon for z-ai

Z.ai: GLM 5.3 Flash

19.5T tokens
Academia (#3)
Finance (#3)
Health (#1)
Legal (#3)
Marketing (#3)

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

by z-ai1.31M context$0.045/M input tokens$0.14/M output tokens50% off
Favicon for openai

OpenAI: GPT-5.6 Luna

8.9T tokens
Academia (#5)
Finance (#5)
Health (#3)
Legal (#2)
Marketing (#4)

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for its price tier.

by openai1.05M context$0.20/M input tokens$1.20/M output tokens
Favicon for xiaomi

Xiaomi: MiMo-V2.5

4.97T tokens
Academia (#10)
Finance (#10)
Health (#8)
Marketing (#29)
SEO (#18)

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding tasks. Its 1M context window supports complete documents, extended conversations, and complex task contexts in a single pass, making it ideal for integration with agent frameworks where strong reasoning, rich perception, and cost efficiency all matter.

by xiaomi1.05M context$0.119/M input tokens$0.238/M output tokens15% off
Favicon for stealth

Space Bunny Alpha

3.68T tokens
Programming (#28)
Science (#31)
Translation (#36)

Space Bunny Alpha is an anonymous large model with blazing-fast inference, strong coding capabilities and native multimodal input support. It delivers adjustable reasoning effort, and a 1M-token context window.

Space Bunny Alpha is a stealth model. It is developed and operated by a third-party provider who has chosen to remain anonymous during this preview. OpenRouter routes requests to it and is not its developer, owner, or provider. Prompts and completions may be retained by the provider but are not used for training; all other use is governed by the Stealth Model Terms.

by stealth1M context$0/M input tokens$0/M output tokens
Favicon for google

Google: Gemini 3.8 Flash

2.19T tokens
Academia (#11)
Finance (#14)
Health (#9)
Legal (#7)
Marketing (#9)

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

by google1.05M context$0.75/M input tokens$3.75/M output tokens50% off
Favicon for xiaomi

Xiaomi: MiMo-V2.6-Flash

2.07T tokens
Finance (#47)
Programming (#6)
Roleplay (#17)
Science (#12)
Technology (#27)

MiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parameters and 15B activated per token, it employs a hybrid attention mechanism for greater computational efficiency. The model features a 1M-token context window and native multimodal capabilities. Optimized for agentic workflows, it delivers strong performance across coding, visual, general, and research scenarios, excelling at complex, long-horizon tasks with robust generalization across a diverse range of agent harnesses.

by xiaomi1.05M context$0.14/M input tokens$0.28/M output tokens
Favicon for meta

Meta: Muse Spark 1.3 Contributor

2.04T tokens
Academia (#32)
Finance (#26)
Health (#42)
Legal (#36)
Marketing (#35)

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information across extended tasks, work through conflicting inputs, and request clarification or confirmation when needed. Prompts and outputs may be used to improve Meta’s products.

by meta1.05M context$0.10/M input tokens$0.20/M output tokens
Favicon for openai

OpenAI: GPT-5.6 Sol

1.78T tokens
Academia (#14)
Finance (#28)
Health (#20)
Legal (#20)
Marketing (#11)

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks and long-horizon problem solving.

by openai1.05M context$2/M input tokens$10/M output tokens50% off
Favicon for anthropic

Anthropic: Claude Sonnet 5

1.5T tokens
Academia (#19)
Finance (#6)
Health (#27)
Legal (#17)
Marketing (#25)

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max, and x-high), a 1M-token context window, and text, image, and file inputs. Sonnet 5 uses an updated tokenizer and includes real-time cyber safeguards that block certain high-risk dual-use activities.

by anthropic1M context$2/M input tokens$10/M output tokens

Explore more collections

  • Free Models
  • Discounted Models
  • Coding
  • Roleplay
  • Tool Calling
  • OpenClaw
  • Image Models
  • Video Models
  • Audio Models
  • Text-to-Speech
  • Speech-to-Text
  • Embedding Models
  • Rerank Models
  • Distillable Models
  • All collections