# OrcaRouter — full model catalog Generated 2026-10-04. One section per model. ## DeepSeek: DeepSeek V4.1 Flash - [DeepSeek: DeepSeek V4.1 Flash](https://www.orcarouter.ai/models/deepseek/deepseek-v4.1-flash) — DeepSeek: DeepSeek V4.1 Flash by deepseek: $0.15/M input, $0.60/M output, 1M context, p50 1751ms, available via OrcaRouter API. DeepSeek-V4.1-Flash is the smallest model in DeepSeek's new architecture family, released September 10, 2026, with native multimodal visual understanding — the production successor to both V4 Flash and the V4 Flash Vision experiment. The new architecture targets a higher capability ceiling, faster inference and higher throughput, and DeepSeek reports V4.1 Flash comprehensively surpasses V4 Pro on performance, cost, speed and total task time. It accepts text and images with text output, serves a 1M-token context window with up to 384K output tokens, and supports thinking (default, with selectable effort low / high / max) and non-thinking modes across the Chat Completions, Responses and Anthropic-compatible APIs, along with JSON output, tool calls and chat prefix completion; FIM works in non-thinking mode only. Model weights are open on Hugging Face (deepseek-ai/DeepSeek-V4.1-Flash). Official benchmarks at release: 90.6 on Terminal-Bench 2.1, 74.2 on DeepSWE v1.1, 65.4 on NL2Repo-Bench, 90.9 on GPQA Diamond, 63.9 on HLE with tools, and strong native-vision agent results (89.6 BabyVision, 78.9 Chartography with tools). Pricing was cut alongside the release: $0.15/M input (cache miss), $0.60/M output and $0.003/M on cache hits at off-peak rates, doubling during weekday peak hours (01:00-04:00 and 06:00-10:00 UTC); all other hours including weekends are off-peak. Canonical URL: https://www.orcarouter.ai/models/deepseek/deepseek-v4.1-flash ## grok/grok-4.3 - [grok/grok-4.3](https://www.orcarouter.ai/models/grok/grok-4.3) — grok/grok-4.3 by grok: $1.25/M input, $2.50/M output, 1M context, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/grok/grok-4.3 ## grok/grok-imagine-image - [grok/grok-imagine-image](https://www.orcarouter.ai/models/grok/grok-imagine-image) — grok/grok-imagine-image by grok: available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/grok/grok-imagine-image ## OpenAI: GPT-4.1 Nano - [OpenAI: GPT-4.1 Nano](https://www.orcarouter.ai/models/openai/gpt-4.1-nano) — OpenAI: GPT-4.1 Nano by openai: $0.10/M input, $0.40/M output, 1M context, available via OrcaRouter API. For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-4.1-nano ## openai/gpt-5-mini-2025-08-07 - [openai/gpt-5-mini-2025-08-07](https://www.orcarouter.ai/models/openai/gpt-5-mini-2025-08-07) — openai/gpt-5-mini-2025-08-07 by openai: $0.25/M input, $2.00/M output, 400K context, p50 919ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5-mini-2025-08-07 ## OpenAI: GPT-5.4 Pro - [OpenAI: GPT-5.4 Pro](https://www.orcarouter.ai/models/openai/gpt-5.4-pro) — OpenAI: GPT-5.4 Pro by openai: $30.00/M input, $180.00/M output, 1M context, available via OrcaRouter API. GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It features a 1M+ token context window (922K input, 128K... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.4-pro ## openai/gpt-image-2 - [openai/gpt-image-2](https://www.orcarouter.ai/models/openai/gpt-image-2) — openai/gpt-image-2 by openai: $8.00/M input, $30.00/M output, available via OrcaRouter API. OpenAI's gpt-image-2 is the next generation of the gpt-image series — a token-billed image model accessed through the standard OpenAI Images API. It's a drop-in upgrade from gpt-image-1: same SDK, same call shape, same parameters. Point your existing OpenAI client at OrcaRouter's base URL and set the model to `openai/gpt-image-2`. Pricing is $5.00 input / $30.00 output per 1M tokens, billed at cost — zero markup, zero migration. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-image-2 ## Qwen: Qwen3.8 Flash - [Qwen: Qwen3.8 Flash](https://www.orcarouter.ai/models/qwen/qwen3.8-flash) — Qwen: Qwen3.8 Flash by qwen: $0.15/M input, $0.47/M output, 1M context, p50 4400ms, available via OrcaRouter API. Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3.8-flash ## Anthropic: Claude Opus 4.5 - [Anthropic: Claude Opus 4.5](https://www.orcarouter.ai/models/anthropic/claude-opus-4.5) — Anthropic: Claude Opus 4.5 by anthropic: $5.00/M input, $25.00/M output, 200K context, p50 1000ms, available via OrcaRouter API. Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and... Canonical URL: https://www.orcarouter.ai/models/anthropic/claude-opus-4.5 ## google/gemini-pro-latest - [google/gemini-pro-latest](https://www.orcarouter.ai/models/google/gemini-pro-latest) — google/gemini-pro-latest by google: $4.00/M input, $18.00/M output, p50 5561ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/google/gemini-pro-latest ## openai/gpt-5.4-nano-2026-03-17 - [openai/gpt-5.4-nano-2026-03-17](https://www.orcarouter.ai/models/openai/gpt-5.4-nano-2026-03-17) — openai/gpt-5.4-nano-2026-03-17 by openai: $0.20/M input, $1.25/M output, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.4-nano-2026-03-17 ## Tencent: Hy3 - [Tencent: Hy3](https://www.orcarouter.ai/models/tencent/hy3) — Tencent: Hy3 by tencent: $0.18/M input, $0.59/M output, 262K context, p50 5003ms, available via OrcaRouter API. Hy3 is Tencent Hunyuan's production-grade Mixture-of-Experts model — 295B total parameters with only 21B active per pass (192 experts, top-8 routing), the upgraded release built on the Hy3-preview line. It expands RL-training scale and post-training data quality for further gains in reasoning, long-context, and agentic tasks, reaching results comparable to flagship models several times its parameter size. It serves a 256K-token context window (text in, text out) with configurable reasoning effort, and is built for real-world coding, tool use, and multi-step agent workflows at a strong quality-to-cost ratio. Canonical URL: https://www.orcarouter.ai/models/tencent/hy3 ## Tencent: Hy4 preview - [Tencent: Hy4 preview](https://www.orcarouter.ai/models/tencent/hy4-preview) — Tencent: Hy4 preview by tencent: $0.83/M input, $2.50/M output, 1M context, p50 5487ms, available via OrcaRouter API. Hy4 preview is Tencent latest Mixture-of-Experts flagship: 770B total parameters with 49B activated per token, 78 layers, 256 routed experts plus 1 shared expert, and a 1M-token context window. It is a text-only model built for coding agents, complex tool-use workflows, and productivity work, with a native MTP layer for speculative decoding. Reasoning is configurable at three levels (high, low, none) and defaults to high, so it can run as a deep chain-of-thought model for math, coding and analysis, or answer directly when latency matters. It supports native tool calling, structured outputs, and the full sampling parameter set; Tencent recommends temperature 0.9 with top_p 1.0. Tencent positions it for software engineering, office and data analysis, game prototyping, and scientific research, and reports an internal blind evaluation (163 experts, 203 engineering tasks) in which Hy4 preview averaged 2.99 against GLM 5.3 at 2.92 and Kimi K3 at 2.94. As a preview release Tencent notes known rough edges, including spending longer than necessary on reasoning and over-verifying its own work. Canonical URL: https://www.orcarouter.ai/models/tencent/hy4-preview ## Tencent: Hy4 preview (Free) - [Tencent: Hy4 preview (Free)](https://www.orcarouter.ai/models/tencent/hy4-preview-free) — Tencent: Hy4 preview (Free) by tencent: 1M context, p50 29199ms, available via OrcaRouter API. Hy4 preview is Tencent latest Mixture-of-Experts flagship: 770B total parameters with 49B activated per token, 78 layers, 256 routed experts plus 1 shared expert, and a 1M-token context window. It is a text-only model built for coding agents, complex tool-use workflows, and productivity work, with a native MTP layer for speculative decoding. Reasoning is configurable at three levels (high, low, none) and defaults to high, so it can run as a deep chain-of-thought model for math, coding and analysis, or answer directly when latency matters. It supports native tool calling, structured outputs, and the full sampling parameter set; Tencent recommends temperature 0.9 with top_p 1.0. Tencent positions it for software engineering, office and data analysis, game prototyping, and scientific research, and reports an internal blind evaluation (163 experts, 203 engineering tasks) in which Hy4 preview averaged 2.99 against GLM 5.3 at 2.92 and Kimi K3 at 2.94. As a preview release Tencent notes known rough edges, including spending longer than necessary on reasoning and over-verifying its own work. Canonical URL: https://www.orcarouter.ai/models/tencent/hy4-preview-free ## TypeSafe: Jev 1.13 - [TypeSafe: Jev 1.13](https://www.orcarouter.ai/models/typesafe/jev-1.13) — TypeSafe: Jev 1.13 by typesafe: $0.04/M input, 66K context, p50 113ms, available via OrcaRouter API. TypeSafe's structured decision and evaluation model. Give it a state and a set of named questions (noul / choice / score) and it returns a structured answer for each. Served via POST /v1/systemone; non-streaming; up to ~64K input tokens; text in, structured JSON out. Canonical URL: https://www.orcarouter.ai/models/typesafe/jev-1.13 ## Anthropic: Claude Sonnet 4.6 - [Anthropic: Claude Sonnet 4.6](https://www.orcarouter.ai/models/anthropic/claude-sonnet-4.6) — Anthropic: Claude Sonnet 4.6 by anthropic: $3.00/M input, $15.00/M output, 1M context, p50 1000ms, available via OrcaRouter API. Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with... Canonical URL: https://www.orcarouter.ai/models/anthropic/claude-sonnet-4.6 ## Google: Gemini 2.5 Flash - [Google: Gemini 2.5 Flash](https://www.orcarouter.ai/models/google/gemini-2.5-flash) — Google: Gemini 2.5 Flash by google: $0.30/M input, $2.50/M output, 1M context, p50 7368ms, available via OrcaRouter API. Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater... Canonical URL: https://www.orcarouter.ai/models/google/gemini-2.5-flash ## OpenAI: GPT-5.1-Codex-Mini - [OpenAI: GPT-5.1-Codex-Mini](https://www.orcarouter.ai/models/openai/gpt-5.1-codex-mini) — OpenAI: GPT-5.1-Codex-Mini by openai: $0.25/M input, $2.00/M output, 400K context, available via OrcaRouter API. GPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.1-codex-mini ## Anthropic: Claude Opus 4.7 - [Anthropic: Claude Opus 4.7](https://www.orcarouter.ai/models/anthropic/claude-opus-4.7) — Anthropic: Claude Opus 4.7 by anthropic: $5.00/M input, $25.00/M output, 1M context, p50 16779ms, available via OrcaRouter API. Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on... Canonical URL: https://www.orcarouter.ai/models/anthropic/claude-opus-4.7 ## Google: Nano Banana (Gemini 2.5 Flash Image) - [Google: Nano Banana (Gemini 2.5 Flash Image)](https://www.orcarouter.ai/models/google/gemini-2.5-flash-image) — Google: Nano Banana (Gemini 2.5 Flash Image) by google: $0.30/M input, $30.00/M output, 33K context, p50 6010ms, available via OrcaRouter API. Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual understanding. It is capable of image generation,... Canonical URL: https://www.orcarouter.ai/models/google/gemini-2.5-flash-image ## openai/gpt-3.5-turbo-1106 - [openai/gpt-3.5-turbo-1106](https://www.orcarouter.ai/models/openai/gpt-3.5-turbo-1106) — openai/gpt-3.5-turbo-1106 by openai: $1.00/M input, $2.00/M output, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-3.5-turbo-1106 ## openai/gpt-4o-mini-tts-2025-12-15 - [openai/gpt-4o-mini-tts-2025-12-15](https://www.orcarouter.ai/models/openai/gpt-4o-mini-tts-2025-12-15) — openai/gpt-4o-mini-tts-2025-12-15 by openai: $0.60/M input, $12.00/M output, p50 646ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-4o-mini-tts-2025-12-15 ## openai/gpt-5.1-chat-latest - [openai/gpt-5.1-chat-latest](https://www.orcarouter.ai/models/openai/gpt-5.1-chat-latest) — openai/gpt-5.1-chat-latest by openai: $1.25/M input, $10.00/M output, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.1-chat-latest ## openai/gpt-5.4-mini-2026-03-17 - [openai/gpt-5.4-mini-2026-03-17](https://www.orcarouter.ai/models/openai/gpt-5.4-mini-2026-03-17) — openai/gpt-5.4-mini-2026-03-17 by openai: $0.75/M input, $4.50/M output, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.4-mini-2026-03-17 ## openai/text-embedding-3-large - [openai/text-embedding-3-large](https://www.orcarouter.ai/models/openai/text-embedding-3-large) — openai/text-embedding-3-large by openai: $0.13/M input, $0.13/M output, p50 137ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/text-embedding-3-large ## openai/tts-1 - [openai/tts-1](https://www.orcarouter.ai/models/openai/tts-1) — openai/tts-1 by openai: $15.00/M characters input, p50 653ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/tts-1 ## Google: Gemini 3.5 Flash-Lite - [Google: Gemini 3.5 Flash-Lite](https://www.orcarouter.ai/models/google/gemini-3.5-flash-lite) — Google: Gemini 3.5 Flash-Lite by google: $0.30/M input, $2.50/M output, 1M context, p50 1027ms, available via OrcaRouter API. Gemini 3.5 Flash-Lite is Google's most cost-efficient Gemini model, purpose-built for very high-volume, latency-sensitive workloads where cost per call dominates. Despite its price point it is natively multimodal — accepting text, images, video, audio, and files with text output — and carries the same 1M-token context window and up to 64K output tokens as the larger Flash tier, so long documents and mixed-media inputs remain in reach. At roughly a fifth of the price of 3.6 Flash, it is the right tool for classification, extraction, routing, summarization, lightweight agents, synthetic-data generation, and any pipeline run millions of times. It still supports configurable reasoning effort, native tool calling, and structured outputs, and punches above its weight on agentic and long-context benchmarks for a lite model — for example 74.0% on OSWorld-Verified and 72.2% on 128k multi-needle retrieval, comfortably ahead of the prior 3.1 Flash-Lite. It speaks both the OpenAI chat-completions format and Gemini's native generateContent API, so you can mix it with heavier models behind a single integration — routing easy, high-volume traffic to Flash-Lite and escalating hard tasks to 3.6 Flash or a Pro model. Canonical URL: https://www.orcarouter.ai/models/google/gemini-3.5-flash-lite ## kling/kling-v2-6 - [kling/kling-v2-6](https://www.orcarouter.ai/models/kling/kling-v2-6) — kling/kling-v2-6 by kling: p50 1000ms, available via OrcaRouter API. Kling 2.6 — text-to-video and image-to-video with motion control + audio control (pro mode), variable duration, 1080p, 24fps. Canonical URL: https://www.orcarouter.ai/models/kling/kling-v2-6 ## Gemma 4 26B A4B - [Gemma 4 26B A4B](https://www.orcarouter.ai/models/obsidian/gemma-4-26b-a4b) — Gemma 4 26B A4B by obsidian: $0.25/M input, $2.90/M output, 262K context, p50 8689ms, available via OrcaRouter API. Gemma 4 26B A4B is a mixture-of-experts build of Google's Gemma 4 family (26B total parameters, about 4B active per token) with a 262K-token context window and text + image input. Preserves the base model's reasoning, multilingual performance, and multimodal understanding. The Balanced variant is recommended for most workloads: consistent sampling, stable long-context behaviour, and reduced topic drift across extended conversations. The model may occasionally frame its reasoning briefly before giving the full response. Well suited for creative writing, roleplay, multilingual assistants, long-context reasoning, and general-purpose AI applications where quality, coherence, and reliability are the primary priorities. Canonical URL: https://www.orcarouter.ai/models/obsidian/gemma-4-26b-a4b ## OpenAI: GPT-5 Pro - [OpenAI: GPT-5 Pro](https://www.orcarouter.ai/models/openai/gpt-5-pro) — OpenAI: GPT-5 Pro by openai: $15.00/M input, $120.00/M output, 400K context, p50 299000ms, available via OrcaRouter API. GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5-pro ## qwen/qwen3-max-preview - [qwen/qwen3-max-preview](https://www.orcarouter.ai/models/qwen/qwen3-max-preview) — qwen/qwen3-max-preview by qwen: $0.86/M input, $3.44/M output, 262K context, p50 858ms, available via OrcaRouter API. Qwen3 Max preview — proprietary chat preview, 256k context, thinking mode + function calling. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3-max-preview ## Z.ai: GLM 5.2 - [Z.ai: GLM 5.2](https://www.orcarouter.ai/models/z-ai/glm-5.2) — Z.ai: GLM 5.2 by z-ai: $1.40/M input, $4.40/M output, 1M context, p50 7339ms, available via OrcaRouter API. GLM-5.2 is Z.ai (Zhipu AI)'s flagship model for the era of long-horizon tasks. It pairs a truly usable 1M-token context window with up to 128K output tokens, letting it hold project-level engineering context, execute long-running tasks more reliably, follow engineering standards more consistently, and carry a task from requirements through to multi-platform deployment in a single run. It is a text-in / text-out model with hybrid reasoning controlled by reasoning_effort (high / max; deep reasoning by default) and native tool calling. Built coding-first as the latest in the GLM-5 line, GLM-5.2 launched on the GLM Coding Plan with standalone API access and MIT-licensed open weights following shortly after. It targets repo-scale agentic coding, autonomous multi-step engineering workflows, and complex long-horizon delivery. Canonical URL: https://www.orcarouter.ai/models/z-ai/glm-5.2 ## google/gemini-3-flash-preview-ks - [google/gemini-3-flash-preview-ks](https://www.orcarouter.ai/models/google/gemini-3-flash-preview-ks) — google/gemini-3-flash-preview-ks by google: $0.16/M input, $0.97/M output, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/google/gemini-3-flash-preview-ks ## google/gemini-2.5-pro-preview-tts - [google/gemini-2.5-pro-preview-tts](https://www.orcarouter.ai/models/google/gemini-2.5-pro-preview-tts) — google/gemini-2.5-pro-preview-tts by google: $1.00/M input, $20.00/M output, p50 3733ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/google/gemini-2.5-pro-preview-tts ## Qwen: Qwen3.5-27B - [Qwen: Qwen3.5-27B](https://www.orcarouter.ai/models/qwen/qwen3.5-27b) — Qwen: Qwen3.5-27B by qwen: $0.09/M input, $0.69/M output, 33K context, p50 2403ms, available via OrcaRouter API. Qwen3.5 27B — open-weight dense multimodal (text/image/video), 27B params, 32k context (vision mode). Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3.5-27b ## Orca: OrcaCyber Zero 1.0 - [Orca: OrcaCyber Zero 1.0](https://www.orcarouter.ai/models/orca/orcacyber-zero-1.0) — Orca: OrcaCyber Zero 1.0 by orca: $3.00/M input, $7.50/M output, 1M context, available via OrcaRouter API. OrcaCyber Zero 1.0 is a cybersecurity post-trained coding model specialized for vulnerability analysis, vulnerability reproduction, red teaming, and offensive-security research. As the frontier-tier model in the OrcaCyber Harness, it achieves 98.07% pass@1 (1,478/1,507) on CyberGym Level 1 — 1,507 real-world vulnerability-reproduction tasks across 188 open-source projects. Zero 1.0 supports a 1M-token context window, native function calling, structured outputs, and configurable reasoning effort. It is served through an OpenAI-compatible API for trusted security researchers, red teams, and authorized security testing. Canonical URL: https://www.orcarouter.ai/models/orca/orcacyber-zero-1.0 ## Google: Gemini 3.1 Pro Preview - [Google: Gemini 3.1 Pro Preview](https://www.orcarouter.ai/models/google/gemini-3.1-pro-preview) — Google: Gemini 3.1 Pro Preview by google: $2.00/M input, $12.00/M output, 1M context, p50 60000ms, available via OrcaRouter API. Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation... Canonical URL: https://www.orcarouter.ai/models/google/gemini-3.1-pro-preview ## Google: Nano Banana Pro (Gemini 3 Pro Image Preview) - [Google: Nano Banana Pro (Gemini 3 Pro Image Preview)](https://www.orcarouter.ai/models/google/gemini-3-pro-image-preview) — Google: Nano Banana Pro (Gemini 3 Pro Image Preview) by google: 66K context, available via OrcaRouter API. Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and... Canonical URL: https://www.orcarouter.ai/models/google/gemini-3-pro-image-preview ## openai/gpt-5.2-pro-2025-12-11 - [openai/gpt-5.2-pro-2025-12-11](https://www.orcarouter.ai/models/openai/gpt-5.2-pro-2025-12-11) — openai/gpt-5.2-pro-2025-12-11 by openai: $21.00/M input, $168.00/M output, 400K context, p50 418ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.2-pro-2025-12-11 ## Qwen: Qwen3.7 Flash - [Qwen: Qwen3.7 Flash](https://www.orcarouter.ai/models/qwen/qwen3.7-flash) — Qwen: Qwen3.7 Flash by qwen: $0.03/M input, $0.13/M output, 1M context, p50 3445ms, available via OrcaRouter API. Qwen3.7 Flash is Alibaba's fast, extremely cost-efficient model in the Qwen3.7 family, sitting below Qwen3.7-Plus and Qwen3.7-Max as the high-throughput tier. It is natively multimodal on the input side — accepting text, images, and video with text output — and serves a 1M-token context window with up to 64K output tokens, so long documents, screenshots, and video frames stay in reach even at Flash-tier pricing. At roughly $0.03 per million input tokens on the base tier it is one of the cheapest 1M-context multimodal models available, making it a natural fit for classification, extraction, routing, summarization, lightweight agents, and any pipeline run at very high volume. It supports reasoning mode, native tool calling, structured outputs via response_format, and the usual sampling controls (temperature / top_p / seed / logprobs). Note that pricing is tiered by prompt length: costs step up above 32K and again above 256K input tokens, so batching short requests is materially cheaper than padding long ones. Use Flash as the volume workhorse and escalate to Qwen3.7-Plus or Max when a task genuinely needs more capability. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3.7-flash ## Qwen: Qwen3.8 27B - [Qwen: Qwen3.8 27B](https://www.orcarouter.ai/models/qwen/qwen3.8-27b) — Qwen: Qwen3.8 27B by qwen: $0.33/M input, $2.40/M output, 262K context, p50 1657ms, available via OrcaRouter API. Qwen3.8-27B is Alibaba's open-weight 27B dense multimodal model, released under Apache-2.0 and self-hosted on OrcaRouter's own infrastructure. It accepts text, images, and video and returns text, with a 256K-token context window (262,144 positions) built on the Qwen3.5 architecture - 64 layers, 5120 hidden size, with a dedicated vision tower. Despite its 27B size it posts unusually strong agentic and coding results on Qwen's own model card, including 90.3 on LiveCodeBench v6, 89.2 on GPQA Diamond, 84.3 on OSWorld-Verified, and 79.0 on QwenSWEBench - competitive with far larger closed models on several axes. It supports native tool calling, structured outputs, reasoning mode, and the full sampling surface. Because the weights are open and we run them ourselves, there is no per-token vendor cost to pass through. That makes Qwen3.8-27B a strong default for high-volume multimodal work: document and screenshot understanding, video frames, coding agents, and computer-use style automation. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3.8-27b ## Z.ai: GLM 4.5 Air - [Z.ai: GLM 4.5 Air](https://www.orcarouter.ai/models/z-ai/glm-4.5-air) — Z.ai: GLM 4.5 Air by z-ai: $0.20/M input, $1.10/M output, 128K context, p50 1000ms, available via OrcaRouter API. Compact MoE sibling of GLM-4.5: 106B total / 12B active. Same hybrid-reasoning and tool-calling stack tuned for high-throughput, low-cost inference. 128K context. Canonical URL: https://www.orcarouter.ai/models/z-ai/glm-4.5-air ## Anthropic: Claude Opus 5.5 - [Anthropic: Claude Opus 5.5](https://www.orcarouter.ai/models/anthropic/claude-opus-5.5) — Anthropic: Claude Opus 5.5 by anthropic: $4.00/M input, $20.00/M output, 1M context, p50 2630ms, available via OrcaRouter API. Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. Particularly strong at multi-step changes in large codebases and code review. Multimodal input (text, image, file), 1M-token context, extended thinking, tools and structured outputs. Canonical URL: https://www.orcarouter.ai/models/anthropic/claude-opus-5.5 ## Anthropic: Claude Sonnet 5 - [Anthropic: Claude Sonnet 5](https://www.orcarouter.ai/models/anthropic/claude-sonnet-5) — Anthropic: Claude Sonnet 5 by anthropic: $2.00/M input, $10.00/M output, 1M context, p50 3632ms, available via OrcaRouter API. Claude Sonnet 5 is Anthropic's most capable Sonnet-class model — frontier-level performance across coding, agentic workflows, and professional knowledge work, at a fraction of the cost of the Opus tier. It serves a 1M-token context window with up to 128K output tokens, accepts text, image, and file inputs with text output, and supports adaptive thinking with selectable reasoning effort (low, medium, high, max) so callers can dial the intelligence / latency / cost tradeoff per request. Built as Anthropic's most agentic Sonnet yet, it posts large gains over Sonnet 4.6 on agentic coding and computer-use and closes much of the gap to Opus 4.8 — 63.2% on SWE-bench Pro, 80.4% on Terminal-Bench 2.1, and 81.2% on OSWorld-Verified — while pricing well below Opus 4.8, GPT-5.5, and Gemini 3.1 Pro. It is a strong default for cost-sensitive agents, coding assistants, and high-volume production workloads that still demand frontier reasoning. Canonical URL: https://www.orcarouter.ai/models/anthropic/claude-sonnet-5 ## Grok 4.7 - [Grok 4.7](https://www.orcarouter.ai/models/grok/grok-4.7) — Grok 4.7 by grok: $2.00/M input, $6.00/M output, 500K context, available via OrcaRouter API. Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. Particularly strong at long-running software engineering and verifying its own work. Multimodal input (text, image, file), 500K-token context, up to 450K output tokens, reasoning with configurable effort, tools and structured outputs. Canonical URL: https://www.orcarouter.ai/models/grok/grok-4.7 ## OpenAI: GPT-5 - [OpenAI: GPT-5](https://www.orcarouter.ai/models/openai/gpt-5) — OpenAI: GPT-5 by openai: $1.25/M input, $10.00/M output, 400K context, p50 818ms, available via OrcaRouter API. GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5 ## openai/gpt-5-2025-08-07 - [openai/gpt-5-2025-08-07](https://www.orcarouter.ai/models/openai/gpt-5-2025-08-07) — openai/gpt-5-2025-08-07 by openai: $1.25/M input, $10.00/M output, 400K context, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5-2025-08-07 ## Google: Gemini 2.5 Pro - [Google: Gemini 2.5 Pro](https://www.orcarouter.ai/models/google/gemini-2.5-pro) — Google: Gemini 2.5 Pro by google: $2.50/M input, $15.00/M output, 1M context, p50 13921ms, available via OrcaRouter API. Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy... Canonical URL: https://www.orcarouter.ai/models/google/gemini-2.5-pro ## OpenAI: GPT-4o-mini (2024-07-18) - [OpenAI: GPT-4o-mini (2024-07-18)](https://www.orcarouter.ai/models/openai/gpt-4o-mini-2024-07-18) — OpenAI: GPT-4o-mini (2024-07-18) by openai: $0.15/M input, $0.60/M output, 128K context, available via OrcaRouter API. GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-4o-mini-2024-07-18 ## OpenAI: GPT-5 Nano - [OpenAI: GPT-5 Nano](https://www.orcarouter.ai/models/openai/gpt-5-nano) — OpenAI: GPT-5 Nano by openai: $0.05/M input, $0.40/M output, 400K context, p50 461ms, available via OrcaRouter API. GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5-nano ## openai/gpt-5-pro-2025-10-06 - [openai/gpt-5-pro-2025-10-06](https://www.orcarouter.ai/models/openai/gpt-5-pro-2025-10-06) — openai/gpt-5-pro-2025-10-06 by openai: $15.00/M input, $120.00/M output, 400K context, p50 1000ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5-pro-2025-10-06 ## openai/gpt-image-1-mini - [openai/gpt-image-1-mini](https://www.orcarouter.ai/models/openai/gpt-image-1-mini) — openai/gpt-image-1-mini by openai: $2.00/M input, $8.00/M output, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-image-1-mini ## google/imagen-4.0-generate-001 - [google/imagen-4.0-generate-001](https://www.orcarouter.ai/models/google/imagen-4.0-generate-001) — google/imagen-4.0-generate-001 by google: available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/google/imagen-4.0-generate-001 ## OpenAI: GPT-4o (2024-08-06) - [OpenAI: GPT-4o (2024-08-06)](https://www.orcarouter.ai/models/openai/gpt-4o-2024-08-06) — OpenAI: GPT-4o (2024-08-06) by openai: $2.50/M input, $10.00/M output, 128K context, p50 707ms, available via OrcaRouter API. The 2024-08-06 version of GPT-4o offers improved performance in structured outputs, with the ability to supply a JSON schema in the respone_format. Read more [here](https://openai.com/index/introducing-structured-outputs-in-the-api/). GPT-4o ("o" for "omni") is... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-4o-2024-08-06 ## OpenAI: GPT-4o (2024-11-20) - [OpenAI: GPT-4o (2024-11-20)](https://www.orcarouter.ai/models/openai/gpt-4o-2024-11-20) — OpenAI: GPT-4o (2024-11-20) by openai: $2.50/M input, $10.00/M output, 128K context, p50 445ms, available via OrcaRouter API. The 2024-11-20 version of GPT-4o offers a leveled-up creative writing ability with more natural, engaging, and tailored writing to improve relevance & readability. It’s also better at working with uploaded... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-4o-2024-11-20 ## Z.ai: GLM 4.7 - [Z.ai: GLM 4.7](https://www.orcarouter.ai/models/z-ai/glm-4.7) — Z.ai: GLM 4.7 by z-ai: $0.60/M input, $2.20/M output, 200K context, p50 6586ms, available via OrcaRouter API. Iteration on GLM-4.6 introducing retained reasoning and round-based reasoning for more stable execution on complex multi-turn tasks. 200K context. Canonical URL: https://www.orcarouter.ai/models/z-ai/glm-4.7 ## kling/kling-v2-1-master - [kling/kling-v2-1-master](https://www.orcarouter.ai/models/kling/kling-v2-1-master) — kling/kling-v2-1-master by kling: available via OrcaRouter API. Kling 2.1 Master — premium text-to-video and image-to-video, 5–10s clips, 1080p, 24fps. Canonical URL: https://www.orcarouter.ai/models/kling/kling-v2-1-master ## OpenAI: GPT-4o (2024-05-13) - [OpenAI: GPT-4o (2024-05-13)](https://www.orcarouter.ai/models/openai/gpt-4o-2024-05-13) — OpenAI: GPT-4o (2024-05-13) by openai: $5.00/M input, $15.00/M output, 128K context, available via OrcaRouter API. GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-4o-2024-05-13 ## Qwen: Qwen3.8 Max - [Qwen: Qwen3.8 Max](https://www.orcarouter.ai/models/qwen/qwen3.8-max) — Qwen: Qwen3.8 Max by qwen: $2.00/M input, $6.00/M output, 1M context, p50 2321ms, available via OrcaRouter API. Qwen3.8-Max is Alibaba's newest flagship model in the Qwen line and its highest-capability tier to date. Alibaba positions it directly against GPT-5.5, Claude Opus 4.7, and Gemini 3.1 Pro in its own migration guide, recommending it whenever a task needs the strongest available reasoning — complex multi-step analysis, deep logical derivation, and demanding agentic work. It is natively multimodal, listed by Alibaba under both text generation and image/video understanding: it accepts text, images, and video and returns text, with a 1M-token context window. The official capability matrix confirms full support for thinking mode, function calling, built-in tools, and structured outputs, making it a complete drop-in for agent frameworks and tool-calling pipelines. Qwen3.8-Max is served through three wire formats — OpenAI-compatible, Anthropic-compatible, and native DashScope — across Beijing, Singapore, Tokyo, Frankfurt, and Virginia. Thinking is toggled with the enable_thinking parameter (or reasoning.effort on the Responses API). It is the premium tier of the family: reach for it when correctness on hard problems outweighs cost, and drop to Qwen3.7-Plus or Qwen3.7-Flash for everyday volume. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3.8-max ## Tencent: Hy3 (Free) - [Tencent: Hy3 (Free)](https://www.orcarouter.ai/models/tencent/hy3-free) — Tencent: Hy3 (Free) by tencent: 262K context, p50 3519ms, available via OrcaRouter API. Hy3 is Tencent Hunyuan's production-grade Mixture-of-Experts model — 295B total parameters with only 21B active per pass (192 experts, top-8 routing), the upgraded release built on the Hy3-preview line. It expands RL-training scale and post-training data quality for further gains in reasoning, long-context, and agentic tasks, reaching results comparable to flagship models several times its parameter size. It serves a 256K-token context window (text in, text out) with configurable reasoning effort, and is built for real-world coding, tool use, and multi-step agent workflows at a strong quality-to-cost ratio. Canonical URL: https://www.orcarouter.ai/models/tencent/hy3-free ## Anthropic: Claude Fable 5 - [Anthropic: Claude Fable 5](https://www.orcarouter.ai/models/anthropic/claude-fable-5) — Anthropic: Claude Fable 5 by anthropic: $10.00/M input, $50.00/M output, 1M context, p50 2654ms, available via OrcaRouter API. Claude Fable 5 is Anthropic's Mythos-class model - a capability tier above the Opus class - made safe for general use. Its capabilities exceed any model Anthropic has previously released broadly, with state-of-the-art results across software engineering, knowledge work, vision, and scientific research; the longer and more complex the task, the larger its lead. It accepts text, image, and file inputs with text output, serves a 1M-token context window with up to 128K output tokens, and supports adaptive reasoning and structured outputs. Fable 5 is built for autonomous, long-horizon work: it stays coherent across millions of tokens, improves its own outputs using file-based memory, and completes complex multi-step tasks with far less scaffolding than prior models. It is the new state of the art for vision-heavy tasks - extracting precise values from scientific figures, rebuilding applications from screenshots alone - and a step change for agentic coding and prototyping, handling codebase-wide migrations and frontier engineering tasks in fewer, more token-efficient turns. This makes it a strong default for AI coding assistants, deep research and analysis pipelines, and long-running autonomous agents where sustained coherence and judgment matter most. Canonical URL: https://www.orcarouter.ai/models/anthropic/claude-fable-5 ## Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview) - [Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview)](https://www.orcarouter.ai/models/google/gemini-3.1-flash-image-preview) — Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview) by google: 66K context, available via OrcaRouter API. Gemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines... Canonical URL: https://www.orcarouter.ai/models/google/gemini-3.1-flash-image-preview ## OpenAI: GPT-4o - [OpenAI: GPT-4o](https://www.orcarouter.ai/models/openai/gpt-4o) — OpenAI: GPT-4o by openai: $2.50/M input, $10.00/M output, 128K context, p50 1368ms, available via OrcaRouter API. GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-4o ## qwen/qwen3.6-plus-2026-04-02 - [qwen/qwen3.6-plus-2026-04-02](https://www.orcarouter.ai/models/qwen/qwen3.6-plus-2026-04-02) — qwen/qwen3.6-plus-2026-04-02 by qwen: $0.50/M input, $3.00/M output, 1M context, p50 938ms, available via OrcaRouter API. Qwen3.6 Plus snapshot 2026-04-02 — frozen version of qwen3.6-plus, same capabilities. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3.6-plus-2026-04-02 ## google/gemini-flash-latest - [google/gemini-flash-latest](https://www.orcarouter.ai/models/google/gemini-flash-latest) — google/gemini-flash-latest by google: $0.50/M input, $3.00/M output, p50 1364ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/google/gemini-flash-latest ## Qwen3.6 35B A3B - [Qwen3.6 35B A3B](https://www.orcarouter.ai/models/obsidian/qwen3.6-35b-a3b) — Qwen3.6 35B A3B by obsidian: $0.31/M input, $4.21/M output, 262K context, p50 3902ms, available via OrcaRouter API. Qwen3.6 35B A3B is a mixture-of-experts build of Alibaba's Qwen3.6 (35B total parameters, about 3B active per token) with a 262K-token context window and text, image, and video input. Preserves the original model's reasoning, coding, multilingual performance, and tool use. Designed to provide direct, complete responses across a wide range of prompts. The model may occasionally append brief informational disclaimers inherited from the base model's training. Ideal for AI research, security testing, red teaming, agent development, coding assistants, and other advanced AI applications that benefit from maximum output flexibility. Canonical URL: https://www.orcarouter.ai/models/obsidian/qwen3.6-35b-a3b ## OpenAI: GPT-4.1 - [OpenAI: GPT-4.1](https://www.orcarouter.ai/models/openai/gpt-4.1) — OpenAI: GPT-4.1 by openai: $2.00/M input, $8.00/M output, 1M context, p50 1796ms, available via OrcaRouter API. GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-4.1 ## openai/gpt-4o-mini-tts - [openai/gpt-4o-mini-tts](https://www.orcarouter.ai/models/openai/gpt-4o-mini-tts) — openai/gpt-4o-mini-tts by openai: $0.60/M input, $12.00/M output, p50 1000ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-4o-mini-tts ## openai/tts-1-hd-1106 - [openai/tts-1-hd-1106](https://www.orcarouter.ai/models/openai/tts-1-hd-1106) — openai/tts-1-hd-1106 by openai: $30.00/M characters input, p50 1000ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/tts-1-hd-1106 ## Qwen: Qwen3.5 397B A17B - [Qwen: Qwen3.5 397B A17B](https://www.orcarouter.ai/models/qwen/qwen3.5-397b-a17b) — Qwen: Qwen3.5 397B A17B by qwen: $0.17/M input, $1.03/M output, 33K context, p50 1258ms, available via OrcaRouter API. Qwen3.5 397B-A17B — open-weight MoE multimodal (text/image/video), 397B total / 17B active params, 32k context (vision mode). Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3.5-397b-a17b ## Qwen: Qwen3.7 Plus - [Qwen: Qwen3.7 Plus](https://www.orcarouter.ai/models/qwen/qwen3.7-plus) — Qwen: Qwen3.7 Plus by qwen: $0.35/M input, $1.42/M output, 1M context, p50 3000ms, available via OrcaRouter API. Qwen3.7-Plus is Alibaba's most capable multimodal agent model, unifying vision and language into a single, versatile agent foundation. Built on the Qwen3.7 text backbone, it delivers a comprehensive upgrade in vision-language understanding while retaining full agentic strength in coding, tool use, and productivity workflows. It accepts text, image, and video inputs with text output, and serves a 1M-token context window with up to 65K output tokens - enough to keep long documents, large codebases, screen recordings, and multi-turn agent sessions coherent without truncation. What sets it apart is its ability to operate as a multimodal interactive hybrid agent: it perceives real-world scenes, reads screens and operates GUIs, writes code from visual references, navigates mobile apps end-to-end, and answers visual questions grounded in web knowledge - blending GUI and CLI interactions within a single agent loop. It generalizes across agent scaffolds, performing consistently whether driven through Claude Code, OpenClaw, Qwen Code, or other frameworks, with native function calling, structured outputs, and controllable reasoning depth. This makes it a dependable default for AI coding assistants, computer-use and browser agents, visual QA pipelines, and long-running automation where perception, reasoning, and execution must stay aligned. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3.7-plus ## DeepSeek: DeepSeek V4 Flash (Free) - [DeepSeek: DeepSeek V4 Flash (Free)](https://www.orcarouter.ai/models/deepseek/deepseek-v4-flash-free) — DeepSeek: DeepSeek V4 Flash (Free) by deepseek: 1M context, p50 2228ms, available via OrcaRouter API. DeepSeek V4 Flash efficient MoE — 284B total / 13B active params, 1M context, optimized for fast everyday workloads. Canonical URL: https://www.orcarouter.ai/models/deepseek/deepseek-v4-flash-free ## DeepSeek: DeepSeek V4 Pro 0813 - [DeepSeek: DeepSeek V4 Pro 0813](https://www.orcarouter.ai/models/deepseek/deepseek-v4-pro-0813) — DeepSeek: DeepSeek V4 Pro 0813 by deepseek: $0.66/M input, $1.98/M output, 1M context, p50 1368ms, available via OrcaRouter API. DeepSeek V4 Pro 0813 is the official release of DeepSeek's flagship V4 Pro model, superseding the April preview and rolled out simultaneously across the app, web, and API. It serves a 1M-token context window with up to 384K output tokens, supports thinking and non-thinking modes with selectable low / high / max reasoning effort, plus JSON output and tool calls, and natively speaks the Responses API with a targeted adaptation for Codex-style coding agents. The 0813 revision is a major step up in agent capability, with DeepSeek reporting especially strong gains in production environments on its published agent benchmarks — 87.9 on Terminal Bench 2.1, 83.3 on Cybergym, 74.1 on Toolathlon Verified, 62.7 on DeepSWE, 61.5 on NL2Repo, and 42.7 / 60.0 on HLE without / with tools. On the DeepSeek first-party API this revision is what the deepseek-v4-pro model id now serves; the dated identifier lists the same revision for callers that reference it by date. It is a strong pick for demanding coding agents, terminal and tool-use workloads, and complex agentic pipelines. Canonical URL: https://www.orcarouter.ai/models/deepseek/deepseek-v4-pro-0813 ## openai/gpt-4-0613 - [openai/gpt-4-0613](https://www.orcarouter.ai/models/openai/gpt-4-0613) — openai/gpt-4-0613 by openai: $30.00/M input, $60.00/M output, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-4-0613 ## OpenAI: GPT-5.3-Codex - [OpenAI: GPT-5.3-Codex](https://www.orcarouter.ai/models/openai/gpt-5.3-codex) — OpenAI: GPT-5.3-Codex by openai: $1.75/M input, $14.00/M output, 400K context, p50 303ms, available via OrcaRouter API. GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with the broader reasoning and professional knowledge capabilities of GPT-5.2. It achieves state-of-the-art results... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.3-codex ## OpenAI: GPT-5.4 Mini - [OpenAI: GPT-5.4 Mini](https://www.orcarouter.ai/models/openai/gpt-5.4-mini) — OpenAI: GPT-5.4 Mini by openai: $0.75/M input, $4.50/M output, 400K context, available via OrcaRouter API. GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.4-mini ## Qwen: Qwen3.6 35B A3B - [Qwen: Qwen3.6 35B A3B](https://www.orcarouter.ai/models/qwen/qwen3.6-35b-a3b) — Qwen: Qwen3.6 35B A3B by qwen: $0.25/M input, $1.49/M output, 262K context, p50 3211ms, available via OrcaRouter API. Qwen3.6 35B-A3B — open-weight MoE multimodal (text/image/video), 35B total / 3B active params, 256k context. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3.6-35b-a3b ## Qwen: Qwen3.6 Flash - [Qwen: Qwen3.6 Flash](https://www.orcarouter.ai/models/qwen/qwen3.6-flash) — Qwen: Qwen3.6 Flash by qwen: $0.25/M input, $1.50/M output, 1M context, p50 1000ms, available via OrcaRouter API. Qwen3.6 Flash — multimodal chat (text/image/video) optimized for cost, 1M context, near-flagship capability. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3.6-flash ## MiniMax: MiniMax-H3 - [MiniMax: MiniMax-H3](https://www.orcarouter.ai/models/minimax/minimax-h3) — MiniMax: MiniMax-H3 by minimax: available via OrcaRouter API. MiniMax-H3 is MiniMax's omni-modal video generation model (the Hailuo 3 generation), released July 31, 2026. It reads text, image, video and audio references as one unified context and generates 4-15 second videos at 768P or 2K with native stereo audio. Billed per second of generated output: $0.08/s at 768P, $0.13/s at 2K. Canonical URL: https://www.orcarouter.ai/models/minimax/minimax-h3 ## kling/kling-v3 - [kling/kling-v3](https://www.orcarouter.ai/models/kling/kling-v3) — kling/kling-v3 by kling: available via OrcaRouter API. Kling 3.0 — flagship text-to-video and image-to-video, multi-shot + subject + motion control, 3–15s clips, up to native 4K. Canonical URL: https://www.orcarouter.ai/models/kling/kling-v3 ## openai/gpt-image-1.5 - [openai/gpt-image-1.5](https://www.orcarouter.ai/models/openai/gpt-image-1.5) — openai/gpt-image-1.5 by openai: $8.00/M input, $32.00/M output, p50 4000ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-image-1.5 ## Anthropic: Claude Opus 4.6 - [Anthropic: Claude Opus 4.6](https://www.orcarouter.ai/models/anthropic/claude-opus-4.6) — Anthropic: Claude Opus 4.6 by anthropic: $5.00/M input, $25.00/M output, 1M context, p50 3790ms, available via OrcaRouter API. Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective... Canonical URL: https://www.orcarouter.ai/models/anthropic/claude-opus-4.6 ## Anthropic: Claude Haiku 4.5 - [Anthropic: Claude Haiku 4.5](https://www.orcarouter.ai/models/anthropic/claude-haiku-4.5) — Anthropic: Claude Haiku 4.5 by anthropic: $1.00/M input, $5.00/M output, 200K context, available via OrcaRouter API. Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance... Canonical URL: https://www.orcarouter.ai/models/anthropic/claude-haiku-4.5 ## DeepSeek: DeepSeek V4 Pro - [DeepSeek: DeepSeek V4 Pro](https://www.orcarouter.ai/models/deepseek/deepseek-v4-pro) — DeepSeek: DeepSeek V4 Pro by deepseek: $0.66/M input, $1.98/M output, 1M context, p50 1931ms, available via OrcaRouter API. DeepSeek V4 Pro flagship MoE — 1.6T total / 49B active params, 1M context, top-tier reasoning + agentic tool use. Canonical URL: https://www.orcarouter.ai/models/deepseek/deepseek-v4-pro ## google/gemini-embedding-2-preview - [google/gemini-embedding-2-preview](https://www.orcarouter.ai/models/google/gemini-embedding-2-preview) — google/gemini-embedding-2-preview by google: $0.20/M input, $0.20/M output, p50 250ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/google/gemini-embedding-2-preview ## kling/kling-v2-5-turbo - [kling/kling-v2-5-turbo](https://www.orcarouter.ai/models/kling/kling-v2-5-turbo) — kling/kling-v2-5-turbo by kling: available via OrcaRouter API. Kling 2.5 Turbo — fast text-to-video and image-to-video with first-last-frame mode, 5–10s clips, 1080p, 24fps. Canonical URL: https://www.orcarouter.ai/models/kling/kling-v2-5-turbo ## MiniMax M2.7 highspeed - [MiniMax M2.7 highspeed](https://www.orcarouter.ai/models/minimax/minimax-m2.7-highspeed) — MiniMax M2.7 highspeed by minimax: $0.60/M input, $2.40/M output, 205K context, p50 23253ms, available via OrcaRouter API. MiniMax M2.7 high-speed — same model + same 200k context as M2.7, faster output (~100 tps vs ~60 tps). Canonical URL: https://www.orcarouter.ai/models/minimax/minimax-m2.7-highspeed ## OpenAI: GPT-4 Turbo - [OpenAI: GPT-4 Turbo](https://www.orcarouter.ai/models/openai/gpt-4-turbo) — OpenAI: GPT-4 Turbo by openai: $10.00/M input, $30.00/M output, 128K context, available via OrcaRouter API. The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December 2023. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-4-turbo ## Anthropic: Claude Opus 5 - [Anthropic: Claude Opus 5](https://www.orcarouter.ai/models/anthropic/claude-opus-5) — Anthropic: Claude Opus 5 by anthropic: $5.00/M input, $25.00/M output, 1M context, p50 2709ms, available via OrcaRouter API. Claude Opus 5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work — the most capable model in the Opus line to date. It is especially strong at end-to-end software engineering, code review and bug finding, and visual analysis, and is built to stay coherent across very long, multi-step agentic sessions with heavy tool use. It accepts text, image, and file inputs with text output, serves a 1M-token context window with up to 128K output tokens, and supports adaptive reasoning with configurable effort. Opus 5 uses Anthropic's adaptive thinking design: instead of sampling parameters like temperature or top_p, callers dial depth through reasoning effort, and the model allocates thinking automatically. It speaks both the OpenAI chat-completions format and Anthropic's native Messages API (/v1/messages), supports structured outputs, verbosity control, and native tool calling, and is a strong default for production coding agents, deep research, and complex automation where correctness and judgment matter most. Canonical URL: https://www.orcarouter.ai/models/anthropic/claude-opus-5 ## Google: Gemma 4 31B - [Google: Gemma 4 31B](https://www.orcarouter.ai/models/google/gemma-4-31b-it) — Google: Gemma 4 31B by google: $0.13/M input, $0.38/M output, available via OrcaRouter API. Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function... Canonical URL: https://www.orcarouter.ai/models/google/gemma-4-31b-it ## SpaceXAI: Grok 4.6 - [SpaceXAI: Grok 4.6](https://www.orcarouter.ai/models/grok/grok-4.6) — SpaceXAI: Grok 4.6 by grok: $2.00/M input, $6.00/M output, 500K context, available via OrcaRouter API. Grok 4.6 is SpaceXAI's smartest model to date, with frontier performance across coding, knowledge work, and STEM. It is the current flagship of the Grok line — listed as "Latest" in xAI's own developer docs — and succeeds Grok 4.5 with the same 500K-token context window, the same text / image / file input surface, and the same base pricing. It supports configurable reasoning effort, native tool calling, structured outputs, and the full sampling surface (temperature / top_p / seed / logprobs / penalties), so it drops into existing integrations unchanged. As a first-class OpenAI Responses model on api.x.ai it plugs directly into agent frameworks and tool-calling loops without a translation layer. Pricing is tiered by prompt length: requests above 200K input tokens bill at double the base rate. Use Grok 4.6 as the high-capability tier for complex coding agents, research, and multi-step automation where quality matters more than cost. Canonical URL: https://www.orcarouter.ai/models/grok/grok-4.6 ## MiniMax: MiniMax M2.5 - [MiniMax: MiniMax M2.5](https://www.orcarouter.ai/models/minimax/minimax-m2.5) — MiniMax: MiniMax M2.5 by minimax: $0.30/M input, $1.20/M output, 205K context, p50 840ms, available via OrcaRouter API. MiniMax M2.5 — SOTA productivity LLM with strong coding + agentic capabilities, 200k context, ~60 tps output. Canonical URL: https://www.orcarouter.ai/models/minimax/minimax-m2.5 ## openai/gpt-5-chat-latest - [openai/gpt-5-chat-latest](https://www.orcarouter.ai/models/openai/gpt-5-chat-latest) — openai/gpt-5-chat-latest by openai: $1.25/M input, $10.00/M output, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5-chat-latest ## OpenAI: GPT-4.1 Mini - [OpenAI: GPT-4.1 Mini](https://www.orcarouter.ai/models/openai/gpt-4.1-mini) — OpenAI: GPT-4.1 Mini by openai: $0.40/M input, $1.60/M output, 1M context, p50 799ms, available via OrcaRouter API. GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-4.1-mini ## openai/gpt-5-search-api - [openai/gpt-5-search-api](https://www.orcarouter.ai/models/openai/gpt-5-search-api) — openai/gpt-5-search-api by openai: $1.25/M input, $10.00/M output, p50 8408ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5-search-api ## OpenAI: GPT-5.2 Pro - [OpenAI: GPT-5.2 Pro](https://www.orcarouter.ai/models/openai/gpt-5.2-pro) — OpenAI: GPT-5.2 Pro by openai: $21.00/M input, $168.00/M output, 400K context, p50 329ms, available via OrcaRouter API. GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It is optimized for complex tasks that require step-by-step reasoning,... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.2-pro ## Anthropic: Claude Opus 4.8 - [Anthropic: Claude Opus 4.8](https://www.orcarouter.ai/models/anthropic/claude-opus-4.8) — Anthropic: Claude Opus 4.8 by anthropic: $5.00/M input, $25.00/M output, 1M context, p50 7044ms, available via OrcaRouter API. Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token context window. It is suited for highly autonomous agents, long-horizon agentic work, knowledge work, and memory-driven tasks where coherence over extended sessions matters. It is particularly strong on multi-step reasoning, complex coding, and end-to-end project orchestration - large codebases, multi-stage debugging, and long-running asynchronous agent pipelines. Beyond coding, it handles knowledge work such as drafting documents, building presentations, and analyzing data, maintaining quality across very long outputs. Canonical URL: https://www.orcarouter.ai/models/anthropic/claude-opus-4.8 ## google/gemini-3.1-flash-lite - [google/gemini-3.1-flash-lite](https://www.orcarouter.ai/models/google/gemini-3.1-flash-lite) — google/gemini-3.1-flash-lite by google: $0.25/M input, $1.50/M output, p50 620ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/google/gemini-3.1-flash-lite ## openai/gpt-4-turbo-2024-04-09 - [openai/gpt-4-turbo-2024-04-09](https://www.orcarouter.ai/models/openai/gpt-4-turbo-2024-04-09) — openai/gpt-4-turbo-2024-04-09 by openai: $10.00/M input, $30.00/M output, 128K context, p50 961ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-4-turbo-2024-04-09 ## Qwen3.7 Max (2026-05-20) - [Qwen3.7 Max (2026-05-20)](https://www.orcarouter.ai/models/qwen/qwen3.7-max-2026-05-20) — Qwen3.7 Max (2026-05-20) by qwen: $1.25/M input, $3.75/M output, 1M context, p50 1466ms, available via OrcaRouter API. Qwen3.7-Max (2026-05-20 snapshot) — Dated checkpoint of Alibaba's flagship proprietary agent-era model, pinned for reproducible production workloads. Native 1M token context window, with an extended thinking mode (and preserve_thinking across turns) tuned for agentic tasks. Frontier-level results on coding (SWE-Verified, SWE-Pro, Terminal-Bench), reasoning (GPQA Diamond, HMMT, IMO), tool use (BFCL, MCP-Mark, MCP-Atlas), and multilingual benchmarks (WMT24++ across 55 languages). Engineered for long-horizon autonomous execution and consistent behavior across agent scaffolds including Claude Code, OpenClaw, and Qwen Code. Use this pinned version when you need stable behavior across releases; use qwen/qwen3.7-max for the rolling alias. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3.7-max-2026-05-20 ## Z.ai: GLM 4.6 - [Z.ai: GLM 4.6](https://www.orcarouter.ai/models/z-ai/glm-4.6) — Z.ai: GLM 4.6 by z-ai: $0.60/M input, $2.20/M output, 200K context, p50 15068ms, available via OrcaRouter API. Successor to GLM-4.5 with the context window extended to 200K, in-thinking tool calls, and stronger agentic / search behavior. Canonical URL: https://www.orcarouter.ai/models/z-ai/glm-4.6 ## Z.ai: GLM 5.3 Flash - [Z.ai: GLM 5.3 Flash](https://www.orcarouter.ai/models/z-ai/glm-5.3-flash) — Z.ai: GLM 5.3 Flash by z-ai: $0.07/M input, $0.25/M output, 1M context, p50 4913ms, available via OrcaRouter API. GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while sharply reducing serving cost. 320B total / 18B active parameters, 1M-token context, text + image + video in, text out. Canonical URL: https://www.orcarouter.ai/models/z-ai/glm-5.3-flash ## Anthropic: Claude Sonnet 4.5 - [Anthropic: Claude Sonnet 4.5](https://www.orcarouter.ai/models/anthropic/claude-sonnet-4.5) — Anthropic: Claude Sonnet 4.5 by anthropic: $3.00/M input, $15.00/M output, 1M context, available via OrcaRouter API. Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with... Canonical URL: https://www.orcarouter.ai/models/anthropic/claude-sonnet-4.5 ## OpenAI: GPT-4 - [OpenAI: GPT-4](https://www.orcarouter.ai/models/openai/gpt-4) — OpenAI: GPT-4 by openai: $30.00/M input, $60.00/M output, 8K context, p50 1000ms, available via OrcaRouter API. OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving difficult problems with greater accuracy than previous models due to its broader general knowledge and advanced reasoning... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-4 ## openai/gpt-4o-mini-tts-2025-03-20 - [openai/gpt-4o-mini-tts-2025-03-20](https://www.orcarouter.ai/models/openai/gpt-4o-mini-tts-2025-03-20) — openai/gpt-4o-mini-tts-2025-03-20 by openai: $0.60/M input, $12.00/M output, p50 783ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-4o-mini-tts-2025-03-20 ## OpenAI: GPT-5 Mini - [OpenAI: GPT-5 Mini](https://www.orcarouter.ai/models/openai/gpt-5-mini) — OpenAI: GPT-5 Mini by openai: $0.25/M input, $2.00/M output, 400K context, available via OrcaRouter API. GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost.... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5-mini ## openai/gpt-5.4-2026-03-05 - [openai/gpt-5.4-2026-03-05](https://www.orcarouter.ai/models/openai/gpt-5.4-2026-03-05) — openai/gpt-5.4-2026-03-05 by openai: $2.50/M input, $15.00/M output, 1M context, p50 968ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.4-2026-03-05 ## OpenAI: GPT-5.5 Pro - [OpenAI: GPT-5.5 Pro](https://www.orcarouter.ai/models/openai/gpt-5.5-pro) — OpenAI: GPT-5.5 Pro by openai: $30.00/M input, $180.00/M output, p50 1000ms, available via OrcaRouter API. GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window (922K input, 128K output) with support for... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.5-pro ## OpenAI: GPT-5.6 Luna - [OpenAI: GPT-5.6 Luna](https://www.orcarouter.ai/models/openai/gpt-5.6-luna) — OpenAI: GPT-5.6 Luna by openai: $0.20/M input, $1.20/M output, 1M context, available via OrcaRouter API. GPT-5.6 Luna is the fast, cost-efficient model in OpenAI's GPT-5.6 series — tuned for high-volume, latency-sensitive workloads while retaining genuinely capable reasoning. It is well suited to chat, classification, extraction, routing, and lightweight agentic workflows where responsiveness and price per call dominate, and still serves a full 1.05M-token context window with up to 128K output tokens. It accepts text, image, and file inputs with text output, supports configurable reasoning effort so you can dial depth up for the occasional hard request, and is a first-class OpenAI Responses model with native tool calling and structured outputs. Reach for Luna at the top of a fan-out, as a cheap first-pass or draft model, or wherever you run millions of calls and every fraction of a cent and millisecond counts. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.6-luna ## OpenAI: GPT-6.1 Sol - [OpenAI: GPT-6.1 Sol](https://www.orcarouter.ai/models/openai/gpt-6.1-sol) — OpenAI: GPT-6.1 Sol by openai: $2.00/M input, $10.00/M output, 1M context, available via OrcaRouter API. GPT-6.1 Sol is an upgrade to GPT-6 Sol from OpenAI, positioned one tier below the flagship GPT-6 Astra in the GPT-6 series. It targets agentic coding, computer use, and document-heavy professional work, pairing a 1M-token context window with 128K max output and multimodal input across text, images, and files. Reasoning is always on, with five effort levels from low through max and a default of medium, so latency and token spend can be tuned per route. It supports native tool calling, structured outputs, response_format, and seed. Sampling parameters such as temperature and top_p are not accepted. Pricing matches the Sonnet tier of the market at 2 USD per million input tokens and 10 USD per million output tokens, with a long-context tier that applies above 272K prompt tokens. On Artificial Analysis it posts an Intelligence Index of 51.8. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-6.1-sol ## Qwen: Qwen3.5-122B-A10B - [Qwen: Qwen3.5-122B-A10B](https://www.orcarouter.ai/models/qwen/qwen3.5-122b-a10b) — Qwen: Qwen3.5-122B-A10B by qwen: $0.12/M input, $0.92/M output, 33K context, p50 656ms, available via OrcaRouter API. Qwen3.5 122B-A10B — open-weight MoE multimodal (text/image/video), 122B total / 10B active params, 32k context (vision mode). Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3.5-122b-a10b ## Meta: Muse Spark 1.2 - [Meta: Muse Spark 1.2](https://www.orcarouter.ai/models/meta/muse-spark-1.2) — Meta: Muse Spark 1.2 by meta: $1.25/M input, $4.25/M output, 1M context, p50 7503ms, available via OrcaRouter API. Muse Spark 1.2 is Meta's reasoning model for complex agentic tasks, and the current checkpoint of the Muse Spark family. Meta describes it as an updated checkpoint over Muse Spark 1.1 with slightly higher performance, served on the same Standard tier at identical pricing — a drop-in upgrade rather than a new tier. It accepts an unusually broad input surface — text, images, video, audio, and PDF documents — and returns text, reasoning natively across modalities within a 1M-token context window. It supports configurable reasoning effort, native tool calling, and structured outputs, making it well suited for multimodal agents, deep research over mixed media, and long-context analysis. Muse Spark 1.2 targets workloads that combine documents, screenshots, recordings, and video with text — from analyzing long reports and media libraries to driving multi-step agentic pipelines that must reason over more than just text. Meta also ships muse-spark-1.2-contributor, the same checkpoint on its Contributor tier. Canonical URL: https://www.orcarouter.ai/models/meta/muse-spark-1.2 ## Orca: OrcaVerify Text 1.0 (Free) - [Orca: OrcaVerify Text 1.0 (Free)](https://www.orcarouter.ai/models/orca/orcaverify-text1.0-free) — Orca: OrcaVerify Text 1.0 (Free) by orca: p50 118ms, available via OrcaRouter API. AI-generated-text detection over the OpenAI Chat Completions format: four calibrated verdicts (AI_GENERATED / AI_ASSISTED / HUMAN / ABSTAIN) with a calibrated probability and paragraph-level localization. English, Chinese, Japanese and Korean. Canonical URL: https://www.orcarouter.ai/models/orca/orcaverify-text1.0-free ## Qwen3.7 Max - [Qwen3.7 Max](https://www.orcarouter.ai/models/qwen/qwen3.7-max) — Qwen3.7 Max by qwen: $1.25/M input, $3.75/M output, 1M context, p50 2817ms, available via OrcaRouter API. Qwen3.7-Max — Alibaba's flagship proprietary model, designed as a foundation for the agent era. Native 1M token context window, with an extended thinking mode (and preserve_thinking across turns) tuned for agentic tasks. Frontier-level results on coding (SWE-Verified, SWE-Pro, Terminal-Bench), reasoning (GPQA Diamond, HMMT, IMO), tool use (BFCL, MCP-Mark, MCP-Atlas), and multilingual benchmarks (WMT24++ across 55 languages). Engineered for long-horizon autonomous execution — sustains coherent strategy across thousands of tool calls and multi-hour sessions — and generalizes consistently across agent scaffolds including Claude Code, OpenClaw, and Qwen Code. Recommended for coding agents, office and workflow automation, long-context RAG, and any system that needs a reliable backbone for sustained tool-use. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3.7-max ## Qwen: Qwen3.8 Max (0902) - [Qwen: Qwen3.8 Max (0902)](https://www.orcarouter.ai/models/qwen/qwen3.8-max-0902) — Qwen: Qwen3.8 Max (0902) by qwen: $2.00/M input, $6.00/M output, 1M context, p50 2808ms, available via OrcaRouter API. Dated September 2, 2026 snapshot of Qwen3.8 Max, Alibaba's flagship multimodal reasoning model: text + image + video in, text out, 1M-token context. Same capabilities and pricing as the base model; pin it for reproducible behaviour. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3.8-max-0902 ## Z.ai: GLM 5.3 Flash (Free) - [Z.ai: GLM 5.3 Flash (Free)](https://www.orcarouter.ai/models/z-ai/glm-5.3-flash-free) — Z.ai: GLM 5.3 Flash (Free) by z-ai: 1M context, p50 5277ms, available via OrcaRouter API. GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while sharply reducing serving cost. 320B total / 18B active parameters, 1M-token context, text + image + video in, text out. Canonical URL: https://www.orcarouter.ai/models/z-ai/glm-5.3-flash-free ## google/gemini-3.1-flash-tts-preview - [google/gemini-3.1-flash-tts-preview](https://www.orcarouter.ai/models/google/gemini-3.1-flash-tts-preview) — google/gemini-3.1-flash-tts-preview by google: $1.00/M input, $20.00/M output, p50 3192ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/google/gemini-3.1-flash-tts-preview ## Google: Gemma 4 26B A4B - [Google: Gemma 4 26B A4B](https://www.orcarouter.ai/models/google/gemma-4-26b-a4b-it) — Google: Gemma 4 26B A4B by google: $0.06/M input, $0.33/M output, 262K context, p50 452ms, available via OrcaRouter API. Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at... Canonical URL: https://www.orcarouter.ai/models/google/gemma-4-26b-a4b-it ## OpenAI: GPT-5.6 Terra - [OpenAI: GPT-5.6 Terra](https://www.orcarouter.ai/models/openai/gpt-5.6-terra) — OpenAI: GPT-5.6 Terra by openai: $2.00/M input, $12.00/M output, 1M context, p50 365ms, available via OrcaRouter API. GPT-5.6 Terra is the balanced model in OpenAI's GPT-5.6 series, sitting between the Sol flagship and the cost-efficient Luna tier to give the best quality-to-cost ratio for everyday production work. It handles day-to-day coding, reasoning, tool use, and agentic workflows with a 1.05M-token context window and up to 128K output tokens, holding up well on long documents, large codebases, and multi-turn sessions. It accepts text, image, and file inputs with text output, supports configurable reasoning effort to tune the intelligence / latency / cost tradeoff, and is a first-class OpenAI Responses model with native tool calling and structured outputs. Terra is a strong default when Sol is more than the task needs but Luna is not quite enough — general assistants, mid-complexity coding, and high-throughput agent pipelines. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.6-terra ## Qwen: Qwen3 VL 235B A22B Instruct - [Qwen: Qwen3 VL 235B A22B Instruct](https://www.orcarouter.ai/models/qwen/qwen3-vl-235b-a22b-instruct) — Qwen: Qwen3 VL 235B A22B Instruct by qwen: $0.40/M input, $1.60/M output, p50 828ms, available via OrcaRouter API. Qwen3-VL 235B-A22B Instruct — open-weight vision-language model, 235B total / 22B active params, 256k context, no thinking mode. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3-vl-235b-a22b-instruct ## Qwen: Qwen3 VL 8B Thinking - [Qwen: Qwen3 VL 8B Thinking](https://www.orcarouter.ai/models/qwen/qwen3-vl-8b-thinking) — Qwen: Qwen3 VL 8B Thinking by qwen: $0.18/M input, $2.10/M output, 131K context, p50 1482ms, available via OrcaRouter API. Qwen3-VL 8B Thinking — open-weight small vision-language reasoning model, 8B params, 128k context. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3-vl-8b-thinking ## Z.ai: GLM 5.1 - [Z.ai: GLM 5.1](https://www.orcarouter.ai/models/z-ai/glm-5.1) — Z.ai: GLM 5.1 by z-ai: $1.40/M input, $4.40/M output, 200K context, p50 3365ms, available via OrcaRouter API. Z.ai's strongest coding-and-agent model in the GLM-5 line; supports streaming tool calls and deep thinking. 200K context. Canonical URL: https://www.orcarouter.ai/models/z-ai/glm-5.1 ## google/imagen-4.0-ultra-generate-001 - [google/imagen-4.0-ultra-generate-001](https://www.orcarouter.ai/models/google/imagen-4.0-ultra-generate-001) — google/imagen-4.0-ultra-generate-001 by google: available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/google/imagen-4.0-ultra-generate-001 ## kimi/kimi-k2.5 - [kimi/kimi-k2.5](https://www.orcarouter.ai/models/kimi/kimi-k2.5) — kimi/kimi-k2.5 by kimi: $0.60/M input, $3.00/M output, 262K context, available via OrcaRouter API. Moonshot Kimi K2 (0905 baseline) — 1T-param MoE chat model with 32B active per pass, 256k context, balanced performance. Canonical URL: https://www.orcarouter.ai/models/kimi/kimi-k2.5 ## openai/gpt-4.1-nano-2025-04-14 - [openai/gpt-4.1-nano-2025-04-14](https://www.orcarouter.ai/models/openai/gpt-4.1-nano-2025-04-14) — openai/gpt-4.1-nano-2025-04-14 by openai: $0.10/M input, $0.40/M output, 1M context, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-4.1-nano-2025-04-14 ## OpenAI: GPT-5.2 - [OpenAI: GPT-5.2](https://www.orcarouter.ai/models/openai/gpt-5.2) — OpenAI: GPT-5.2 by openai: $1.75/M input, $14.00/M output, 400K context, available via OrcaRouter API. GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.2 ## openai/gpt-5.2-chat-latest - [openai/gpt-5.2-chat-latest](https://www.orcarouter.ai/models/openai/gpt-5.2-chat-latest) — openai/gpt-5.2-chat-latest by openai: $1.75/M input, $14.00/M output, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.2-chat-latest ## OpenAI: GPT-5.5 - [OpenAI: GPT-5.5](https://www.orcarouter.ai/models/openai/gpt-5.5) — OpenAI: GPT-5.5 by openai: $5.00/M input, $30.00/M output, p50 8396ms, available via OrcaRouter API. GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.5 ## openai/gpt-5.5-pro-2026-04-23 - [openai/gpt-5.5-pro-2026-04-23](https://www.orcarouter.ai/models/openai/gpt-5.5-pro-2026-04-23) — openai/gpt-5.5-pro-2026-04-23 by openai: $30.00/M input, $180.00/M output, p50 450ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.5-pro-2026-04-23 ## OpenAI: GPT-5.6 Sol - [OpenAI: GPT-5.6 Sol](https://www.orcarouter.ai/models/openai/gpt-5.6-sol) — OpenAI: GPT-5.6 Sol by openai: $4.00/M input, $20.00/M output, 1M context, available via OrcaRouter API. GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series — the tier built for the hardest work: deep multi-step reasoning, large-scale software engineering, and long-horizon agentic workflows. It is especially strong at command-line and multi-file coding tasks, planning and executing across many tool calls while staying coherent over a 1.05M-token context window, and can emit up to 128K output tokens in a single response. It accepts text, image, and file inputs with text output, and exposes configurable reasoning effort so callers can trade latency and cost against depth per request. As a first-class OpenAI Responses model it plugs directly into agent frameworks, structured-output pipelines, and tool-calling loops. Use Sol when correctness on complex, high-value tasks matters more than cost — production coding agents, research and analysis, and multi-step automation that must not drift. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.6-sol ## DeepSeek: DeepSeek V4 Flash - [DeepSeek: DeepSeek V4 Flash](https://www.orcarouter.ai/models/deepseek/deepseek-v4-flash) — DeepSeek: DeepSeek V4 Flash by deepseek: $0.22/M input, $0.66/M output, 1M context, p50 787ms, available via OrcaRouter API. DeepSeek V4 Flash efficient MoE — 284B total / 13B active params, 1M context, optimized for fast everyday workloads. Canonical URL: https://www.orcarouter.ai/models/deepseek/deepseek-v4-flash ## DeepSeek: DeepSeek V4 Flash 0731 - [DeepSeek: DeepSeek V4 Flash 0731](https://www.orcarouter.ai/models/deepseek/deepseek-v4-flash-0731) — DeepSeek: DeepSeek V4 Flash 0731 by deepseek: $0.22/M input, $0.66/M output, 1M context, p50 3271ms, available via OrcaRouter API. DeepSeek V4 Flash 0731 is the re-post-trained official release of DeepSeek's efficient V4 Flash mixture-of-experts model — 284B total / 13B active parameters, identical in architecture and size to the April V4 Flash release, with the post-training redone from scratch. It serves a 1M-token context window with up to 384K output tokens, supports both thinking and non-thinking modes (thinking on by default) plus JSON output and tool calls, and natively speaks the Responses API with a targeted adaptation for Codex-style coding agents. The 0731 revision is a major step up in agent capability, outscoring even DeepSeek V4 Pro Preview on DeepSeek's published agent benchmarks — 82.7 on Terminal Bench 2.1, 76.7 on Cybergym, 70.3 on Toolathlon Verified, 54.4 on DeepSWE, and 54.2 on NL2Repo. On the DeepSeek first-party API this revision is what the deepseek-v4-flash model id now serves; the dated identifier lists the same revision for callers that reference it by date. It is a strong pick for coding agents, terminal and tool-use workloads, and high-volume agentic pipelines at flash-tier cost. Canonical URL: https://www.orcarouter.ai/models/deepseek/deepseek-v4-flash-0731 ## google/gemini-2.5-flash-preview-tts - [google/gemini-2.5-flash-preview-tts](https://www.orcarouter.ai/models/google/gemini-2.5-flash-preview-tts) — google/gemini-2.5-flash-preview-tts by google: $0.50/M input, $10.00/M output, p50 5991ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/google/gemini-2.5-flash-preview-tts ## google/imagen-4.0-fast-generate-001 - [google/imagen-4.0-fast-generate-001](https://www.orcarouter.ai/models/google/imagen-4.0-fast-generate-001) — google/imagen-4.0-fast-generate-001 by google: available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/google/imagen-4.0-fast-generate-001 ## xAI: Grok 4.5 - [xAI: Grok 4.5](https://www.orcarouter.ai/models/grok/grok-4.5) — xAI: Grok 4.5 by grok: $2.00/M input, $6.00/M output, 500K context, available via OrcaRouter API. Grok 4.5 is xAI's flagship model — its smartest to date, with frontier performance across coding, knowledge work, and STEM. Built on the 1.5-trillion-parameter V9 foundation and trained alongside the Cursor coding editor, it serves a 500K-token context window and accepts text, image, and file inputs with text output. It emphasizes strong agentic coding with notable token efficiency — resolving software-engineering tasks in far fewer output tokens than comparable frontier models — and is priced aggressively for high-volume production use. Canonical URL: https://www.orcarouter.ai/models/grok/grok-4.5 ## Kling: Kling 3.0 Turbo - [Kling: Kling 3.0 Turbo](https://www.orcarouter.ai/models/kling/kling-3-turbo) — Kling: Kling 3.0 Turbo by kling: available via OrcaRouter API. Kling 3.0 Turbo is Kuaishou's speed-optimized model in the Kling 3.0 generation, built for fast, cost-efficient video with native audio bundled in. It handles both text-to-video and image-to-video with strong prompt adherence, stable motion, and multi-shot storyboarding (up to six shots in a single generation), and brings a standout lip-sync improvement for natural talking-head and dialogue clips — all at noticeably faster output than the Standard and Pro variants in the same generation. It outputs 720p or 1080p, produces 3–15 second clips, and bills per second with audio included. Canonical URL: https://www.orcarouter.ai/models/kling/kling-3-turbo ## OpenAI: GPT-3.5 Turbo - [OpenAI: GPT-3.5 Turbo](https://www.orcarouter.ai/models/openai/gpt-3.5-turbo) — OpenAI: GPT-3.5 Turbo by openai: $0.50/M input, $1.50/M output, 16K context, p50 828ms, available via OrcaRouter API. GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-3.5-turbo ## OpenAI: GPT-4o-mini - [OpenAI: GPT-4o-mini](https://www.orcarouter.ai/models/openai/gpt-4o-mini) — OpenAI: GPT-4o-mini by openai: $0.15/M input, $0.60/M output, 128K context, p50 539ms, available via OrcaRouter API. GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-4o-mini ## Google: Gemini 3.1 Flash Lite Preview - [Google: Gemini 3.1 Flash Lite Preview](https://www.orcarouter.ai/models/google/gemini-3.1-flash-lite-preview) — Google: Gemini 3.1 Flash Lite Preview by google: $0.25/M input, $1.50/M output, 1M context, p50 7414ms, available via OrcaRouter API. Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall quality and approaches Gemini 2.5 Flash performance across... Canonical URL: https://www.orcarouter.ai/models/google/gemini-3.1-flash-lite-preview ## DeepSeek: DeepSeek V3 - [DeepSeek: DeepSeek V3](https://www.orcarouter.ai/models/deepseek/deepseek-chat) — DeepSeek: DeepSeek V3 by deepseek: $0.22/M input, $0.66/M output, 1M context, p50 362ms, available via OrcaRouter API. DeepSeek alias for V4 Flash non-thinking mode — 1M context, strong instruction following and coding (legacy alias, slated for deprecation). Canonical URL: https://www.orcarouter.ai/models/deepseek/deepseek-chat ## Google: Gemini 3.1 Pro Preview Custom Tools - [Google: Gemini 3.1 Pro Preview Custom Tools](https://www.orcarouter.ai/models/google/gemini-3.1-pro-preview-customtools) — Google: Gemini 3.1 Pro Preview Custom Tools by google: $4.00/M input, $18.00/M output, 1M context, p50 35534ms, available via OrcaRouter API. Gemini 3.1 Pro Preview Custom Tools is a variant of Gemini 3.1 Pro that improves tool selection behavior by preventing overuse of a general bash tool when more efficient third-party... Canonical URL: https://www.orcarouter.ai/models/google/gemini-3.1-pro-preview-customtools ## openai/gpt-4.1-2025-04-14 - [openai/gpt-4.1-2025-04-14](https://www.orcarouter.ai/models/openai/gpt-4.1-2025-04-14) — openai/gpt-4.1-2025-04-14 by openai: $2.00/M input, $8.00/M output, 1M context, p50 379ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-4.1-2025-04-14 ## OpenAI: GPT-5.4 Nano - [OpenAI: GPT-5.4 Nano](https://www.orcarouter.ai/models/openai/gpt-5.4-nano) — OpenAI: GPT-5.4 Nano by openai: $0.20/M input, $1.25/M output, 400K context, available via OrcaRouter API. GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.4-nano ## openai/gpt-5.4-pro-2026-03-05 - [openai/gpt-5.4-pro-2026-03-05](https://www.orcarouter.ai/models/openai/gpt-5.4-pro-2026-03-05) — openai/gpt-5.4-pro-2026-03-05 by openai: $30.00/M input, $180.00/M output, 1M context, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.4-pro-2026-03-05 ## openai/gpt-5.5-2026-04-23 - [openai/gpt-5.5-2026-04-23](https://www.orcarouter.ai/models/openai/gpt-5.5-2026-04-23) — openai/gpt-5.5-2026-04-23 by openai: $5.00/M input, $30.00/M output, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.5-2026-04-23 ## OpenAI: GPT-6 Sol - [OpenAI: GPT-6 Sol](https://www.orcarouter.ai/models/openai/gpt-6-sol) — OpenAI: GPT-6 Sol by openai: $2.00/M input, $10.00/M output, 1M context, p50 501ms, available via OrcaRouter API. GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier. It is suited for demanding professional and agentic work. Multimodal input (text, image, file), 1.05M-token context, configurable reasoning effort, tools and structured outputs. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-6-sol ## kling/kling-v2-master - [kling/kling-v2-master](https://www.orcarouter.ai/models/kling/kling-v2-master) — kling/kling-v2-master by kling: available via OrcaRouter API. Kling 2.0 Master — early flagship text-to-video and image-to-video, 5–10s clips, 720p, 24fps. Canonical URL: https://www.orcarouter.ai/models/kling/kling-v2-master ## kling/kling-v3-omni - [kling/kling-v3-omni](https://www.orcarouter.ai/models/kling/kling-v3-omni) — kling/kling-v3-omni by kling: available via OrcaRouter API. Kling 3.0 Omni — unified text-to-video and image-to-video API with multi-shot, subject control and video-reference, 3–15s clips, up to native 4K. Canonical URL: https://www.orcarouter.ai/models/kling/kling-v3-omni ## Meta: Muse Spark 1.1 - [Meta: Muse Spark 1.1](https://www.orcarouter.ai/models/meta/muse-spark-1.1) — Meta: Muse Spark 1.1 by meta: $1.25/M input, $4.25/M output, 1M context, p50 3847ms, available via OrcaRouter API. Muse Spark 1.1 is Meta's multimodal reasoning model, built for agentic tasks. It accepts an unusually broad range of inputs — text, images, video, audio, and PDF documents — and returns text, reasoning natively across modalities within a 1M-token context window. It supports configurable reasoning effort and native tool calling, making it well suited for multimodal agents, deep research over mixed media, and long-context analysis. With efficient inference and a broad input surface, Muse Spark 1.1 targets workloads that combine documents, screenshots, recordings, and video with text — from analyzing long reports and media libraries to driving multi-step agentic pipelines that must reason over more than just text. Canonical URL: https://www.orcarouter.ai/models/meta/muse-spark-1.1 ## OpenAI: GPT-3.5 Turbo 16k - [OpenAI: GPT-3.5 Turbo 16k](https://www.orcarouter.ai/models/openai/gpt-3.5-turbo-16k) — OpenAI: GPT-3.5 Turbo 16k by openai: $3.00/M input, $4.00/M output, 16K context, p50 9221ms, available via OrcaRouter API. This model offers four times the context length of gpt-3.5-turbo, allowing it to support approximately 20 pages of text in a single request at a higher cost. Training data: up... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-3.5-turbo-16k ## OpenAI: GPT-6 Luna - [OpenAI: GPT-6 Luna](https://www.orcarouter.ai/models/openai/gpt-6-luna) — OpenAI: GPT-6 Luna by openai: $0.10/M input, $0.50/M output, 1M context, available via OrcaRouter API. GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive workloads such as chat, classification, and lightweight agentic tasks. Multimodal input (text, image, file), 1.05M-token context, configurable reasoning effort, tools and structured outputs. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-6-luna ## openai/text-embedding-3-small - [openai/text-embedding-3-small](https://www.orcarouter.ai/models/openai/text-embedding-3-small) — openai/text-embedding-3-small by openai: $0.02/M input, $0.02/M output, p50 119ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/text-embedding-3-small ## Z.ai: GLM 4.5 - [Z.ai: GLM 4.5](https://www.orcarouter.ai/models/z-ai/glm-4.5) — Z.ai: GLM 4.5 by z-ai: $0.60/M input, $2.20/M output, 128K context, p50 2000ms, available via OrcaRouter API. Zhipu (Z.ai) flagship open-source MoE: 355B total / 32B active. Hybrid reasoning (thinking / non-thinking modes), native tool calling and agentic surface, 128K context. Canonical URL: https://www.orcarouter.ai/models/z-ai/glm-4.5 ## Anthropic: Claude Sonnet 5.5 - [Anthropic: Claude Sonnet 5.5](https://www.orcarouter.ai/models/anthropic/claude-sonnet-5.5) — Anthropic: Claude Sonnet 5.5 by anthropic: $2.00/M input, $10.00/M output, 1M context, available via OrcaRouter API. Claude Sonnet 5.5 is the current Sonnet-class model from Anthropic and a direct upgrade over Claude Sonnet 5, aimed at well-scoped everyday work: building features, fixing bugs, and producing reliable output on agentic coding and enterprise tasks. It pairs a 1M-token context window with 128K max output and accepts text, image, and file input. Reasoning is always on, with five effort levels (low through max, default high) that trade thoroughness against token spend. The effort levels were recalibrated relative to Sonnet 5, so an existing integration should re-tune them rather than carry its old setting over. Native tool calling, structured outputs, and response_format are all supported. It speaks the Anthropic Messages API natively and is also reachable through the OpenAI chat-completions format, making it a drop-in upgrade for existing Claude integrations at the same price as Sonnet 5. Canonical URL: https://www.orcarouter.ai/models/anthropic/claude-sonnet-5.5 ## OpenAI: GPT-5.2-Codex - [OpenAI: GPT-5.2-Codex](https://www.orcarouter.ai/models/openai/gpt-5.2-codex) — OpenAI: GPT-5.2-Codex by openai: $1.75/M input, $14.00/M output, 400K context, available via OrcaRouter API. GPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks.... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.2-codex ## openai/tts-1-hd - [openai/tts-1-hd](https://www.orcarouter.ai/models/openai/tts-1-hd) — openai/tts-1-hd by openai: $30.00/M characters input, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/tts-1-hd ## Qwen: Qwen3 VL 8B Instruct - [Qwen: Qwen3 VL 8B Instruct](https://www.orcarouter.ai/models/qwen/qwen3-vl-8b-instruct) — Qwen: Qwen3 VL 8B Instruct by qwen: $0.18/M input, $0.70/M output, 131K context, p50 20755ms, available via OrcaRouter API. Qwen3-VL 8B Instruct — open-weight small vision-language model, 8B params, 128k context, no thinking mode. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3-vl-8b-instruct ## qwen/qwen3.5-flash-2026-02-23 - [qwen/qwen3.5-flash-2026-02-23](https://www.orcarouter.ai/models/qwen/qwen3.5-flash-2026-02-23) — qwen/qwen3.5-flash-2026-02-23 by qwen: $0.10/M input, $0.40/M output, 1M context, p50 1061ms, available via OrcaRouter API. Qwen3.5 Flash snapshot 2026-02-23 — frozen version of qwen3.5-flash, same capabilities. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3.5-flash-2026-02-23 ## qwen/qwen3.5-plus - [qwen/qwen3.5-plus](https://www.orcarouter.ai/models/qwen/qwen3.5-plus) — qwen/qwen3.5-plus by qwen: $0.40/M input, $2.40/M output, 1M context, p50 1207ms, available via OrcaRouter API. Qwen3.5 Plus — multimodal chat (text/image/video), 1M context, strong coding + agent capability. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3.5-plus ## Z.ai: GLM 5 - [Z.ai: GLM 5](https://www.orcarouter.ai/models/z-ai/glm-5) — Z.ai: GLM 5 by z-ai: $1.00/M input, $3.20/M output, 200K context, p50 5805ms, available via OrcaRouter API. Next-generation Zhipu flagship with multiple thinking modes and strong tool calling. 200K context / 128K max output. Canonical URL: https://www.orcarouter.ai/models/z-ai/glm-5 ## deepseek/deepseek-reasoner - [deepseek/deepseek-reasoner](https://www.orcarouter.ai/models/deepseek/deepseek-reasoner) — deepseek/deepseek-reasoner by deepseek: $0.24/M input, $0.73/M output, 1M context, p50 371ms, available via OrcaRouter API. DeepSeek alias for V4 Flash thinking mode — 1M context, open chain-of-thought reasoning (legacy alias, slated for deprecation). Canonical URL: https://www.orcarouter.ai/models/deepseek/deepseek-reasoner ## MoonshotAI: Kimi K3 - [MoonshotAI: Kimi K3](https://www.orcarouter.ai/models/kimi/kimi-k3) — MoonshotAI: Kimi K3 by kimi: $3.00/M input, $15.00/M output, 1M context, p50 8398ms, available via OrcaRouter API. Kimi K3 is Moonshot AI's flagship model and its most capable release to date — a 2.8-trillion-parameter Mixture-of-Experts model built for long-horizon coding and end-to-end knowledge work. It pairs a 1M-token context window with native visual understanding, accepting text and image input with text output, and is designed to stay coherent across very long agentic sessions. Moonshot positions K3 for programming-agent scenarios such as Codex, Claude Code, Cline, and RooCode, as well as deep reasoning and knowledge work. It exposes a top-level reasoning_effort control and native tool calling, and speaks the OpenAI API format for drop-in integration. Note that K3 does not accept sampling parameters (no temperature / top_p / seed) — reasoning depth is controlled through reasoning_effort instead. Canonical URL: https://www.orcarouter.ai/models/kimi/kimi-k3 ## MiniMax M2.5 highspeed - [MiniMax M2.5 highspeed](https://www.orcarouter.ai/models/minimax/minimax-m2.5-highspeed) — MiniMax M2.5 highspeed by minimax: $0.60/M input, $2.40/M output, 205K context, p50 1218ms, available via OrcaRouter API. MiniMax M2.5 high-speed — same model + same 200k context as M2.5, faster output (~100 tps vs ~60 tps). Canonical URL: https://www.orcarouter.ai/models/minimax/minimax-m2.5-highspeed ## MiniMax: MiniMax M3 - [MiniMax: MiniMax M3](https://www.orcarouter.ai/models/minimax/minimax-m3) — MiniMax: MiniMax M3 by minimax: $0.30/M input, $1.20/M output, 1M context, p50 13755ms, available via OrcaRouter API. MiniMax-M3 is MiniMax's flagship open-weight foundation model and the first to combine three frontier capabilities at once: frontier-level coding and agentic performance, a 1M-token context window, and native multimodality. It accepts text, image, and video inputs with text output, and is powered by the proprietary MiniMax Sparse Attention (MSA) architecture, which sustains up to 1M tokens of context (with a guaranteed minimum of 512K) - the foundation for long-range agent tasks, long-horizon coding, and long-video understanding. Multimodality is a native core capability rather than an add-on: the data pipeline was rebuilt to scale pretraining to 100T+ tokens with multimodal training from step zero, deeply aligning textual and visual semantic spaces. M3 achieves top-tier results across coding and agentic benchmarks spanning software engineering, terminal execution, and autonomous browsing (scoring 83.5 on BrowseComp), with autonomous task decomposition, tool invocation, and multi-step reasoning. It is well suited to AI coding assistants, automated workflows, and long-running asynchronous agent pipelines where coherence over extended sessions matters. Canonical URL: https://www.orcarouter.ai/models/minimax/minimax-m3 ## openai/gpt-3.5-turbo-0125 - [openai/gpt-3.5-turbo-0125](https://www.orcarouter.ai/models/openai/gpt-3.5-turbo-0125) — openai/gpt-3.5-turbo-0125 by openai: $0.50/M input, $1.50/M output, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-3.5-turbo-0125 ## openai/gpt-5.1-2025-11-13 - [openai/gpt-5.1-2025-11-13](https://www.orcarouter.ai/models/openai/gpt-5.1-2025-11-13) — openai/gpt-5.1-2025-11-13 by openai: $1.25/M input, $10.00/M output, 400K context, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.1-2025-11-13 ## OpenAI: GPT-6 Astra - [OpenAI: GPT-6 Astra](https://www.orcarouter.ai/models/openai/gpt-6-astra) — OpenAI: GPT-6 Astra by openai: $10.00/M input, $50.00/M output, 1M context, available via OrcaRouter API. GPT-6 Astra is OpenAI's flagship model for demanding, long-horizon end-to-end work — advanced analysis, software engineering, deep research, scientific work, and document creation. It pairs a 1M-token context window with multimodal input (text, images, files) and always-on reasoning, with configurable effort from low through max for latency/quality trade-offs. It is a native tool-use and structured-output model: function calling, response_format / structured outputs, seed, and web search as a tool are all supported. Astra speaks both the OpenAI chat-completions format and the native Responses API, so it drops into existing OpenAI-compatible integrations and returns full reasoning traces where the Responses surface is used. On Artificial Analysis it posts an Intelligence Index of 54.7, a Coding Index of 76.9, and an Agentic Index of 51.6. Note tiered pricing: requests above 272K prompt tokens bill at the long-context rate. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-6-astra ## qwen/qwen3.6-flash-2026-04-16 - [qwen/qwen3.6-flash-2026-04-16](https://www.orcarouter.ai/models/qwen/qwen3.6-flash-2026-04-16) — qwen/qwen3.6-flash-2026-04-16 by qwen: $0.25/M input, $1.50/M output, 1M context, p50 2679ms, available via OrcaRouter API. Qwen3.6 Flash snapshot 2026-04-16 — frozen version of qwen3.6-flash, same capabilities. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3.6-flash-2026-04-16 ## Google: Gemini 3 Flash Preview - [Google: Gemini 3 Flash Preview](https://www.orcarouter.ai/models/google/gemini-3-flash-preview) — Google: Gemini 3 Flash Preview by google: $0.50/M input, $3.00/M output, 1M context, p50 10890ms, available via OrcaRouter API. Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool... Canonical URL: https://www.orcarouter.ai/models/google/gemini-3-flash-preview ## Gemini 3.5 Flash - [Gemini 3.5 Flash](https://www.orcarouter.ai/models/google/gemini-3.5-flash) — Gemini 3.5 Flash by google: $1.50/M input, $9.00/M output, 1M context, p50 5263ms, available via OrcaRouter API. Google Gemini 3.5 Flash — Google's strongest Flash-tier model, sitting just under 3.1-pro-preview on intelligence (AA Intel 55.3 vs 57.2) at Flash latency. 1M token context window (1,048,576), 64K max output (65,536). Full multimodal input across text, image, audio, video, and file; text output only. AA Coding 45, GPQA Diamond 92.2%, IFBench 76.3% (top-tier instruction-following), Long-Context Recall 69.3%, SciCode 53.1%, τ²-Bench 95.3% (agentic tool-use), HLE 41%. Supports reasoning (include_reasoning / reasoning), structured outputs, tool calls, response_format, seed, stop, and standard sampling controls. Endpoints: chat_completions, gemini_generate. Best fit: high-volume agentic loops, long-document analysis, and multimodal extraction where Pro-tier latency or cost is overkill. Canonical URL: https://www.orcarouter.ai/models/google/gemini-3.5-flash ## google/gemini-embedding-001 - [google/gemini-embedding-001](https://www.orcarouter.ai/models/google/gemini-embedding-001) — google/gemini-embedding-001 by google: $0.15/M input, p50 277ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/google/gemini-embedding-001 ## google/gemini-robotics-er-1.6-preview - [google/gemini-robotics-er-1.6-preview](https://www.orcarouter.ai/models/google/gemini-robotics-er-1.6-preview) — google/gemini-robotics-er-1.6-preview by google: $1.00/M input, $5.00/M output, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/google/gemini-robotics-er-1.6-preview ## Qwen3.8 27B - [Qwen3.8 27B](https://www.orcarouter.ai/models/obsidian/qwen3.8-27b) — Qwen3.8 27B by obsidian: $0.40/M input, $4.21/M output, 262K context, p50 2827ms, available via OrcaRouter API. Qwen3.8 27B served in block-FP8 for higher throughput and lower memory footprint, with the vision tower kept at full precision. Preserves the original model's reasoning, coding, multilingual performance, and tool use. Designed to provide direct, complete responses across a wide range of prompts. The model may occasionally append brief informational disclaimers inherited from the base model's training. Ideal for AI research, security testing, red teaming, agent development, coding assistants, and other advanced AI applications that benefit from maximum output flexibility. Canonical URL: https://www.orcarouter.ai/models/obsidian/qwen3.8-27b ## OpenAI: GPT-5.1 - [OpenAI: GPT-5.1](https://www.orcarouter.ai/models/openai/gpt-5.1) — OpenAI: GPT-5.1 by openai: $1.25/M input, $10.00/M output, 400K context, available via OrcaRouter API. GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.1 ## OpenAI: GPT-5.1-Codex - [OpenAI: GPT-5.1-Codex](https://www.orcarouter.ai/models/openai/gpt-5.1-codex) — OpenAI: GPT-5.1-Codex by openai: $1.25/M input, $10.00/M output, 400K context, available via OrcaRouter API. GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks.... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.1-codex ## openai/text-embedding-ada-002 - [openai/text-embedding-ada-002](https://www.orcarouter.ai/models/openai/text-embedding-ada-002) — openai/text-embedding-ada-002 by openai: $0.10/M input, $0.10/M output, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/text-embedding-ada-002 ## Z.ai: GLM 5.3 - [Z.ai: GLM 5.3](https://www.orcarouter.ai/models/z-ai/glm-5.3) — Z.ai: GLM 5.3 by z-ai: $1.26/M input, $3.96/M output, 1M context, p50 3688ms, available via OrcaRouter API. GLM-5.3 is Z.ai (Zhipu AI)'s latest flagship model for complex software engineering and long-horizon agentic tasks. It delivers roughly a 50% improvement in coding experience over GLM-5.2, matches Mythos 5 on selected cybersecurity capabilities, and strikes a better balance between raw performance and token efficiency. It is a text-in / text-out model built for repo-scale coding, autonomous multi-step engineering, and agent workflows that must stay coherent over long horizons. GLM-5.3 uses the same API surface as the GLM-5 line with two changes callers must handle: thinking is always on (thinking.type only accepts enabled; passing disabled now fails the request), and reasoning depth is controlled by reasoning_effort with values low / high / max, defaulting to max. It supports native tool calling and structured JSON output, and speaks the OpenAI-compatible chat-completions format. Canonical URL: https://www.orcarouter.ai/models/z-ai/glm-5.3 ## google/gemini-flash-lite-latest - [google/gemini-flash-lite-latest](https://www.orcarouter.ai/models/google/gemini-flash-lite-latest) — google/gemini-flash-lite-latest by google: $0.25/M input, $1.50/M output, p50 977ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/google/gemini-flash-lite-latest ## OpenAI: GPT-5.4 - [OpenAI: GPT-5.4](https://www.orcarouter.ai/models/openai/gpt-5.4) — OpenAI: GPT-5.4 by openai: $2.50/M input, $15.00/M output, 1M context, available via OrcaRouter API. GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for... Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.4 ## openai/gpt-image-1 - [openai/gpt-image-1](https://www.orcarouter.ai/models/openai/gpt-image-1) — openai/gpt-image-1 by openai: $5.00/M input, $40.00/M output, p50 7000ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-image-1 ## OrcaDub: OrcaDub 1.0 - [OrcaDub: OrcaDub 1.0](https://www.orcarouter.ai/models/orca/dub) — OrcaDub: OrcaDub 1.0 by orca: available via OrcaRouter API. OrcaDub 1.0 is an end-to-end, prosody-aware dubbing and video-translation model. Given a source video, it produces a dubbed video in a target language in the original speakers' cloned voices, preserving delivery — emphasis, pauses, intonation — and word-level timing. It covers 28 languages in any direction, clones a separate voice per speaker in multi-speaker footage, preserves background music and ambience, and aligns the dub to the source at word level. It is reachable from the free web Studio and from an OpenAI-compatible API. Canonical URL: https://www.orcarouter.ai/models/orca/dub ## Qwen: Qwen3.5-35B-A3B - [Qwen: Qwen3.5-35B-A3B](https://www.orcarouter.ai/models/qwen/qwen3.5-35b-a3b) — Qwen: Qwen3.5-35B-A3B by qwen: $0.06/M input, $0.46/M output, 33K context, p50 1395ms, available via OrcaRouter API. Qwen3.5 35B-A3B — open-weight MoE multimodal (text/image/video), 35B total / 3B active params, 32k context (vision mode). Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3.5-35b-a3b ## qwen/qwen3.5-flash - [qwen/qwen3.5-flash](https://www.orcarouter.ai/models/qwen/qwen3.5-flash) — qwen/qwen3.5-flash by qwen: $0.10/M input, $0.40/M output, 1M context, p50 15370ms, available via OrcaRouter API. Qwen3.5 Flash — multimodal chat (text/image/video) optimized for cost, 1M context. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3.5-flash ## Anthropic: Claude Fable 5.1 - [Anthropic: Claude Fable 5.1](https://www.orcarouter.ai/models/anthropic/claude-fable-5.1) — Anthropic: Claude Fable 5.1 by anthropic: $10.00/M input, $50.00/M output, 1M context, p50 3191ms, available via OrcaRouter API. Claude Fable 5.1 is Anthropic's Mythos-class model — a capability tier above the Opus class — made safe for broad use, and the successor to Claude Fable 5. It improves on Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual engineering, and multi-step tasks where sustained coherence and judgment matter most. It accepts text, image, and file inputs with text output, serves a 1M-token context window with up to 128K output tokens, and supports adaptive reasoning and structured outputs. Fable 5.1 keeps the autonomous, long-horizon posture of the Fable line: it stays coherent across millions of tokens, improves its own outputs using file-based memory, and completes complex multi-step work with far less scaffolding than prior models. It is a strong default for AI coding assistants, deep research and analysis pipelines, and long-running autonomous agents. Canonical URL: https://www.orcarouter.ai/models/anthropic/claude-fable-5.1 ## Google: Gemini 2.5 Flash Lite - [Google: Gemini 2.5 Flash Lite](https://www.orcarouter.ai/models/google/gemini-2.5-flash-lite) — Google: Gemini 2.5 Flash Lite by google: $0.10/M input, $0.40/M output, 1M context, p50 1204ms, available via OrcaRouter API. Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance... Canonical URL: https://www.orcarouter.ai/models/google/gemini-2.5-flash-lite ## DeepSeek: DeepSeek V4 Flash Vision (Exp) - [DeepSeek: DeepSeek V4 Flash Vision (Exp)](https://www.orcarouter.ai/models/deepseek/deepseek-v4-flash-vision-exp) — DeepSeek: DeepSeek V4 Flash Vision (Exp) by deepseek: $0.22/M input, $0.66/M output, 1M context, p50 765ms, available via OrcaRouter API. Experimental vision-enabled variant of DeepSeek V4 Flash: text + image in, text out, 1M context, 384K max output, thinking and non-thinking modes. Pure-text capability on par with V4 Flash; multimodal agent capability approaching Opus 4.8. Canonical URL: https://www.orcarouter.ai/models/deepseek/deepseek-v4-flash-vision-exp ## MoonshotAI: Kimi K2.7 Code - [MoonshotAI: Kimi K2.7 Code](https://www.orcarouter.ai/models/kimi/kimi-k2.7-code) — MoonshotAI: Kimi K2.7 Code by kimi: $0.95/M input, $4.00/M output, 262K context, p50 2475ms, available via OrcaRouter API. Kimi K2.7 Code is Moonshot AI's strongest coding model to date - a coding-focused agentic model built on Kimi K2.6, with a 1T-parameter Mixture-of-Experts architecture (32B active per pass) and a 256K-token context window. It delivers substantial gains on real-world long-horizon coding: stronger end-to-end task completion across complex software-engineering workflows, while cutting thinking-token usage by roughly 30% versus K2.6 for better token efficiency. The model is natively multimodal via the MoonViT vision encoder, accepting text, image, and video inputs with text output, and runs with thinking always on for deep agentic reasoning. It is purpose-built for autonomous coding agents and long-horizon software work - large multi-file changes, production-grade engineering tasks, and multi-step tool use - and integrates cleanly with agent harnesses and coding-agent frameworks. Canonical URL: https://www.orcarouter.ai/models/kimi/kimi-k2.7-code ## MiniMax: MiniMax M2.7 - [MiniMax: MiniMax M2.7](https://www.orcarouter.ai/models/minimax/minimax-m2.7) — MiniMax: MiniMax M2.7 by minimax: $0.30/M input, $1.20/M output, 205K context, p50 1178ms, available via OrcaRouter API. MiniMax M2.7 — next-gen agentic LLM optimized for autonomous workflows and continuous self-improvement, 200k context, ~60 tps output. Canonical URL: https://www.orcarouter.ai/models/minimax/minimax-m2.7 ## openai/gpt-4.1-mini-2025-04-14 - [openai/gpt-4.1-mini-2025-04-14](https://www.orcarouter.ai/models/openai/gpt-4.1-mini-2025-04-14) — openai/gpt-4.1-mini-2025-04-14 by openai: $0.40/M input, $1.60/M output, 1M context, p50 282ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-4.1-mini-2025-04-14 ## openai/gpt-5.2-2025-12-11 - [openai/gpt-5.2-2025-12-11](https://www.orcarouter.ai/models/openai/gpt-5.2-2025-12-11) — openai/gpt-5.2-2025-12-11 by openai: $1.75/M input, $14.00/M output, 400K context, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5.2-2025-12-11 ## Qwen: Qwen3 Max - [Qwen: Qwen3 Max](https://www.orcarouter.ai/models/qwen/qwen3-max) — Qwen: Qwen3 Max by qwen: $0.36/M input, $1.43/M output, 262K context, p50 2152ms, available via OrcaRouter API. Qwen3 Max — proprietary flagship chat model, 256k context, thinking mode + function calling. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3-max ## Google: Gemini 3.8 Flash - [Google: Gemini 3.8 Flash](https://www.orcarouter.ai/models/google/gemini-3.8-flash) — Google: Gemini 3.8 Flash by google: $0.75/M input, $3.75/M output, 1M context, p50 4118ms, available via OrcaRouter API. Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning. Canonical URL: https://www.orcarouter.ai/models/google/gemini-3.8-flash ## kling/kling-video-o1 - [kling/kling-video-o1](https://www.orcarouter.ai/models/kling/kling-video-o1) — kling/kling-video-o1 by kling: available via OrcaRouter API. Kling O1 — next-gen video generation with text-to-video, image-to-video, subject control and video-reference editing, 5–10s clips, std/pro modes. Canonical URL: https://www.orcarouter.ai/models/kling/kling-video-o1 ## openai/gpt-5-nano-2025-08-07 - [openai/gpt-5-nano-2025-08-07](https://www.orcarouter.ai/models/openai/gpt-5-nano-2025-08-07) — openai/gpt-5-nano-2025-08-07 by openai: $0.05/M input, $0.40/M output, 400K context, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5-nano-2025-08-07 ## openai/gpt-5-search-api-2025-10-14 - [openai/gpt-5-search-api-2025-10-14](https://www.orcarouter.ai/models/openai/gpt-5-search-api-2025-10-14) — openai/gpt-5-search-api-2025-10-14 by openai: $1.25/M input, $10.00/M output, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/gpt-5-search-api-2025-10-14 ## Orca: OrcaVerify Text 1.0 - [Orca: OrcaVerify Text 1.0](https://www.orcarouter.ai/models/orca/orcaverify-text1.0) — Orca: OrcaVerify Text 1.0 by orca: $2.00/M input, p50 115ms, available via OrcaRouter API. AI-generated-text detection over the OpenAI Chat Completions format: four calibrated verdicts (AI_GENERATED / AI_ASSISTED / HUMAN / ABSTAIN) with a calibrated probability and paragraph-level localization. English, Chinese, Japanese and Korean. Canonical URL: https://www.orcarouter.ai/models/orca/orcaverify-text1.0 ## Qwen: Qwen3 VL 235B A22B Thinking - [Qwen: Qwen3 VL 235B A22B Thinking](https://www.orcarouter.ai/models/qwen/qwen3-vl-235b-a22b-thinking) — Qwen: Qwen3 VL 235B A22B Thinking by qwen: $0.40/M input, $4.00/M output, 131K context, p50 13328ms, available via OrcaRouter API. Qwen3-VL 235B-A22B Thinking — open-weight vision-language reasoning model, 235B total / 22B active params, 128k context. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3-vl-235b-a22b-thinking ## qwen/qwen3.5-plus-2026-02-15 - [qwen/qwen3.5-plus-2026-02-15](https://www.orcarouter.ai/models/qwen/qwen3.5-plus-2026-02-15) — qwen/qwen3.5-plus-2026-02-15 by qwen: $0.40/M input, $2.40/M output, 1M context, p50 1118ms, available via OrcaRouter API. Qwen3.5 Plus snapshot 2026-02-15 — frozen version of qwen3.5-plus, same capabilities. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3.5-plus-2026-02-15 ## Google: Gemini 3.6 Flash - [Google: Gemini 3.6 Flash](https://www.orcarouter.ai/models/google/gemini-3.6-flash) — Google: Gemini 3.6 Flash by google: $0.75/M input, $3.75/M output, 1M context, p50 2828ms, available via OrcaRouter API. Gemini 3.6 Flash is Google's latest fast, cost-efficient model in the Gemini 3 family — a natively multimodal model that reads text, images, video, audio, and files and responds in text, with a 1M-token context window and up to 64K output tokens. It is built for teams that need frontier-adjacent quality at Flash-tier speed and price, and is a strong default across coding agents, document and media understanding, retrieval-augmented generation, and high-throughput automation. Against the previous 3.5 Flash it is markedly more token-efficient and less verbose — reaching the same or better quality while emitting fewer output tokens, which directly lowers cost and latency on long agentic runs. It supports configurable reasoning effort (so callers can dial depth against speed per request), native tool calling, and structured outputs, and it holds up strongly on long-context and agentic-coding evaluations for its tier, including 91.8% on 128k-average multi-needle retrieval and competitive scores on SWE-Bench Pro, Terminal-Bench, and OSWorld. Gemini 3.6 Flash speaks both the OpenAI chat-completions format and Gemini's native generateContent API, making it a drop-in upgrade for existing Gemini and OpenAI-compatible integrations. Use it as the everyday workhorse when you want most of the intelligence of a frontier model without paying frontier prices. Canonical URL: https://www.orcarouter.ai/models/google/gemini-3.6-flash ## kimi/kimi-k2.6 - [kimi/kimi-k2.6](https://www.orcarouter.ai/models/kimi/kimi-k2.6) — kimi/kimi-k2.6 by kimi: $0.95/M input, $4.00/M output, 262K context, p50 3297ms, available via OrcaRouter API. Moonshot Kimi K2 Thinking — most advanced open reasoning model in the K2 series, agentic long-horizon tasks, 256k context. Canonical URL: https://www.orcarouter.ai/models/kimi/kimi-k2.6 ## openai/tts-1-1106 - [openai/tts-1-1106](https://www.orcarouter.ai/models/openai/tts-1-1106) — openai/tts-1-1106 by openai: $15.00/M characters input, p50 889ms, available via OrcaRouter API. Canonical URL: https://www.orcarouter.ai/models/openai/tts-1-1106 ## Qwen: Qwen3.6 Plus - [Qwen: Qwen3.6 Plus](https://www.orcarouter.ai/models/qwen/qwen3.6-plus) — Qwen: Qwen3.6 Plus by qwen: $0.28/M input, $1.65/M output, 1M context, p50 3561ms, available via OrcaRouter API. Qwen3.6 Plus — flagship multimodal chat (text/image/video), 1M context, Vibe Coding + function calling. Canonical URL: https://www.orcarouter.ai/models/qwen/qwen3.6-plus