Model catalogue
429 models from OpenRouter. Any of them can back a token.
GPT-6 Astra
OpenAI
GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...
- Context
- 1.1M
- In
- $10.00/M
- Out
- $50.00/M
Claude Fable 5.1
Anthropic
Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...
- Context
- 1.0M
- In
- $10.00/M
- Out
- $50.00/M
Gemma 4 26B A4B
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...
- Context
- 262K
- In
- $0.04/M
- Out
- $0.22/M
Grok 4.6
xAI
Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.
- Context
- 500K
- In
- $2.00/M
- Out
- $6.00/M
DeepSeek V4.1 Flash
DeepSeek
DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...
- Context
- 1.0M
- In
- $0.15/M
- Out
- $0.60/M
Llama Guard 4 12B
Meta
Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM...
- Context
- 164K
- In
- $0.18/M
- Out
- $0.18/M
Qwen3.8 Flash
Qwen
Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.
- Context
- 1.0M
- In
- $0.15/M
- Out
- $0.47/M
Mistral Medium 3.5
Mistral
Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...
- Context
- 262K
- In
- $1.50/M
- Out
- $7.50/M
Kimi K3
Moonshot
Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...
- Context
- 1.0M
- In
- $2.65/M
- Out
- $13.28/M
GLM 5.3 Flash
Z.ai
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
- Context
- 1.3M
- In
- $0.15/M
- Out
- $0.50/M
MiniMax M3
MiniMax
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
- Context
- 1.0M
- In
- $0.30/M
- Out
- $1.20/M
Nova Premier 1.0
Amazon
Amazon Nova Premier is the most capable of Amazon’s multimodal models for complex reasoning tasks and for use as the best teacher for distilling custom models.
- Context
- 1.0M
- In
- $2.50/M
- Out
- $12.50/M
Phi 4
Microsoft
[Microsoft Research](/microsoft) Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or where quick responses are needed. At 14 billion...
- Context
- 16K
- In
- $0.07/M
- Out
- $0.14/M
Nemotron 3.5 Lightning
NVIDIA
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...
- Context
- 262K
- In
- $0.08/M
- Out
- $0.20/M
Command A
Cohere
Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual, and coding use cases. Compared to other leading proprietary...
- Context
- 256K
- In
- $2.50/M
- Out
- $10.00/M
Sonar Pro Search
Perplexity
Exclusively available on the OpenRouter API, Sonar Pro's new Pro Search mode is Perplexity's most advanced agentic search system. It is designed for deeper reasoning and analysis. Pricing is based...
- Context
- 200K
- In
- $3.00/M
- Out
- $15.00/M
Seed 2.1 Turbo
ByteDance
Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step task execution, and understanding visual and...
- Context
- 262K
- In
- $0.50/M
- Out
- $2.50/M
Hy-MT2-1.8B
Tencent
Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided...
- Context
- 8K
- In
- $0.04/M
- Out
- $0.18/M
Schematron V2 Turbo
Inference.net
Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extraction instructions must be supplied through a JSON schema in response_format rather...
- Context
- 128K
- In
- $0.03/M
- Out
- $0.15/M
Fugu Ultra v2
Sakana
Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to...
- Context
- 1.0M
- In
- $5.00/M
- Out
- $30.00/M
Ling 3.0 Flash VL
Inclusion AI
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
- Context
- 131K
- In
- $0.06/M
- Out
- $0.18/M
Mercury 2.5
Inception
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...
- Context
- 260K
- In
- $0.04/M
- Out
- $0.15/M
Muse Spark 1.3 Contributor
Meta
Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...
- Context
- 1.0M
- In
- $0.10/M
- Out
- $0.20/M
Granite 4.2 8B
Ibm Granite
Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-step reasoning. It supports full, low-effort,...
- Context
- 131K
- In
- $0.06/M
- Out
- $0.25/M
Solar Pro 4
Upstage
Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with strong capabilities in office productivity, document-intensive...
- Context
- 524K
- In
- $0.09/M
- Out
- $0.36/M
Laguna S 2.1
Poolside
Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...
- Context
- 1.0M
- In
- $0.09/M
- Out
- $0.18/M
LongCat 2.0
Meituan
LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, repository-level changes, long-horizon problem solving, and agentic...
- Context
- 1.0M
- In
- $0.30/M
- Out
- $1.20/M
Inkling
Thinking Machines
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
- Context
- 1.0M
- In
- $1.00/M
- Out
- $4.05/M
KAT-Coder-Pro V2.5
Kwaipilot
KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...
- Context
- 262K
- In
- $0.74/M
- Out
- $2.96/M
Aion-3.0
Aion Labs
Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation process in which multiple specialized models each contribute...
- Context
- 131K
- In
- $3.00/M
- Out
- $6.00/M
Fusion
Openrouter
Fusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel with web search and web fetch enabled, then a...
- Context
- 1.0M
- In
- $-1000000.0000/M
- Out
- $-1000000.0000/M
Step 3.7 Flash
Stepfun
Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...
- Context
- 262K
- In
- $0.20/M
- Out
- $1.15/M
Perceptron Mk1
Perceptron
Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and video inputs paired with natural language queries, and produces detailed visual understanding...
- Context
- 33K
- In
- $0.15/M
- Out
- $1.50/M
MiMo-V2.5-Pro
Xiaomi
MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro....
- Context
- 1.1M
- In
- $0.43/M
- Out
- $0.87/M
Trinity Large Thinking
Arcee Ai
Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic workloads, and reasoning tasks. Launch video: https://youtu.be/Gc82AXLa0Rg?si=4RLn6WBz33qT--B7...
- Context
- 262K
- In
- $0.25/M
- Out
- $0.80/M
Reka Edge
Rekaai
Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. This model is optimized specifically to deliver industry-leading performance in image understanding,...
- Context
- 16K
- In
- $0.10/M
- Out
- $0.10/M
Palmyra X5
Writer
Palmyra X5 is Writer's most advanced model, purpose-built for building and scaling AI agents across the enterprise. It delivers industry-leading speed and efficiency on context windows up to 1 million...
- Context
- 1.0M
- In
- $0.60/M
- Out
- $6.00/M
Relace Search
Relace
The relace-search model uses 4-12 `view_file` and `grep` tools in parallel to explore a codebase and return relevant files to the user request. In contrast to RAG, relace-search performs agentic...
- Context
- 256K
- In
- $1.00/M
- Out
- $3.00/M
Cydonia 24B V4.1
Thedrummer
Uncensored and creative writing model based on Mistral Small 3.2 24B with good recall, prompt adherence, and intelligence.
- Context
- 131K
- In
- $0.30/M
- Out
- $0.50/M
Hermes 4 405B
Nous Research
Hermes 4 is a large-scale reasoning model built on Meta-Llama-3.1-405B and released by Nous Research. It introduces a hybrid reasoning mode, where the model can choose to deliberate internally with...
- Context
- 131K
- In
- $1.00/M
- Out
- $3.00/M
UI-TARS 7B
Bytedance
UI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, including desktop interfaces, web browsers, mobile systems, and games. Built by ByteDance, it builds upon the UI-TARS framework with reinforcement...
- Context
- 128K
- In
- $0.10/M
- Out
- $0.20/M
Uncensored
Cognitivecomputations
Venice Uncensored Dolphin Mistral 24B Venice Edition is a fine-tuned variant of Mistral-Small-24B-Instruct-2501, developed by dphn.ai in collaboration with Venice.ai. This model is designed as an “uncensored” instruct-tuned LLM, preserving...
- Context
- 128K
- In
- $0.20/M
- Out
- $0.90/M
Morph V3 Large
Morph
Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations. The model requires the prompt to be in the following format: <instruction>{instruction}</instruction> <code>{initial_code}</code>...
- Context
- 262K
- In
- $0.90/M
- Out
- $1.90/M
ERNIE 4.5 VL 424B A47B
Baidu
ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu’s ERNIE 4.5 series, featuring 424B total parameters with 47B active per token. It is trained jointly on text and image data...
- Context
- 123K
- In
- $0.42/M
- Out
- $1.25/M
Llama 3.3 Euryale 70B
Sao10k
Euryale L3.3 70B is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.2](/models/sao10k/l3-euryale-70b).
- Context
- 131K
- In
- $0.65/M
- Out
- $0.75/M
Magnum v4 72B
Anthracite Org
This is a series of models designed to replicate the prose quality of the Claude 3 models, specifically Sonnet(https://openrouter.ai/anthropic/claude-3.5-sonnet) and Opus(https://openrouter.ai/anthropic/claude-3-opus). The model is fine-tuned on top of [Qwen2.5 72B](https://openrouter.ai/qwen/qwen-2.5-72b-instruct).
- Context
- 33K
- In
- $2.50/M
- Out
- $5.00/M
Weaver (alpha)
Mancer
An attempt to recreate Claude-style verbosity, but don't expect the same level of coherence or memory. Meant for use in roleplay/narrative situations.
- Context
- 8K
- In
- $0.40/M
- Out
- $0.75/M
ReMM SLERP 13B
Undi95
A recreation trial of the original MythoMax-L2-B13 but with updated models. #merge
- Context
- 6K
- In
- $0.35/M
- Out
- $0.65/M
48 of 429