Model catalogue

429 models from OpenRouter. Any of them can back a token.

Showing 429 of 429

GPT-6 Astra

OpenAI

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...

Context
1.1M
In
$10.00/M
Out
$50.00/M

Claude Fable 5.1

Anthropic

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

Context
1.0M
In
$10.00/M
Out
$50.00/M

Gemma 4 26B A4B

Google

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

Context
262K
In
$0.04/M
Out
$0.22/M

Grok 4.6

xAI

Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

Context
500K
In
$2.00/M
Out
$6.00/M

DeepSeek V4.1 Flash

DeepSeek

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...

Context
1.0M
In
$0.15/M
Out
$0.60/M

Llama Guard 4 12B

Meta

Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM...

Context
164K
In
$0.18/M
Out
$0.18/M

Qwen3.8 Flash

Qwen

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

Context
1.0M
In
$0.15/M
Out
$0.47/M

Mistral Medium 3.5

Mistral

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...

Context
262K
In
$1.50/M
Out
$7.50/M

Kimi K3

Moonshot

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

Context
1.0M
In
$2.65/M
Out
$13.28/M

GLM 5.3 Flash

Z.ai

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

Context
1.3M
In
$0.15/M
Out
$0.50/M

MiniMax M3

MiniMax

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

Context
1.0M
In
$0.30/M
Out
$1.20/M

Nova Premier 1.0

Amazon

Amazon Nova Premier is the most capable of Amazon’s multimodal models for complex reasoning tasks and for use as the best teacher for distilling custom models.

Context
1.0M
In
$2.50/M
Out
$12.50/M

Phi 4

Microsoft

[Microsoft Research](/microsoft) Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or where quick responses are needed. At 14 billion...

Context
16K
In
$0.07/M
Out
$0.14/M

Nemotron 3.5 Lightning

NVIDIA

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

Context
262K
In
$0.08/M
Out
$0.20/M

Command A

Cohere

Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual, and coding use cases. Compared to other leading proprietary...

Context
256K
In
$2.50/M
Out
$10.00/M

Sonar Pro Search

Perplexity

Exclusively available on the OpenRouter API, Sonar Pro's new Pro Search mode is Perplexity's most advanced agentic search system. It is designed for deeper reasoning and analysis. Pricing is based...

Context
200K
In
$3.00/M
Out
$15.00/M

Seed 2.1 Turbo

ByteDance

Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step task execution, and understanding visual and...

Context
262K
In
$0.50/M
Out
$2.50/M

Hy-MT2-1.8B

Tencent

Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided...

Context
8K
In
$0.04/M
Out
$0.18/M

Schematron V2 Turbo

Inference.net

Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extraction instructions must be supplied through a JSON schema in response_format rather...

Context
128K
In
$0.03/M
Out
$0.15/M

Fugu Ultra v2

Sakana

Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to...

Context
1.0M
In
$5.00/M
Out
$30.00/M

Ling 3.0 Flash VL

Inclusion AI

Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...

Context
131K
In
$0.06/M
Out
$0.18/M
IN

Mercury 2.5

Inception

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...

Context
260K
In
$0.04/M
Out
$0.15/M

Muse Spark 1.3 Contributor

Meta

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...

Context
1.0M
In
$0.10/M
Out
$0.20/M
IB

Granite 4.2 8B

Ibm Granite

Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-step reasoning. It supports full, low-effort,...

Context
131K
In
$0.06/M
Out
$0.25/M
UP

Solar Pro 4

Upstage

Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with strong capabilities in office productivity, document-intensive...

Context
524K
In
$0.09/M
Out
$0.36/M

Laguna S 2.1

Poolside

Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...

Context
1.0M
In
$0.09/M
Out
$0.18/M
ME

LongCat 2.0

Meituan

LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, repository-level changes, long-horizon problem solving, and agentic...

Context
1.0M
In
$0.30/M
Out
$1.20/M

Inkling

Thinking Machines

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

Context
1.0M
In
$1.00/M
Out
$4.05/M
KW

KAT-Coder-Pro V2.5

Kwaipilot

KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...

Context
262K
In
$0.74/M
Out
$2.96/M
AI

Aion-3.0

Aion Labs

Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation process in which multiple specialized models each contribute...

Context
131K
In
$3.00/M
Out
$6.00/M

Fusion

Openrouter

Fusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel with web search and web fetch enabled, then a...

Context
1.0M
In
$-1000000.0000/M
Out
$-1000000.0000/M

Step 3.7 Flash

Stepfun

Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...

Context
262K
In
$0.20/M
Out
$1.15/M
PE

Perceptron Mk1

Perceptron

Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and video inputs paired with natural language queries, and produces detailed visual understanding...

Context
33K
In
$0.15/M
Out
$1.50/M
XI

MiMo-V2.5-Pro

Xiaomi

MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro....

Context
1.1M
In
$0.43/M
Out
$0.87/M
AR

Trinity Large Thinking

Arcee Ai

Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic workloads, and reasoning tasks. Launch video: https://youtu.be/Gc82AXLa0Rg?si=4RLn6WBz33qT--B7...

Context
262K
In
$0.25/M
Out
$0.80/M
RE

Reka Edge

Rekaai

Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. This model is optimized specifically to deliver industry-leading performance in image understanding,...

Context
16K
In
$0.10/M
Out
$0.10/M
WR

Palmyra X5

Writer

Palmyra X5 is Writer's most advanced model, purpose-built for building and scaling AI agents across the enterprise. It delivers industry-leading speed and efficiency on context windows up to 1 million...

Context
1.0M
In
$0.60/M
Out
$6.00/M
RE

Relace Search

Relace

The relace-search model uses 4-12 `view_file` and `grep` tools in parallel to explore a codebase and return relevant files to the user request. In contrast to RAG, relace-search performs agentic...

Context
256K
In
$1.00/M
Out
$3.00/M
TH

Cydonia 24B V4.1

Thedrummer

Uncensored and creative writing model based on Mistral Small 3.2 24B with good recall, prompt adherence, and intelligence.

Context
131K
In
$0.30/M
Out
$0.50/M

Hermes 4 405B

Nous Research

Hermes 4 is a large-scale reasoning model built on Meta-Llama-3.1-405B and released by Nous Research. It introduces a hybrid reasoning mode, where the model can choose to deliberate internally with...

Context
131K
In
$1.00/M
Out
$3.00/M

UI-TARS 7B

Bytedance

UI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, including desktop interfaces, web browsers, mobile systems, and games. Built by ByteDance, it builds upon the UI-TARS framework with reinforcement...

Context
128K
In
$0.10/M
Out
$0.20/M
CO

Uncensored

Cognitivecomputations

Venice Uncensored Dolphin Mistral 24B Venice Edition is a fine-tuned variant of Mistral-Small-24B-Instruct-2501, developed by dphn.ai in collaboration with Venice.ai. This model is designed as an “uncensored” instruct-tuned LLM, preserving...

Context
128K
In
$0.20/M
Out
$0.90/M
MO

Morph V3 Large

Morph

Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations. The model requires the prompt to be in the following format: <instruction>{instruction}</instruction> <code>{initial_code}</code>...

Context
262K
In
$0.90/M
Out
$1.90/M

ERNIE 4.5 VL 424B A47B

Baidu

ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu’s ERNIE 4.5 series, featuring 424B total parameters with 47B active per token. It is trained jointly on text and image data...

Context
123K
In
$0.42/M
Out
$1.25/M
SA

Llama 3.3 Euryale 70B

Sao10k

Euryale L3.3 70B is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.2](/models/sao10k/l3-euryale-70b).

Context
131K
In
$0.65/M
Out
$0.75/M
AN

Magnum v4 72B

Anthracite Org

This is a series of models designed to replicate the prose quality of the Claude 3 models, specifically Sonnet(https://openrouter.ai/anthropic/claude-3.5-sonnet) and Opus(https://openrouter.ai/anthropic/claude-3-opus). The model is fine-tuned on top of [Qwen2.5 72B](https://openrouter.ai/qwen/qwen-2.5-72b-instruct).

Context
33K
In
$2.50/M
Out
$5.00/M
MA

Weaver (alpha)

Mancer

An attempt to recreate Claude-style verbosity, but don't expect the same level of coherence or memory. Meant for use in roleplay/narrative situations.

Context
8K
In
$0.40/M
Out
$0.75/M
UN

ReMM SLERP 13B

Undi95

A recreation trial of the original MythoMax-L2-B13 but with updated models. #merge

Context
6K
In
$0.35/M
Out
$0.65/M

48 of 429