apitokens.rent
                                                          .....::::------------------------------------------------------------------::::.....                                                          
                                                         .....:::--******************************************************************--:::.....                                                         
                                                         .....::--****............................................................****--::.....                                                         
                                                         .....::--****............................................................****--::.....                                                         
                                                         .....::--****............................................................****--::.....                                                         
                                                         .....::--****..$ apitokens.rent                                       ...****--::.....                                                         
                                                         .....::--****..                                                       ...****--::.....             ====---:::::                                
                                                         .....::--****..  422 models        credits: $--                       ...****--::.....            =====---:::::.                               
                                                        .....::--=****..  rewards claimed   holders: --                        ...****=--::.....           ==@@@@--:::::.                               
             :  :   :  :  :                             .....::--=****..                                                       ...****=--::.....           ==****--:::::.                               
             :--:-  -  - --             .               .....::--=***+..$ run claude-opus-5 _                                  ...+***=--::.....           =====---:::::.                               
               -- -  - - -                ..            .....::--=***+............................................................+***=--::.....           =====-------:.                               
                = -= = ==-                 ::           .....::--=***+............................................................+***=--::.....           =====-------:.                               
                 = ======                 --            .....::--=****............................................................****=--::.....           =====-------:.                               
                 =++++=                 --              .....::--=********************************************************************=--::.....           =====---:::::.                               
                ===--:::              +===--:-----      .....::--==******************************************************************==--::.....           =====---:::::.        ==================     
                ===--:::              +===---:   --     ......:::---======+==++==++==++==++===+===+===+===+===+===+===+===+=========---:::......           =====---:::::.         ----------------      
                ===--:::              +===--:-----       .......::::::::::+--++--++--++=-++=-=+=-=+===+=-=+=-=+=-=+=-=+=-=+=-=::::::::::.......             ====---:::::         ::::::::::::::::::     
::::::::::::::::::::::::::::::-------------------------------------------======================================================-------------------------------------------::::::::::::::::::::::::::::::
                ........................................................................................................................................................................                
                ........................................................................................................................................................................                
Tokens
Requests
Rewards
Cost
hourly · UTC
Market cap
$217.2K
Liquidity
$44.1K
24h volume
$1.62M
Holders
1.4K
10.0K to get access
Creator rewards
143.689 SOL
$13,511.21
Inference served
$368.44
24 wallets
Treasury
$13,947.99
148.334 SOL
Runway
265d
tick 0m ago
01
Every trade pays a fee

$API generates creator fees on every trade, 30 basis points of it, forever. Those fees pay for inference.

02
Holding is the subscription

Hold 10,000 $API and the API turns on. Sell, and it turns off. There is nothing to buy and no balance to top up.

03
Then use it, unmetered

Not a quota and not a share. Every holder gets the same access to all 422 models, whether they hold the minimum or a thousand times it.

Every model, unmetered

Prices are what each one costs the treasury, not what it costs you.

422

Claude Opus 5 (Fast)

anthropic/claude-opus-5-fast

Fast-mode variant of [Opus 5](/anthropic/claude-opus-5) - identical capabilities with higher output speed at 2x pricing relative to regular

1M31 runs$10/M in · $50/M outRun now

Claude Opus 5

anthropic/claude-opus-5

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end

1M519 runs$5/M in · $25/M outRun now

MoonshotAI: Kimi K3

moonshotai/kimi-k3

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and lo

1.0M2 runs$3/M in · $15/M outRun now

Claude Opus 5 (batch)

anthropic/claude-opus-5:batch

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end

1M$2.5/M in · $12.5/M outRun now

Qwen: Qwen3.8 2.4T A95B

qwen/qwen3.8-2.4t-a95b

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max

1.0M2 runs$2/M in · $6/M outRun now

Qwen: Qwen3.8 Max

qwen/qwen3.8-max

Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to the Qwen3.8 Max Preview. It is a multim

1M1 runs$2/M in · $6/M outRun now

SpaceXAI: Grok 4.6

x-ai/grok-4.6

Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

500K33 runs$2/M in · $6/M outRun now

Z.ai: GLM Latest

~z-ai/glm-latest

This model always redirects to the latest GLM model from Z.ai.

1.0M$1.4/M in · $4.4/M outRun now

Z.ai: GLM 5.3

z-ai/glm-5.3

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text in

1.0M$1.4/M in · $4.4/M outRun now

Meta: Muse Spark 1.1

meta/muse-spark-1.1

Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents

1.0M$1.25/M in · $4.25/M outRun now

Meta: Muse Spark 1.2

meta/muse-spark-1.2

Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents,

1.0M$1.25/M in · $4.25/M outRun now

Thinking Machines: Inkling

thinkingmachines/inkling

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It i

1.0M1 runs$1/M in · $4.05/M outRun now

Thinking Machines: Inkling (batch)

thinkingmachines/inkling:batch

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It i

524K$1/M in · $4.05/M outRun now

Sakana: Sakana Namazu

sakana/sakana-namazu

Sakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language an

262K$0.95/M in · $4/M outRun now

Google: Gemini 3.6 Flash

google/gemini-3.6-flash

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produc

1.0M2 runs$0.75/M in · $3.75/M outRun now

DeepSeek: DeepSeek V4 Pro 0813

deepseek/deepseek-v4-pro-0813

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

1.0M2 runs$1.12/M in · $3.37/M outRun now

Kwaipilot: KAT-Coder-Pro V2.5

kwaipilot/kat-coder-pro-v2.5

KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it

256K1 runs$0.74/M in · $2.96/M outRun now

ByteDance Seed: Seed-2.0-Code

bytedance-seed/seed-2.0-code

Seed 2.0 Code is a model from ByteDance Seed optimized for agentic coding. It is suited for frontend development, multilingual programming t

262K$0.5/M in · $3/M outRun now

Qwen: Qwen3.8 27B

qwen/qwen3.8-27b

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal i

1M$0.4/M in · $3/M outRun now

ByteDance Seed: Seed 2.1 Turbo

bytedance-seed/seed-2-1-turbo

Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software d

262K1 runs$0.5/M in · $2.5/M outRun now

Google: Gemini 3.5 Flash Lite

google/gemini-3.5-flash-lite

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute foc

1.0M$0.3/M in · $2.5/M outRun now

Google: Gemini 3.6 Flash (batch)

google/gemini-3.6-flash:batch

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produc

1.0M$0.375/M in · $1.88/M outRun now

Google: Gemini 3.7 Flash

google/gemini-3.7-flash

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for

1.0M1 runs$0.375/M in · $1.88/M outRun now

Meta: Muse Glimmer 30B

meta/muse-glimmer-30b

Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for auto

131K$0.35/M in · $1.5/M outRun now

Thinking Machines: Inkling Small

thinkingmachines/inkling-small

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total

1.0M$0.45/M in · $1.2/M outRun now

Meituan: LongCat 2.0

meituan/longcat-2.0

LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for codin

1.0M$0.3/M in · $1.2/M outRun now

Google: Gemini 3.5 Flash Lite (batch)

google/gemini-3.5-flash-lite:batch

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute foc

1.0M$0.15/M in · $1.25/M outRun now

Google: Gemini 3.7 Flash (batch)

google/gemini-3.7-flash:batch

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for

1.0M$0.188/M in · $0.938/M outRun now

DeepSeek: DeepSeek V4 Flash Vision Exp

deepseek/deepseek-v4-flash-vision-exp

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v

1.0M$0.22/M in · $0.66/M outRun now

Kwaipilot: KAT-Coder-Air V2.5

kwaipilot/kat-coder-air-v2.5

KAT-Coder-Air V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it

256K$0.15/M in · $0.6/M outRun now

DeepSeek: DeepSeek V4 Flash 0731

deepseek/deepseek-v4-flash-0731

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-traine

1.3M$0.14/M in · $0.28/M outRun now

Tencent: Hy-MT2-30B-A3B

tencent/hy-mt2-30b-a3b

Hy-MT2-30B-A3B is Tencent's flagship translation model in the Hy-MT2 family. It supports 33 language pairs and five Chinese dialect and mino

8K$0.074/M in · $0.295/M outRun now

Tencent: Hy-MT2-7B

tencent/hy-mt2-7b

Hy-MT2-7B is a 7B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pai

8K$0.074/M in · $0.295/M outRun now

Meta: Muse Spark 1.2 Contributor

meta/muse-spark-1.2-contributor

Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’

1.0M$0.1/M in · $0.2/M outRun now

NVIDIA: Nemotron 3.5 Lightning

nvidia/nemotron-3.5-lightning

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for

262K1 runs$0.08/M in · $0.2/M outRun now

Poolside: Laguna S 2.1

poolside/laguna-s-2.1

Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B

1.0M$0.09/M in · $0.18/M outRun now

Tencent: Hy-MT2-1.8B

tencent/hy-mt2-1.8b

Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-

8K$0.044/M in · $0.177/M outRun now

Qwen: Qwen3.7 Flash

qwen/qwen3.7-flash

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer int

1M1 runs$0.03/M in · $0.13/M outRun now

Upstage: Solar Pro 4

upstage/solar-pro4

Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agenti

524K$0.03/M in · $0.12/M outRun now

DeepSeek V4 Flash Latest

~deepseek/deepseek-v4-flash-latest

This model always redirects to the latest model in the DeepSeek V4 Flash family.

1.3M$0.04/M in · $0.08/M outRun now

Ling-3.0-flash

inclusionai/ling-3.0-flash

*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model i

262K$0.021/M in · $0.063/M outRun now

Dots Studio: Dots3-Note Preview (free)

dots-studio/dots-3-note-preview:free

Dots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio, with 16B active parameters out of 280B total. It is the ligh

512KFree in · Free outRun now

LiquidAI: LFM2.5-2.6B (free)

liquid/lfm-2.5-2.6b:free

LFM2.5-2.6B is a compact reasoning model from Liquid AI. It is suited for agent workflows, data extraction, RAG, and long-context processing

66KFree in · Free outRun now

NVIDIA: Nemotron 3.5 Lightning (free)

nvidia/nemotron-3.5-lightning:free

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for

1MFree in · Free outRun now

Poolside: Laguna S 2.1 (free)

poolside/laguna-s-2.1:free

Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B

262KFree in · Free outRun now

Ox Alpha

stealth/ox-alpha

Ox Alpha is a reasoning model designed for coding, sustained agentic work, and production workloads. It is suited for long-horizon software

1.0MFree in · Free outRun now

Thinking Machines: Inkling Small (free)

thinkingmachines/inkling-small:free

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total

262KFree in · Free outRun now

Thinking Machines: Inkling (free)

thinkingmachines/inkling:free

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It i

262KFree in · Free outRun now

Auto Router (Beta)

openrouter/auto-beta

Auto Router (Beta) is a task-aware router from OpenRouter. It classifies each request, then routes it the [most popular model](/rankings#tas

2MFree in · Free outRun now

OpenAI: o1-pro

openai/o1-pro

The o1 series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o1-pro model

200K6 runs$150/M in · $600/M outRun now

OpenAI: o1-pro (batch)

openai/o1-pro:batch

The o1 series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o1-pro model

200K$75/M in · $300/M outRun now

OpenAI: GPT-5.4 Pro

openai/gpt-5.4-pro

GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, hi

1.1M1 runs$30/M in · $180/M outRun now

OpenAI: GPT-5.5 Pro

openai/gpt-5.5-pro

GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+

1.1M4 runs$30/M in · $180/M outRun now

OpenAI: GPT-5.2 Pro

openai/gpt-5.2-pro

GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It i

400K$21/M in · $168/M outRun now

Anthropic: Claude Opus 4.7 (Fast)

anthropic/claude-opus-4.7-fast

Fast-mode variant of [Opus 4.7](/anthropic/claude-opus-4.7) - identical capabilities with higher output speed at premium 6x pricing. Learn

1M$30/M in · $150/M outRun now

OpenAI: GPT-5 Pro

openai/gpt-5-pro

GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for

400K$15/M in · $120/M outRun now

OpenAI: GPT-5.4 Pro (batch)

openai/gpt-5.4-pro:batch

GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, hi

1.1M$15/M in · $90/M outRun now

OpenAI: GPT-5.5 Pro (batch)

openai/gpt-5.5-pro:batch

GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+

1.1M$15/M in · $90/M outRun now

OpenAI: o3 Pro

openai/o3-pro

The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model u

200K$20/M in · $80/M outRun now

OpenAI: GPT-5.2 Pro (batch)

openai/gpt-5.2-pro:batch

GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It i

400K$10.5/M in · $84/M outRun now

Anthropic: Claude Opus 4

anthropic/claude-opus-4

Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-running t

200K$15/M in · $75/M outRun now

Anthropic: Claude Opus 4.1

anthropic/claude-opus-4.1

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks.

200K$15/M in · $75/M outRun now

OpenAI: GPT-4

openai/gpt-4

OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving difficult problems with greater accuracy than p

8K$30/M in · $60/M outRun now

OpenAI: o1

openai/o1

The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trai

200K$15/M in · $60/M outRun now

OpenAI: GPT-5 Pro (batch)

openai/gpt-5-pro:batch

GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for

400K$7.5/M in · $60/M outRun now

Anthropic: Claude Fable Latest

~anthropic/claude-fable-latest

This model always redirects to the latest model in the Claude Fable family.

1M$10/M in · $50/M outRun now

Anthropic: Claude Fable 5

anthropic/claude-fable-5

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inp

1M214 runs$10/M in · $50/M outRun now

Anthropic: Claude Opus 4.8 (Fast)

anthropic/claude-opus-4.8-fast

Fast-mode variant of [Opus 4.8](/anthropic/claude-opus-4.8) - identical capabilities with higher output speed at 2x pricing relative to regu

1M$10/M in · $50/M outRun now

OpenAI: o3 Pro (batch)

openai/o3-pro:batch

The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model u

200K$10/M in · $40/M outRun now

Anthropic: Claude Opus 4.1 (batch)

anthropic/claude-opus-4.1:batch

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks.

200K$7.5/M in · $37.5/M outRun now

OpenAI: GPT-4 Turbo

openai/gpt-4-turbo

The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to Dec

128K$10/M in · $30/M outRun now

OpenAI: GPT-4 Turbo Preview

openai/gpt-4-turbo-preview

The preview GPT-4 model with improved instruction following, JSON mode, reproducible outputs, parallel function calling, and more. Training

128K$10/M in · $30/M outRun now

OpenAI: o1 (batch)

openai/o1:batch

The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trai

200K$7.5/M in · $30/M outRun now

OpenAI: GPT-5.5

openai/gpt-5.5

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliabil

1.1M$5/M in · $30/M outRun now

OpenAI: GPT Chat Latest

openai/gpt-chat-latest

GPT Chat Latest points to OpenAI's stable API alias `chat-latest` that always resolves to the latest Instant chat model used in ChatGPT. As

400K$5/M in · $30/M outRun now

Sakana: Fugu Ultra

sakana/fugu-ultra

Fugu Ultra is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent

1M$5/M in · $30/M outRun now

Anthropic: Claude Opus Latest

~anthropic/claude-opus-latest

This model always redirects to the latest model in the Claude Opus family.

1M$5/M in · $25/M outRun now

Anthropic: Claude Fable 5 (batch)

anthropic/claude-fable-5:batch

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inp

1M$5/M in · $25/M outRun now

Anthropic: Claude Opus 4.5

anthropic/claude-opus-4.5

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon comp

200K$5/M in · $25/M outRun now

Anthropic: Claude Opus 4.6

anthropic/claude-opus-4.6

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire wo

1M$5/M in · $25/M outRun now

Anthropic: Claude Opus 4.7

anthropic/claude-opus-4.7

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic

1M$5/M in · $25/M outRun now

Anthropic: Claude Opus 4.8

anthropic/claude-opus-4.8

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text

1M$5/M in · $25/M outRun now

OpenAI: GPT-5.4 Image 2

openai/gpt-5.4-image-2

[GPT-5.4](https://openrouter.ai/openai/gpt-5.4) Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities

272K5 runs$8/M in · $15/M outRun now

OpenAI: GPT-4 Turbo (batch)

openai/gpt-4-turbo:batch

The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to Dec

128K$5/M in · $15/M outRun now

OpenAI: GPT-4o (2024-05-13)

openai/gpt-4o-2024-05-13

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence

128K$5/M in · $15/M outRun now

OpenAI: GPT-5 Image

openai/gpt-5-image

[GPT-5](https://openrouter.ai/openai/gpt-5) Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It offe

400K2 runs$10/M in · $10/M outRun now

Anthropic: Claude Sonnet 4

anthropic/claude-sonnet-4

Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with im

1M$3/M in · $15/M outRun now

Anthropic: Claude Sonnet 4.5

anthropic/claude-sonnet-4.5

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state

1M$3/M in · $15/M outRun now

Anthropic: Claude Sonnet 4.6

anthropic/claude-sonnet-4.6

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It ex

1M$3/M in · $15/M outRun now

Perplexity: Sonar Pro

perplexity/sonar-pro

Note: Sonar Pro pricing includes Perplexity search pricing. See [details here](https://docs.perplexity.ai/guides/pricing#detailed-pricing-br

200K$3/M in · $15/M outRun now

Perplexity: Sonar Pro Search

perplexity/sonar-pro-search

Exclusively available on the OpenRouter API, Sonar Pro's new Pro Search mode is Perplexity's most advanced agentic search system. It is desi

200K$3/M in · $15/M outRun now

OpenAI: GPT-5.4

openai/gpt-5.4

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (92

1.1M$2.5/M in · $15/M outRun now

OpenAI: GPT-5.5 (batch)

openai/gpt-5.5:batch

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliabil

1.1M$2.5/M in · $15/M outRun now

OpenAI: GPT-5.2

openai/gpt-5.2

GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. I

400K$1.75/M in · $14/M outRun now

OpenAI: GPT-5.2 Chat

openai/gpt-5.2-chat

GPT-5.2 Chat (AKA Instant) is the fast, lightweight member of the 5.2 family, optimized for low-latency chat while retaining strong general

128K$1.75/M in · $14/M outRun now

OpenAI: GPT-5.2-Codex

openai/gpt-5.2-codex

GPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows. It is designed for both inter

400K$1.75/M in · $14/M outRun now

OpenAI: GPT-5.3-Codex

openai/gpt-5.3-codex

GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with

400K1 runs$1.75/M in · $14/M outRun now

MoonshotAI Kimi Latest

~moonshotai/kimi-latest

This model always redirects to the latest model in the MoonshotAI Kimi family.

1.0M$2.6/M in · $13/M outRun now

Amazon: Nova Premier 1.0

amazon/nova-premier-v1

Amazon Nova Premier is the most capable of Amazon’s multimodal models for complex reasoning tasks and for use as the best teacher for distil

1M$2.5/M in · $12.5/M outRun now

Anthropic: Claude Opus 4.5 (batch)

anthropic/claude-opus-4.5:batch

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon comp

200K$2.5/M in · $12.5/M outRun now

Anthropic: Claude Opus 4.6 (batch)

anthropic/claude-opus-4.6:batch

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire wo

1M$2.5/M in · $12.5/M outRun now

Anthropic: Claude Opus 4.7 (batch)

anthropic/claude-opus-4.7:batch

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic

1M$2.5/M in · $12.5/M outRun now

Anthropic: Claude Opus 4.8 (batch)

anthropic/claude-opus-4.8:batch

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text

1M$2.5/M in · $12.5/M outRun now

Google Gemini Pro Latest

~google/gemini-pro-latest

This model always redirects to the latest model in the Google Gemini Pro family.

1.0M$2/M in · $12/M outRun now

Google: Nano Banana Pro (Gemini 3 Pro Image)

google/gemini-3-pro-image

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana wit

131K16 runs$2/M in · $12/M outRun now

Google: Nano Banana Pro (Gemini 3 Pro Image Preview)

google/gemini-3-pro-image-preview

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana wit

66K$2/M in · $12/M outRun now

Google: Gemini 3.1 Pro Preview

google/gemini-3.1-pro-preview

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliabil

1.0M$2/M in · $12/M outRun now

Google: Gemini 3.1 Pro Preview Custom Tools

google/gemini-3.1-pro-preview-customtools

Gemini 3.1 Pro Preview Custom Tools is a variant of Gemini 3.1 Pro that improves tool selection behavior by preventing overuse of a general

1.0M$2/M in · $12/M outRun now

OpenAI: GPT-5.6 Terra

openai/gpt-5.6-terra

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It

1.1M$2/M in · $12/M outRun now

OpenAI: GPT-5.6 Terra Pro

openai/gpt-5.6-terra-pro

GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode`

1.1M9 runs$2/M in · $12/M outRun now

Cohere: Command A

cohere/command-a

Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multili

256K$2.5/M in · $10/M outRun now

Cohere: Command R+ (08-2024)

cohere/command-r-plus-08-2024

command-r-plus-08-2024 is an update of the [Command R+](/models/cohere/command-r-plus) with roughly 50% higher throughput and 25% lower late

128K$2.5/M in · $10/M outRun now

OpenAI: GPT-4o

openai/gpt-4o

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence

128K$2.5/M in · $10/M outRun now

OpenAI: GPT-4o (2024-08-06)

openai/gpt-4o-2024-08-06

The 2024-08-06 version of GPT-4o offers improved performance in structured outputs, with the ability to supply a JSON schema in the respone_

128K$2.5/M in · $10/M outRun now

OpenAI: GPT-4o (2024-11-20)

openai/gpt-4o-2024-11-20

The 2024-11-20 version of GPT-4o offers a leveled-up creative writing ability with more natural, engaging, and tailored writing to improve r

128K$2.5/M in · $10/M outRun now

OpenAI: GPT Audio

openai/gpt-audio

The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural soundi

128K$2.5/M in · $10/M outRun now

Anthropic Claude Sonnet Latest

~anthropic/claude-sonnet-latest

This model always redirects to the latest model in the Anthropic Claude Sonnet family.

1M$2/M in · $10/M outRun now

OpenAI GPT Latest

~openai/gpt-latest

This model always redirects to the latest model in the OpenAI GPT family.

1.1M$2/M in · $10/M outRun now

Anthropic: Claude Sonnet 5

anthropic/claude-sonnet-5

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports

1M1 runs$2/M in · $10/M outRun now

OpenAI: GPT-5.6 Sol

openai/gpt-5.6-sol

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is part

1.1M$2/M in · $10/M outRun now

OpenAI: GPT-5.6 Sol Pro

openai/gpt-5.6-sol-pro

GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to

1.1M$2/M in · $10/M outRun now

Google: Gemini 2.5 Pro

google/gemini-2.5-pro

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs

1.0M$1.25/M in · $10/M outRun now

Google: Gemini 2.5 Pro Preview 06-05

google/gemini-2.5-pro-preview

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs

1.0M$1.25/M in · $10/M outRun now

Google: Gemini 2.5 Pro Preview 05-06

google/gemini-2.5-pro-preview-05-06

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs

1.0M$1.25/M in · $10/M outRun now

OpenAI: GPT-5

openai/gpt-5

GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for comp

400K$1.25/M in · $10/M outRun now

OpenAI: GPT-5.1

openai/gpt-5.1

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence,

400K$1.25/M in · $10/M outRun now

OpenAI: GPT-5.1-Codex

openai/gpt-5.1-codex

GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows. It is designed for both interacti

400K$1.25/M in · $10/M outRun now

OpenAI: GPT-5.1-Codex-Max

openai/gpt-5.1-codex-max

GPT-5.1-Codex-Max is OpenAI’s latest agentic coding model, designed for long-running, high-context software development tasks. It is based o

400K$1.25/M in · $10/M outRun now

Google: Gemini 3.5 Flash

google/gemini-3.5-flash

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It

1.0M$1.5/M in · $9/M outRun now

OpenAI: GPT-4.1

openai/gpt-4.1

GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context r

1.0M$2/M in · $8/M outRun now

OpenAI: o3

openai/o3

o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It als

200K$2/M in · $8/M outRun now

Perplexity: Sonar Deep Research

perplexity/sonar-deep-research

Sonar Deep Research is a research-focused model designed for multi-step retrieval, synthesis, and reasoning across complex topics. It autono

128K$2/M in · $8/M outRun now

Perplexity: Sonar Reasoning Pro

perplexity/sonar-reasoning-pro

Note: Sonar Pro pricing includes Perplexity search pricing. See [details here](https://docs.perplexity.ai/guides/pricing#detailed-pricing-br

128K$2/M in · $8/M outRun now

AionLabs: Aion-3.0

aion-labs/aion-3.0

Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative gene

131K$3/M in · $6/M outRun now

Anthropic: Claude Sonnet 4.5 (batch)

anthropic/claude-sonnet-4.5:batch

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state

1M$1.5/M in · $7.5/M outRun now

Anthropic: Claude Sonnet 4.6 (batch)

anthropic/claude-sonnet-4.6:batch

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It ex

1M$1.5/M in · $7.5/M outRun now

Mistral: Mistral Medium 3.5

mistralai/mistral-medium-3-5

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is d

262K$1.5/M in · $7.5/M outRun now

OpenAI: GPT-5.4 (batch)

openai/gpt-5.4:batch

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (92

1.1M$1.25/M in · $7.5/M outRun now

xAI: Grok Latest

~x-ai/grok-latest

This model always redirects to the latest Grok model from xAI.

500K$2/M in · $6/M outRun now

Magnum v4 72B

anthracite-org/magnum-v4-72b

This is a series of models designed to replicate the prose quality of the Claude 3 models, specifically Sonnet(https://openrouter.ai/anthrop

33K$3/M in · $5/M outRun now

Mistral Large

mistralai/mistral-large

This is Mistral AI's flagship model, Mistral Large 2 (version `mistral-large-2407`). It's a proprietary weights-available model and excels a

128K$2/M in · $6/M outRun now

Mistral Large 2407

mistralai/mistral-large-2407

This is Mistral AI's flagship model, Mistral Large 2 (version mistral-large-2407). It's a proprietary weights-available model and excels at

131K$2/M in · $6/M outRun now

Mistral: Mixtral 8x22B Instruct

mistralai/mixtral-8x22b-instruct

Mistral's official instruct fine-tuned version of [Mixtral 8x22B](/models/mistralai/mixtral-8x22b). It uses 39B active parameters out of 141

66K$2/M in · $6/M outRun now

SpaceXAI: Grok 4.5

x-ai/grok-4.5

Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM.

500K$2/M in · $6/M outRun now

OpenAI: GPT-5.2 (batch)

openai/gpt-5.2:batch

GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. I

400K$0.875/M in · $7/M outRun now

Qwen: Qwen3.6 Max Preview

qwen/qwen3.6-max-preview

Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately

262K$1.03/M in · $6.16/M outRun now

Google: Gemini 3.1 Pro Preview (batch)

google/gemini-3.1-pro-preview:batch

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliabil

1.0M$1/M in · $6/M outRun now

OpenAI: GPT-3.5 Turbo 16k

openai/gpt-3.5-turbo-16k

This model offers four times the context length of gpt-3.5-turbo, allowing it to support approximately 20 pages of text in a single request

16K$3/M in · $4/M outRun now

OpenAI: GPT-5.6 Terra Pro (batch)

openai/gpt-5.6-terra-pro:batch

GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode`

1.1M$1/M in · $6/M outRun now

OpenAI: GPT-5.6 Terra (batch)

openai/gpt-5.6-terra:batch

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It

1.1M$1/M in · $6/M outRun now

Writer: Palmyra X5

writer/palmyra-x5

Palmyra X5 is Writer's most advanced model, purpose-built for building and scaling AI agents across the enterprise. It delivers industry-lea

1.0M$0.6/M in · $6/M outRun now

OpenAI: GPT-4o (batch)

openai/gpt-4o:batch

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence

128K$1.25/M in · $5/M outRun now

Anthropic Claude Haiku Latest

~anthropic/claude-haiku-latest

This model always redirects to the latest model in the Anthropic Claude Haiku family.

200K$1/M in · $5/M outRun now

Anthropic: Claude Haiku 4.5

anthropic/claude-haiku-4.5

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latenc

200K$1/M in · $5/M outRun now

Anthropic: Claude Sonnet 5 (batch)

anthropic/claude-sonnet-5:batch

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports

1M$1/M in · $5/M outRun now

OpenAI: GPT-5.6 Sol Pro (batch)

openai/gpt-5.6-sol-pro:batch

GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to

1.1M$1/M in · $5/M outRun now

OpenAI: GPT-5.6 Sol (batch)

openai/gpt-5.6-sol:batch

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is part

1.1M$1/M in · $5/M outRun now

Qwen: Qwen3.7 Max

qwen/qwen3.7-max

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads

1M$1.48/M in · $4.42/M outRun now

Z.ai: GLM 5.2 (batch)

z-ai/glm-5.2:batch

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long

1.0M$1.4/M in · $4.4/M outRun now

Google: Gemini 2.5 Pro (batch)

google/gemini-2.5-pro:batch

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs

1.0M$0.625/M in · $5/M outRun now

OpenAI: GPT-5 Codex (batch)

openai/gpt-5-codex:batch

GPT-5-Codex is a specialized version of GPT-5 optimized for software engineering and coding workflows. It is designed for both interactive d

400K$0.625/M in · $5/M outRun now

OpenAI: GPT-5 (batch)

openai/gpt-5:batch

GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for comp

400K$0.625/M in · $5/M outRun now

OpenAI: GPT-5.1 (batch)

openai/gpt-5.1:batch

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence,

400K$0.625/M in · $5/M outRun now

OpenAI: o3 Mini

openai/o3-mini

OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and co

200K$1.1/M in · $4.4/M outRun now

OpenAI: o3 Mini High

openai/o3-mini-high

OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language

200K$1.1/M in · $4.4/M outRun now

OpenAI: o4 Mini

openai/o4-mini

OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimoda

200K$1.1/M in · $4.4/M outRun now

OpenAI: o4 Mini High

openai/o4-mini-high

OpenAI o4-mini-high is the same model as [o4-mini](/openai/o4-mini) with reasoning_effort set to high. OpenAI o4-mini is a compact reasoning

200K$1.1/M in · $4.4/M outRun now

OpenAI GPT Mini Latest

~openai/gpt-mini-latest

This model always redirects to the latest model in the OpenAI GPT Mini family.

400K$0.75/M in · $4.5/M outRun now

Google: Gemini 3.5 Flash (batch)

google/gemini-3.5-flash:batch

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It

1.0M$0.75/M in · $4.5/M outRun now

OpenAI: GPT-5.4 Mini

openai/gpt-5.4-mini

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports

400K$0.75/M in · $4.5/M outRun now

Z.ai: GLM 5 Turbo

z-ai/glm-5-turbo

GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenar

203K$1.2/M in · $4/M outRun now

Z.ai: GLM 5V Turbo

z-ai/glm-5v-turbo

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively han

203K$1.2/M in · $4/M outRun now

OpenAI: GPT-4.1 (batch)

openai/gpt-4.1:batch

GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context r

1.0M$1/M in · $4/M outRun now

OpenAI: o3 (batch)

openai/o3:batch

o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It als

200K$1/M in · $4/M outRun now

MoonshotAI: Kimi K2.6

moonshotai/kimi-k2.6

Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-age

262K1 runs$0.95/M in · $4/M outRun now

MoonshotAI: Kimi K2.7 Code (batch)

moonshotai/kimi-k2.7-code:batch

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliabl

262K$0.95/M in · $4/M outRun now

Qwen: Qwen3 Max

qwen/qwen3-max

Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual sup

262K$0.78/M in · $3.9/M outRun now

Qwen: Qwen3 Max Thinking

qwen/qwen3-max-thinking

Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-st

262K$0.78/M in · $3.9/M outRun now

OpenAI: GPT-5 Image Mini

openai/gpt-5-image-mini

GPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by [GPT-5 Mini](https://openrouter.ai/openai/gpt-5-mini), with GP

400K2 runs$2.5/M in · $2/M outRun now

Qwen: Qwen3 VL 235B A22B Thinking

qwen/qwen3-vl-235b-a22b-thinking

Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The

131K$0.4/M in · $4/M outRun now

NVIDIA: Nemotron 3 Ultra

nvidia/nemotron-3-ultra-550b-a55b

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE

512K$0.6/M in · $3.6/M outRun now

NVIDIA: Nemotron 3 Ultra (batch)

nvidia/nemotron-3-ultra-550b-a55b:batch

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE

512K$0.6/M in · $3.6/M outRun now

Qwen: Qwen3.5 397B A17B

qwen/qwen3.5-397b-a17b

The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism wit

262K$0.5/M in · $3.6/M outRun now

MoonshotAI: Kimi K2.7 Code

moonshotai/kimi-k2.7-code

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliabl

262K$0.67/M in · $3.4/M outRun now

Z.ai: GLM 5.1

z-ai/glm-5.1

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous mode

205K$0.966/M in · $3.04/M outRun now

Z.ai: GLM 5.2

z-ai/glm-5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long

1.0M$0.966/M in · $3.04/M outRun now

Nous: Hermes 4 405B

nousresearch/hermes-4-405b

Hermes 4 is a large-scale reasoning model built on Meta-Llama-3.1-405B and released by Nous Research. It introduces a hybrid reasoning mode,

131K$1/M in · $3/M outRun now

Relace: Relace Search

relace/relace-search

The relace-search model uses 4-12 `view_file` and `grep` tools in parallel to explore a codebase and return relevant files to the user reque

256K$1/M in · $3/M outRun now

Amazon: Nova Pro 1.0

amazon/nova-pro-v1

Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination of accuracy, speed, and cost for a wide ran

300K$0.8/M in · $3.2/M outRun now

Qwen: Qwen3 Coder Plus

qwen/qwen3-coder-plus

Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializing

1M$0.65/M in · $3.25/M outRun now

SpaceXAI: Grok 4.20

x-ai/grok-4.20

Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallu

2M$1.25/M in · $2.5/M outRun now

SpaceXAI: Grok 4.20 Multi-Agent

x-ai/grok-4.20-multi-agent

Grok 4.20 Multi-Agent is a variant of SpaceXAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in par

2M$1.25/M in · $2.5/M outRun now

SpaceXAI: Grok 4.3

x-ai/grok-4.3

Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruc

1M$1.25/M in · $2.5/M outRun now

Qwen: Qwen3.6 27B

qwen/qwen3.6-27b

Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimo

262K$0.32/M in · $3.2/M outRun now

Google: Gemini 3 Flash Preview

google/gemini-3-flash-preview

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It

1.0M$0.5/M in · $3/M outRun now

Google: Nano Banana 2 (Gemini 3.1 Flash Image)

google/gemini-3.1-flash-image

Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level

131K10 runs$0.5/M in · $3/M outRun now

Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview)

google/gemini-3.1-flash-image-preview

Gemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering P

66K$0.5/M in · $3/M outRun now

OpenAI: GPT-3.5 Turbo Instruct

openai/gpt-3.5-turbo-instruct

This model is a variant of GPT-3.5 Turbo tuned for instructional prompts and omitting chat-related optimizations. Training data: up to Sep 2

4K$1.5/M in · $2/M outRun now

DeepSeek: R1

deepseek/deepseek-r1

DeepSeek R1 is here: Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B param

64K$0.7/M in · $2.5/M outRun now

MoonshotAI: Kimi K2 0905

moonshotai/kimi-k2-0905

Kimi K2 0905 is the September update of [Kimi K2 0711](moonshotai/kimi-k2). It is a large-scale Mixture-of-Experts (MoE) language model deve

262K$0.6/M in · $2.5/M outRun now

MoonshotAI: Kimi K2 Thinking

moonshotai/kimi-k2-thinking

Kimi K2 Thinking is Moonshot AI’s most advanced open reasoning model to date, extending the K2 series into agentic, long-horizon reasoning.

262K$0.6/M in · $2.5/M outRun now

Anthropic: Claude Haiku 4.5 (batch)

anthropic/claude-haiku-4.5:batch

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latenc

200K$0.5/M in · $2.5/M outRun now

OpenAI: GPT-3.5 Turbo (older v0613)

openai/gpt-3.5-turbo-0613

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional

4K$1/M in · $2/M outRun now

OpenAI: GPT Audio Mini

openai/gpt-audio-mini

A cost-efficient version of GPT Audio. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better v

128K$0.6/M in · $2.4/M outRun now

SpaceXAI: Grok Build 0.1

x-ai/grok-build-0.1

Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image i

256K$1/M in · $2/M outRun now

MoonshotAI: Kimi K2 0711

moonshotai/kimi-k2

Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters wi

131K$0.57/M in · $2.3/M outRun now

Z.ai: GLM 4.5

z-ai/glm-4.5

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) archite

131K$0.6/M in · $2.2/M outRun now

Amazon: Nova 2 Lite

amazon/nova-2-lite-v1

Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nov

1M$0.3/M in · $2.5/M outRun now

Google: Gemini 2.5 Flash

google/gemini-2.5-flash

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scient

1.0M$0.3/M in · $2.5/M outRun now

Google: Nano Banana (Gemini 2.5 Flash Image)

google/gemini-2.5-flash-image

Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual un

33K$0.3/M in · $2.5/M outRun now

Morph: Morph V3 Large

morph/morph-v3-large

Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations. The model re

262K$0.9/M in · $1.9/M outRun now

MiniMax: MiniMax M1

minimax/minimax-m1

MiniMax-M1 is a large-scale, open-weight reasoning model designed for extended context and high-efficiency inference. It leverages a hybrid

1M$0.55/M in · $2.2/M outRun now

OpenAI: o3 Mini High (batch)

openai/o3-mini-high:batch

OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language

200K$0.55/M in · $2.2/M outRun now

OpenAI: o3 Mini (batch)

openai/o3-mini:batch

OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and co

200K$0.55/M in · $2.2/M outRun now

OpenAI: o4 Mini High (batch)

openai/o4-mini-high:batch

OpenAI o4-mini-high is the same model as [o4-mini](/openai/o4-mini) with reasoning_effort set to high. OpenAI o4-mini is a compact reasoning

200K$0.55/M in · $2.2/M outRun now

OpenAI: o4 Mini (batch)

openai/o4-mini:batch

OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimoda

200K$0.55/M in · $2.2/M outRun now

MoonshotAI: Kimi K2.5

moonshotai/kimi-k2.5

Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm par

262K$0.45/M in · $2.25/M outRun now

DeepSeek: R1 0528

deepseek/deepseek-r1-0528

May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and w

164K$0.5/M in · $2.15/M outRun now

OpenAI: GPT-5.4 Mini (batch)

openai/gpt-5.4-mini:batch

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports

400K$0.375/M in · $2.25/M outRun now

Qwen: Qwen3 30B A3B Thinking 2507

qwen/qwen3-30b-a3b-thinking-2507

Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step

82K$0.2/M in · $2.4/M outRun now

Qwen: Qwen3 VL 30B A3B Thinking

qwen/qwen3-vl-30b-a3b-thinking

Qwen3-VL-30B-A3B-Thinking is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Thi

262K$0.2/M in · $2.4/M outRun now

Qwen: Qwen3 235B A22B Thinking 2507

qwen/qwen3-235b-a22b-thinking-2507

Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tas

262K$0.23/M in · $2.3/M outRun now

Z.ai: GLM 5

z-ai/glm-5

GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expe

205K$0.6/M in · $1.92/M outRun now

Z.ai: GLM 4.6

z-ai/glm-4.6

Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128

205K$0.5/M in · $2/M outRun now

AionLabs: Aion-2.0

aion-labs/aion-2.0

Aion-2.0 is a variant of DeepSeek V3.2 optimized for immersive roleplaying and storytelling. It is particularly strong at introducing tensio

131K$0.8/M in · $1.6/M outRun now

AionLabs: Aion-RP 1.0 (8B)

aion-labs/aion-rp-llama-3.1-8b

Aion-RP-Llama-3.1-8B ranks the highest in the character evaluation portion of the RPBench-Auto benchmark, a roleplaying-specific variant of

33K$0.8/M in · $1.6/M outRun now

Mistral: Mistral Medium 3

mistralai/mistral-medium-3

Mistral Medium 3 is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly redu

131K$0.4/M in · $2/M outRun now

Mistral: Mistral Medium 3.1

mistralai/mistral-medium-3.1

Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to delive

131K$0.4/M in · $2/M outRun now

Z.ai: GLM 4.5V

z-ai/glm-4.5v

GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B

66K$0.6/M in · $1.8/M outRun now

Qwen: Qwen3.5-122B-A10B

qwen/qwen3.5-122b-a10b

The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a spa

262K$0.26/M in · $2.08/M outRun now

Qwen: Qwen3 VL 8B Thinking

qwen/qwen3-vl-8b-thinking

Qwen3-VL-8B-Thinking is the reasoning-optimized variant of the Qwen3-VL-8B multimodal model, designed for advanced visual and textual reason

131K$0.18/M in · $2.1/M outRun now

Qwen: Qwen3 235B A22B

qwen/qwen3-235b-a22b

Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass. It support

131K$0.455/M in · $1.82/M outRun now

Qwen: Qwen3.6 Plus

qwen/qwen3.6-plus

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling stro

1M$0.325/M in · $1.95/M outRun now

Google Gemini Flash Latest

~google/gemini-flash-latest

This model always redirects to the latest model in the Google Gemini Flash family.

1.0M$0.375/M in · $1.88/M outRun now

ByteDance Seed: Seed 1.6

bytedance-seed/seed-1.6

Seed 1.6 is a general-purpose model released by the ByteDance Seed team. It incorporates multimodal capabilities and adaptive deep thinking

262K$0.25/M in · $2/M outRun now

ByteDance Seed: Seed-2.0-Lite

bytedance-seed/seed-2.0-lite

Seed-2.0-Lite is a versatile, cost‑efficient enterprise workhorse that delivers strong multimodal and agent capabilities while offering noti

262K$0.25/M in · $2/M outRun now

OpenAI: GPT-5 Mini

openai/gpt-5-mini

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and

400K$0.25/M in · $2/M outRun now

OpenAI: GPT-5.1-Codex-Mini

openai/gpt-5.1-codex-mini

GPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex

400K$0.25/M in · $2/M outRun now

DeepSeek: DeepSeek V3.1

deepseek/deepseek-chat-v3.1

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt

164K$0.55/M in · $1.65/M outRun now

Z.ai: GLM 4.7

z-ai/glm-4.7

GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step r

205K$0.4/M in · $1.75/M outRun now

Qwen: Qwen3 VL 235B A22B Instruct

qwen/qwen3-vl-235b-a22b-instruct

Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across images a

262K$0.21/M in · $1.9/M outRun now

Relace: Relace Apply 3

relace/relace-apply-3

Relace Apply 3 is a specialized code-patching LLM that merges AI-suggested edits straight into your source files. It can apply updates from

256K$0.85/M in · $1.25/M outRun now

AionLabs: Aion-3.0-Mini

aion-labs/aion-3.0-mini

Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. It uses a collabor

131K$0.7/M in · $1.4/M outRun now

Qwen: Qwen3.5 Plus 2026-04-20

qwen/qwen3.5-plus-20260420

Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text

1M1 runs$0.3/M in · $1.8/M outRun now

Mistral: Mistral Large 3 2512

mistralai/mistral-large-2512

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters

262K$0.5/M in · $1.5/M outRun now

Morph: Morph V3 Fast

morph/morph-v3-fast

Morph's fastest apply model for code edits. ~10,500 tokens/sec with 96% accuracy for rapid code transformations. The model requires the prom

82K$0.8/M in · $1.2/M outRun now

Nous: Hermes 3 405B Instruct

nousresearch/hermes-3-llama-3.1-405b

Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplayi

131K$1/M in · $1/M outRun now

OpenAI: GPT-3.5 Turbo

openai/gpt-3.5-turbo

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional

16K$0.5/M in · $1.5/M outRun now

Perplexity: Sonar

perplexity/sonar

Sonar is lightweight, affordable, fast, and simple to use — now featuring citations and the ability to customize sources. It is designed for

127K$1/M in · $1/M outRun now

OpenAI: GPT-4.1 Mini

openai/gpt-4.1-mini

GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 mil

1.0M$0.4/M in · $1.6/M outRun now

Arcee AI: Virtuoso Large

arcee-ai/virtuoso-large

Virtuoso‑Large is Arcee's top‑tier general‑purpose LLM at 72 B parameters, tuned to tackle cross‑domain reasoning, creative writing and ente

131K$0.75/M in · $1.2/M outRun now

Qwen: Qwen3.5 Plus 2026-02-15

qwen/qwen3.5-plus-02-15

The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sp

1M$0.26/M in · $1.56/M outRun now

Qwen: Qwen2.5 VL 72B Instruct

qwen/qwen2.5-vl-72b-instruct

Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly capable of analyzing tex

128K$0.8/M in · $1/M outRun now

Qwen: Qwen3.5-27B

qwen/qwen3.5-27b

The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing

262K$0.195/M in · $1.56/M outRun now

Google: Gemini 3 Flash Preview (batch)

google/gemini-3-flash-preview:batch

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It

1.0M$0.25/M in · $1.5/M outRun now

Google: Gemini 3.1 Flash Lite

google/gemini-3.1-flash-lite

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, im

1.0M$0.25/M in · $1.5/M outRun now

Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)

google/gemini-3.1-flash-lite-image

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity develo

66K$0.25/M in · $1.5/M outRun now

Google: Gemini 3.1 Flash Lite Preview

google/gemini-3.1-flash-lite-preview

Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on

1.0M$0.25/M in · $1.5/M outRun now

Sao10K: Llama 3.1 Euryale 70B v2.2

sao10k/l3.1-euryale-70b

Euryale L3.1 70B v2.2 is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70

131K$0.85/M in · $0.85/M outRun now

Baidu: ERNIE 4.5 VL 424B A47B

baidu/ernie-4.5-vl-424b-a47b

ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu’s ERNIE 4.5 series, featuring 424B total parameters with 47

123K$0.42/M in · $1.25/M outRun now

Qwen2.5 Coder 32B Instruct

qwen/qwen-2.5-coder-32b-instruct

Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). Qwen2.5-Coder brings the follow

33K$0.66/M in · $1/M outRun now

Perceptron: Perceptron Mk1

perceptron/perceptron-mk1

Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and vid

33K$0.15/M in · $1.5/M outRun now

Qwen: Qwen3.7 Plus

qwen/qwen3.7-plus

Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the serie

1M1 runs$0.32/M in · $1.28/M outRun now

DeepSeek: R1 Distill Llama 70B

deepseek/deepseek-r1-distill-llama-70b

DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), usi

8K$0.8/M in · $0.8/M outRun now

DeepSeek: DeepSeek V4 Pro 0423

deepseek/deepseek-v4-pro

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting

1.0M$0.526/M in · $1.05/M outRun now

Anthropic: Claude 3 Haiku

anthropic/claude-3-haiku

Claude 3 Haiku is Anthropic's fastest and most compact model for near-instant responsiveness. Quick and accurate targeted performance. See

200K$0.25/M in · $1.25/M outRun now

Kwaipilot: KAT-Coder-Pro V2

kwaipilot/kat-coder-pro-v2

KAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engineer

262K$0.3/M in · $1.2/M outRun now

MiniMax: MiniMax M2-her

minimax/minimax-m2-her

MiniMax M2-her is a dialogue-first large language model built for immersive roleplay, character-driven chat, and expressive multi-turn conve

66K$0.3/M in · $1.2/M outRun now

MiniMax: MiniMax M2.1

minimax/minimax-m2.1

MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application develop

205K$0.3/M in · $1.2/M outRun now

MiniMax: MiniMax M3

minimax/minimax-m3

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context win

1.0M4 runs$0.3/M in · $1.2/M outRun now

MiniMax: MiniMax M3 (batch)

minimax/minimax-m3:batch

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context win

524K$0.3/M in · $1.2/M outRun now

Qwen: Qwen3.5-35B-A3B

qwen/qwen3.5-35b-a3b

The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms

262K$0.25/M in · $1.25/M outRun now

OpenAI: GPT-5.4 Nano

openai/gpt-5.4-nano

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. I

400K$0.2/M in · $1.25/M outRun now

Google: Gemini 2.5 Flash (batch)

google/gemini-2.5-flash:batch

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scient

1.0M$0.15/M in · $1.25/M outRun now

Nous: Hermes 3 70B Instruct

nousresearch/hermes-3-llama-3.1-70b

Hermes 3 is a generalist language model with many improvements over [Hermes 2](/models/nousresearch/nous-hermes-2-mistral-7b-dpo), including

131K$0.7/M in · $0.7/M outRun now

OpenAI: GPT-5.6 Luna

openai/gpt-5.6-luna

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat,

1.1M$0.2/M in · $1.2/M outRun now

OpenAI: GPT-5.6 Luna Pro

openai/gpt-5.6-luna-pro

GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set

1.1M$0.2/M in · $1.2/M outRun now

Sao10K: Llama 3.3 Euryale 70B

sao10k/l3.3-euryale-70b

Euryale L3.3 70B is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.

131K$0.65/M in · $0.75/M outRun now

MiniMax: MiniMax M2.5

minimax/minimax-m2.5

MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital w

205K$0.27/M in · $1.08/M outRun now

TheDrummer: Skyfall 36B V2

thedrummer/skyfall-36b-v2

Skyfall 36B v2 is an enhanced iteration of Mistral Small 2501, specifically fine-tuned for improved creativity, nuanced writing, role-playin

33K$0.55/M in · $0.8/M outRun now

Qwen: Qwen3 Next 80B A3B Thinking

qwen/qwen3-next-80b-a3b-thinking

Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’

262K$0.15/M in · $1.2/M outRun now

StepFun: Step 3.7 Flash

stepfun/step-3.7-flash

Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a v

262K$0.2/M in · $1.15/M outRun now

Qwen: Qwen3.6 Flash

qwen/qwen3.6-flash

Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token c

1M$0.188/M in · $1.13/M outRun now

Xiaomi: MiMo-V2.5-Pro

xiaomi/mimo-v2.5-pro

MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and l

1.1M$0.435/M in · $0.87/M outRun now

Google: Gemma 2 27B

google/gemma-2-27b-it

Gemma 2 27B by Google is an open model built from the same research and technology used to create the [Gemini models](/models?q=gemini). Gem

8K$0.65/M in · $0.65/M outRun now

MiniMax: MiniMax-01

minimax/minimax-01

MiniMax-01 is a combines MiniMax-Text-01 for text generation and MiniMax-VL-01 for image understanding. It has 456 billion parameters, with

1.0M$0.2/M in · $1.1/M outRun now

Qwen: Qwen3 Coder 480B A35B

qwen/qwen3-coder

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic c

262K$0.3/M in · $1/M outRun now

DeepSeek: DeepSeek V3

deepseek/deepseek-chat

DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous version

164K$0.257/M in · $1.03/M outRun now

MiniMax: MiniMax M2

minimax/minimax-m2

MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activat

205K$0.255/M in · $1.02/M outRun now

DeepSeek: DeepSeek V3.1 Terminus

deepseek/deepseek-v3.1-terminus

DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while

164K$0.27/M in · $1/M outRun now

DeepSeek: DeepSeek V3 0324

deepseek/deepseek-chat-v3-0324

DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. I

164K$0.25/M in · $1/M outRun now

Mancer: Weaver (alpha)

mancer/weaver

An attempt to recreate Claude-style verbosity, but don't expect the same level of coherence or memory. Meant for use in roleplay/narrative s

8K$0.5/M in · $0.75/M outRun now

Nex AGI: Nex-N2-Pro

nex-agi/nex-n2-pro

Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total. Built on the Qwen3.5 architect

262K$0.25/M in · $1/M outRun now

WizardLM-2 8x22B

microsoft/wizardlm-2-8x22b

WizardLM-2 8x22B is Microsoft AI's most advanced Wizard model. It demonstrates highly competitive performance compared to leading proprietar

66K$0.62/M in · $0.62/M outRun now

Qwen: Qwen3 Next 80B A3B Instruct

qwen/qwen3-next-80b-a3b-instruct

Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinki

262K$0.1/M in · $1.1/M outRun now

MiniMax: MiniMax M2.7

minimax/minimax-m2.7

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to

205K$0.24/M in · $0.96/M outRun now

Mistral: Codestral 2508

mistralai/codestral-2508

Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such

256K$0.3/M in · $0.9/M outRun now

Z.ai: GLM 4.6V

z-ai/glm-4.6v

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, an

131K$0.3/M in · $0.9/M outRun now

Qwen: Qwen3 Coder Flash

qwen/qwen3-coder-flash

Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus. It is a powerful coding agent model sp

1M$0.195/M in · $0.975/M outRun now

Qwen: Qwen3.6 35B A3B

qwen/qwen3.6-35b-a3b

Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per t

262K$0.14/M in · $1/M outRun now

OpenAI: GPT-5 Mini (batch)

openai/gpt-5-mini:batch

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and

400K$0.125/M in · $1/M outRun now

ReMM SLERP 13B

undi95/remm-slerp-l2-13b

A recreation trial of the original MythoMax-L2-B13 but with updated models. #merge

6K$0.45/M in · $0.65/M outRun now

Venice: Uncensored

cognitivecomputations/dolphin-mistral-24b-venice-edition

Venice Uncensored Dolphin Mistral 24B Venice Edition is a fine-tuned variant of Mistral-Small-24B-Instruct-2501, developed by dphn.ai in col

128K$0.2/M in · $0.9/M outRun now

Arcee AI: Trinity Large Thinking

arcee-ai/trinity-large-thinking

Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agent

262K$0.22/M in · $0.85/M outRun now

Qwen: Qwen-Plus

qwen/qwen-plus

Qwen-Plus, based on the Qwen2.5 foundation model, is a 131K context model with a balanced performance, speed, and cost combination.

1M$0.26/M in · $0.78/M outRun now

Qwen: Qwen Plus 0728

qwen/qwen-plus-2025-07-28

Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, and c

1M$0.26/M in · $0.78/M outRun now

Qwen: Qwen Plus 0728 (thinking)

qwen/qwen-plus-2025-07-28:thinking

Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, and c

1M$0.26/M in · $0.78/M outRun now

Inception: Mercury 2

inception/mercury-2

Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercu

128K$0.25/M in · $0.75/M outRun now

OpenAI: GPT-3.5 Turbo (batch)

openai/gpt-3.5-turbo:batch

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional

16K$0.25/M in · $0.75/M outRun now

Meta: Llama 4 Maverick

meta-llama/llama-4-maverick

Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architectur

1.0M$0.2/M in · $0.8/M outRun now

OpenAI: GPT-4.1 Mini (batch)

openai/gpt-4.1-mini:batch

GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 mil

1.0M$0.2/M in · $0.8/M outRun now

Z.ai: GLM 4.5 Air

z-ai/glm-4.5-air

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5,

131K$0.13/M in · $0.85/M outRun now

Qwen: Qwen3 Coder Next

qwen/qwen3-coder-next

Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse MoE d

262K$0.12/M in · $0.8/M outRun now

Mistral: Mistral Small 3.1 24B

mistralai/mistral-small-3.1-24b-instruct

Mistral Small 3.1 24B Instruct is an upgraded variant of Mistral Small 3 (2501), featuring 24 billion parameters with advanced multimodal ca

128K$0.351/M in · $0.555/M outRun now

Google: Gemini 3.1 Flash Lite (batch)

google/gemini-3.1-flash-lite:batch

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, im

1.0M$0.125/M in · $0.75/M outRun now

TheDrummer: Cydonia 24B V4.1

thedrummer/cydonia-24b-v4.1

Uncensored and creative writing model based on Mistral Small 3.2 24B with good recall, prompt adherence, and intelligence.

131K$0.3/M in · $0.5/M outRun now

Meta: Llama 3.1 70B Instruct

meta-llama/llama-3.1-70b-instruct

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 70B instruct-tuned version is optimized for high q

131K$0.4/M in · $0.4/M outRun now

Mistral: Saba

mistralai/mistral-saba

Mistral Saba is a 24B-parameter language model specifically designed for the Middle East and South Asia, delivering accurate and contextuall

33K$0.2/M in · $0.6/M outRun now

TheDrummer: UnslopNemo 12B

thedrummer/unslopnemo-12b

UnslopNemo v4.1 is the latest addition from the creator of Rocinante, designed for adventure writing and role-play scenarios.

1.0M$0.4/M in · $0.4/M outRun now

Tencent: Hy3 preview

tencent/hy3-preview

Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports config

262K$0.18/M in · $0.6/M outRun now

Qwen2.5 72B Instruct

qwen/qwen-2.5-72b-instruct

Qwen2.5 72B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more k

33K$0.36/M in · $0.4/M outRun now

Cohere: Command R (08-2024)

cohere/command-r-08-2024

command-r-08-2024 is an update of the [Command R](/models/cohere/command-r) with improved performance for multilingual retrieval-augmented g

128K$0.15/M in · $0.6/M outRun now

Mistral: Mistral Small 4

mistralai/mistral-small-2603

Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a s

262K$0.15/M in · $0.6/M outRun now

OpenAI: GPT-4o-mini

openai/gpt-4o-mini

GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As

128K1 runs$0.15/M in · $0.6/M outRun now

OpenAI: GPT-4o-mini (2024-07-18)

openai/gpt-4o-mini-2024-07-18

GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As

128K$0.15/M in · $0.6/M outRun now

TheDrummer: Rocinante 12B

thedrummer/rocinante-12b

Rocinante 12B is designed for engaging storytelling and rich prose. Early testers have reported: - Expanded vocabulary with unique and expre

66K$0.25/M in · $0.5/M outRun now

Upstage: Solar Pro 3

upstage/solar-pro-3

Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward

131K$0.15/M in · $0.6/M outRun now

OpenAI: GPT-5.4 Nano (batch)

openai/gpt-5.4-nano:batch

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. I

400K$0.1/M in · $0.625/M outRun now

Tencent: Hunyuan A13B Instruct

tencent/hunyuan-a13b-instruct

Hunyuan-A13B is a 13B active parameter Mixture-of-Experts (MoE) language model developed by Tencent, with a total parameter count of 80B and

131K$0.14/M in · $0.57/M outRun now

inclusionAI: Ling-2.6-1T

inclusionai/ling-2.6-1t

Ling-2.6-1T is an instant (instruct) model from inclusionAI and the company’s trillion-parameter flagship, designed for real-world agents th

262K$0.075/M in · $0.625/M outRun now

inclusionAI: Ring-2.6-1T

inclusionai/ring-2.6-1t

Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both strong

262K$0.075/M in · $0.625/M outRun now

OpenAI: GPT-5.6 Luna Pro (batch)

openai/gpt-5.6-luna-pro:batch

GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set

1.1M$0.1/M in · $0.6/M outRun now

OpenAI: GPT-5.6 Luna (batch)

openai/gpt-5.6-luna:batch

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat,

1.1M$0.1/M in · $0.6/M outRun now

DeepSeek: DeepSeek V3.2 Exp

deepseek/deepseek-v3.2-exp

DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures

164K$0.27/M in · $0.41/M outRun now

Tencent: Hy3

tencent/hy3

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic work

262K$0.132/M in · $0.528/M outRun now

AllenAI: Olmo 3 32B Think

allenai/olmo-3-32b-think

Olmo 3 32B Think is a large-scale, 32-billion-parameter model purpose-built for deep reasoning, complex logic chains and advanced instructio

66K$0.15/M in · $0.5/M outRun now

Qwen: Qwen3 VL 30B A3B Instruct

qwen/qwen3-vl-30b-a3b-instruct

Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Ins

262K$0.13/M in · $0.52/M outRun now

DeepSeek: DeepSeek V3.2

deepseek/deepseek-v3.2

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use perfo

164K$0.26/M in · $0.38/M outRun now

Qwen: Qwen3 235B A22B Instruct 2507

qwen/qwen3-235b-a22b-2507

Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, w

262K$0.09/M in · $0.55/M outRun now

Qwen: Qwen3 30B A3B

qwen/qwen3-30b-a3b

Qwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to exce

131K$0.12/M in · $0.5/M outRun now

Qwen: Qwen3 8B

qwen/qwen3-8b

Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialog

131K$0.117/M in · $0.455/M outRun now

Qwen: Qwen3 VL 8B Instruct

qwen/qwen3-vl-8b-instruct

Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning acr

262K$0.117/M in · $0.455/M outRun now

Nous: Hermes 4 70B

nousresearch/hermes-4-70b

Hermes 4 70B is a hybrid reasoning model from Nous Research, built on Meta-Llama-3.1-70B. It introduces the same hybrid mode as the larger 4

131K$0.13/M in · $0.4/M outRun now

Google: Gemma 3 27B

google/gemma-3-27b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understan

262K$0.08/M in · $0.45/M outRun now

Qwen: Qwen3 VL 32B Instruct

qwen/qwen3-vl-32b-instruct

Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text,

131K$0.104/M in · $0.416/M outRun now

ByteDance Seed: Seed-2.0-Mini

bytedance-seed/seed-2.0-mini

Seed-2.0-mini targets latency-sensitive, high-concurrency, and cost-sensitive scenarios, emphasizing fast response and flexible inference de

262K$0.1/M in · $0.4/M outRun now

Google: Gemini 2.5 Flash Lite

google/gemini-2.5-flash-lite

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It off

1.0M1 runs$0.1/M in · $0.4/M outRun now

OpenAI: GPT-4.1 Nano

openai/gpt-4.1-nano

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance

1.0M$0.1/M in · $0.4/M outRun now

NVIDIA: Nemotron 3 Super

nvidia/nemotron-3-super-120b-a12b

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accurac

1M$0.085/M in · $0.4/M outRun now

Z.ai: GLM 4.7 Flash

z-ai/glm-4.7-flash

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic c

203K$0.06/M in · $0.4/M outRun now

OpenAI: GPT-5 Nano

openai/gpt-5-nano

GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency

400K$0.05/M in · $0.4/M outRun now

Google: Gemma 4 31B

google/gemma-4-31b-it

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K tok

262K$0.1/M in · $0.34/M outRun now

Xiaomi: MiMo-V2.5

xiaomi/mimo-v2.5

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpass

1.1M$0.14/M in · $0.28/M outRun now

Meta: Llama 3.3 70B Instruct

meta-llama/llama-3.3-70b-instruct

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out).

131K$0.1/M in · $0.32/M outRun now

Google: Gemma 4 26B A4B

google/gemma-4-26b-a4b-it

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B ac

262K$0.07/M in · $0.34/M outRun now

Meta: Llama 4 Scout

meta-llama/llama-4-scout

Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a t

1.3M$0.1/M in · $0.3/M outRun now

Mistral: Ministral 3 14B 2512

mistralai/ministral-14b-2512

The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral S

262K$0.2/M in · $0.2/M outRun now

Mistral: Voxtral Small 24B 2507

mistralai/voxtral-small-24b-2507

Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class te

32K$0.1/M in · $0.3/M outRun now

StepFun: Step 3.5 Flash

stepfun/step-3.5-flash

Step 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it selective

262K$0.1/M in · $0.3/M outRun now

Meta: Llama 3.2 3B Instruct

meta-llama/llama-3.2-3b-instruct

Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialo

131K$0.05/M in · $0.33/M outRun now

ByteDance Seed: Seed 1.6 Flash

bytedance-seed/seed-1.6-flash

Seed 1.6 Flash is an ultra-fast multimodal deep thinking model by ByteDance Seed, supporting both text and visual understanding. It features

262K$0.075/M in · $0.3/M outRun now

OpenAI: GPT-4o-mini (batch)

openai/gpt-4o-mini:batch

GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As

128K$0.075/M in · $0.3/M outRun now

OpenAI: gpt-oss-safeguard-20b

openai/gpt-oss-safeguard-20b

gpt-oss-safeguard-20b is a safety reasoning model from OpenAI built upon gpt-oss-20b. This open-weight, 21B-parameter Mixture-of-Experts (Mo

131K$0.075/M in · $0.3/M outRun now

Qwen: Qwen3 32B

qwen/qwen3-32b

Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogu

131K$0.08/M in · $0.28/M outRun now

Meta: Llama Guard 4 12B

meta-llama/llama-guard-4-12b

Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous vers

1.0M$0.18/M in · $0.18/M outRun now

Qwen: Qwen3 14B

qwen/qwen3-14b

Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue

131K$0.12/M in · $0.24/M outRun now

Qwen: Qwen3 Coder 30B A3B Instruct

qwen/qwen3-coder-30b-a3b-instruct

Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for

262K$0.07/M in · $0.28/M outRun now

Qwen: Qwen3.5-Flash

qwen/qwen3.5-flash-02-23

The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a spars

1M$0.065/M in · $0.26/M outRun now

Amazon: Nova Lite 1.0

amazon/nova-lite-v1

Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to gen

300K$0.06/M in · $0.24/M outRun now

ByteDance: UI-TARS 7B

bytedance/ui-tars-1.5-7b

UI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, including desktop interfaces, web browsers, mobile s

128K$0.1/M in · $0.2/M outRun now

Mistral: Ministral 3 8B 2512

mistralai/ministral-8b-2512

A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.

262K$0.15/M in · $0.15/M outRun now

Qwen: Qwen2.5 7B Instruct

qwen/qwen-2.5-7b-instruct

Qwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more kn

33K$0.1/M in · $0.2/M outRun now

Reka Flash 3

rekaai/reka-flash-3

Reka Flash 3 is a general-purpose, instruction-tuned large language model with 21 billion parameters, developed by Reka. It excels at genera

66K$0.1/M in · $0.2/M outRun now

Mistral: Mistral Small 3.2 24B

mistralai/mistral-small-3.2-24b-instruct

Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduction

131K$0.075/M in · $0.2/M outRun now

Qwen: Qwen3.5-9B

qwen/qwen3.5-9b

Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding

262K$0.1/M in · $0.15/M outRun now

Google: Gemini 2.5 Flash Lite (batch)

google/gemini-2.5-flash-lite:batch

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It off

1.0M$0.05/M in · $0.2/M outRun now

NVIDIA: Nemotron 3 Nano 30B A3B

nvidia/nemotron-3-nano-30b-a3b

NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialize

262K$0.05/M in · $0.2/M outRun now

OpenAI: GPT-4.1 Nano (batch)

openai/gpt-4.1-nano:batch

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance

1.0M$0.05/M in · $0.2/M outRun now

Qwen: Qwen3 30B A3B Instruct 2507

qwen/qwen3-30b-a3b-instruct-2507

Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. It

262K$0.048/M in · $0.193/M outRun now

Meta: Llama 3.2 1B Instruct

meta-llama/llama-3.2-1b-instruct

Llama 3.2 1B is a 1-billion-parameter language model focused on efficiently performing natural language tasks, such as summarization, dialog

60K$0.027/M in · $0.201/M outRun now

OpenAI: GPT-5 Nano (batch)

openai/gpt-5-nano:batch

GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency

400K$0.025/M in · $0.2/M outRun now

Mistral: Ministral 8B

mistralai/ministral-8b

Ministral 8B is an 8B parameter model featuring a unique interleaved sliding-window attention pattern for faster, memory-efficient inference

128K$0.11/M in · $0.11/M outRun now

Microsoft: Phi 4

microsoft/phi-4

[Microsoft Research](/microsoft) Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with

16K$0.07/M in · $0.14/M outRun now

OpenAI: gpt-oss-120b

openai/gpt-oss-120b

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and

131K$0.037/M in · $0.17/M outRun now

Google: Gemma 3 12B

google/gemma-3-12b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understan

131K$0.05/M in · $0.15/M outRun now

Mistral: Ministral 3 3B 2512

mistralai/ministral-3b-2512

The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.

131K$0.1/M in · $0.1/M outRun now

Reka Edge

rekaai/reka-edge

Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. Thi

16K$0.1/M in · $0.1/M outRun now

Cohere: Command R7B (12-2024)

cohere/command-r7b-12-2024

Command R7B (12-2024) is a small, fast update of the Command R+ model, delivered in December 2024. It excels at RAG, tool use, agents, and s

128K$0.037/M in · $0.15/M outRun now

Google: Gemma 3n 4B

google/gemma-3n-e4b-it

Gemma 3n E4B-it is optimized for efficient execution on mobile and low-resource devices, such as phones, laptops, and tablets. It supports m

33K$0.06/M in · $0.12/M outRun now

Poolside: Laguna XS 2.1

poolside/laguna-xs-2.1

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their L

262K$0.06/M in · $0.12/M outRun now

Amazon: Nova Micro 1.0

amazon/nova-micro-v1

Amazon Nova Micro 1.0 is a text-only model that delivers the lowest latency responses in the Amazon Nova family of models at a very low cost

128K$0.035/M in · $0.14/M outRun now

DeepSeek: DeepSeek V4 Flash 0423

deepseek/deepseek-v4-flash

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters,

1.0M$0.057/M in · $0.115/M outRun now

OpenAI: gpt-oss-20b

openai/gpt-oss-20b

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) archit

131K$0.03/M in · $0.13/M outRun now

Google: Gemma 3 4B

google/gemma-3-4b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understan

131K$0.05/M in · $0.1/M outRun now

IBM: Granite 4.1 8B

ibm-granite/granite-4.1-8b

Granite 4.1 8B is a dense, decoder-only 8-billion-parameter language model from IBM, part of the Granite 4.1 family. It supports a 131K-toke

131K$0.05/M in · $0.1/M outRun now

Meta: Llama 3.1 8B Instruct

meta-llama/llama-3.1-8b-instruct

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. I

131K$0.05/M in · $0.08/M outRun now

Mistral: Mistral Small 3

mistralai/mistral-small-24b-instruct-2501

Mistral Small 3 is a 24B-parameter language model optimized for low-latency performance across common AI tasks. Released under the Apache 2.

33K$0.05/M in · $0.08/M outRun now

IBM: Granite 4.0 Micro

ibm-granite/granite-4.0-h-micro

Granite-4.0-H-Micro is a 3B parameter from the Granite 4 family of models. These models are the latest in a series of models released by IBM

131K$0.017/M in · $0.112/M outRun now

Nex AGI: Nex-N2-Mini

nex-agi/nex-n2-mini

Nex-N2-Mini is an open-source agentic mixture-of-experts model from Nex AGI, the smaller sibling in the Nex-N2 series. It accepts text and i

262K$0.025/M in · $0.1/M outRun now

MythoMax 13B

gryphe/mythomax-l2-13b

One of the highest performing and most popular fine-tunes of Llama 2 13B, with rich descriptions and roleplay. #merge

8K$0.06/M in · $0.06/M outRun now

Sao10K: Llama 3 8B Lunaris

sao10k/l3-lunaris-8b

Lunaris 8B is a versatile generalist and roleplaying model based on Llama 3. It's a strategic merge of multiple models, designed to balance

8K$0.04/M in · $0.05/M outRun now

Mistral: Mistral Nemo

mistralai/mistral-nemo

A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. The model is multilingual, supporting

131K$0.019/M in · $0.03/M outRun now

inclusionAI: Ling-2.6-flash

inclusionai/ling-2.6-flash

Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-worl

262K$0.01/M in · $0.03/M outRun now

Cohere: North Mini Code (free)

cohere/north-mini-code:free

North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total p

256KFree in · Free outRun now

Google: Gemma 4 26B A4B (free)

google/gemma-4-26b-a4b-it:free

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B ac

262KFree in · Free outRun now

Google: Gemma 4 31B (free)

google/gemma-4-31b-it:free

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K tok

262KFree in · Free outRun now

Google: Lyria 3 Clip Preview

google/lyria-3-clip-preview

30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemini A

1.0MFree in · Free outRun now

Google: Lyria 3 Pro Preview

google/lyria-3-pro-preview

Full-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. Wit

1.0MFree in · Free outRun now

NVIDIA: Nemotron 3 Nano 30B A3B (free)

nvidia/nemotron-3-nano-30b-a3b:free

NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialize

256KFree in · Free outRun now

NVIDIA: Nemotron 3 Nano Omni (free)

nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free

NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise age

256KFree in · Free outRun now

NVIDIA: Nemotron 3 Super (free)

nvidia/nemotron-3-super-120b-a12b:free

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accurac

262KFree in · Free outRun now

NVIDIA: Nemotron 3 Ultra (free)

nvidia/nemotron-3-ultra-550b-a55b:free

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE

1MFree in · Free outRun now

NVIDIA: Nemotron 3.5 Content Safety (free)

nvidia/nemotron-3.5-content-safety:free

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It m

128KFree in · Free outRun now

NVIDIA: Nemotron Nano 12B 2 VL (free)

nvidia/nemotron-nano-12b-v2-vl:free

NVIDIA Nemotron Nano 2 VL is a 12-billion-parameter open multimodal reasoning model designed for video understanding and document intelligen

128KFree in · Free outRun now

NVIDIA: Nemotron Nano 9B V2 (free)

nvidia/nemotron-nano-9b-v2:free

NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasonin

128KFree in · Free outRun now

Free Models Router

openrouter/free

The simplest way to get free inference. openrouter/free is a router that selects free models at random from the models available on OpenRout

200KFree in · Free outRun now

Poolside: Laguna XS 2.1 (free)

poolside/laguna-xs-2.1:free

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their L

262KFree in · Free outRun now

Z.ai: GLM 5.2 (free)

z-ai/glm-5.2:free

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long

256KFree in · Free outRun now

Auto Router

openrouter/auto

Your prompt will be processed by a meta-model and routed to one of dozens of models (see below), optimizing for the best possible output. To

2MFree in · Free outRun now

Body Builder (beta)

openrouter/bodybuilder

Transform your natural language requests into structured OpenRouter API request objects. Describe what you want to accomplish with AI models

128KFree in · Free outRun now

OpenRouter: Fusion

openrouter/fusion

Fusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel with w

1MFree in · Free outRun now

Pareto Code Router

openrouter/pareto-code

The Pareto Router maintains a tiered shortlist of strong coding models, ranked by [Artificial Analysis](https://artificialanalysis.ai/) codi

2MFree in · Free outRun now