Every model, explained and priced.
Filter by what you’re building, search by capability or tag, and compare cost, speed, and benchmarks — on one OpenAI-compatible key.
What are you building?
Pick a job to narrow the catalog and surface leaders.
Fast, low-cost Nova for high-volume text and light multimodal tasks.
Smallest Nova — text-only ultra-cheap option for classification and routing.
Bedrock Nova Pro — multimodal workhorse with 300K context.
Anthropic's next-generation model for complex knowledge work and coding, built for sustained autonomous operation across multi-day tasks — plans across stages, delegates to sub-agents, and self-verifies its work.
Fastest and most compact Claude model. Ideal for classification, extraction, and high-volume tasks where cost and latency matter most.
Most capable Claude model. Sets the bar on advanced reasoning, research synthesis, and nuanced long-form content.
The latest Opus 4.x model — advanced reasoning and research synthesis, with fast mode available for premium-priced low-latency responses.
Balanced intelligence and speed. Handles complex coding, analysis, and multi-step reasoning without the cost of Opus.
Balanced Claude 5.x model with the full 1M-token context window at standard pricing. Handles complex coding, analysis, and multi-step reasoning at Sonnet-tier cost.
Mistral’s code-specialized model — fill-in-the-middle and multi-language coding.
Cohere Command R — efficient RAG and tool use at budget rates.
Cohere’s RAG-oriented flagship — strong retrieval-augmented generation and tool use.
OpenAI image generation — priced per image upstream.
DeepSeek’s chat alias — V3-class coding and general chat at low cost.
Open-source reasoning model rivalling o1. Publishes its chain of thought — useful for auditability and research workflows.
Mixture-of-experts model with strong coding and math benchmarks at a fraction of frontier model prices.
Bedrock-hosted DeepSeek V3.2 — coding and math at competitive MoE rates.
DeepSeek's flagship (Pro tier) — 1M-token context with strong reasoning and coding at a fraction of frontier-tier cost.
Mistral Devstral 2 — large coding-focused model on Bedrock.
Prior Gemini Flash — cheap long-context option.
Prior Gemini Pro generation — 1M context for long documents.
Low-latency workhorse with a 1M context window and multimodal inputs.
Best-value Gemini model. Handles massive context at low cost with strong performance on coding and document analysis.
Gemini 2.5 Flash with native audio — preview long-context multimodal.
Google's highest-capability model with native multimodal reasoning across text, image, video, and audio — all within a 1M-token window.
Previous Flash tier — prefer 3.6 Flash for new workloads.
High-throughput Flash-Lite for extraction, classification, and subagent loops.
July 2026 Flash workhorse — stronger coding/agentic loops than 3.5 at lower output cost, 1M context.
Open Gemma 3 12B instruct — compact open-weight generalist on Bedrock.
Open Gemma 3 27B instruct — strong open-weight generalist on Bedrock.
Smallest Gemma 3 instruct — ultra-cheap classification and light chat.
Shared community capacity for Google Gemma 4 26B A4B IT (multimodal).
Shared community capacity for Google Gemma 4 31B IT. Rate-limited free pool.
Prior GLM-4.7 generation — capable bilingual generalist.
Fast GLM-4.7 tier for high-throughput bilingual workloads.
Zhipu GLM-5 — strong Chinese/English bilingual reasoning and coding.
Lightweight 4.1 variant with the same 1M-token window at a fraction of the cost. Ideal for agentic pipelines that need long context cheaply.
Cheapest 4.1-family model — high-throughput extraction and subagent loops with 1M context.
Flagship multimodal model optimised for real-time dialogue. Natively processes text, images, and structured data.
GPT-4o with native audio input/output (preview).
Affordable GPT-4o sibling that handles everyday tasks with excellent speed. Strong default for high-volume deployments.
OpenAI's most capable model — advanced coding, research, analysis, software operation, document workflows, and long-running agentic tasks with less orchestration.
Successor to GPT-5.5 (Sol tier) — OpenAI's current flagship for advanced coding, research, analysis, and long-running agentic tasks.
Groq-hosted open-weight flagship replacing Llama 3.3 70B. Strong reasoning at low latency.
Open GPT-OSS 120B via Bedrock — large open-weight generalist.
Groq-hosted open-weight model replacing Llama 3.1 8B. Fast, cheap, and ideal for high-volume prototyping.
Open GPT-OSS 20B via Bedrock — compact open-weight option.
Shared community capacity for GPT-OSS 20B. Strong default for light prototyping.
Safety-tuned GPT-OSS 120B on Bedrock for moderation-style workloads.
Safety-tuned GPT-OSS 20B on Bedrock — compact moderation helper.
Prior Grok flagship — strong reasoning and coding with a 131K context window.
Lightweight Grok 3 — fast and cheap for lighter agent loops.
Long-context Grok — 1M tokens with strong reasoning at a lower rate than Grok 4.5.
xAI's current frontier model — strong reasoning and agentic coding with a 500K-token context window.
Thinking-mode Kimi for multi-step reasoning and research synthesis.
General Kimi K2.5 — long-context reasoning and coding at Bedrock rates.
Moonshot coding-focused Kimi — 256K context for large repos and agent loops.
Shared community capacity for Poolside Laguna S 2.1.
Shared community capacity for Poolside Laguna XS 2.1 — compact coding model.
Shared community capacity for InclusionAI Ling 3.0 Flash (MoE).
Meta Llama 3 70B instruct on Bedrock — classic open-weight workhorse.
Meta Llama 3 8B instruct — small, fast open-weight option on Bedrock.
Meta Llama 3.1 405B instruct via Fireworks.
Meta Llama 3.1 70B instruct via Fireworks.
Groq speculative-decoding Llama 3.3 70B — faster throughput variant.
Mistral Magistral Small — reasoning-oriented small model on Bedrock.
MiniMax M2 — long-context agentic model via Bedrock.
MiniMax M2.1 — long-context agentic model via Bedrock.
MiniMax M2.5 — long-context agentic model via Bedrock.
Ministral 3 14B — mid-size open instruct model on Bedrock.
Tiny Ministral 3 — ultra-cheap classification and light chat on Bedrock.
Compact Ministral 3 8B — small open instruct model on Bedrock.
Classic Mistral 7B instruct on Bedrock.
Flagship Mistral model with top-tier coding, reasoning, and multilingual performance. Native function calling support.
Mistral's current flagship (675B) — top-tier coding, reasoning, and multilingual performance, routed via Bedrock.
Efficient European model for classification, summarisation, and structured extraction. Strong on French and other EU languages.
Mixtral MoE instruct on Bedrock — classic open MoE generalist.
Shared community capacity for NVIDIA Nemotron 3 Nano 30B A3B.
Shared community capacity for NVIDIA Nemotron 3 Nano Omni — multimodal reasoning.
Shared community capacity for NVIDIA Nemotron 3 Super.
Shared community capacity for NVIDIA Nemotron 3 Ultra (MoE). Large free context window.
Shared community capacity for NVIDIA Nemotron 3.5 content-safety classifier.
NVIDIA Nemotron Nano 12B — slightly larger nano tier for general chat.
Shared community capacity for NVIDIA Nemotron Nano 12B V2 vision-language.
NVIDIA Nemotron Nano 3 30B — open reasoning-capable mid-size model.
NVIDIA Nemotron Nano 9B — compact open model for light workloads.
Shared community capacity for NVIDIA Nemotron Nano 9B v2. Good for low-stakes tasks.
NVIDIA Nemotron Super 3 120B — open reasoning and coding at Bedrock rates.
Shared community capacity for Cohere North Mini Code.
OpenAI's most powerful reasoning model. Spends more compute thinking before answering — ideal for math, science, and strategic planning.
Compact reasoning model tuned for STEM tasks. Faster and cheaper than o3 while retaining strong chain-of-thought capabilities.
Next-generation compact reasoning model. Faster than o3-mini with improved accuracy on code, math, and agentic workflows.
Dense Qwen3 32B — solid generalist at low cost on Bedrock.
Qwen3 MoE coder — strong coding at low cost with 256K context.
Latest Qwen3 coding model — strong agentic coding with 256K context.
Qwen3 Next MoE — long-context generalist for analysis and agents.
Large vision-language Qwen3 — document and image understanding at MoE scale.
OpenAI embedding model — higher-dimensional vectors.
OpenAI embedding model — small, cheap vectors for retrieval.
Mistral Voxtral Mini — compact audio-aware model on Bedrock.
Mistral Voxtral Small — larger audio-aware model on Bedrock.
OpenAI speech-to-text — priced per minute upstream.
Three steps to your first call
Getting startedCreate a free account
Email and a password. No card, no sales call. Community models are free from the first minute.
Create an inference key
Name it, scope it to the models you want, and set a spend ceiling. The cap is enforced before each request, not reconciled after.
Change one line
Point your existing OpenAI-compatible client at the Relixr base URL. Same SDK, same request shape, same response shape.