Models

OpenAI: GPT-6 Astra Pro

NewT
55.3M tokens

OpenAI's most capable frontier reasoning and multimodal model for highly complex workflows, software engineering, and scientific tasks.

by openaiSep 9, 20261.05M context$10/M input tokens$50/M output tokens

OpenAI: GPT-6 Astra

NewT
1.1M tokens

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, and long-horizon tasks.

by openaiSep 4, 20261.05M context$10/M input tokens$50/M output tokens

OpenAI: GPT-5.6 Sol

T
18.2M tokens

Frontier balance between deep reasoning, extensive context, and real-time generation speed for everyday production workflows.

by openaiAug 28, 20261.05M context$5/M input tokens$25/M output tokens

OpenAI: GPT-5.6 Terra

T
32.8M tokens

High-efficiency frontier model optimized for enterprise data pipelines, document extraction, and long-form analytical synthesis.

by openaiAug 20, 20261.05M context$2.5/M input tokens$12.5/M output tokens

OpenAI: GPT-5.6 Luna

50% offT
89.4M tokens

Ultra-fast low-cost model with full 1M+ context window for high-volume automation and classification.

by openaiAug 15, 20261.05M context$0.5/M input tokens$2/M output tokens

OpenAI: GPT-5

T
24.1M tokens

Powerful reasoning and general intelligence engine with multimodal vision support for creative and analytical generation.

by openaiAug 10, 2026400K context$1.25/M input tokens$10/M output tokens

OpenAI: GPT-5 mini

T
142.0M tokens

Cost-effective, low-latency reasoning and conversational intelligence for consumer-facing apps.

by openaiAug 10, 2026400K context$0.25/M input tokens$2/M output tokens

OpenAI: GPT-5 nano

T
61.3M tokens

Ultra-compact sub-second execution model for edge tasks, summarization, and content classification.

by openaiAug 5, 2026400K context$0.05/M input tokens$0.4/M output tokens

OpenAI: GPT-4.1

T
74.5M tokens

Reliable workhorse model with massive 1M token context for long document synthesis and legal review.

by openaiJul 12, 20261M context$2/M input tokens$8/M output tokens

OpenAI: GPT-4.1 mini

T
115.8M tokens

Affordable 1M-context model for processing large codebases and complex multi-document repositories.

by openaiJul 12, 20261M context$0.4/M input tokens$1.6/M output tokens

OpenAI: GPT-4.1 nano

T
44.1M tokens

Lowest latency 1M-context variant for lightweight summarization and routing pipelines.

by openaiJul 8, 20261M context$0.1/M input tokens$0.4/M output tokens

OpenAI: GPT-4o

T
310.2M tokens

Flagship omni-channel model for reasoning across text, code, and vision with low latency.

by openaiMay 13, 2026128K context$2.5/M input tokens$10/M output tokens

OpenAI: GPT-4o mini

T
580.4M tokens

High-speed, cost-efficient model for lightweight multimodal operations and structured outputs.

by openaiMay 13, 2026128K context$0.15/M input tokens$0.6/M output tokens

OpenAI: o3

T
98.7M tokens

Advanced deliberate reasoning model tailored for science, math, and competitive programming benchmarks.

by openaiJun 1, 2026200K context$2/M input tokens$8/M output tokens

OpenAI: o4-mini

T
82.4M tokens

Fast and affordable deep reasoning model designed for STEM, theorem proving, and analytical logic.

by openaiJun 10, 2026200K context$1.1/M input tokens$4.4/M output tokens

OpenAI: GPT OSS 120B

T
38.6M tokens

Open-weights foundation model offering transparency, fine-tunability, and rapid throughput.

by openaiMay 20, 2026128K context$0.15/M input tokens$0.6/M output tokens

Anthropic: Claude Opus 4.7

T
76.4M tokens

Anthropic's pinnacle intelligence for deep analysis, creative synthesis, and multi-step complex coding.

by anthropicSep 7, 2026200K context$5/M input tokens$25/M output tokens

Anthropic: Claude Sonnet 4

T
245.9M tokens

Top-tier balance of frontier intelligence, speed, and vision reasoning for daily software engineering.

by anthropicSep 1, 2026200K context$3/M input tokens$15/M output tokens

Anthropic: Claude 3.5 Sonnet v2

T
190.1M tokens

Industry-leading model for computer use, code generation, and visual understanding.

by anthropicAug 14, 2026200K context$6/M input tokens$30/M output tokens

Amazon: Amazon Nova Premier

T
42.0M tokens

Amazon's most capable multimodal model for complex reasoning across text, image, and high-framerate video.

by amazonAug 29, 20261M context$2.5/M input tokens$12.5/M output tokens

Amazon: Amazon Nova Pro

T
58.1M tokens

High-accuracy multimodal model tailored for enterprise visual reasoning, documents, and agentic tasks.

by amazonAug 20, 2026300K context$0.8/M input tokens$3.2/M output tokens

Amazon: Amazon Nova Lite

50% offT
128.5M tokens

Very low-cost, lightning-fast multimodal processing across text, image, and video media streams.

by amazonAug 18, 2026300K context$0.06/M input tokens$0.24/M output tokens

Amazon: Amazon Nova Micro

T
210.0M tokens

Text-only model offering lowest latency and rock-bottom inference costs for classification and intent routing.

by amazonAug 15, 2026128K context$0.035/M input tokens$0.14/M output tokens
G

Google: Gemma 4 31B

T
64.2M tokens

Google's high-efficiency open model engineered for versatile linguistic and code understanding.

by googleJul 25, 2026128K context$0.14/M input tokens$0.4/M output tokens
G

Google: Gemma 4 E2B

T
49.8M tokens

Ultra-compact open model optimized for lightweight tasks and responsive chat completions.

by googleJul 20, 2026128K context$0.04/M input tokens$0.08/M output tokens
DS

DeepSeek: DeepSeek R1

Free tierT
412.5M tokens

Open-architecture reasoning model matching top frontier models on math, code, and logic benchmarks.

by deepseekAug 19, 2026163K context$1.35/M input tokens$5.4/M output tokens
DS

DeepSeek: DeepSeek V3.2

T
295.3M tokens

Next-iteration general LLM with enhanced instruction following and agentic workflows.

by deepseekAug 12, 2026163K context$1.14/M input tokens$4.56/M output tokens
DS

DeepSeek: DeepSeek V3.1

T
340.1M tokens

High-performance Mixture-of-Experts architecture delivering frontier output at extreme value.

by deepseekAug 2, 2026128K context$0.2987/M input tokens$0.8652/M output tokens
∞

Meta: Llama 4 Maverick 17B

T
185.0M tokens

Meta's flagship open multimodal model with 1M context window and vision comprehension.

by metaSep 2, 20261M context$0.24/M input tokens$0.97/M output tokens
∞

Meta: Llama 3.3 70B

T
270.8M tokens

Proven open-weights model delivering industry benchmark performance in coding and reasoning.

by metaJul 30, 2026128K context$0.71/M input tokens$0.71/M output tokens
𝕏

xAI: Grok 4

T
95.6M tokens

Frontier model with extensive world knowledge, sharp reasoning, and unfiltered analytical depth.

by xaiAug 24, 2026256K context$3/M input tokens$15/M output tokens
𝕏

xAI: Grok 4 Fast

T
168.2M tokens

Huge 2M token context window paired with lightning-fast token generation speed.

by xaiAug 20, 20262M context$0.2/M input tokens$0.5/M output tokens

Mistral AI: Mistral Large 3

T
87.3M tokens

Mistral's flagship multilingual model with state-of-the-art reasoning and native function calling.

by mistral-aiAug 16, 2026256K context$0.5/M input tokens$1.5/M output tokens

Mistral AI: Devstral 2 123B

T
104.9M tokens

Engineered specifically for complex software architecture, multi-file refactoring, and code review.

by mistral-aiAug 10, 2026256K context$0.4/M input tokens$2/M output tokens

Mistral AI: Mistral Medium 3.5

T
52.4M tokens

Well-rounded performance across European languages, logic reasoning, and enterprise formatting.

by mistral-aiJul 28, 2026128K context$0.4/M input tokens$2/M output tokens

Microsoft: Phi-4

T
41.9M tokens

High-density small language model with state-of-the-art synthetic-data reasoning abilities.

by microsoftJul 15, 2026128K context$0.125/M input tokens$0.5/M output tokens

Microsoft: Phi-4 mini

T
63.7M tokens

Lightweight SLM engineered for rapid text processing, categorization, and routing.

by microsoftJul 10, 2026128K context$0.075/M input tokens$0.3/M output tokens

Microsoft: Phi-4 multimodal

T
38.2M tokens

Compact multi-modal model handling simultaneous text, visual diagrams, and speech audio.

by microsoftJul 5, 2026128K context$0.08/M input tokens$0.32/M output tokens

Microsoft: MAI-DS-R1

T
28.4M tokens

Microsoft-tuned DeepSeek R1 reasoning distribution for enterprise analytical tasks.

by microsoftAug 1, 2026163K context$1.35/M input tokens$5.4/M output tokens
MO

Moonshot AI: Kimi K2.5

T
49.0M tokens

Specialized long-context model capable of lossless document search and summarization.

by moonshot-aiJul 18, 2026256K context$0.6/M input tokens$3/M output tokens
MO

Moonshot AI: Kimi K2.5 Thinking

T
34.6M tokens

Deliberate chain-of-thought architecture for intricate logic, research, and coding tasks.

by moonshot-aiJul 22, 2026256K context$0.6/M input tokens$3/M output tokens
MI

MiniMax AI: MiniMax M2.5

T
22.3M tokens

Fluid conversational and creative writing engine with high token generation throughput.

by minimax-aiJun 20, 2026200K context$0.3/M input tokens$1.2/M output tokens
NV

NVIDIA: Nemotron 3 Super 120B

T
31.5M tokens

NVIDIA-accelerated LLM built for high-throughput enterprise synthesis and data curation.

by nvidiaJun 15, 2026256K context$0.15/M input tokens$0.65/M output tokens
QW

Qwen: Qwen3 Next 80B A3B

T
77.8M tokens

Next-generation sparse MoE architecture with expansive multilingual understanding.

by qwenJul 11, 2026256K context$0.15/M input tokens$1.2/M output tokens
QW

Qwen: Qwen3 VL 235B A22B

T
45.9M tokens

Advanced visual-language model adept at document OCR, charts, and spatial understanding.

by qwenJul 8, 2026256K context$0.53/M input tokens$2.66/M output tokens
Z

Z AI: GLM 5

T
19.8M tokens

Bilingual powerhouse for Chinese and English reasoning, translation, and code synthesis.

by z-aiJun 2, 2026200K context$1/M input tokens$3.2/M output tokens
Z

Z AI: GLM 4.7

T
14.2M tokens

High-speed bilingual conversational model for production latency requirements.

by z-aiMay 25, 2026200K context$0.6/M input tokens$2.2/M output tokens
CO

Cohere: Cohere Command A

T
36.4M tokens

Enterprise-optimized LLM specifically tuned for Retrieval-Augmented Generation (RAG) and tool use.

by cohereJun 28, 2026256K context$2.5/M input tokens$10/M output tokens