{"data":[{"id":"anthropic-claude-fable-5","name":"Anthropic: Claude Fable 5","description":"Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...","contextLength":1000000,"pricing":{"input":12,"output":60,"cacheInput":1.2}},{"id":"anthropic-claude-opus-4-5","name":"Anthropic: Claude Opus 4.5","description":"Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...","contextLength":198000,"pricing":{"input":6,"output":30,"cacheInput":0.6}},{"id":"anthropic-claude-opus-4-6","name":"Anthropic: Claude Opus 4.6","description":"Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...","contextLength":1000000,"pricing":{"input":6,"output":30,"cacheInput":0.6}},{"id":"anthropic-claude-opus-4-6-fast","name":"Anthropic: Claude Opus 4.6 (Fast)","description":"Fast-mode variant of [Opus 4.6](/anthropic/claude-opus-4.6) - identical capabilities with higher output speed at premium 6x pricing.\n\nLearn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode","contextLength":1000000,"pricing":{"input":36,"output":180,"cacheInput":3.6}},{"id":"anthropic-claude-opus-4-7","name":"Anthropic: Claude Opus 4.7","description":"Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...","contextLength":1000000,"pricing":{"input":6,"output":30,"cacheInput":0.6}},{"id":"anthropic-claude-opus-4-7-fast","name":"Anthropic: Claude Opus 4.7 (Fast)","description":"Fast-mode variant of [Opus 4.7](/anthropic/claude-opus-4.7) - identical capabilities with higher output speed at premium 6x pricing.\n\nLearn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode","contextLength":1000000,"pricing":{"input":36,"output":180,"cacheInput":3.6}},{"id":"anthropic-claude-opus-4-8","name":"Anthropic: Claude Opus 4.8","description":"Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...","contextLength":1000000,"pricing":{"input":6,"output":30,"cacheInput":0.6}},{"id":"anthropic-claude-opus-4-8-fast","name":"Anthropic: Claude Opus 4.8 (Fast)","description":"Fast-mode variant of [Opus 4.8](/anthropic/claude-opus-4.8) - identical capabilities with higher output speed at 2x pricing relative to regular Opus 4.8.\n\nLearn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode","contextLength":1000000,"pricing":{"input":12,"output":60,"cacheInput":1.2}},{"id":"anthropic-claude-opus-5","name":"Claude Opus 5","description":"Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...","contextLength":1000000,"pricing":{"input":6,"output":30,"cacheInput":0.6}},{"id":"anthropic-claude-opus-5-fast","name":"Claude Opus 5 (Fast)","description":"Fast-mode variant of [Opus 5](/anthropic/claude-opus-5) - identical capabilities with higher output speed at 2x pricing relative to regular Opus 5.\n\nLearn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode","contextLength":1000000,"pricing":{"input":12,"output":60,"cacheInput":1.2}},{"id":"anthropic-claude-sonnet-4-5","name":"Anthropic: Claude Sonnet 4.5","description":"Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...","contextLength":198000,"pricing":{"input":3.75,"output":18.75,"cacheInput":0.375}},{"id":"anthropic-claude-sonnet-4-6","name":"Anthropic: Claude Sonnet 4.6","description":"Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...","contextLength":1000000,"pricing":{"input":3.6,"output":18,"cacheInput":0.36}},{"id":"anthropic-claude-sonnet-5","name":"Anthropic: Claude Sonnet 5","description":"Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...","contextLength":1000000,"pricing":{"input":3,"output":15,"cacheInput":0.3}},{"id":"deepseek-deepseek-v3-2","name":"DeepSeek: DeepSeek V3.2","description":"DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...","contextLength":160000,"pricing":{"input":0.33,"output":0.48,"cacheInput":0.16}},{"id":"deepseek-deepseek-v4-flash","name":"DeepSeek: DeepSeek V4 Flash","description":"DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...","contextLength":1000000,"pricing":{"input":0.138,"output":0.275,"cacheInput":0.028}},{"id":"deepseek-deepseek-v4-flash-0731","name":"DeepSeek: DeepSeek V4 Flash 0731","description":"DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.","contextLength":1000000,"pricing":{"input":0.175,"output":0.35,"cacheInput":0.035}},{"id":"deepseek-deepseek-v4-pro","name":"DeepSeek: DeepSeek V4 Pro","description":"DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...","contextLength":1000000,"pricing":{"input":1.65,"output":3.301,"cacheInput":0.33}},{"id":"deepseek-deepseek-v4-pro-0813","name":"DeepSeek: DeepSeek V4 Pro 0813","description":"DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.","contextLength":1000000,"pricing":{"input":1.73,"output":4.95,"cacheInput":0.33}},{"id":"deepseek-v4-flash-0731-fast","name":"DeepSeek V4 Flash 0731 Fast","description":"DeepSeek V4 Flash is an efficiency-optimized 284B-parameter Mixture-of-Experts model with 13B active parameters and a 1M-token context window. Tuned for fast inference and high-throughput workloads while maintaining strong reasoning and coding performance.","contextLength":1000000,"pricing":{"input":0.35,"output":0.7,"cacheInput":0.0875}},{"id":"e2ee-deepseek-v4-flash","name":"DeepSeek V4 Flash","description":"DeepSeek V4 Flash running in a Trusted Execution Environment (TEE). Hardware attestation evidence is available for independent verification of enclave identity and configuration.","contextLength":1000000,"pricing":{"input":0.182,"output":0.373,"cacheInput":0.038}},{"id":"google-gemini-3-1-pro-preview","name":"Google: Gemini 3.1 Pro Preview","description":"Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...","contextLength":1000000,"pricing":{"input":2.5,"output":15,"cacheInput":0.5}},{"id":"google-gemini-3-5-flash","name":"Google: Gemini 3.5 Flash","description":"Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...","contextLength":1000000,"pricing":{"input":1.55,"output":9.45,"cacheInput":0.155}},{"id":"google-gemini-3-5-flash-lite","name":"Google: Gemini 3.5 Flash-Lite","description":"Gemini 3.5 Flash-Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.","contextLength":1000000,"pricing":{"input":0.375,"output":3.125,"cacheInput":0.0375}},{"id":"google-gemini-3-6-flash","name":"Google: Gemini 3.6 Flash","description":"Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...","contextLength":1000000,"pricing":{"input":1.875,"output":9.375,"cacheInput":0.1875}},{"id":"google-gemini-3-7-flash","name":"Google: Gemini 3.7 Flash","description":"Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...","contextLength":1000000,"pricing":{"input":1.875,"output":9.375,"cacheInput":0.1875}},{"id":"google-gemini-3-flash-preview","name":"Google: Gemini 3 Flash Preview","description":"Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...","contextLength":256000,"pricing":{"input":0.7,"output":3.75,"cacheInput":0.07}},{"id":"kimi-k3-fast-api","name":"Kimi K3 Fast","description":"Kimi K3 is an ultra-large-scale, open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at navigating large repositories, using tools, debugging, and iterating against images, logs, tests, and runtime feedback.","contextLength":1000000,"pricing":{"input":4.5,"output":22.5,"cacheInput":0.45}},{"id":"minimax-m3-preview","name":"MiniMax M3 Preview","description":"MiniMax-M3 preview is a 1.4T-parameter frontier model from MiniMax for coding, agentic workflows, and complex reasoning, served at fp8 with a 512K context window.","contextLength":524288,"pricing":{"input":0.3,"output":1.2,"cacheInput":0.06}},{"id":"minimax-minimax-m2-5","name":"MiniMax: MiniMax M2.5","description":"MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1...","contextLength":198000,"pricing":{"input":0.27,"output":0.95,"cacheInput":0.03}},{"id":"minimax-minimax-m2-7","name":"MiniMax: MiniMax M2.7","description":"MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...","contextLength":198000,"pricing":{"input":0.375,"output":1.5,"cacheInput":0.06875}},{"id":"minimax-minimax-m3","name":"MiniMax: MiniMax M3","description":"MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...","contextLength":500000,"pricing":{"input":0.3,"output":1.2,"cacheInput":0.06}},{"id":"moonshotai-kimi-k2-5","name":"MoonshotAI: Kimi K2.5","description":"Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm. Built on Kimi K2 with continued pretraining over approximately 15T mixed...","contextLength":256000,"pricing":{"input":0.56,"output":3.5,"cacheInput":0.22}},{"id":"moonshotai-kimi-k2-6","name":"MoonshotAI: Kimi K2.6","description":"Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and...","contextLength":256000,"pricing":{"input":0.75,"output":3.5,"cacheInput":0.16}},{"id":"moonshotai-kimi-k2-7-code","name":"MoonshotAI: Kimi K2.7 Code","description":"MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...","contextLength":256000,"pricing":{"input":0.75,"output":3.5,"cacheInput":0.16}},{"id":"moonshotai-kimi-k3","name":"MoonshotAI: Kimi K3","description":"Kimi K3 is an ultra-large-scale, open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at navigating...","contextLength":1000000,"pricing":{"input":3.75,"output":18.75,"cacheInput":0.375}},{"id":"openai-gpt-52","name":"GPT-5.2","description":"GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context performance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly to simple queries while spending more depth on complex tasks.","contextLength":256000,"pricing":{"input":2.19,"output":17.5,"cacheInput":0.219}},{"id":"openai-gpt-52-codex","name":"GPT-5.2 Codex","description":"GPT-5.2 Codex is OpenAI specialized coding model built on GPT-5.2, optimized for advanced software development, code generation, and technical problem-solving.","contextLength":256000,"pricing":{"input":2.19,"output":17.5,"cacheInput":0.219}},{"id":"openai-gpt-53-codex","name":"GPT-5.3 Codex","description":"GPT-5.3 Codex is OpenAI specialized coding model built on GPT-5.3, optimized for advanced software development, code generation, and technical problem-solving.","contextLength":400000,"pricing":{"input":2.19,"output":17.5,"cacheInput":0.219}},{"id":"openai-gpt-54","name":"GPT-5.4","description":"GPT-5.4 is the latest frontier model in the GPT-5 series with a 1M+ context window, offering improved agentic and long context performance. It uses adaptive reasoning to dynamically allocate computation across tasks.","contextLength":1000000,"pricing":{"input":3.13,"output":18.8,"cacheInput":0.313}},{"id":"openai-gpt-54-mini","name":"GPT-5.4 Mini","description":"GPT-5.4 Mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding, and tool use.","contextLength":400000,"pricing":{"input":0.9375,"output":5.625,"cacheInput":0.09375}},{"id":"openai-gpt-54-pro","name":"GPT-5.4 Pro","description":"GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning for complex, high-stakes tasks. It provides a 1M+ token context window (922K input, 128K output) and supports text and image inputs.","contextLength":1000000,"pricing":{"input":37.5,"output":225,"cacheInput":null}},{"id":"openai-gpt-55","name":"GPT-5.5","description":"GPT-5.5 is the latest frontier model in the GPT-5 series with a 1M+ context window, offering improved agentic and long context performance. It uses adaptive reasoning to dynamically allocate computation across tasks.","contextLength":1000000,"pricing":{"input":6.25,"output":37.5,"cacheInput":0.625}},{"id":"openai-gpt-55-pro","name":"GPT-5.5 Pro","description":"GPT-5.5 Pro is OpenAI's most advanced model, building on GPT-5.5's unified architecture with enhanced reasoning for complex, high-stakes tasks. It provides a 1M+ token context window (922K input, 128K output) and supports text and image inputs.","contextLength":1000000,"pricing":{"input":37.5,"output":225,"cacheInput":null}},{"id":"openai-gpt-56-luna","name":"GPT-5.6 Luna","description":"GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for its price tier.","contextLength":1000000,"pricing":{"input":0.26666667,"output":1.6,"cacheInput":0.02666667}},{"id":"openai-gpt-56-luna-pro","name":"GPT-5.6 Luna Pro","description":"GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks.","contextLength":1000000,"pricing":{"input":1.25,"output":7.5,"cacheInput":0.125}},{"id":"openai-gpt-56-sol","name":"GPT-5.6 Sol","description":"GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks and long-horizon problem solving.","contextLength":1000000,"pricing":{"input":6.25,"output":37.5,"cacheInput":0.625}},{"id":"openai-gpt-56-sol-pro","name":"GPT-5.6 Sol Pro","description":"GPT-5.6 Sol Pro is the same underlying model as GPT-5.6 Sol, served with reasoning.mode set to pro for higher-quality responses on complex tasks.","contextLength":1000000,"pricing":{"input":6.25,"output":37.5,"cacheInput":0.625}},{"id":"openai-gpt-56-terra","name":"GPT-5.6 Terra","description":"GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic tasks where capability and cost need to be balanced.","contextLength":1000000,"pricing":{"input":3.125,"output":18.75,"cacheInput":0.3125}},{"id":"openai-gpt-56-terra-pro","name":"GPT-5.6 Terra Pro","description":"GPT-5.6 Terra Pro is the same underlying model as GPT-5.6 Terra, served with reasoning.mode set to pro for higher-quality responses on complex tasks.","contextLength":1000000,"pricing":{"input":3.125,"output":18.75,"cacheInput":0.3125}},{"id":"x-ai-grok-4-20","name":"xAI: Grok 4.20","description":"Grok 4.20 is a reasoning model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...","contextLength":2000000,"pricing":{"input":1.42,"output":2.83,"cacheInput":0.23}},{"id":"x-ai-grok-4-20-multi-agent","name":"xAI: Grok 4.20 Multi-Agent","description":"Grok 4.20 Multi-Agent is a variant of xAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct deep research, coordinate tool use, and synthesize information...","contextLength":2000000,"pricing":{"input":1.42,"output":2.83,"cacheInput":0.23}},{"id":"x-ai-grok-4-3","name":"xAI: Grok 4.3","description":"Grok 4.3 is a reasoning model from xAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...","contextLength":1000000,"pricing":{"input":1.42,"output":2.83,"cacheInput":0.23}},{"id":"x-ai-grok-4-5","name":"xAI: Grok 4.5","description":"Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.","contextLength":500000,"pricing":{"input":2.27,"output":6.8,"cacheInput":0.34}},{"id":"x-ai-grok-4-6","name":"SpaceXAI: Grok 4.6","description":"Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.","contextLength":500000,"pricing":{"input":2.27,"output":6.8,"cacheInput":0.57}},{"id":"x-ai-grok-build-0-1","name":"xAI: Grok Build 0.1","description":"Grok Build 0.1 is xAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...","contextLength":256000,"pricing":{"input":1,"output":2,"cacheInput":0.2}},{"id":"z-ai-glm-4-6","name":"Z.ai: GLM 4.6","description":"Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...","contextLength":198000,"pricing":{"input":0.43,"output":1.75,"cacheInput":0.08}},{"id":"z-ai-glm-4-7","name":"Z.ai: GLM 4.7","description":"GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...","contextLength":198000,"pricing":{"input":0.55,"output":2.65,"cacheInput":0.11}},{"id":"z-ai-glm-4-7-flash","name":"Z.ai: GLM 4.7 Flash","description":"As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...","contextLength":128000,"pricing":{"input":0.06,"output":0.4,"cacheInput":0.01}},{"id":"z-ai-glm-5","name":"Z.ai: GLM 5","description":"GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading...","contextLength":198000,"pricing":{"input":1,"output":3.2,"cacheInput":0.2}},{"id":"z-ai-glm-5-1","name":"Z.ai: GLM 5.1","description":"GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...","contextLength":200000,"pricing":{"input":1.54,"output":4.84,"cacheInput":0.286}},{"id":"z-ai-glm-5-2","name":"Z.ai: GLM 5.2","description":"GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...","contextLength":1000000,"pricing":{"input":1.4,"output":4.4,"cacheInput":0.26}},{"id":"z-ai-glm-5-turbo","name":"Z.ai: GLM 5 Turbo","description":"GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows...","contextLength":200000,"pricing":{"input":1.2,"output":4,"cacheInput":0.24}},{"id":"z-ai-glm-5v-turbo","name":"GLM 5V Turbo","description":"GLM-5V-Turbo is Z.ai's first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks with image, video, and text inputs.","contextLength":200000,"pricing":{"input":1.5,"output":5,"cacheInput":0.3}}]}