
·
19 min read
TL;DR
Current flagship: Grok 4.7, released September 21, 2026, costs $2.00 per million input tokens, $0.50 per million cached input tokens, and $6.00 per million output tokens for prompts below 200K tokens. It has a 500K-token context window and is xAI's recommended model for coding, agents, and general text work.
Same price as Grok 4.6: Grok 4.7 keeps Grok 4.6's rates at both context tiers. Upgrading does not change your per-token bill.
Long-context pricing: Once a prompt reaches 200K tokens, xAI bills the entire request at $4.00/M input, $1.00/M cached input, and $12.00/M output.
Cheapest options: Grok Build 0.1 is the lowest-priced text model at $1.00/M input and $2.00/M output. For general-purpose work, Grok 4.3 and the three Grok 4.20 variants cost $1.25/M input and $2.50/M output.
Grok 4.7 Fast is not on the public API: It runs at 2x standard rates below 200K tokens and is available only through Cursor and Grok Build.
Check before deploying: Model availability can vary by account and region. Verify it in the xAI Console.
See the live Grok 4.7 API pricing page or use the LLM cost calculator to estimate your own workload.
Grok Model API Pricing at a Glance
All prices are in USD per 1 million tokens. Updated September 25, 2026.
Model | Context window | Standard input | Cached input | Standard output | Long-context input | Long-context cached | Long-context output | Best for |
|---|---|---|---|---|---|---|---|---|
| 500K | $2.00 | $0.50 | $6.00 | $4.00 | $1.00 | $12.00 | Latest flagship for coding, agents, and knowledge work |
| 500K | $2.00 | $0.50 | $6.00 | $4.00 | $1.00 | $12.00 | Previous flagship, existing agent deployments |
| 500K | $2.00 | $0.30 | $6.00 | $4.00 | $0.60 | $12.00 | Cache-heavy workloads (lowest flagship-tier cached rate) |
| 256K | $1.00 | $0.20 | $2.00 | $2.00 | $0.40 | $4.00 | Cost-efficient coding specialist |
| 1M | $1.25 | $0.20 | $2.50 | $2.50 | $0.40 | $5.00 | General-purpose, cost-efficient reasoning |
| 1M | $1.25 | $0.20 | $2.50 | $2.50 | $0.40 | $5.00 | Multi-agent orchestration and long context |
| 1M | $1.25 | $0.20 | $2.50 | $2.50 | $0.40 | $5.00 | Reasoning-heavy tasks |
| 1M | $1.25 | $0.20 | $2.50 | $2.50 | $0.40 | $5.00 | Lower-latency standard completions |
Important: Long-context rates apply to all tokens in a request once its prompt reaches 200K tokens. A 210K-token prompt is not billed as 200K tokens at the standard rate plus 10K at the higher rate. The full request receives the higher rate.
xAI also offers Priority Processing at 2x standard token prices. This is a service tier for higher scheduling priority, not a separate Grok model. You are charged the priority rate only when the response confirms "service_tier": "priority".
How to Calculate Grok 4.7 API Cost
For a request below the 200K-token threshold:
Total cost = (uncached input tokens / 1,000,000 × $2.00)
+ (cached input tokens / 1,000,000 × $0.50)
+ (output tokens / 1,000,000 × $6.00)
For example, 1 million standard input tokens plus 250,000 output tokens cost:
Input: 1,000,000 ÷ 1,000,000 × $2.00 = $2.00
Output: 250,000 ÷ 1,000,000 × $6.00 = $1.50
Total: $3.50
If 600,000 of those input tokens are served from cache, the same workload costs less:
Uncached input: 400,000 ÷ 1,000,000 × $2.00 = $0.80
Cached input: 600,000 ÷ 1,000,000 × $0.50 = $0.30
Output: 250,000 ÷ 1,000,000 × $6.00 = $1.50
Total: $2.60
These examples assume each prompt stays below 200K tokens. Here is what happens when a single prompt crosses the threshold:
Single request: 210,000 input tokens + 5,000 output tokens
At long-context rates: (210,000 × $4.00 + 5,000 × $12.00) ÷ 1,000,000 = $0.90
If it had stayed below: (210,000 × $2.00 + 5,000 × $6.00) ÷ 1,000,000 = $0.45
Crossing the threshold by 10K tokens doubles the cost of the entire request.
What is the Grok API?
The Grok API is xAI's pay-per-token developer interface for building applications, agents, and integrations on top of Grok models, billed separately from any consumer Grok subscription. You don't need SuperGrok, X Premium, or any other consumer plan to use it: you generate an API key through the xAI Console, add a payment method, and pay only for the tokens you actually consume, at the per-model rates listed above.
This is a different product from the Grok chat experience inside the X app or at grok.com. Those are subscription-based, flat-rate consumer products. The API is usage-based and built for developers integrating Grok into their own software, whether that's a chatbot, a coding agent, or an automated workflow.
Grok Subscription Plans (At a Glance)
Plan | Price | Best for |
|---|---|---|
Free | $0 | Casual use with usage limits |
X Premium | $8/month | Basic Grok access bundled with X platform features |
SuperGrok Lite | $10/month | Light standalone Grok access |
SuperGrok | $30/month (or $300/year) | Daily consumer use with DeepSearch and Big Brain mode |
X Premium+ | $40/month | X power users who also want meaningful Grok access |
SuperGrok Heavy | $300/month | Heavy power users, developers, high-volume media generation |
These are consumer subscriptions for the chat interface, not the API. If you're building a product or integration, use the API pricing above instead.
Why Grok API Costs Scale Faster Than Expected
The model's list price is only one part of the bill. Multi-turn agents frequently send conversation history, system prompts, retrieved documents, and tool results back to the model on every request.
For example, a 20-turn agent that repeatedly sends 30K tokens of context can process roughly 600K input tokens across the session before counting its generated output. If an individual prompt crosses 200K tokens, Grok 4.7's long-context rate applies to the entire request.
A memory layer helps by storing durable facts and retrieving only what matters for the current request. Instead of repeatedly sending the complete transcript, the application can provide a smaller set of relevant memories.
Reduce repeated context in Grok applications
Mem0 stores and retrieves relevant memory instead of forcing an application to resend its full conversation history on every call. Read the LLM token-cost reduction guide to see where memory, caching, and model routing can reduce unnecessary input tokens.
Grok 4.7
Grok 4.7 is xAI's current flagship model, released September 21, 2026. xAI positions it as its most capable model for coding and knowledge work. It is designed to stay on difficult tasks for longer and verify its own output more carefully than Grok 4.6. xAI's documentation now recommends Grok 4.7 for everything except dedicated image, video, and voice workloads.
Grok 4.7 accepts text and image inputs and returns text. It supports configurable reasoning effort (low, medium, high, and xhigh, with high as the default), function calling, and structured outputs, through both the Responses and Chat Completions APIs. Its knowledge cutoff is May 2026.
Grok 4.7 pricing
Usage | Below 200K prompt tokens | At or above 200K prompt tokens |
|---|---|---|
Input | $2.00/M | $4.00/M |
Cached input | $0.50/M | $1.00/M |
Output | $6.00/M | $12.00/M |
Grok 4.7 is available through the xAI API, Cursor, Grok Build, and gateway partners including OpenRouter, Vercel, and Cloudflare. Grok Build now uses Grok 4.7 as its default model. Partner prices, promotions, and included usage may differ from direct API rates.
Grok 4.7 Fast
Grok 4.7 Fast is the same model served on faster infrastructure. It is not a separate API model: it is available only through Cursor and Grok Build, and Grok Build's free tier does not include it.
Prompt tokens | Input | Cached input | Output |
|---|---|---|---|
Below 200K | $4.00/M | $1.00/M | $12.00/M |
Above 200K | $6.00/M | $1.50/M | $18.00/M |
US regional endpoint
Requests sent to https://us.api.x.ai/v1 run inference in the United States and are billed at 1.1x global rates. For Grok 4.7, that is $2.20/M input, $0.55/M cached input, and $6.60/M output below 200K prompt tokens, and $4.40 / $1.10 / $13.20 above. The regional endpoint currently serves only Grok 4.7 and Grok 4.6.
What to watch in multi-turn agents
On the Responses API, Grok 4.7 always returns encrypted reasoning content, and xAI asks developers to pass those reasoning items back unchanged in multi-turn work. That keeps the agent coherent, but it also adds tokens to every subsequent request. xAI recommends using a prompt cache key and context compaction for long agent loops.
Keep long-running Grok 4.7 agents below the long-context threshold
Mem0 stores durable facts from earlier turns and retrieves only what the current request needs, so your agent doesn't resend its full history on every call. See the LLM token-cost reduction guide for how memory, caching, and routing work together.
Use Grok 4.7 for new deployments that need xAI's strongest coding and agentic performance. Because it costs the same as Grok 4.6, there is no pricing reason to stay on 4.6 once you've tested 4.7 on representative tasks.
Grok 4.6
Grok 4.6 was released August 12, 2026, and is now the previous-generation flagship. It remains available at the same rates as Grok 4.7, with a 500K-token context window. It is also one of only two models served on the US regional endpoint.
According to xAI, it builds on Grok 4.5 with a focus on long-running agents and more ambitious interactive and visual work. The company positions it for multi-step tasks such as research, codebase analysis, application development, and the production of polished work artifacts.
Grok 4.6 accepts text and image inputs and returns text. It supports reasoning, function calling, and structured outputs.
Grok 4.6 pricing
Usage | Below 200K prompt tokens | At or above 200K prompt tokens |
|---|---|---|
Input | $2.00/M | $4.00/M |
Cached input | $0.50/M | $1.00/M |
Output | $6.00/M | $12.00/M |
The model has a 500K-token context window. Grok 4.6 is available through the xAI API and partners including Cursor, Grok Build, OpenRouter, Vercel, and Cloudflare. Partner prices, promotions, and included usage may differ from direct API rates.
Artificial Analysis scores Grok 4.6 (high) at 44 on its Intelligence Index v4.3. The index was rebalanced in September 2026, so scores from earlier versions are not directly comparable.
A composite benchmark does not make the models interchangeable for every workload, but it provides a more useful capability reference than price alone. Compare the live token rates and context limits on Grok 4.6 vs GPT-5.6 Sol.
Existing Grok 4.6 deployments can keep running, but new work should start with Grok 4.7, which costs the same.
Grok 4.5
Grok 4.5 was released July 8, 2026, and is now two generations behind the current flagship. It remains available with a 500K-token context window and the same standard input and output prices as Grok 4.7 and 4.6.
The main pricing difference is prompt caching. Grok 4.5 cached input costs $0.30/M below the long-context threshold, compared with $0.50/M for Grok 4.7 and 4.6. Above the threshold, Grok 4.5 cached input costs $0.60/M.
Existing applications can continue using Grok 4.5. New deployments should test both models on representative tasks before migrating, since a newer model does not automatically produce a lower total cost for every workload.
Grok 4.3
Grok 4.3 sits in xAI's lowest-priced active general-purpose tier, tied with the three Grok 4.20 variants. It costs $1.25/M input, $0.20/M cached input, and $2.50/M output below the long-context threshold.
It also has a larger 1M-token context window. However, requests reaching 200K prompt tokens are billed at $2.50/M input, $0.40/M cached input, and $5.00/M output.
Use Grok 4.3 for general production workloads where cost efficiency and a larger context window matter more than Grok 4.7's newest agentic capabilities.
Grok 4.20
Grok 4.20 is available in three variants with the same published token rates:
grok-4.20-multi-agent-0309for multi-agent orchestration.grok-4.20-0309-reasoningfor reasoning-heavy workloads.grok-4.20-0309-non-reasoningfor standard completions without the reasoning variant.
Each variant has a 1M-token context window. Standard pricing is $1.25/M input, $0.20/M cached input, and $2.50/M output. At or above 200K prompt tokens, pricing increases to $2.50/M input, $0.40/M cached input, and $5.00/M output.
Grok Server-Side Tools Pricing
Server-side tool usage is billed in addition to the model's token costs. Agentic requests can therefore cost more than text-only calls even when they use the same model.
Tool | Cost |
|---|---|
Web Search | $5.00 per 1K calls |
X Search | $5.00 per 1K posts returned, $10.00 per 1K profiles returned |
Code Execution | $5.00 per 1K calls |
Image Generation | Billed at the applicable Imagine API rates |
File Attachments Search | $10.00 per 1K calls |
Collections Search | $2.50 per 1K calls |
Image Understanding from search results | Token-based |
X Video Understanding | Token-based |
Remote MCP Tools | Token-based |
Tool costs can multiply during autonomous agent runs because the model decides how many searches or executions to perform. Set tool limits, cache reusable results, and monitor invocation counts separately from token usage.
Usage Guidelines Violation Fee
xAI charges $0.05 per request when a usage-guideline violation is detected before generation in the Responses API. When a violation is detected during generation, the request is still charged for the tokens generated. Treat policy failures as both a safety issue and a measurable cost item when monitoring production traffic.
Grok Voice and Imagine API Pricing
Voice API
Model or service | Cost |
|---|---|
| $0.05/minute or $3.00/hour for audio, plus $0.004 for text input |
| $0.08/minute or $4.80/hour for audio, plus $0.004 for text input |
Speech to Text, REST | $0.10/hour |
Speech to Text, streaming | $0.20/hour |
Text to Speech | $15.00 per 1M characters |
Imagine API
Model | Media input | Output pricing |
|---|---|---|
| $0.01/image | $0.05/image at 1K, $0.07/image at 2K |
| $0.002/image | $0.02/image at 1K or 2K |
| $0.01/image | $0.04 at 1K Low, $0.06 at 2K Low, $0.06 at 1K Medium, $0.08 at 2K Medium |
| $0.01/image | $0.08/second at 480p, $0.14/second at 720p, $0.25/second at 1080p |
| $0.01/second for video or $0.002/image | $0.05/second at 480p, $0.07/second at 720p |
Files and Collections Pricing
Resource | Rate |
|---|---|
File storage | $0.025 per GiB per day |
Collection storage | $0.10 per GiB per day |
File downloads | $0.20 per GiB downloaded |
Collection downloads | $0.20 per GiB downloaded |
Collection storage costs four times as much as raw file storage. For large retrieval pipelines, remove obsolete documents and avoid maintaining duplicate indexes.
Batch and Priority Processing
The Batch API processes asynchronous jobs, typically within 24 hours. xAI currently lists a 20% batch discount for Grok 4.3 and the three Grok 4.20 variants. Models that are not listed in xAI's batch-discount table do not currently receive a published batch discount.
Priority Processing provides higher scheduling priority at 2x standard token rates. It applies to input, output, cached, and reasoning tokens, with prompt-caching discounts applied before the 2x multiplier. It is available for Chat Completions and Responses endpoints and is not supported for image generation, video generation, or Batch API requests.
API vs Subscription: Which Should You Use?
Use the API when you are building a product, backend integration, chatbot, coding agent, or automated workflow. API billing is usage-based and does not require a consumer Grok subscription.
Use a consumer subscription when you primarily want the Grok chat interface rather than programmatic model access. Consumer plan names, model access, and usage allowances change more frequently than API token prices, so verify the current offer on Grok before buying.
Use case | Recommended option |
|---|---|
Latest xAI model for long-running agents, coding, and research | API with Grok 4.7 |
Existing coding or agentic deployment that does not yet need 4.7 | API with Grok 4.6 or 4.5 |
Cost-efficient general-purpose reasoning | API with Grok 4.3 |
Lower-cost coding-specific work | API with Grok Build 0.1 |
Multi-agent orchestration | API with Grok 4.20 Multi-Agent |
Personal chat usage | Current Grok consumer subscription |
Enterprise requirements, SLAs, or custom terms | Contact xAI sales |
Grok 4.7 Pricing vs Other Frontier Models
The comparison below uses each provider's listed standard rates. All three labs released new models in the week of September 21, 2026, and prices can change. Caching, batch, and priority tiers can also change the effective bill.
Model | Positioning | Input $/M | Output $/M | Context window | Long-context rule |
|---|---|---|---|---|---|
Grok 4.7 | xAI flagship | $2.00 | $6.00 | 500K | Whole request at 2x input and 2x output once the prompt reaches 200K |
GPT-6 Sol | OpenAI mid-frontier model | $2.00 | $10.00 | 1.05M | Whole request at $4 input and $15 output above 272K |
Claude Opus 5.5 | Anthropic frontier model | $4.00 | $20.00 | 1M | No published long-context surcharge |
GPT-6 Luna | OpenAI low-cost tier | $0.10 | $0.50 | 1M | NA |
At standard rates, Grok 4.7 matches GPT-6 Sol on input and costs 40% less on output. Compared with Claude Opus 5.5, it costs 50% less on input and 70% less on output.
Long prompts change the ranking. Between 200K and 272K prompt tokens, Grok 4.7 has already switched to $4/$12, while GPT-6 Sol is still billed at $2/$10. In that range, GPT-6 Sol is cheaper on both input and output. Above 272K, both charge $4 for input, and Grok 4.7 is cheaper on output ($12 vs $15). Claude Opus 5.5 bills its full 1M-token context window at standard rates, so a 900K-token request costs the same per token as a 9K-token one. Above 200K prompt tokens, its $4/M input rate matches Grok 4.7's long-context input rate, while Grok 4.7 remains cheaper on output ($12 vs $20).
Caching can narrow the gap. Grok 4.7's cached input is $0.50/M, a 75% discount. GPT-6 Sol's cached input is $0.20/M. For workloads where most input is served from cache, compare cached rates rather than list prices.
These percentages compare token prices, not model quality. On the Artificial Analysis Intelligence Index (v4.3), Grok 4.7 at xhigh reasoning effort scores 46, up two points from Grok 4.6. That places it one point behind GPT-5.6 Sol (47) and behind Claude Opus 5 (51), Claude Fable 5.1 (53), and GPT-6 Astra (53). Artificial Analysis reports that GPT-6 Sol scores level with GPT-5.6 Sol. Claude Opus 5.5 had not been scored at the time of writing.
Grok 4.7 performs better on agentic coding than its headline score suggests. Paired with Grok Build, it scores 56 on the Artificial Analysis Coding Agent Index, ranking fourth among models in their native harnesses and ahead of GPT-5.6 Sol.
Token usage matters as much as token price. Grok 4.7 (xhigh) used roughly 81K output tokens per Intelligence Index task, more than double the roughly 36K used by Grok 4.6 (high). A model with a lower per-token rate can still cost more per completed task if it generates more tokens. Test each model on representative workloads and compare the total cost of accepted results, along with latency, tool support, context requirements, and provider availability.
How to Reduce Grok API Costs
1. Route requests by difficulty
Do not send every request to Grok 4.7. Use Grok 4.3 for cost-sensitive general work and Grok Build 0.1 for lower-cost coding workloads. Reserve Grok 4.7 for tasks that need its strongest agentic and coding capabilities.
2. Keep individual prompts below the long-context threshold
Crossing 200K prompt tokens increases Grok 4.7 rates from $2/$0.50/$6 to $4/$1/$12. Summarize old context, retrieve only relevant documents, and split independent work into smaller requests where appropriate.
3. Take advantage of prompt caching
xAI automatically reports cached token usage in the API response. Keep stable instructions and repeated context in consistent prompt prefixes so the provider can reuse them when eligible.
4. Limit autonomous tool calls
Web Search, X Search, Code Execution, File Attachments, and Collections Search are billed separately. Set clear budgets and stopping conditions for autonomous agents.
5. Replace repeated history with retrieved memory
Long-running agents often resend information that has already been established. Store durable facts once and retrieve a small relevant set for each request. This can reduce both standard input usage and the chance of crossing the 200K long-context threshold.
Useful Sources
Frequently Asked Questions
Q. What does Grok 4.7 cost?
For prompts below 200K tokens, Grok 4.7 costs $2 per million input tokens, $0.50 per million cached input tokens, and $6 per million output tokens. At or above 200K prompt tokens, the entire request is billed at $4/M input, $1/M cached input, and $12/M output. Tool calls are billed separately.
Q. Is Grok 4.7 more expensive than Grok 4.6?
No. Grok 4.7 and Grok 4.6 have identical standard, cached, and long-context rates, and both have a 500K-token context window. Grok 4.5 has a lower cached-input rate ($0.30/M) but the same standard input and output prices.
Q. What is Grok 4.7 Fast, and can I use it through the API?
Grok 4.7 Fast is Grok 4.7 served on faster infrastructure at 2x standard rates below 200K tokens. It is available only in Cursor and Grok Build, not on the public xAI API.
Q. What is the latest Grok API model?
As of September 2026, the latest Grok API model is grok-4.7. xAI's docs recommend it for code, chat, and general text work, with separate models for image, video, and voice.
Q. Does Grok 4.7 get a Batch API discount?
Not currently. xAI lists a 20% batch discount only for Grok 4.3 and the three Grok 4.20 variants.
Q. What is the cheapest active Grok API model?
As of September 2026, Grok Build 0.1 has the lowest listed text-token rate at $1/M input and $2/M output, and it is designed for coding work. For general-purpose use, Grok 4.3 is tied with the three Grok 4.20 variants at $1.25/M input and $2.50/M output. Grok 4.3 and Grok 4.20 also receive a 20% Batch API discount.
Q. Is Grok 4.7 cheaper than GPT-6 Sol and Claude Opus 5.5?
Yes, for prompts below 200K tokens. Grok 4.7 costs $2/M input and $6/M output. That matches GPT-6 Sol on input ($2/M) and is 40% cheaper on output ($10/M). Compared with Claude Opus 5.5 ($4/M input, $20/M output), Grok 4.7 is 50% cheaper on input and 70% cheaper on output.
Token price is not the same as cost per task. Models differ in how many tokens they use to finish the same job, how often they need retries, and when long-context surcharges apply. Test each model on representative workloads and compare the total cost of accepted results, not just the rate card.
Q. Does Grok 4.7 support prompt caching?
Yes. Cached input costs $0.50/M below the 200K prompt threshold and $1/M at or above it, a 75% discount off standard input. Cached token usage is shown in the API response. For multi-turn agents, xAI recommends setting a prompt cache key to improve cache hit rates. If caching makes up most of your input, Grok 4.5 has a lower cached rate at $0.30/M.
Q. Does xAI offer free API credits?
xAI has offered promotional or data-sharing credits, but availability and amounts can change. Check the xAI Console rather than treating a historical credit amount as guaranteed.
Q. Should I use the Grok API or a Grok subscription?
Use the API for applications and automated workflows. Use a subscription when you primarily want access through the consumer chat interface. API access does not require a consumer subscription.
Q. How can I reduce Grok API costs at scale?
Use a cheaper model for routine requests, keep prompts below 200K tokens, reuse cacheable prompt prefixes, limit autonomous tool calls, and retrieve relevant memory instead of repeatedly sending complete conversation history.

Aashi Dutt
She is a senior technical content writer at Mem0. She covers agent memory architecture and the engineering decisions behind building agents that actually remember. She experiments with new features and turns research into posts developers can put straight to use.
Start building with memory
Free tier, no card









