The cheapest Gemini API is still Gemini 2.5 Flash-Lite: $0.10 per million input tokens and $0.40 per million output tokens on the Standard paid tier, or $0.05 / $0.20 with Batch or Flex, as of October 8, 2026 (Gemini API pricing). Against that model, only one alternative in this comparison is cheaper on list price: Groq's GPT OSS 20B at $0.075 / $0.30. GPT-6 Luna and Claude Haiku 5.5 match Flash-Lite's $0.10 input but charge $0.50 for output, so they cost slightly more on the same text job.
Whether switching saves money depends mostly on which Gemini model you pay for today. If you're on Gemini 3.1 Flash-Lite, 3.5 Flash-Lite or a 3.x Flash model, moving the same text work to GPT-6 Luna, Claude Haiku 5.5 or Groq cuts the bill by roughly 60–90% in the worked example below. If you're already on 2.5 Flash-Lite, the savings are small and only Groq delivers them, unless cached input or a quality difference changes the math.
Gemini API alternatives compared: prices per 1M tokens (October 8, 2026)
All prices are Standard paid rates per 1 million tokens, without caching, as listed on each provider's own pricing page on October 8, 2026. The third column compares each rate with Gemini 2.5 Flash-Lite.
| Model (price source) | Input / output per 1M | vs Gemini 2.5 Flash-Lite | Conditions that change the price |
|---|---|---|---|
| Gemini 2.5 Flash-Lite | $0.10 / $0.40 | Baseline | Batch and Flex $0.05 / $0.20; audio input $0.30; free tier available |
| Gemini 3.1 Flash-Lite | $0.25 / $1.50 | Higher on both | Free tier available; Batch and Flex half price |
| Gemini 3.5 Flash-Lite | $0.30 / $2.50 | Higher on both | Free tier available |
| Gemini 3.6, 3.7 and 3.8 Flash | $0.75 / $3.75 | Higher on both | Rises to $1.50 / $7.50 on January 1, 2027 |
| GPT-6 Luna (OpenAI) | $0.10 / $0.50 | Same input, higher output | Requests over 272K input tokens: $0.20 / $0.75; cached input $0.01 |
| Claude Haiku 5.5 (Anthropic) | $0.10 / $0.50 | Same input, higher output | Prompts over 100,000 tokens: every rate ×5 ($0.50 / $2.50); cache read $0.01 |
| DeepSeek V4.1-Flash | $0.15–$0.30 / $0.60–$1.20 | Higher on both | Off-peak is the lower end; cache-hit input $0.003–$0.006 |
| Mistral Small 4 | $0.15 / $0.60 | Higher on both | Replaces Mistral Small 3.2, retired July 31, 2026 |
| Groq GPT OSS 20B | $0.075 / $0.30 | Lower on both | Text model; GPT OSS 120B costs $0.15 / $0.60 |
Two things stand out. First, every newer Gemini text model costs more than 2.5 Flash-Lite, so "newer Gemini" and "cheaper Gemini" point in opposite directions. Second, the non-Google options here sit between $0.075 and $0.15 for input and between $0.30 and $0.60 for output at their lowest rates, which is far closer to 2.5 Flash-Lite than to any Gemini 3.x model.
Token counts aren't identical across providers, because each model uses its own tokenizer. Treat the table as a comparison of rates, then measure your own prompts on any model you shortlist.
Cheapest Gemini API: 2.5 Flash-Lite at $0.10 / $0.40, not Gemini 3.x
Gemini 2.5 Flash-Lite is the lowest-priced paid Gemini text model on Google's price list as of October 8, 2026. It accepts text, image and video input at $0.10 per million tokens (audio input is $0.30) and returns output at $0.40.
The newer Flash-Lite generations are not budget upgrades. Gemini 3.1 Flash-Lite costs $0.25 / $1.50 and Gemini 3.5 Flash-Lite costs $0.30 / $2.50. Gemini 2.5 Flash sits at $0.30 / $2.50, the same as 3.5 Flash-Lite. If a cheap Gemini model is what you want, the first saving usually comes from moving down to 2.5 Flash-Lite, not from leaving Google.
Gemini 3.6, 3.7 and 3.8 Flash need a budget note. Google lists them at $0.75 / $3.75 through December 31, 2026, then at $1.50 / $7.50 from January 1, 2027. A workload priced on today's rate will cost twice as much in January without any change on your side. Gemini 3.5 Flash is already at $1.50 / $9.00.
Google's own discounts matter before you compare vendors:
- Batch and Flex cost half of Standard. Gemini 2.5 Flash-Lite drops to $0.05 / $0.20, the lowest rate in the table above. Jobs that can wait for a delayed result, such as nightly classification or bulk summarization, should be priced on Batch or Flex first.
- Priority costs 1.8× Standard, so use it only for traffic that truly needs it.
- Gemma 4 runs on the Gemini API only on the free tier. Google lists no paid tier for it, so it can't be your production lane on the Gemini API.
When switching saves money: it depends on your current Gemini model
The cost of a text job on any model is the same formula:
cost = (input tokens ÷ 1,000,000) × input price + (output tokens ÷ 1,000,000) × output price
Run that formula with the rates above and a clear rule falls out. Compare against the Gemini model you actually use, not against Gemini in general.

If you're on Gemini 2.5 Flash-Lite, switching rarely saves money:
- Groq GPT OSS 20B is about 25% cheaper on both input and output, so it costs less on any mix of tokens.
- GPT-6 Luna and Claude Haiku 5.5 cost the same for input and $0.10 more per million output tokens. They are never cheaper at uncached Standard rates. The gap grows with output volume: 200,000 output tokens add $0.02, while 10 million add $1.00.
- Mistral Small 4 and DeepSeek V4.1-Flash cost more on both sides at uncached rates, even at DeepSeek's off-peak price.
- If your job qualifies for Batch or Flex, Gemini 2.5 Flash-Lite at $0.05 / $0.20 is the lowest list price in this comparison.
If you're on Gemini 3.1 Flash-Lite, 3.5 Flash-Lite or a 3.x Flash model, almost every alternative here is cheaper for text. GPT-6 Luna, Claude Haiku 5.5 (on prompts up to 100,000 tokens), Mistral Small 4 and Groq are cheaper on both input and output. DeepSeek V4.1-Flash at peak hours ($0.30 / $1.20) costs more than 3.1 Flash-Lite for input but less for output, so it comes out ahead only when output is a meaningful share of the job. Before you switch vendors, check whether Gemini 2.5 Flash-Lite already handles the task, because that move needs no new account or SDK.
Cached input can flip the result. GPT-6 Luna and Claude Haiku 5.5 charge $0.01 per million cached input tokens, and DeepSeek charges $0.003–$0.006 for cache hits. If most of each prompt is a repeated system prompt or document, the effective input price of those models falls far below $0.10. Gemini also offers context caching, so compare against Google's own cached rate on its pricing page before deciding that caching justifies a move.
Long prompts flip it the other way. Claude Haiku 5.5 charges five times every rate when a prompt is over 100,000 tokens, and GPT-6 Luna charges $0.20 / $0.75 once a request passes 272,000 input tokens. For long-document work, check prompt length before you trust the $0.10 / $0.50 headline.
Worked example: 1M input + 200K output tokens on each model
This example uses 1 million input tokens and 200,000 output tokens, Standard rates and no caching. It's an input-heavy text job, similar to summarizing or classifying documents.

| Model | Input cost | Output cost | Total |
|---|---|---|---|
| Gemini 2.5 Flash-Lite, Batch or Flex | $0.05 | $0.04 | $0.09 |
| Groq GPT OSS 20B | $0.075 | $0.06 | $0.135 |
| Gemini 2.5 Flash-Lite | $0.10 | $0.08 | $0.18 |
| GPT-6 Luna | $0.10 | $0.10 | $0.20 |
| Claude Haiku 5.5 (prompts up to 100K tokens) | $0.10 | $0.10 | $0.20 |
| Mistral Small 4 | $0.15 | $0.12 | $0.27 |
| DeepSeek V4.1-Flash, off-peak | $0.15 | $0.12 | $0.27 |
| DeepSeek V4.1-Flash, peak | $0.30 | $0.24 | $0.54 |
| Gemini 3.1 Flash-Lite | $0.25 | $0.30 | $0.55 |
| Gemini 3.8 Flash (2026 rate) | $0.75 | $0.75 | $1.50 |
Read against the Gemini model you use today:
- From Gemini 3.8 Flash, GPT-6 Luna costs $0.20 instead of $1.50, about 87% less. After the January 1, 2027 price change, the same 3.8 Flash job costs $3.00.
- From Gemini 3.1 Flash-Lite, GPT-6 Luna costs $0.20 instead of $0.55, about 64% less.
- From Gemini 2.5 Flash-Lite, GPT-6 Luna and Claude Haiku 5.5 add $0.02, and Groq GPT OSS 20B saves $0.045.
If every prompt in this example were over 100,000 tokens, the Claude Haiku 5.5 total would rise from $0.20 to $1.00 ($0.50 input + $0.50 output). DeepSeek's peak window runs 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday, so a US-hours workload mostly lands on off-peak pricing.
Price doesn't settle quality. No first-hand quality test sits behind these numbers, so run your own prompts on the two or three cheapest models that fit your job before you move traffic.
GPT-6 Luna and Claude Haiku 5.5: same $0.10 / $0.50, different caps
GPT-6 Luna and Claude Haiku 5.5 have the same base price: $0.10 per million input tokens and $0.50 per million output tokens. Both cost $0.01 per million cached input tokens, and both offer Batch at half price. Among OpenAI and Anthropic models, they are the closest in price to Gemini 2.5 Flash-Lite.
They split on long prompts:
- GPT-6 Luna keeps its base price up to 272,000 input tokens per request. Above that, the whole request is billed at $0.20 / $0.75.
- Claude Haiku 5.5 keeps its base price for prompts up to 100,000 tokens. Above that, every rate is multiplied by five, to $0.50 / $2.50.
For short chat, extraction and classification, the two cost the same. For prompts between 100,000 and 272,000 tokens, Luna stays at its base price while Haiku costs five times as much. The older Claude Haiku 4.5 ($1 / $5) is now a legacy model and is no longer Anthropic's low-cost option. For the full tier math, see Claude Haiku 5.5 vs GPT-6 Luna Pricing: Same Until 100K, Then 5x. If you're deciding between OpenAI's cheaper and larger GPT-6 models, GPT-6 Luna vs. Sol pricing works through workload costs.
DeepSeek, Mistral and Groq: which old cheap prices no longer apply
Three low prices from earlier in 2026 no longer match the providers' own pages.
DeepSeek. The deepseek-flash model is now DeepSeek V4.1-Flash. As of October 8, 2026, it costs $0.15–$0.30 per million input tokens on a cache miss, $0.003–$0.006 on a cache hit, and $0.60–$1.20 per million output tokens. Off-peak rates are half of peak. It offers a 1M-token context window and up to 384K output tokens, and DeepSeek now serves legacy model names with V4.1-Flash. Older "DeepSeek-V3.2 at $0.28 / $0.42" figures no longer apply. DeepSeek V4 Pro costs $0.66–$1.32 / $1.98–$3.96, so it isn't a budget pick.
Mistral. Mistral Small 3.2 was retired on July 31, 2026, and Mistral Small 4 replaced it at $0.15 / $0.60. The $0.10 / $0.30 price belonged to the retired Small 3.2.
Groq. GPT OSS 20B ($0.075 / $0.30) and GPT OSS 120B ($0.15 / $0.60) have public prices. Llama 3.1 8B and Llama 3.3 70B are now listed as Enterprise models with "Contact Sales" instead of a public price, so the once-quoted $0.05 / $0.08 Llama 3.1 8B rate is no longer something you can plan around. Groq's GPT OSS models are text models, so they replace Gemini only for text traffic.
Is there a free Gemini API? Free tier, Gemma 4 and free alternatives
Yes. Google lists a free tier for Gemini 2.5 Flash-Lite, 2.5 Flash, 3.1 Flash-Lite, 3.5 Flash-Lite, 3.5 Flash, 3.6–3.8 Flash and 2.5 Pro. Gemini 3.1 Pro Preview has no free tier. Gemma 4 is free-tier only on the Gemini API.
Free-tier rate limits vary by model and aren't part of the price list. Before you build on free requests, read Gemini API Free Tier in 2026: What Is Free and What to Check. If you're ready to pay, Do You Buy a Gemini API Key? The Safe Way to Pay for API Access explains how paid access is set up.
Among the alternatives here, Gemini's price page is the one that marks free-tier availability model by model. If a free tier from Groq or Mistral is part of your plan, read that provider's current terms directly, because the paid rates above don't tell you what is free.
What price doesn't settle: multimodal input, regions and migration
Multimodal input. Gemini 2.5 Flash-Lite takes text, image and video input at the same $0.10 rate. Groq's GPT OSS models are text models. If your app sends images or video, a cheaper text model only saves money when you split traffic: text jobs to the cheaper model, media jobs to Gemini.
Migration effort. Gemini models are available through the OpenAI Python and JavaScript libraries by changing three lines of code and pointing the base URL at https://generativelanguage.googleapis.com/v1beta/openai/ (OpenAI compatibility). Using OpenAI's SDK isn't a reason to leave Gemini, and moving between 2.5 Flash-Lite and a 3.x model is a model-name change.
Regions. As of October 8, 2026, the Gemini API available regions list excludes Russia, mainland China and Hong Kong, and the OpenAI and Anthropic API country lists exclude the same three. If you build from one of those places, the cheapest official rate may not be one you can use. laozhang.ai offers a single API key for several of these models, but its listed prices are equal to or higher than the official ones (for example gemini-2.5-flash-lite at $0.10 / $0.40 and deepseek-v4-flash at $0.44 / $1.32 on its model list), and it listed no Claude models at the time. It's an access route, not a way to pay less.
For a wider view that adds premium models and context-length trade-offs, see Gemini API vs OpenAI vs Claude: Complete 2026 Cost Decision Guide.
FAQ
What is the cheapest Gemini API version?
Gemini 2.5 Flash-Lite, at $0.10 input / $0.40 output per million tokens on Standard and $0.05 / $0.20 on Batch or Flex, as of October 8, 2026. Newer Gemini 3.x Flash-Lite and Flash models cost more.
Is Gemini 3.1 Flash-Lite cheaper than Gemini 2.5 Flash-Lite?
No. Gemini 3.1 Flash-Lite costs $0.25 / $1.50 per million tokens, against $0.10 / $0.40 for 2.5 Flash-Lite. On 1 million input plus 200,000 output tokens, that's $0.55 versus $0.18.
What is the cheapest Gemini API alternative for text?
At list price, Groq GPT OSS 20B at $0.075 / $0.30 per million tokens, the only option in this comparison that costs less than Gemini 2.5 Flash-Lite on both input and output. GPT-6 Luna and Claude Haiku 5.5 ($0.10 / $0.50) are the low-cost picks if you prefer OpenAI or Anthropic models.
Is GPT-6 Luna cheaper than Gemini?
It's cheaper than every Gemini 3.x Flash and Flash-Lite model, but not cheaper than Gemini 2.5 Flash-Lite at uncached Standard rates. It matches Flash-Lite's $0.10 input and charges $0.50 instead of $0.40 for output. Heavy use of cached input ($0.01 per million) can change that.
What is the best free Gemini API model, and is there a paid Gemma 4 API?
Most Gemini Flash and Flash-Lite models, plus Gemini 2.5 Pro, have a free tier. Gemma 4 is free-tier only on the Gemini API, and Google lists no paid tier for it. Free-tier rate limits vary by model, so check them before you depend on free requests.
Do I need to leave Gemini to keep using OpenAI libraries?
No. Google's OpenAI compatibility endpoint lets you call Gemini models from the OpenAI libraries by changing the API key, the base URL and the model name.



