Skip to content
Xavi Creus

AI

How to choose an AI model: frontier, open-weight or local

Choose an AI model by task: frontier for judgement, mid-tier for volume, open-weight or local for privacy and EU residency. 2026 prices and framework.

By Xavi Creus7 min read

The right AI model for your company is rarely 1 model. In September 2026 the sensible answer is a portfolio: a frontier model for the 10% of tasks that need judgement, a mid-tier model for the 80% that need volume, and an open-weight or local model wherever the data cannot leave your control. I am Xavi Creus, CEO and CTO of +10 SaaS and AI companies in Barcelona, and this is the framework we use to decide.

I compare the Claude 5 family, OpenAI's GPT-6 and GPT-5.6, Google's Gemini 3 and the main open-weight models on quality, cost per million tokens, privacy, latency and EU data residency, and finish with a simple decision framework. Every price comes from the vendors' pricing pages on 18 September 2026 and will change, so the method is the durable part.

Key takeaways

  • Choose by task, not by brand: frontier models for judgement and long agentic work, mid-tier models for volume, open-weight or local models where data cannot leave your control.
  • The price gap is more than 10 to 1: Claude Fable 5.1 costs $10 per million input tokens and $50 output, Claude Sonnet 5 costs $2 and $10, and Gemini 3.8 Flash costs $0.75 and $3.75, according to the vendors' pricing pages.
  • Open-weight models are serious business options: Mistral Small 4 and DeepSeek V4 ship under Apache 2.0 and MIT licences with 256K to 1M token context windows and can run on hardware you own.
  • EU data residency is a configuration decision: Claude through Bedrock or Vertex EU endpoints, OpenAI's European region and Mistral's EU hosting keep inference in Europe, and the AI Act's GPAI obligations have applied since 2 August 2025.

What is the difference between frontier, open-weight and local models?

A frontier model is the most capable model a lab offers, available only through its API or a cloud partner; an open-weight model is one whose trained parameters are published so anyone can download and run it; a local model is any model, usually open-weight, running on hardware you own or rent exclusively.

The 3 categories trade capability for control. Frontier models such as Claude Fable 5.1, GPT-6 Astra or Gemini 3.8 Flash are the strongest and the simplest to use, but your prompts go to the vendor's servers. Open-weight models such as Mistral Small 4, DeepSeek V4, Qwen 3.8 or Kimi K3 can be served by a cloud provider of your choice or by you, and the licence decides what you may do with them. Local means the data never leaves the building.

Which frontier models exist in September 2026 and what do they cost?

The frontier tier in September 2026 is Anthropic's Claude Fable 5.1 and Opus 5, OpenAI's GPT-6 Astra and GPT-5.6, and Google's Gemini 3.8 Flash and 3.1 Pro, with list prices between $0.75 and $10 per million input tokens and between $3.75 and $50 per million output tokens.

A token is roughly 0.75 words, so 1 million tokens is about 700,000 words. According to Anthropic's pricing page, Claude Fable 5.1 costs $10 per million input tokens and $50 per million output tokens, Claude Opus 5 costs $5 and $25, Claude Sonnet 5 costs $2 and $10 and Claude Haiku 4.5 costs $1 and $5, with a 1 million token context window on the Claude 5 models.

OpenAI's pricing page lists GPT-6 Astra, its flagship released in September 2026, at $10 and $50; GPT-5.6 Sol at $4 and $20 under promotional pricing through at least 21 November 2026; GPT-5.6 Terra at $2 and $12; and GPT-5.6 Luna at $0.20 and $1.20. Google's Gemini API pricing lists Gemini 3.8 Flash at $0.75 and $3.75 through 31 December 2026, rising to $1.50 and $7.50 afterwards, and Gemini 3.1 Pro Preview at $2 and $12 up to 200K tokens. Batch processing is 50% off at all 3 vendors, and cached input costs 10% of the base price or less at Anthropic and OpenAI.

Which open-weight models are worth considering?

The open-weight models worth considering in 2026 are Mistral Small 4 and Large 3 (Apache 2.0), DeepSeek V4 Pro and V4 Flash (MIT), Qwen 3.8 (Apache 2.0 for the 27B model) and Kimi K3 (custom licence), because each has a permissive licence and a context window of 256K tokens or more.

Mistral is the European option. Mistral Small 4, released in March 2026, has 119 billion total parameters with 6 billion active per token, a 256K context window, an Apache 2.0 licence and an API price of $0.15 per million input tokens and $0.60 output, according to Mistral's announcement. DeepSeek V4 Pro is a 1.6 trillion parameter mixture-of-experts model with 49 billion active parameters, a 1 million token context and an MIT licence, according to its Hugging Face model card. Check the licences of the rest: the largest Qwen 3.8 model and Kimi K3 carry attribution clauses above 100 million monthly users or $20 million in monthly revenue, and Meta's newest Muse Spark models are API-only. Read the licence before you build, not after.

  • Mistral Small 4: 119B total / 6B active, 256K context, Apache 2.0, $0.15 / $0.60 on Mistral's API.
  • DeepSeek V4 Pro: 1.6T total / 49B active, 1M context, MIT; V4 Flash: 284B / 13B, MIT.
  • Qwen 3.8: 27B model under Apache 2.0, the 2.4T flagship under a custom licence.
  • Kimi K3: 2.8T total / 104B active, 1M context, custom licence with attribution clauses.

When does a local model make sense?

A local model makes sense when the data must not leave your control, when you need predictable latency without network round trips, or when your volume is high enough that owning hardware beats paying per token. If none of the 3 applies, use an API.

The hardware has caught up faster than most CEOs realise: a workstation with 512 GB of unified memory can now load the largest open-weight models entirely on device, and a single 32 GB consumer GPU runs a 27B model comfortably. In my companies local models handle classification, extraction and first drafts on confidential documents. Anything that needs judgement across a long context still goes to a frontier model, because a wrong answer costs more than the tokens.

  • Use local when: regulated or confidential data, on-premise requirements, offline use, very high steady volume.
  • Do not use local when: fewer than 10 million tokens a month, no engineer to own it, tasks that need frontier judgement.

How do you keep AI data in the EU?

You keep AI data in the EU by choosing an EU inference endpoint, which every major vendor now offers through its own API or a cloud partner: Claude through the EU regions of Amazon Bedrock or Google Cloud, OpenAI through its European API region, Mistral hosted in the EU by default, and any open-weight model on servers you choose.

According to Anthropic's documentation, Claude on Amazon Bedrock supports EU inference profiles in Frankfurt, Zurich, Stockholm, Milan, Spain, Ireland, London and Paris with a 10% premium for regional endpoints, and on Google Cloud the Claude 5 models use the "eu" multi-region endpoint. Anthropic's first-party API offers a US-only option but no EU-only option yet, so European residency for Claude goes through the cloud partners. OpenAI's data residency guide states that its Europe region performs inference in region, and Mistral's help centre states that customer data is hosted in the European Union by default.

Then there is the law. The EU AI Act's obligations for general-purpose AI models have applied since 2 August 2025, its transparency rules came into effect in August 2026, and the AI Omnibus, in force since 27 July 2026, moved most high-risk obligations to 2 December 2027, according to the European Commission. For a company using models in support, back office or sales, the practical duties today are transparency and data protection, not high-risk conformity.

What is a simple framework for choosing an AI model?

Split your AI work into 3 buckets and assign a tier to each: judgement tasks go to a frontier model, volume tasks go to a mid-tier model, and confidential or regulated tasks go to an open-weight model on EU or local infrastructure. Then route every request to the cheapest tier that passes your own evaluation.

Your own evaluation is the key phrase. Leaderboards such as Arena and Artificial Analysis are useful to shortlist, and in September 2026 the Artificial Analysis Intelligence Index shows Claude Fable 5.1, GPT-6 Astra and Claude Opus 5 at the top with small gaps. But a model that wins a coding benchmark may lose on your Spanish invoices. Build 50 to 100 real examples with correct answers, run every candidate and measure cost per completed task, including retries and human corrections.

This is how my companies run it: roughly 10% of requests go to Fable 5.1 or Opus 5, around 80% to Sonnet 5, Gemini 3.8 Flash or GPT-5.6 Terra, and the rest to an open-weight model on infrastructure we control. Re-evaluate every quarter: Sonnet 5 already had a planned price increase cancelled and Gemini 3.8 Flash doubles in price in January 2027.

Choosing an AI model in 2026 is a routing problem, not a loyalty problem. The frontier tier, Claude Fable 5.1, Opus 5 and GPT-6 Astra, is for tasks where judgement is worth $50 per million output tokens. The mid tier, Sonnet 5, Gemini 3.8 Flash and GPT-5.6 Terra, is where most of your volume should live at a fifth of the price or less. Open-weight models and local hardware are for the data you cannot let go. Put your own 50 examples in front of every candidate, measure cost per completed task, keep inference in the EU when the data requires it, and revisit the decision every quarter. That is what I do in the companies I run.

Frequently asked questions

Which AI model is best for a company in 2026?
There is no single best model. Claude Fable 5.1, Claude Opus 5 and GPT-6 Astra lead on judgement and long agentic work; Claude Sonnet 5, Gemini 3.8 Flash and GPT-5.6 Terra cover volume at a fraction of the price; open-weight models such as Mistral Small 4 and DeepSeek V4 cover privacy. Test each on your own tasks.
How much does it cost to use a frontier AI model?
According to the vendors' pricing pages in September 2026, Claude Fable 5.1 and GPT-6 Astra cost $10 per million input tokens and $50 per million output tokens, Claude Opus 5 costs $5 and $25, and Claude Sonnet 5 costs $2 and $10. Batch processing halves those prices and cached input costs 10% or less.
Can a company use Claude or GPT and keep its data in the EU?
Yes. Claude runs in EU regions of Amazon Bedrock and through Google Cloud's EU multi-region endpoint, and OpenAI offers a European API region with in-region inference. Mistral hosts in the EU by default and its open-weight models can run on your own servers.

Sources

  1. 01Anthropic: Claude API pricing
  2. 02OpenAI: API pricing
  3. 03Google AI for Developers: Gemini Developer API pricing
  4. 04Mistral AI: Mistral Small 4
  5. 05Hugging Face: deepseek-ai/DeepSeek-V4-Pro model card
  6. 06European Commission: Regulatory framework for AI (AI Act timeline)

Want to apply this to your company?

Book an hour, a morning or a day with me and we will turn the article into decisions.

See the sessions