open-weightlocal-ailanguage-modelsfrontier-modelsai-hardware

Open-Weight Models in 2026: How Far Behind Are They?

October 5, 20267 min readPIXEL MANAGEMENT

This article is also available in Dutch

Anyone who tried a local language model in 2024 mostly remembers that it was worse than ChatGPT. That verdict was fair at the time. It no longer is.

Open-weight models are language models whose makers publish the weights, so you can run them on your own hardware without sending data to a vendor. In 2026, the best open-weight models trail the most expensive closed models from OpenAI, Anthropic and Google by about four months on average. What the frontier could do in spring, you run yourself by autumn.

How far behind are open-weight models now?

The gap between open-weight and closed models has been measured, and it is smaller than most decision-makers assume. Research institute Epoch AI compares models on a composite capabilities index, and concluded in May 2026 that since January 2026 the best open-weight models have lagged the best closed models by an average of four months. The average gap on that index is roughly the difference between GPT-5 and GPT-5.5.

Four months. In 2024, the same institute put the lag at five to twenty-two months, depending on which benchmark you picked, which in practice meant a local model was a generation behind for any serious work.

The gap doesn't close neatly in one direction. Stanford HAI's AI Index Report 2026 measured on Chatbot Arena that the top closed model led the top open model by 3.3 percent in March 2026, up from 0.5 percent in August 2024. Six of the ten highest-ranked models were closed. The distance breathes with every major release: a new frontier model pulls the gap open, the next open release closes it again.

MeasureSourceResult
Lag in timeEpoch AI, May 20264 months on average
Lag in 2024Epoch AI, November 20245 to 22 months
Gap on Chatbot ArenaStanford AI Index 20263.3% (March 2026)
Models on one consumer GPUEpoch AI, September 2025catch the frontier within 6 to 12 months

That last row matters most for a business. According to Epoch AI, models that fit on a single consumer GPU consistently catch the frontier within six to twelve months. What needs a subscription to the most expensive model today will run on a workstation in a year. Under your desk.

Which open-weight models can you run yourself today?

Open-weight models come in every size in 2026, and nearly all of them under a licence that allows commercial use. The selection below is based on what a company can realistically run in its own environment.

ModelMakerParameters (active)LicenceMinimum hardware
gpt-oss-20bOpenAI21B (3.6B)Apache 2.016 GB of memory
Qwen3.6-27BAlibaba27BApache 2.0one workstation GPU
Gemma 4 31BGoogle31BApache 2.0one H100 (80 GB) unquantised
gpt-oss-120bOpenAI117B (5.1B)Apache 2.0one 80 GB GPU
Mistral Small 4Mistral AI119B (6B)Apache 2.0four H100s
DeepSeek V4-FlashDeepSeek284B (13B)MITa multi-GPU server

Two things stand out. First, many of the strongest open models come from the same companies that sell closed ones: OpenAI's gpt-oss-120b runs on a single 80 GB GPU according to its own model card, and Google released Gemma 4 in April 2026 under Apache 2.0.

Second, the name tells you nothing about the hardware. Mistral Small 4 is "small" within its own family, but it has 119 billion parameters and according to Mistral needs at least four NVIDIA H100s. The rule of thumb: total parameters decide how much memory you need, active parameters decide how fast the model answers.

If you're after models of 1 to 15 billion parameters, we cover those separately in small language models for business. Our older overview of open-source AI models shows how quickly this list dates.

How good are open-weight models in languages other than English?

For a European business, the English benchmark matters less than performance in the language your documents are in. The independent EuroEval project tests exactly that per language, and its Dutch leaderboard covers 170 models, which makes it a good test case for a mid-sized European language.

The result on 5 October 2026: the top spot goes to the newest GPT-6 models. Next, the open GLM-5.3-Flash sits level with Google's Gemini 3.7 Flash. Qwen3.6-27B and Gemma 4 31B, both small enough for a single workstation, score at the level of Claude Sonnet 4.5 and GPT-4.1. A year ago, many companies were paying per item for exactly that level.

An honest caveat: the most expensive flagships, such as Claude Opus and Gemini Pro, aren't on this leaderboard. "Just as good in Dutch" here means just as good as the mid-tier of the closed vendors. For most business work, that's the relevant comparison.

What hardware do you need to run open-weight models?

Hardware for open-weight models has become a purchase in 2026, not a construction project. Three classes cover almost every business need.

The first is a compact AI workstation. The NVIDIA DGX Spark has 128 GB of unified memory and, according to NVIDIA, runs models up to 200 billion parameters. Its list price was raised from $3,999 to $4,699 in February 2026 because of memory shortages. An Apple Mac Studio with M5 Ultra and up to 512 GB of memory falls in the same category.

The second is a workstation or server with one heavy GPU, such as the RTX PRO 6000 Blackwell with 96 GB, or an 80 GB H100. That runs gpt-oss-120b and Gemma 4 31B for a department.

The third is a multi-GPU server for the largest open models, or for many concurrent users. Take that step only once the first two demonstrably fall short. Our guide to AI hardware for business works through the choice per class.

[ TIME SAVED ]

Save 6 hours per week on manually summarising and routing internal documents with a local model

Where is a frontier model still worth the money?

A frontier model still pays off on three kinds of work, and they share one trait: many steps in a row have to go right.

Long agent tasks come first. An agent that carries out twenty steps on its own multiplies every small error: a model that gets each step right 97 percent of the time completes a twenty-step chain only 54 percent of the time, and at 99 percent per step that becomes 82 percent. That's arithmetic. The hardest reasoning work is second, such as legal or financial analysis where the question itself is unclear. Writing or understanding complex software is third.

For everything else, the reverse holds. Classifying, extracting, summarising, translating and answering from your own knowledge base sit comfortably within what open-weight models can do now. And when a task is narrow enough, a distilled model can even match the large model. How that works, and which licences allow it, is in AI model distillation.

How do you plan around a four-month lag?

The measured lag of open-weight models gives you a planning rule we use in projects. What a frontier model can do today, you can run in your own environment in about a quarter. In a year, it will probably fit on one workstation.

That changes how you build.

Prototype with the best model, but build model-agnostic. A prototype on a frontier API shows whether the task can be automated at all. Put the model behind your own abstraction layer, and switching to a local model becomes a configuration change rather than a rebuild.

Don't sign multi-year API commitments for work that can run locally within a year. The price per item is also falling hard: the AI Index Report 2025 measured that the cost of querying a model at GPT-3.5 level fell more than 280-fold between November 2022 and October 2024.

And keep your own test set of real cases from your business. You can check every new open release against it in an afternoon. That's all it takes. That way you see for yourself when your task crosses the line, instead of waiting for a vendor to tell you. Whether the moment has come to switch is something you judge with the trade-offs in local AI vs cloud AI and the broader guide to local AI for business.

[ SERVICE ]

Learn more about local AI?

View service →

Curious how much time you could save?

Request a free efficiency audit. We'll analyze your processes and show you where the gains are, no strings attached.