On September 11 the GOLD table at the top of octobench's BENCHMARK.md refreshed with a new run. Twenty real bug fixes and features, taken from merged pull requests in projects like Redis, Fastify, Guzzle and ripgrep, each graded by the project's own hidden tests. One column on the board cost $1.64. Not per task. For all twenty. It landed 19 of them.

The column next to it was Claude Opus 5 running in Claude Code. Also 19 of 20. That one cost $41.34.

The $1.64 column was GLM-5.3 Flash, an open weight model anyone can download from Hugging Face under the MIT license. So here's the short answer: open weight models publish their trained parameters, so you can download, run and modify them yourself, and the best of them now hold up on real coding work. On our real-PR benchmark, GLM-5.3 solved 20 of 20 and GLM-5.3 Flash 19 of 20, at a fraction of closed-model cost.

Below: what "open weight" actually gets you, how we tested, the full scoreboard, and the places where open weights still lose.

What is an open weight model?

An open weight model is one whose trained parameters, the weights, are published for anyone to download. You can run it on your own hardware, host it with any provider, or fine-tune it, within the terms of its license. You get the finished model, not the recipe that produced it.

"Open weight" says nothing about size, so I checked what downloading one actually means. On October 2 I asked the Hugging Face API for the file lists of the three open models in this post and summed the weight shards:

bash
$ for m in zai-org/GLM-5.3 zai-org/GLM-5.3-Flash deepseek-ai/DeepSeek-V4-Flash-0731; do
    curl -s "https://huggingface.co/api/models/$m?blobs=true" | jq -r --arg m "$m" \
      '[.siblings[] | select(.rfilename | endswith(".safetensors")) | .size]
       | "\($m): \(length) shards, \(add / 1e9 | floor) GB"'
  done
zai-org/GLM-5.3: 141 shards, 755 GB
zai-org/GLM-5.3-Flash: 62 shards, 328 GB
deepseek-ai/DeepSeek-V4-Flash-0731: 48 shards, 166 GB

Open does not mean it fits on your laptop. GLM-5.3 Flash is a mixture-of-experts model with 320B total parameters and 18B active per token, according to its model card. Only the active slice does work on each token, which keeps serving cheap, but 328 GB of weights is a multi-GPU server, not a MacBook. If a laptop is the goal, our guide to running a coding agent locally covers the smaller models that fit.

ModelLicenseWeightsOctomind hub price (in / out per 1M tokens)
GLM-5.3GLM-5.3 License (MIT plus one clause)755 GB$1.40 / $4.40
GLM-5.3 FlashMIT328 GB$0.15 / $0.50
DeepSeek V4 Flash (0731)MIT166 GB$0.44 / $1.32

Open weight vs open source: not the same thing

Open weight means the parameters are public. Open source, for AI, asks for much more. The Open Source Initiative's Open Source AI Definition requires the parameters, the complete code used to train and run the system, and data information detailed enough that "a skilled person can build a substantially equivalent system", all under OSI-approved terms.

None of the models in this post meets that bar. Nobody can rebuild GLM-5.3 from what Z.ai published; you can only use, modify and redistribute the result. So when someone asks for the best open source model for coding, the honest answer is that the strong contenders are open weight. The open source vs open weight difference matters in two practical places.

The license is the real contract. DeepSeek V4 Flash and GLM-5.3 Flash are MIT: commercial use, fine-tuning and redistribution, no strings. GLM-5.3 ships under its own GLM-5.3 License, which reads like MIT except for one clause: a model-as-a-service operator whose revenue exceeds $10 billion over any 12 months must pass Z.ai's security review before commercial use. That clause won't touch most of us, but it's exactly why you read the license instead of trusting the word "open."

You can audit behavior, not provenance. You can probe, quantize and fine-tune open weights as much as you like. You still can't inspect what they were trained on.

Open weight vs closed weight models: what you actually trade

Closed weight models such as Claude, GPT and Gemini exist only behind their vendor's API. You rent them per token, on the vendor's terms, for as long as the vendor keeps them running.

That last part is concrete for us. Our own hub retired Claude Opus 5, the closed model in the benchmark below, on October 4; Opus 5.5 is the current one. GLM-5.2, the open model that topped our first real-PR benchmark in July, left the hub on September 17. Its weights are still on Hugging Face today. When a closed model retires, the exact version you tuned your prompts and workflows against is gone. With open weights, the checkpoint you tested is the checkpoint you can keep running.

The trade runs the other way too. Closed frontier models are still faster per task, as the numbers below show, and someone else carries the GPUs.

How we tested open weight models on real pull requests

We tested on octobench, our public benchmark built from merged pull requests. Each case checks out the repository at the commit before the fix, hands the agent the request, then grades the result two ways: the fix's own held-out tests, which the agent never sees, and a three-model judge panel scoring the work from 0 to 100.

This post draws on the GOLD set: 20 one-shot cases across C++, JavaScript, PHP, Python and Rust, plus 10 long-run sequences that add up to 90 consecutive feature turns on repos like DuckDB, cargo, CPython and ESLint. Every client got the same system prompt. Web search and fetch were disabled, and GitHub was unreachable once setup finished, so nobody could read the merged fix. We learned that one the hard way; the DeepSeek V4 Flash post has the story.

One caveat up front, because it's our benchmark. The open models ran in Octomind; the closed ones ran in their vendors' own clients, Claude Code and Codex. Octobench compares client-plus-model pairs on purpose. To check that the model isn't just riding our harness, the same GLM-5.3 Flash also ran in opencode: 18 of 20 for $2.05.

The scoreboard: 20 real fixes, open vs closed weights

On the one-shot cases, the only perfect score belongs to an open weight model, and the cheapest model to match Opus is open weight too:

ModelWeightsClientSolvedJudge avgTotal costCost per solveMedian time per case
GLM-5.3openOctomind20/2092.92$21.14$1.0615.0 min
GLM-5.3 FlashopenOctomind19/2089.47$1.64$0.0917.7 min
Claude Opus 5closedClaude Code19/2090.10$41.34$2.184.7 min
GPT-5.6 SolclosedCodex19/2089.60$15.35$0.814.0 min
GPT-5.6 LunaclosedCodex16/2082.08$0.88$0.063.4 min

Cost per solve is total cost divided by solved cases. Every per-case result, with traces and judge verdicts, is in BENCHMARK.md, pinned to the commit.

Three things stand out. GLM-5.3 is the only model that landed all twenty. GLM-5.3 Flash matched Opus's solved count at 4% of its cost, which is the number that changes budgets. And cheap is not exclusive to open weights: GPT-5.6 Luna had the lowest cost per solve, but it solved 16. Four failed fixes are a cost too, because someone has to redo them.

The long runs, 90 turns of building on the same codebase, tell a similar story:

ModelWeightsClientTurns passedTotal costCost per passed turnAgent time
GPT-5.6 SolclosedCodex82/90$116.18$1.424.9 h
Claude Opus 5closedClaude Code81/90$363.77$4.4915.2 h
GLM-5.3openOctomind80/90$150.44$1.8817.9 h
GLM-5.3 FlashopenOctomind80/90$10.36$0.1325.7 h
GPT-5.6 LunaclosedCodex75/90$6.32$0.085.2 h

An honesty note on the Flash row: three of its ten sequences come from a second run, and BENCHMARK.md flags them. On its first pass it passed 75 of 90 turns. Either way, it delivered long-run work in the same band as the frontier models for a few percent of the Opus bill. It also needed 25.7 hours of agent time to do it, which brings us to the costs that aren't dollars.

Where open weight models still lose

Open weight models lose on speed and hosting, and they share the hardest failures with everyone else.

Speed. Median time per one-shot case was 17.7 minutes for Flash and 15.0 for GLM-5.3, against 4.7 for Opus and 4.0 for Sol. Over the long runs, Flash needed about five times Sol's agent time. Part of that is the model reasoning more, and part is provider throughput: when we moved DeepSeek V4 Flash from a third-party endpoint to DeepSeek's official API, requests got roughly three times faster with nothing else changed. If you sit and watch the agent work, this is the number you'll feel. If it runs in the background while you do something else, it matters much less.

Hosting. Open weights are free to download and expensive to serve. 328 GB of Flash weights means a GPU server, and at that point you're running infrastructure, not an agent. We run open weights through hosted providers and keep self-hosting as the exit, which exists only because the weights are public.

The hardest walls are shared. Flash's one failed one-shot case was yaml-cpp's octal scalars, the same case Opus failed. React's hidden hydration hang has beaten every model we've run on it, open or closed, including a 271-minute DeepSeek attempt that passed 66 of 67 hidden tests and still counts as a fail. At the top end, the gap between open and closed weights is now smaller than the gap between easy and hard tasks.

Which open weight model should you use for coding?

Of the best open weight models we've measured, I'd start with GLM-5.3 Flash for most coding work and step up to GLM-5.3 when a task needs the ceiling. Here are my picks from our own numbers:

  • Default: GLM-5.3 Flash. MIT license, $0.15 in and $0.50 out per million tokens on our hub, 19 of 20 one-shot fixes and 80 of 90 long-run turns. It's what I'd give a background agent that runs all day.
  • When you want the ceiling: GLM-5.3. The only 20 of 20 on the board, at about 13 times Flash's cost on the one-shot set. Read the license if you run a model-as-a-service business.
  • The cheapest strong option on an older set: DeepSeek V4 Flash. 45 of 50 in our August run at about $0.035 per solve, MIT. That was a different 50-case set, so don't line it up against the tables above. DeepSeek has since shipped V4.1 Flash, which we haven't benchmarked yet.

For how complete agents compare, rather than models, see our ranking of the best AI coding agents. Every model above also has a page with current prices, such as GLM-5.3 Flash.

How to run an open weight model in Octomind

Switching Octomind to an open weight model takes one flag. Through the Octomind hub, you pass the hub's model name:

bash
$ octomind run developer --format jsonl -m octohub:glm-5.3-flash < q.txt > out.jsonl
$ jq -r 'select(.type=="assistant") | .content' out.jsonl
`dict.setdefault(key, default)` returns the existing value for `key` without inserting the default, and only sets/returns the default when the key is absent.
$ jq -c 'select(.type=="cost") | {session_tokens, input_tokens, session_cost}' out.jsonl
{"session_tokens":13162,"input_tokens":13118,"session_cost":0.0020111619999999995}

That's a real run from October 2, with a one-line Python question in q.txt. It cost two tenths of a cent, and almost all of the 13,118 input tokens were the system prompt and tool definitions. Inside an interactive session, /model octohub:glm-5.3-flash switches without restarting.

For weights you host yourself behind Ollama or any OpenAI-compatible server, use the ollama: or local: provider with your endpoint URL. The providers guide lists the exact variables. If nothing should leave your network, point the supervisor and compression models at your server too.

FAQ

What is an open weight model?

An open weight model is an AI model whose trained parameters are published for download. You can run it on your own hardware, host it with any provider and fine-tune it, within the terms of its license. You get the finished model, but usually not its training data or full training code.

What is the difference between open weight and open source?

Open weight means the parameters are public. Open source AI, under OSI's definition, also requires the complete training and inference code and data information detailed enough to rebuild a substantially equivalent system, all under OSI-approved terms. Popular "open" models such as GLM and DeepSeek are open weight, not open source.

What is the best open source model for coding?

Strictly by OSI's definition, the strongest coding models aren't open source; they're open weight. The best open weight coding model in our octobench GOLD run was GLM-5.3, which solved 20 of 20 real pull requests, and the MIT-licensed GLM-5.3 Flash solved 19 of 20 for $1.64 in total.

Are open weight models as good as closed weight models for coding?

On our benchmark, the best ones are. GLM-5.3 Flash matched Claude Opus 5's 19 of 20 solved at about 4% of the cost, and GLM-5.3 solved all 20. Closed models are still roughly three to four times faster per task, and both camps fail the same hardest cases.

Run it on your own backlog

A year ago, picking an open weight model for coding meant accepting worse results to save money or keep code private. On real pull requests, that trade is gone; what's left is wall-clock time against a model you can keep. Octobench is public, so run it against your own repos before you believe my numbers.