The size difference made the comparison interesting before the first result was read: a local model recorded at 27 billion parameters against a cloud model recorded at 284 billion. We called it David versus Goliath, and the name has held up rather better than our expectations did.
The two setups were Qwen3.8 27B, running locally, and DeepSeek-V4-Flash 284B, in the cloud. Model names and sizes are reproduced as recorded in the experiment notes. The two jobs were business-value analysis, across six cases, and daily summaries, across twelve. The question was not which model is better in general. It was how each setup performed on the actual work, across speed, output quality, validation, and repeatability.
Every figure below is an aggregate reported by our own internal benchmark on 1 September 2026. The complete tables, the derivation arithmetic, the open CSV, and a full list of what we could not verify from our own records are published in the evidence note. Read that before citing any of this.
01The obvious result, and it went to the cloud
DeepSeek was faster, and not by a margin that invites argument. Mean business-analysis runtime was 17.6 seconds in the cloud against 347.8 seconds locally, about 19.8× on elapsed time. For summaries the reported means were 3.4 seconds and 32.6 seconds.
Both models completed all eighteen reported business runs. If successful completion and elapsed time were the only criteria, this comparison ends in the first paragraph and the larger model wins.
Which is close to how model selection actually gets decided. Latency is easy to measure, easy to feel, and it shows up in a demo. The properties that took us the rest of the exercise to measure do not show up anywhere until a person is reading the output and deciding whether to act on it.
02The result that changed the story
The business-value quality assessment favoured Qwen across all six cases. Case wins: Qwen 6, DeepSeek 0, ties 0. Ensemble overall quality 77.4 against 63.9.
What makes that worth writing down is not the headline gap. It is where the gap sits.
| Quality metric | Qwen3.8 local | DeepSeek V4 cloud |
|---|---|---|
| Ensemble overall quality | 77.4 | 63.9 |
| Factual grounding | 81.3 | 80.6 |
| Material coverage | 72.4 | 49.1 |
| Prioritisation and usefulness | 77.3 | 56.4 |
| Calibration | 76.4 | 71.5 |
| Clarity and concision | 76.7 | 76.5 |
Judge direction agreement: 5 of 6 cases.
Clarity and concision differed by 0.2 points. Both reports read well. This was not a prose-polish win, which is what we half expected to find once we started looking past the runtime.
The two large gaps were material coverage, 23.3 points, and prioritisation and usefulness, 20.9 points. These are reported score points on a scale whose endpoints our notes do not record, so read them as ordered, not as percentages.
03The row that turned into the whole argument
Missed major facts per output: 3.89 local against 10.54 cloud, about 63% fewer.
That single row did more to change how we evaluate models than the six case wins did, because it names a failure mode almost nobody measures. Every evaluation rubric we have read grades what the answer says. Very few grade what the answer left out.
The reason that matters is unglamorous: a fluent report which omits three material facts is harder to catch than a badly written one. The reviewer's attention goes to the prose. If there is no omission metric on your scorecard, you are grading writing.
We are not claiming this cost anyone anything. No specific missed fact, and no resulting business loss, was recorded in the experiment notes, so we are not going to invent an example.
Also on the record, and easy to overstate: the local model accepted all 186 of 186 reported candidates against 136 of 144 for the cloud model, and the narrative reports no validator errors locally where the cloud table records eight duplicates. A completed run is not the same as an accepted candidate, and an accepted candidate is not a unique insight. Both distinctions are easy to blur into a better-sounding claim, and blurring them is how a benchmark stops being useful to the only people worth convincing.
04The results that keep it honest
Two results go the other way, and they belong in the same size type as the wins.
Unsupported claims per output: 1.78 local against 0.69 cloud. The larger model asserted less than its source supported. That is a real advantage and it is a headline, not a footnote. Coverage and unsupported claims are separate properties, and in our test they pointed in opposite directions. That is the part worth taking away: a model can preserve more of what matters and still over-claim while doing it.
Then we ran the second task, and the ranking flipped.
| Measure | Qwen3.8 local | DeepSeek V4 cloud |
|---|---|---|
| Ensemble overall quality | 77.3 | 80.0 |
| Factual grounding | 76.6 | 81.5 |
| Clarity and concision | 79.6 | 82.5 |
| Unsupported claims per output | 1.44 | 0.92 |
| Missed major facts per output | 0.92 | 1.25 |
| Case wins | 5 | 6 |
One tie. Judge direction agreement: 3 of 12 cases.
DeepSeek scored 80.0 against 77.3 and won six cases to five with one tie. Judge direction agreement was three of twelve, which is low enough that the summary result should be read as evaluator uncertainty rather than a settled outcome. We report it because it is what the test returned, not because it helps.
05Repeatability, and what we can honestly say about it
The local model repeated its outputs exactly in every tested case: 6 of 6 business cases and 12 of 12 summary cases. The cloud model repeated none of them: 0 of 6 and 0 of 12.
The caveat goes first, because it is load-bearing. Our notes do not record what exact match compared. Raw string, normalised output, structured fields, semantic equivalence: we do not know. So this is exact repeatability in the cases we tested, not a property of either model.
Why it still matters. A report regenerated weekly and reviewed by a person costs review time on every cycle when the output changes between runs, whether or not the change is meaningful. We did not measure that cost, and no saved-hours figure should be read into this.
And the line that stops it becoming a sales point: the local model repeated its summary outputs exactly, 12 of 12, and still lost the summary task on quality, 77.3 against 80.0. A repeatable output can be a consistently weaker one.
06What we cannot tell you
This is published with its gaps listed, because a benchmark whose gaps are not listed is unusable by anyone careful.
- No per-run logs, so no variability, medians or uncertainty.
- No judge identities, prompts or blinding procedure, and no scoring rubric or ensemble weights.
- No exact-match rule, which is why the repeatability result is scoped to the tested cases.
- No model digests, quantisation, sampling settings or seeds, and no record of the local September hardware.
- No per-case scores, and no original screenshots.
What that costs us. With six business cases, twelve summary cases and no per-case scores, there is no variance, no confidence interval and no inferential test available. We make no claim of statistical significance and none should be read into these numbers.
The complete missing-field list, and the per-run record schema we would use to run this properly next time, are both published on the evidence note. Take the schema.
07What we actually changed
The finding is not that small models win, and it is not that local deployment beats cloud. Parameter count predicted neither of our two results, and the two results disagree with each other.
What changed is the question we ask at the start of an engagement. Not which model is best, but what this workflow can absorb. A 347-second run is fine for a weekly report and useless inside a live tool. Ten missed major facts is survivable in a first draft and not survivable in a board pack. Same two models, opposite answers, depending on the job.
In practice that means scoring the workflow on five things before choosing anything: output quality, omissions, unsupported claims, latency tolerance, and repeatability need. Most teams we speak to have scored the first and the fourth, and never the middle three.
There is a separate engineering story sitting behind this one, and it should not be merged into it. In a different experiment, on a different machine, in August 2026, we ran a 284B-parameter model locally on a desktop workstation at about 3.5 tokens per second. That is a feasibility result about hardware. It says nothing about the quality comparison above, and we do not report the two together.
Frequently asked questions
Did the smaller local model win?
On business-value analysis, yes, in the six cases we tested: 6 case wins to 0, with ensemble overall quality 77.4 against 63.9. On daily summaries it lost, 77.3 against 80.0, and the cloud model was about 19.8× faster throughout and produced fewer unsupported claims per output. There is no overall winner in this test, and we have not invented one by combining the two task families into a single score.
What does the missed-major-facts figure measure?
It is the reported average number of material facts absent from an output, scored by the same evaluation ensemble that produced the quality figures. Business outputs averaged 3.89 locally and 10.54 in the cloud, about 63% fewer. It is an omission measure, and it is separate from unsupported claims, where the cloud model was ahead. We do not have the rubric that defined a major fact, and that limitation is published with the figures.
Can these numbers support a significance claim?
No, and we do not claim it. Six business cases, twelve summary cases, and no per-case scores means no variance, no confidence interval and no inferential test. The reason to read the benchmark is not our result. It is that the ranking reversed between two tasks, which is an argument for testing your own workflow rather than trusting ours.
Can we see the underlying data?
Yes. The reported aggregates are published as an open CSV, alongside every table, the derivation arithmetic, and the complete list of experiment artefacts we could not recover. Nothing is gated and no email address is asked for.
Every figure above, the full tables, the derivations, and what we could not verify: read the complete evidence note, or take the aggregates directly as an open CSV. If you want the same five-dimension read on one of your own workflows, start a conversation.
Internal benchmark, reported 1 September 2026. Qwen3.8 27B local against DeepSeek-V4-Flash 284B cloud, on six business-value analysis cases and twelve daily-summary cases. Model names and sizes as recorded in the experiment notes.