The finding is not that small models win. It is that the ranking changed when the job changed. On business-value analysis, a locally run Qwen3.8 27B beat cloud-hosted DeepSeek-V4-Flash, recorded at 284B, on all six cases. On daily summaries the cloud model won. Both results are published here at the same size.
This page is the evidence, not the argument. The narrative version is a separate post. What follows is the reported performance and repeatability data, the business-value quality data, the daily-summary quality data, the derivation arithmetic, the claim boundaries, and the complete list of experiment artifacts we do not have.
Read this before the tables. This is not a rerun. Every figure below is a transcription of reported aggregate results from three contemporaneous experiment records, not a result recomputed from individual runs. The per-run logs, judge identities, scoring rubric, ensemble weights and exact-match rule were not recovered. Six business cases and twelve summary cases, with no per-case scores, means no variance, no confidence interval and no inferential test. We make no significance claim. Section 06 lists every missing field.
Two separate experiments, and they are not merged anywhere on this page. The model comparison here was reported on 1 September 2026. A local-hardware feasibility run on a different machine was reported on 27 August 2026. No hardware figure from the August run appears on this page, and no claim here draws on it.
01Performance and repeatability
Source: S1. Reconstructed from flattened screenshot text and its accompanying narrative.
| Metric | Qwen3.8 Local | DeepSeek V4 Cloud | Reading |
|---|---|---|---|
| Business runs successful | 18/18 | 18/18 | Both completed all reported business runs |
| Mean business runtime | 347.8 s | 17.6 s | Cloud substantially faster |
| Accepted candidates | 186/186 | 136/144 | Local accepted every reported candidate |
| Validation errors | Narrative reports none | Table text reports “8 duplicates” | Exact error taxonomy needs original screenshot/logs |
| Exact business repeatability | 6/6 cases | 0/6 cases | Local repeated exactly across all tested business cases |
| Mean summary runtime | 32.6 s | 3.4 s | Cloud substantially faster |
| Required section coverage | 100% | 100% | Both met reported section coverage |
| Numeric grounding precision | 94.0% | 93.1% | Local ahead by 0.9 percentage points |
| Exact summary repeatability | 12/12 cases | 0/12 cases | Local repeated exactly across all tested summary cases |
Details that matter
Completion and acceptance are different. DeepSeek completed 18/18 runs even though only 136/144 candidates were accepted. A completed request does not mean every candidate passed validation.
Candidate count is not unique business value. These may include repeated outputs across repeated cases. Do not describe 186 accepted candidates as 186 unique discoveries or customer outcomes.
Repeatability is scoped. The narrative says Qwen was “completely deterministic”. The narrower supported statement is exact repeatability in the reported cases. The metric’s comparison rule, raw string, normalised output, structured fields or semantic equivalence, is not supplied. No universal claim about all Qwen outputs follows from this test.
Grounding precision is a separate measure. The 94.0% numeric grounding precision is not the same as the factual-grounding score in either quality table. Their denominators and scoring procedures are not documented here.
Validation-row ambiguity is preserved. The extracted table includes the header “Fatal validation errors”, omits a clear local numeric cell, and includes “8 duplicates” for DeepSeek. The prose independently says Qwen produced no validator errors. That is the wording we use; we have not silently converted the table into a fully verified error taxonomy.
02Business-value quality
Source: S2. Score units and scale endpoints were not supplied. These are reported score points, without a percent sign and without assuming a 100-point rubric.
| Quality metric | Qwen3.8 Local | DeepSeek V4 Cloud | Qwen minus DeepSeek | Interpretation |
|---|---|---|---|---|
| Ensemble overall quality | 77.4 | 63.9 | +13.5 | Local higher |
| Factual grounding | 81.3 | 80.6 | +0.7 | Small local advantage |
| Material coverage | 72.4 | 49.1 | +23.3 | Largest score gap in local’s favour |
| Prioritisation/usefulness | 77.3 | 56.4 | +20.9 | Substantial local advantage |
| Calibration | 76.4 | 71.5 | +4.9 | Local higher |
| Clarity/concision | 76.7 | 76.5 | +0.2 | Nearly equal reported scores |
| Unsupported claims/output | 1.78 | 0.69 | +1.09 | Cloud better; lower is preferable |
| Missed major facts/output | 3.89 | 10.54 | −6.65 | Local better; lower is preferable |
| Case-level result | Reported result |
|---|---|
| Qwen wins | 6 |
| DeepSeek wins | 0 |
| Ties | 0 |
| Judge direction agreement | 5/6 cases |
The local result was not principally a prose-polish win. Clarity/concision differed by only 0.2 points. The striking gaps were material coverage and prioritisation/usefulness. That supports a statement about preserving decision-relevant information, and it supports no causal claim about why Qwen performed better.
The missed-major-facts measure gives the result an accessible business question: what did the report leave out? Local outputs reportedly missed 3.89 major facts on average compared with 10.54 for cloud outputs.
The local model also generated more unsupported claims per output. The larger model’s lower count is a real advantage and it belongs in the main result, not in a footnote. These are counts per output, not normalised rates. Output lengths and total factual-claim counts are unavailable.
Five-of-six judge direction agreement is not unanimous agreement across all cases. The source does not identify the judges, the ensemble rule, the blinding procedure, or the tie rules.
03Daily-summary quality
Source: S3. Again, values are reported score points, not percentages unless explicitly stated elsewhere.
| Quality metric | Qwen3.8 Local | DeepSeek V4 Cloud | Qwen minus DeepSeek | Interpretation |
|---|---|---|---|---|
| Ensemble overall quality | 77.3 | 80.0 | −2.7 | Cloud higher |
| Factual grounding | 76.6 | 81.5 | −4.9 | Cloud higher |
| Material coverage | 79.8 | 79.5 | +0.3 | Nearly equal |
| Prioritisation/usefulness | 78.3 | 79.9 | −1.6 | Cloud higher |
| Calibration | 74.3 | 78.9 | −4.6 | Cloud higher |
| Clarity/concision | 79.6 | 82.5 | −2.9 | Cloud higher |
| Unsupported claims/output | 1.44 | 0.92 | +0.52 | Cloud better |
| Missed major facts/output | 0.92 | 1.25 | −0.33 | Local better |
| Case-level result | Reported result |
|---|---|
| Qwen wins | 5 |
| DeepSeek wins | 6 |
| Ties | 1 |
| Judge direction agreement | 3/12 cases |
The business-analysis win does not transfer to every task. Daily-summary results show a small cloud aggregate lead, a split case result, and low reported judge direction agreement. This is a mixed outcome with evaluator uncertainty, not a decisive universal ranking.
Exact repeatability and quality are distinct. Qwen repeated its summary outputs exactly in 12/12 cases, and DeepSeek still had the higher aggregate summary quality score. A repeatable output can be a consistently weaker one.
The two task families are not combined into a single overall score anywhere on this page. Their scoring weights, use-case importance, and aggregation rules are unknown.
04Calculations and reusable data
All calculations use rounded values from the source. The extra decimal places below explain the arithmetic. Copy should retain sensible rounding.
| Derived quantity | Formula | Result | Suggested wording |
|---|---|---|---|
| Reported total-parameter ratio | 284 / 27 | 10.52× | Approximately 10.5× the reported total parameters |
| Business runtime ratio | 347.8 / 17.6 | 19.76× | Cloud was about 19.8× faster by mean elapsed runtime |
| Business time difference | 347.8 − 17.6 | 330.2 s | About 5 minutes 30 seconds less per reported mean run |
| Business relative elapsed-time reduction | (347.8 − 17.6) / 347.8 | 94.94% | Cloud mean elapsed time was about 94.9% lower |
| Summary runtime ratio | 32.6 / 3.4 | 9.59× | About 9.6× from displayed means; source narrative says 9.5× |
| Summary time difference | 32.6 − 3.4 | 29.2 s | About 29 seconds less per reported mean summary |
| Local candidate acceptance | 186 / 186 | 100% | All reported local candidates accepted |
| Cloud candidate acceptance | 136 / 144 | 94.44% | About 94.4% of reported cloud candidates accepted |
| Acceptance-rate gap | 100 − 94.4444 | 5.56 pp | Local ahead by about 5.6 percentage points |
| Accepted-candidate count difference | 186 − 136 | 50 | 50 more accepted candidates across the reported runs |
| Relative accepted-candidate difference | (186 − 136) / 136 | 36.76% | About 36.8% more accepted candidates; not a recall metric |
| Business overall score gap | 77.4 − 63.9 | 13.5 points | Local scored 13.5 points higher |
| Business material-coverage gap | 72.4 − 49.1 | 23.3 points | Local scored 23.3 points higher on coverage |
| Business usefulness gap | 77.3 − 56.4 | 20.9 points | Local scored 20.9 points higher on usefulness |
| Reduction in missed major facts, business | (10.54 − 3.89) / 10.54 | 63.09% | About 63.1% fewer missed major facts per output |
| Business unsupported-claim difference | 1.78 − 0.69 | 1.09/output | Local had 1.09 more unsupported claims per output |
| Summary overall score gap, cloud lead | 80.0 − 77.3 | 2.7 points | Cloud scored 2.7 points higher |
| Numeric grounding precision gap | 94.0 − 93.1 | 0.9 pp | Local ahead by 0.9 percentage points |
| Business judge direction agreement | 5 / 6 | 83.33% | Prefer the original 5/6 cases |
| Summary judge direction agreement | 3 / 12 | 25% | Prefer the original 3/12 cases |
The aggregates as an open CSV
The same figures, machine readable. The CSV excludes the ambiguous validation row. Ties and judge agreement are experiment-level values rather than separate model values: business ties 0 and agreement 5/6; summary ties 1 and agreement 3/12.
Download benchmark-aggregates.csv
task,metric,qwen_local,deepseek_cloud,unit,source business,successful_runs,18,18,count,S1 business,attempted_runs,18,18,count,S1 business,mean_runtime,347.8,17.6,seconds,S1 business,accepted_candidates,186,136,count,S1 business,generated_candidates,186,144,count,S1 business,exact_repeatability_cases,6,0,count,S1 business,repeatability_cases_total,6,6,count,S1 business,ensemble_overall_quality,77.4,63.9,reported_score,S2 business,factual_grounding,81.3,80.6,reported_score,S2 business,material_coverage,72.4,49.1,reported_score,S2 business,prioritization_usefulness,77.3,56.4,reported_score,S2 business,calibration,76.4,71.5,reported_score,S2 business,clarity_concision,76.7,76.5,reported_score,S2 business,unsupported_claims_per_output,1.78,0.69,count_per_output,S2 business,missed_major_facts_per_output,3.89,10.54,count_per_output,S2 business,case_wins,6,0,count,S2 summary,mean_runtime,32.6,3.4,seconds,S1 summary,required_section_coverage,100,100,percent,S1 summary,numeric_grounding_precision,94.0,93.1,percent,S1 summary,exact_repeatability_cases,12,0,count,S1 summary,repeatability_cases_total,12,12,count,S1 summary,ensemble_overall_quality,77.3,80.0,reported_score,S3 summary,factual_grounding,76.6,81.5,reported_score,S3 summary,material_coverage,79.8,79.5,reported_score,S3 summary,prioritization_usefulness,78.3,79.9,reported_score,S3 summary,calibration,74.3,78.9,reported_score,S3 summary,clarity_concision,79.6,82.5,reported_score,S3 summary,unsupported_claims_per_output,1.44,0.92,count_per_output,S3 summary,missed_major_facts_per_output,0.92,1.25,count_per_output,S3 summary,case_wins,5,6,count,S3
05Claim bank and boundaries
The left column includes claims we do not make. They are published so that nobody has to guess where the boundary sits, and so that a reader can check any sentence we publish elsewhere against this table.
| Proposed claim | Status | Basis / required context |
|---|---|---|
| “Our smaller local model won all six business-analysis cases.” | Supported as reported | S2; name both models and retain six-case scope |
| “27B local versus 284B cloud.” | Source-recorded labels | S1; verify manifests for independently certified specifications |
| “13.5 points higher overall business-analysis quality.” | Supported arithmetic | S2; reported score points, not percentage points |
| “About 63% fewer missed major facts per business output.” | Supported arithmetic | S2; aggregate means, not a claim about every output |
| “100% exact repeatability in the tested local cases.” | Supported with scope | S1; include 6 business and 12 summary cases |
| “Cloud was about 20× faster on business analysis.” | Supported rounded comparison | S1; elapsed runtime under the tested setups |
| “The local model was more accurate in every way.” | Contradicted | Cloud had fewer unsupported claims and stronger summary scores |
| “The local model hallucinated less.” | Not supported | Reported unsupported-claim counts favour cloud |
| “No validator errors means no incorrect claims.” | Invalid inference | Validation and factual support are different checks |
| “186 unique business insights.” | Not established | Candidate uniqueness across runs is unknown |
| “63% less business risk.” | Not measured | Missed facts are not a business-risk metric |
| “A 10× smaller model uses 10× less compute.” | Not established | Total parameters do not establish active compute or cost |
| “The local setup saved a measured amount of money.” | Not measured | No energy, hardware amortisation, or billing comparison |
| “The experiment proves local AI is always better.” | Contradicted by scope | Different task outcomes and a small case set |
| “These results are statistically significant.” | Not established | No per-case variance, inferential test, or confidence interval |
| “The cloud test exercised a million-token context.” | Not established | A reported capability label is not actual input length |
| “We ran DeepSeek locally at ~3.5 tokens/second on the existing PC.” | Separately reported | S4/S5; do not merge with the September cloud experiment |
The result in one paragraph
We compared a locally run Qwen3.8 27B model with cloud-hosted DeepSeek-V4-Flash, recorded as 284B, on business-value analysis and daily summaries. Qwen won all six business-analysis cases and scored 77.4 versus 63.9 overall, with stronger material coverage and fewer missed major facts. DeepSeek was about 19.8× faster on business analysis, produced fewer unsupported claims, and scored higher overall on summaries. Our result supports evaluating models against the specific workflow and its trade-offs.
06Missing experiment artifacts and reproducibility fields
The following gaps limit independent verification. They do not erase the recorded result, but they define what this evidence note can and cannot substantiate.
| Missing field | Why it matters |
|---|---|
| Exact model artifact/tag/digest | Distinguishes labels from reproducible model identity |
| Quantization and precision | Can materially change local quality and performance |
| Local September hardware | Needed to interpret local runtime |
| Cloud provider, endpoint, and serving configuration | Needed to interpret the cloud setup |
| Prompt templates and input packets | Needed to verify input parity |
| Token counts and output budgets | Needed to distinguish content length from throughput |
| Sampling settings, seeds, and thinking settings | Needed to interpret repeatability and quality |
| Retry policy and timeout handling | Needed to interpret successful-run counts |
| Case selection and case IDs | Needed to assess representativeness and cherry-picking |
| Per-run runtime values | Needed for variability, medians, and uncertainty |
| Exact-match rule | Needed to reproduce repeatability |
| Candidate validation rules | Needed to interpret duplicates and acceptance |
| Grounding definitions and denominators | Needed to interpret precision and factual scores |
| Judge models/people, prompts, and blinding | Needed to understand evaluation bias |
| Ensemble weights and score scale | Needed to reproduce overall quality |
| Per-case scores | Needed to verify wins and uncertainty |
| Billing, power, and amortization | Needed for a cost comparison |
| Production follow-up | Needed for claims about deployment outcomes |
Suggested per-run record format for recovered logs
This is the schema we would run the next benchmark against. Take it.
{
"experiment_id": null,
"task_family": null,
"case_id": null,
"repeat_index": null,
"executed_at": null,
"model_label": null,
"model_digest": null,
"deployment": null,
"hardware_or_endpoint": null,
"quantization": null,
"prompt_hash": null,
"input_hash": null,
"temperature": null,
"seed": null,
"thinking_mode": null,
"context_limit": null,
"input_tokens": null,
"output_tokens": null,
"runtime_seconds": null,
"retry_count": null,
"completed": null,
"generated_candidates": null,
"accepted_candidates": null,
"validation_errors": null,
"output_hash": null,
"judge_ids": null,
"rubric_version": null,
"quality_scores": null,
"unsupported_claims": null,
"missed_major_facts": null,
"source_artifact": null
}
Null means unknown, not zero. This is a collection template, not an actual run record.
Publishing a benchmark with its own gaps listed is not modesty. It is the only version of it a technical buyer can use. If you want the same five dimensions run on one of your own workflows, start a conversation, or read how we build AI in production under AI solutions.
Internal benchmark, reported 1 September 2026. Figures are transcriptions of reported aggregates across six business-value analysis cases and twelve daily-summary cases. Evidence records S1, S2 and S3, reported 1 September 2026; S4 and S5 relate to a separate hardware experiment reported 27 August 2026 and are not used on this page.