Home  ›  Insights  ›  Local versus cloud model benchmark
AI & ML

Local versus cloud model benchmark: the complete figures

A 27B model running locally against a cloud model recorded at 284B, over two jobs, with opposite rankings. Every reported aggregate, every derivation, the open CSV, and a complete list of what we could not verify.

The finding is not that small models win. It is that the ranking changed when the job changed. On business-value analysis, a locally run Qwen3.8 27B beat cloud-hosted DeepSeek-V4-Flash, recorded at 284B, on all six cases. On daily summaries the cloud model won. Both results are published here at the same size.

This page is the evidence, not the argument. The narrative version is a separate post. What follows is the reported performance and repeatability data, the business-value quality data, the daily-summary quality data, the derivation arithmetic, the claim boundaries, and the complete list of experiment artifacts we do not have.

Read this before the tables. This is not a rerun. Every figure below is a transcription of reported aggregate results from three contemporaneous experiment records, not a result recomputed from individual runs. The per-run logs, judge identities, scoring rubric, ensemble weights and exact-match rule were not recovered. Six business cases and twelve summary cases, with no per-case scores, means no variance, no confidence interval and no inferential test. We make no significance claim. Section 06 lists every missing field.

Two separate experiments, and they are not merged anywhere on this page. The model comparison here was reported on 1 September 2026. A local-hardware feasibility run on a different machine was reported on 27 August 2026. No hardware figure from the August run appears on this page, and no claim here draws on it.

01Performance and repeatability

Source: S1. Reconstructed from flattened screenshot text and its accompanying narrative.

MetricQwen3.8 LocalDeepSeek V4 CloudReading
Business runs successful18/1818/18Both completed all reported business runs
Mean business runtime347.8 s17.6 sCloud substantially faster
Accepted candidates186/186136/144Local accepted every reported candidate
Validation errorsNarrative reports noneTable text reports “8 duplicates”Exact error taxonomy needs original screenshot/logs
Exact business repeatability6/6 cases0/6 casesLocal repeated exactly across all tested business cases
Mean summary runtime32.6 s3.4 sCloud substantially faster
Required section coverage100%100%Both met reported section coverage
Numeric grounding precision94.0%93.1%Local ahead by 0.9 percentage points
Exact summary repeatability12/12 cases0/12 casesLocal repeated exactly across all tested summary cases

Details that matter

Completion and acceptance are different. DeepSeek completed 18/18 runs even though only 136/144 candidates were accepted. A completed request does not mean every candidate passed validation.

Candidate count is not unique business value. These may include repeated outputs across repeated cases. Do not describe 186 accepted candidates as 186 unique discoveries or customer outcomes.

Repeatability is scoped. The narrative says Qwen was “completely deterministic”. The narrower supported statement is exact repeatability in the reported cases. The metric’s comparison rule, raw string, normalised output, structured fields or semantic equivalence, is not supplied. No universal claim about all Qwen outputs follows from this test.

Grounding precision is a separate measure. The 94.0% numeric grounding precision is not the same as the factual-grounding score in either quality table. Their denominators and scoring procedures are not documented here.

Validation-row ambiguity is preserved. The extracted table includes the header “Fatal validation errors”, omits a clear local numeric cell, and includes “8 duplicates” for DeepSeek. The prose independently says Qwen produced no validator errors. That is the wording we use; we have not silently converted the table into a fully verified error taxonomy.

02Business-value quality

Source: S2. Score units and scale endpoints were not supplied. These are reported score points, without a percent sign and without assuming a 100-point rubric.

Quality metricQwen3.8 LocalDeepSeek V4 CloudQwen minus DeepSeekInterpretation
Ensemble overall quality77.463.9+13.5Local higher
Factual grounding81.380.6+0.7Small local advantage
Material coverage72.449.1+23.3Largest score gap in local’s favour
Prioritisation/usefulness77.356.4+20.9Substantial local advantage
Calibration76.471.5+4.9Local higher
Clarity/concision76.776.5+0.2Nearly equal reported scores
Unsupported claims/output1.780.69+1.09Cloud better; lower is preferable
Missed major facts/output3.8910.54−6.65Local better; lower is preferable
Case-level resultReported result
Qwen wins6
DeepSeek wins0
Ties0
Judge direction agreement5/6 cases

The local result was not principally a prose-polish win. Clarity/concision differed by only 0.2 points. The striking gaps were material coverage and prioritisation/usefulness. That supports a statement about preserving decision-relevant information, and it supports no causal claim about why Qwen performed better.

The missed-major-facts measure gives the result an accessible business question: what did the report leave out? Local outputs reportedly missed 3.89 major facts on average compared with 10.54 for cloud outputs.

The local model also generated more unsupported claims per output. The larger model’s lower count is a real advantage and it belongs in the main result, not in a footnote. These are counts per output, not normalised rates. Output lengths and total factual-claim counts are unavailable.

Five-of-six judge direction agreement is not unanimous agreement across all cases. The source does not identify the judges, the ensemble rule, the blinding procedure, or the tie rules.

03Daily-summary quality

Source: S3. Again, values are reported score points, not percentages unless explicitly stated elsewhere.

Quality metricQwen3.8 LocalDeepSeek V4 CloudQwen minus DeepSeekInterpretation
Ensemble overall quality77.380.0−2.7Cloud higher
Factual grounding76.681.5−4.9Cloud higher
Material coverage79.879.5+0.3Nearly equal
Prioritisation/usefulness78.379.9−1.6Cloud higher
Calibration74.378.9−4.6Cloud higher
Clarity/concision79.682.5−2.9Cloud higher
Unsupported claims/output1.440.92+0.52Cloud better
Missed major facts/output0.921.25−0.33Local better
Case-level resultReported result
Qwen wins5
DeepSeek wins6
Ties1
Judge direction agreement3/12 cases

The business-analysis win does not transfer to every task. Daily-summary results show a small cloud aggregate lead, a split case result, and low reported judge direction agreement. This is a mixed outcome with evaluator uncertainty, not a decisive universal ranking.

Exact repeatability and quality are distinct. Qwen repeated its summary outputs exactly in 12/12 cases, and DeepSeek still had the higher aggregate summary quality score. A repeatable output can be a consistently weaker one.

The two task families are not combined into a single overall score anywhere on this page. Their scoring weights, use-case importance, and aggregation rules are unknown.

04Calculations and reusable data

All calculations use rounded values from the source. The extra decimal places below explain the arithmetic. Copy should retain sensible rounding.

Derived quantityFormulaResultSuggested wording
Reported total-parameter ratio284 / 2710.52×Approximately 10.5× the reported total parameters
Business runtime ratio347.8 / 17.619.76×Cloud was about 19.8× faster by mean elapsed runtime
Business time difference347.8 − 17.6330.2 sAbout 5 minutes 30 seconds less per reported mean run
Business relative elapsed-time reduction(347.8 − 17.6) / 347.894.94%Cloud mean elapsed time was about 94.9% lower
Summary runtime ratio32.6 / 3.49.59×About 9.6× from displayed means; source narrative says 9.5×
Summary time difference32.6 − 3.429.2 sAbout 29 seconds less per reported mean summary
Local candidate acceptance186 / 186100%All reported local candidates accepted
Cloud candidate acceptance136 / 14494.44%About 94.4% of reported cloud candidates accepted
Acceptance-rate gap100 − 94.44445.56 ppLocal ahead by about 5.6 percentage points
Accepted-candidate count difference186 − 1365050 more accepted candidates across the reported runs
Relative accepted-candidate difference(186 − 136) / 13636.76%About 36.8% more accepted candidates; not a recall metric
Business overall score gap77.4 − 63.913.5 pointsLocal scored 13.5 points higher
Business material-coverage gap72.4 − 49.123.3 pointsLocal scored 23.3 points higher on coverage
Business usefulness gap77.3 − 56.420.9 pointsLocal scored 20.9 points higher on usefulness
Reduction in missed major facts, business(10.54 − 3.89) / 10.5463.09%About 63.1% fewer missed major facts per output
Business unsupported-claim difference1.78 − 0.691.09/outputLocal had 1.09 more unsupported claims per output
Summary overall score gap, cloud lead80.0 − 77.32.7 pointsCloud scored 2.7 points higher
Numeric grounding precision gap94.0 − 93.10.9 ppLocal ahead by 0.9 percentage points
Business judge direction agreement5 / 683.33%Prefer the original 5/6 cases
Summary judge direction agreement3 / 1225%Prefer the original 3/12 cases

The aggregates as an open CSV

The same figures, machine readable. The CSV excludes the ambiguous validation row. Ties and judge agreement are experiment-level values rather than separate model values: business ties 0 and agreement 5/6; summary ties 1 and agreement 3/12.

Download benchmark-aggregates.csv

task,metric,qwen_local,deepseek_cloud,unit,source
business,successful_runs,18,18,count,S1
business,attempted_runs,18,18,count,S1
business,mean_runtime,347.8,17.6,seconds,S1
business,accepted_candidates,186,136,count,S1
business,generated_candidates,186,144,count,S1
business,exact_repeatability_cases,6,0,count,S1
business,repeatability_cases_total,6,6,count,S1
business,ensemble_overall_quality,77.4,63.9,reported_score,S2
business,factual_grounding,81.3,80.6,reported_score,S2
business,material_coverage,72.4,49.1,reported_score,S2
business,prioritization_usefulness,77.3,56.4,reported_score,S2
business,calibration,76.4,71.5,reported_score,S2
business,clarity_concision,76.7,76.5,reported_score,S2
business,unsupported_claims_per_output,1.78,0.69,count_per_output,S2
business,missed_major_facts_per_output,3.89,10.54,count_per_output,S2
business,case_wins,6,0,count,S2
summary,mean_runtime,32.6,3.4,seconds,S1
summary,required_section_coverage,100,100,percent,S1
summary,numeric_grounding_precision,94.0,93.1,percent,S1
summary,exact_repeatability_cases,12,0,count,S1
summary,repeatability_cases_total,12,12,count,S1
summary,ensemble_overall_quality,77.3,80.0,reported_score,S3
summary,factual_grounding,76.6,81.5,reported_score,S3
summary,material_coverage,79.8,79.5,reported_score,S3
summary,prioritization_usefulness,78.3,79.9,reported_score,S3
summary,calibration,74.3,78.9,reported_score,S3
summary,clarity_concision,79.6,82.5,reported_score,S3
summary,unsupported_claims_per_output,1.44,0.92,count_per_output,S3
summary,missed_major_facts_per_output,0.92,1.25,count_per_output,S3
summary,case_wins,5,6,count,S3

05Claim bank and boundaries

The left column includes claims we do not make. They are published so that nobody has to guess where the boundary sits, and so that a reader can check any sentence we publish elsewhere against this table.

Proposed claimStatusBasis / required context
“Our smaller local model won all six business-analysis cases.”Supported as reportedS2; name both models and retain six-case scope
“27B local versus 284B cloud.”Source-recorded labelsS1; verify manifests for independently certified specifications
“13.5 points higher overall business-analysis quality.”Supported arithmeticS2; reported score points, not percentage points
“About 63% fewer missed major facts per business output.”Supported arithmeticS2; aggregate means, not a claim about every output
“100% exact repeatability in the tested local cases.”Supported with scopeS1; include 6 business and 12 summary cases
“Cloud was about 20× faster on business analysis.”Supported rounded comparisonS1; elapsed runtime under the tested setups
“The local model was more accurate in every way.”ContradictedCloud had fewer unsupported claims and stronger summary scores
“The local model hallucinated less.”Not supportedReported unsupported-claim counts favour cloud
“No validator errors means no incorrect claims.”Invalid inferenceValidation and factual support are different checks
“186 unique business insights.”Not establishedCandidate uniqueness across runs is unknown
“63% less business risk.”Not measuredMissed facts are not a business-risk metric
“A 10× smaller model uses 10× less compute.”Not establishedTotal parameters do not establish active compute or cost
“The local setup saved a measured amount of money.”Not measuredNo energy, hardware amortisation, or billing comparison
“The experiment proves local AI is always better.”Contradicted by scopeDifferent task outcomes and a small case set
“These results are statistically significant.”Not establishedNo per-case variance, inferential test, or confidence interval
“The cloud test exercised a million-token context.”Not establishedA reported capability label is not actual input length
“We ran DeepSeek locally at ~3.5 tokens/second on the existing PC.”Separately reportedS4/S5; do not merge with the September cloud experiment

The result in one paragraph

We compared a locally run Qwen3.8 27B model with cloud-hosted DeepSeek-V4-Flash, recorded as 284B, on business-value analysis and daily summaries. Qwen won all six business-analysis cases and scored 77.4 versus 63.9 overall, with stronger material coverage and fewer missed major facts. DeepSeek was about 19.8× faster on business analysis, produced fewer unsupported claims, and scored higher overall on summaries. Our result supports evaluating models against the specific workflow and its trade-offs.

06Missing experiment artifacts and reproducibility fields

The following gaps limit independent verification. They do not erase the recorded result, but they define what this evidence note can and cannot substantiate.

Missing fieldWhy it matters
Exact model artifact/tag/digestDistinguishes labels from reproducible model identity
Quantization and precisionCan materially change local quality and performance
Local September hardwareNeeded to interpret local runtime
Cloud provider, endpoint, and serving configurationNeeded to interpret the cloud setup
Prompt templates and input packetsNeeded to verify input parity
Token counts and output budgetsNeeded to distinguish content length from throughput
Sampling settings, seeds, and thinking settingsNeeded to interpret repeatability and quality
Retry policy and timeout handlingNeeded to interpret successful-run counts
Case selection and case IDsNeeded to assess representativeness and cherry-picking
Per-run runtime valuesNeeded for variability, medians, and uncertainty
Exact-match ruleNeeded to reproduce repeatability
Candidate validation rulesNeeded to interpret duplicates and acceptance
Grounding definitions and denominatorsNeeded to interpret precision and factual scores
Judge models/people, prompts, and blindingNeeded to understand evaluation bias
Ensemble weights and score scaleNeeded to reproduce overall quality
Per-case scoresNeeded to verify wins and uncertainty
Billing, power, and amortizationNeeded for a cost comparison
Production follow-upNeeded for claims about deployment outcomes

Suggested per-run record format for recovered logs

This is the schema we would run the next benchmark against. Take it.

{
  "experiment_id": null,
  "task_family": null,
  "case_id": null,
  "repeat_index": null,
  "executed_at": null,
  "model_label": null,
  "model_digest": null,
  "deployment": null,
  "hardware_or_endpoint": null,
  "quantization": null,
  "prompt_hash": null,
  "input_hash": null,
  "temperature": null,
  "seed": null,
  "thinking_mode": null,
  "context_limit": null,
  "input_tokens": null,
  "output_tokens": null,
  "runtime_seconds": null,
  "retry_count": null,
  "completed": null,
  "generated_candidates": null,
  "accepted_candidates": null,
  "validation_errors": null,
  "output_hash": null,
  "judge_ids": null,
  "rubric_version": null,
  "quality_scores": null,
  "unsupported_claims": null,
  "missed_major_facts": null,
  "source_artifact": null
}

Null means unknown, not zero. This is a collection template, not an actual run record.

Publishing a benchmark with its own gaps listed is not modesty. It is the only version of it a technical buyer can use. If you want the same five dimensions run on one of your own workflows, start a conversation, or read how we build AI in production under AI solutions.

Internal benchmark, reported 1 September 2026. Figures are transcriptions of reported aggregates across six business-value analysis cases and twelve daily-summary cases. Evidence records S1, S2 and S3, reported 1 September 2026; S4 and S5 relate to a separate hardware experiment reported 27 August 2026 and are not used on this page.