BIGMODEL?CHECK A TASK
PUBLIC EVAL · VERSION 0.1 + SUPPLEMENT · UPDATED 10 AUG 2026

INVOICE FIELD
EXTRACTION.

On this 12-record synthetic fixture, Luna matched Sol exactly at one twenty-fifth of the cost. A later DeepSeek V4 Flash supplement was even cheaper and faster—but returned only 6 of 12 records exactly.

STATUSRUNS COMPLETE · FOUR MODELS SCOREDLUNA IS THE HOSTED DEFAULT CANDIDATE
12/12LUNA · WHOLE RECORD EXACT
12/12SOL · WHOLE RECORD EXACT
11/12LOCAL 12B · WHOLE RECORD EXACT
6/12DEEPSEEK V4 FLASH · LATER SUPPLEMENT

CHEAPEST WAS FASTEST.
IT ALSO BROKE THE CONTRACT.

Luna and Sol returned all five fields correctly on all 12 records. Supplemental DeepSeek V4 Flash returned valid JSON every time, but only seven outputs used the exact five-field contract and only six records were fully correct.

ROUTEMODELEXACTFIELD CONTRACTMEDIANMEASURED COST
POST-RESULT SUPPLEMENTDEEPSEEK V4 FLASH6/127/120.28s$0.00046
FRONTIER HOSTEDGPT-5.6 SOL12/1212/121.58s$0.03070
LOCAL OPEN WEIGHTGEMMA4 12B · Q411/1212/123.79s$0 API*

LATER SUPPLEMENT: DeepSeek was locked and run after the original results were known. It used the same cases and scorer, one attempt per case, thinking disabled, and no retries. JSON-object mode produced 12/12 parseable JSON; “field contract” requires the exact five property names.

* The local route had no per-call API bill. Hardware, energy, maintenance, and concurrency were not measured. Its first request also incurred a 67.45-second cold load.

Rates locked 10 Aug 2026 per 1M tokens: Luna $0.20 input / $1.20 output; Sol $5 input / $30 output; DeepSeek V4 Flash $0.0028 cached input / $0.14 uncached input / $0.28 output. Luna source ↗ Sol source ↗ DeepSeek source ↗

ROUTING DECISION

START WITH LUNA.

It tied Sol without breaking the field contract. DeepSeek was 63% cheaper than Luna and much faster, but the six incorrect records make that saving a false economy under this prompt.

THE DEEPSEEK FAILURE

FIVE WRONG KEYS. ONE WRONG NAME.

Five outputs renamed total_amount to amount_due or amount_payable. One selected the Japanese legal name instead of the expected English trading name. The values themselves were otherwise correct.

THE LOCAL FAILURE

GEMMA4 FLIPPED A DATE.

Gemma4 read the spaced UK date 03/08/2026 as 8 March instead of 3 August. Supplier, invoice number, currency, and amount were correct.

MUSE GLIMMER
DID NOT RUN HERE.

The official 17 GB Muse Glimmer 30B build targets 24 GB of VRAM. This laptop has 8,151 MiB, and its installed Ollama 0.32.7 runtime predates the non-MLX release path.

That is exactly why “open weights” is not the same as “runs well on my laptop.” The model was excluded before download; a 16.76 GB transfer outside the vendor hardware target would not be a credible scored comparison.

Meta GGUF model card ↗Ollama release-status issue ↗

REVISIT CONDITION

Run Muse Glimmer when a supported runtime and at least the vendor-target 24 GB memory envelope are available. Until then, it stays a candidate—not a result.

CHEAP MODEL,
OR FALSE ECONOMY?

The workload is deliberately narrow: return supplier name, invoice number, issue date, currency, and total amount from OCR-like text as valid JSON.

The fixture excludes image rendering and OCR quality. All invoices are synthetic. Twelve cases are enough to test the harness and expose repeated DeepSeek contract failures plus a local-model date failure—not enough to establish production readiness.

NOT A PRODUCTION PASS

No operational error budget or pass threshold exists yet. The result identifies Luna as the next hosted candidate for this narrow fixture; it does not prove that Luna, Sol, DeepSeek, or the local model is reliable on production invoices.

REPRODUCE
THE FIXTURE.

The original package and later DeepSeek supplement retain every request, response, prediction, usage count, latency, and failure—not only the aggregate scores.

RESULTS.JSONThe comparison, recommendation, measured costs, latency, failure, and links to every run artifact.DOWNLOAD ↗PREREGISTRATION.JSONCandidates, prompt, schema, prices, hashes, hardware, metrics, and decision rule locked before the first scored call.DOWNLOAD ↗PROTOCOL.JSONThe locked scope, metrics, normalisation, and run requirements.DOWNLOAD ↗CASES.JSONLTwelve entirely synthetic OCR-like invoice records with gold outputs.DOWNLOAD ↗SCORER.MJSA deterministic scorer for record match, field match, schema validity, and missing cases.DOWNLOAD ↗DEEPSEEK SUPPLEMENT RESULTSThe later candidate's measured score, cost, latency, six failures, comparison boundary, and retained zero-spend infrastructure attempt.DOWNLOAD ↗DEEPSEEK SUPPLEMENT PREREGISTRATIONDeepSeek settings, current prices, hashes, no-retry rule, and $0.01 ceiling locked before the first provider call.DOWNLOAD ↗DEEPSEEK RAW JSONLAll 12 sanitised requests and responses, including exact outputs, usage, measured cost, and latency.DOWNLOAD ↗VERIFY DEEPSEEK RESULTA deterministic reconciliation of raw records, saved score, token cost, median latency, and reported failure groups.DOWNLOAD ↗LUNA RAW JSONLAll 12 sanitised requests, full API responses, parsed outputs, usage, cost, and latency.DOWNLOAD ↗SOL RAW JSONLAll 12 sanitised requests, full API responses, parsed outputs, usage, cost, and latency.DOWNLOAD ↗LOCAL RAW JSONLAll 12 local requests and responses, including the retained inv-010 date failure and hardware timing.DOWNLOAD ↗
PRIMARY

WHOLE-RECORD EXACT MATCH

All five normalised fields must match the gold record.

SECONDARY

FIELD + CONTRACT DIAGNOSTICS

Field exact match, exact-property validity, JSON syntax, and missing predictions expose how a model failed.

RUN LOG

MODEL, PRICE, LATENCY, HARDWARE

Candidate version, provider or runtime, quantisation, raw output, retry, token cost, and latency must travel with the result.

RUN LOCALLYnode score.mjs cases.jsonl predictions.jsonl

SPREAD THE
MEASURED RESULT.

This project grows when one useful result reaches another person choosing an AI model—not when a sponsor buys a verdict.