No —1 start small.
FOR: Extract fields from invoices
“The expensive model is mostly waiting for the schema.”
Stop paying frontier prices for routine work — and stop cheaping out when failure is expensive. Pick a workload. Get a blunt verdict, the cost maths, and a permanent result you can send to the person choosing your AI stack.
FOR: Extract fields from invoices
“The expensive model is mostly waiting for the schema.”
Em dash set by hand. Human generated, not AI generated.
There are 74 editorial classifications and one completed workload evaluation. Invoice extraction has a locked protocol, raw outputs, deterministic scores, measured cost, and a documented hardware gate. The other verdicts remain guidance for what to test.
1 completed eval. Four candidates scored on one 12-case synthetic invoice fixture; DeepSeek V4 Flash was added as a separately preregistered supplement. No result is described as production proof.
Invoice field extraction. Constrained, cheap to run, and objectively scoreable. The first package separates OCR quality from field extraction and retains every failure.
LUNA 12/12 · SOL 12/12 · LOCAL 12B 11/12 · DEEPSEEK 6/12. Luna tied Sol at one twenty-fifth of the measured API cost. Supplemental DeepSeek was cheaper and faster but broke the field contract on five records and chose the wrong supplier-name variant on one.
The recommender sets the capability tier, filters candidates by deployment and hardware, then ranks only models that pass the workload eval.
Shortlist hosted models in the required capability tier, then choose the cheapest one that clears the workload evaluation and latency target.
Meta targets 24GB and 32GB hardware; the remaining memory must cover KV cache, the perception encoder and speculative drafter. Fit still does not establish quality on your workload.
META HARDWARE NOTES ↗K3 ships with MXFP4 weights and MXFP8 activations after quantization-aware training. The artifact is cluster or hosted class for almost everyone.
READ THE KIMI K3 LICENSE ↗Published artifact size, licence and memory fit are filters — not proof of useful speed or task quality. New releases remain candidates until a reproducible workload eval and hardware probe say otherwise.
Replace the assumptions with your real traffic and token usage. We show the arithmetic — no invented savings counter.
Estimate only. Cached reads bill at 10% of standard input; batch halves rates. Excludes tools, retries, reasoning-token behaviour, hosting and the cost of errors.
Editorial starting points for eval design — not universal claims about every prompt, dataset or model release.
Does the task require a long chain of dependent decisions — or mainly classification, extraction and transformation?
Are the goal, inputs and acceptable answers clear? Bigger models help more when several interpretations remain live.
Can you detect and retry a failure cheaply, or could one plausible mistake create material harm?
Does success depend on connecting evidence across many sources, tools, steps or modalities?
Use the cheapest setup that clears a task-specific evaluation at the reliability you need. Small model still means prompts, schemas, tests, fallbacks and escalation.
The frontier model is a tool.
Not a default.
The utility is free. If it earns a meaningful audience of people choosing AI models and infrastructure, clearly labelled sponsor placements may fund new evaluations. Sponsors will never buy rankings, scores or recommendations.
Know someone paying frontier prices by default? Send them the permanent result for Extract fields from invoices.