BIGMODEL?Model-sizing utility
AN INDEPENDENT MODEL-SIZING UTILITY

Do you really need the big model?

Stop paying frontier prices for routine work — and stop cheaping out when failure is expensive. Pick a workload. Get a blunt verdict, the cost maths, and a permanent result you can send to the person choosing your AI stack.

OR TRY
74SHAREABLE WORKLOADS
4DECISION FACTORS
1REPRODUCIBLE EVAL
0PAID RANKINGS
02 / THE VERDICTExtraction · invoice-fields
Small-model job · hostedBIGMODEL? / INVOICE-FIELDS

No —1 start small.

FOR: Extract fields from invoices

The expensive model is mostly waiting for the schema.

RECOMMENDED SETUPSmall, fast model + schema + deterministic checks
Permanent link · /tasks/invoice-fields
1

Em dash set by hand. Human generated, not AI generated.

00 / EVIDENCE LEDGER

The first result is now live.

There are 74 editorial classifications and one completed workload evaluation. Invoice extraction has a locked protocol, raw outputs, deterministic scores, measured cost, and a documented hardware gate. The other verdicts remain guidance for what to test.

CURRENT VERIFIED RESULT COUNT

1 completed eval. Four candidates scored on one 12-case synthetic invoice fixture; DeepSeek V4 Flash was added as a separately preregistered supplement. No result is described as production proof.

FIRST WORKLOAD

Invoice field extraction. Constrained, cheap to run, and objectively scoreable. The first package separates OCR quality from field extraction and retains every failure.

PACKAGE STATUS

LUNA 12/12 · SOL 12/12 · LOCAL 12B 11/12 · DEEPSEEK 6/12. Luna tied Sol at one twenty-fifth of the measured API cost. Supplemental DeepSeek was cheaper and faster but broke the field contract on five records and chose the wrong supplier-name variant on one.

02B / PICK THE DEPLOYMENT ROUTE

Open weights doesn’t mean laptop-local.

The recommender sets the capability tier, filters candidates by deployment and hardware, then ranks only models that pass the workload eval.

1 · TASK TIER2 · DEPLOYMENT FIT3 · EVAL GATE4 · LOWEST TOTAL COST
WHERE SHOULD IT RUN?
CURRENT RECOMMENDATION

SMALL-MODEL JOB · HOSTED

Shortlist hosted models in the required capability tier, then choose the cheapest one that clears the workload evaluation and latency target.

TASK PASS RATEP95 LATENCYTOTAL COSTLICENCEHARDWARE FIT
MODEL ADDED · 10 AUG 2026LAPTOP / WORKSTATION CLASS

MUSE GLIMMER 30B

29.6B PARAMETERS<20GB QUANTIZED MODELAPACHE-2.0 WEIGHTS

Meta targets 24GB and 32GB hardware; the remaining memory must cover KV cache, the perception encoder and speculative drafter. Fit still does not establish quality on your workload.

META HARDWARE NOTES ↗
FRONTIER OPEN WEIGHTSCUSTOM KIMI K3 LICENSE

KIMI K3

2.8T TOTAL PARAMETERS104B ACTIVE / TOKEN1.561TB 96 SHARDS

K3 ships with MXFP4 weights and MXFP8 activations after quantization-aware training. The artifact is cluster or hosted class for almost everyone.

READ THE KIMI K3 LICENSE ↗

Published artifact size, licence and memory fit are filters — not proof of useful speed or task quality. New releases remain candidates until a reproducible workload eval and hardware probe say otherwise.

03 / RUN YOUR NUMBERS

Prestige isn’t a pricing strategy.

Replace the assumptions with your real traffic and token usage. We show the arithmetic — no invented savings counter.

PRICE ANCHOR · EFFECTIVE 30 JUL 2026
LUNA $0.20 / $1.20 → SOL $5 / $30
OpenAI price source ↗
MODEL TIERINPUT / 1MOUTPUT / 1M
SMALL / FAST
FRONTIER
ALL SMALL$54/ MONTH
ALL FRONTIER$1,350/ MONTH
ESTIMATED MONTHLY SAVING$1,29696% VS ALL-FRONTIER

Estimate only. Cached reads bill at 10% of standard input; batch halves rates. Excludes tools, retries, reasoning-token behaviour, hosting and the cost of errors.

04 / THE METHOD

Four questions before you upgrade.

Editorial starting points for eval design — not universal claims about every prompt, dataset or model release.

01

REASONING

Does the task require a long chain of dependent decisions — or mainly classification, extraction and transformation?

02

AMBIGUITY

Are the goal, inputs and acceptable answers clear? Bigger models help more when several interpretations remain live.

03

ERROR COST

Can you detect and retry a failure cheaply, or could one plausible mistake create material harm?

04

CONTEXT

Does success depend on connecting evidence across many sources, tools, steps or modalities?

THE RULE THAT MATTERS

Use the cheapest setup that clears a task-specific evaluation at the reliability you need. Small model still means prompts, schemas, tests, fallbacks and escalation.

05 / BROWSE THE CATALOGUE

74 workloads. No leaderboard theatre.

Coding

12 TASKS

Extraction

10 TASKS

Writing

10 TASKS

Research

8 TASKS

Analysis

10 TASKS

Support

8 TASKS

Vision

8 TASKS

Agents

8 TASKS

The frontier model is a tool.
Not a default.