# One-Page Model Picker

From Distru's No Bullshit AI Course, module 02 (lessons 01 and 02, and module 01 lesson 06 for decision models). Prices and names change; the tiers do not. Re-check prices on the vendor's page before budgeting. Last verified: September 2026.

## The four tiers

| Tier | Nickname | What it is for | Anthropic example (price per 1M tokens in / out) | Others in this tier |
|---|---|---|---|---|
| Frontier | The expert | Hard, one-off reasoning. Long regulations. Multi-step agent work. Errors are expensive. | Claude Opus 5 · `claude-opus-5` · $5 / $25 · 1M context | OpenAI's top GPT model; Google Gemini Pro |
| Mid | The workhorse | Daily driver. Extraction, normalization, drafting, code generation. | Claude Sonnet 5 · `claude-sonnet-5` · $2 / $10 · 1M context | OpenAI's mid GPT model; Gemini Flash (upper) |
| Small | The intern | High-volume, simple, repeated. Classification, templated generation, routing. | Claude Haiku 4.5 · `claude-haiku-4-5` · $1 / $5 · 200K context | OpenAI "mini"/"nano" models; Gemini Flash-Lite |
| Local | The closet | Data that cannot leave the building. Very high volume, low stakes. | n/a | Llama, Qwen, Gemma, Mistral via Ollama or LM Studio |

Output tokens cost ~5x input. Reading is cheap; writing is where tier matters.

## The fifth option: a model that never writes

If the answer your software needs is a label, a score, or a yes/no, and nobody reads prose, you may not want a text model at all. TypeSafe's Jev is a "System One" decision model: send the text (the state) plus typed questions (Choice / Score / Noul) and get typed answers with calibrated probabilities and a confidence number. Per TypeSafe's docs (September 2026): text-only input, $0.042 per million input tokens, output free, `POST /v1/systemone`, `jev-latest`. One vendor and one term so far; calibration holds across many answers, not any single one; test on your own data and set thresholds per action.

| Cannabis task | Use | Note |
|---|---|---|
| Route a buyer text to the right rep | Decision model (Choice) | Below your confidence threshold, a person picks |
| Flag a health claim in menu copy | Decision model (Noul) | A guardrail on LLM output; state ad rules set the criteria |
| Score urgency of a vendor email | Decision model (Score) | Describe situations for each level, not degrees |

## Task map

| Cannabis task | Tier | Note |
|---|---|---|
| Read a state bulletin, tell me what changed for us | Frontier | Once. Quote the source text. |
| Buyer spreadsheet → my SKUs | Mid | Ask for UNKNOWN on low confidence, never a guess |
| Collections / vendor / restock emails | Mid | Paste one email you like as the voice |
| Explain a Metrc error or mismatch | Mid | Give it the row, not the whole export |
| Sort 2,000 expenses into COGS / not (first pass) | Small | Batch it. Accountant reviews. |
| 300 menu descriptions from spec sheets | Small | Spot-check 10% with Mid |
| Anything with customer names, IDs, PII | Local, or redact first | |
| Totals, counts, averages | **No AI** | Spreadsheet formula. Or have AI write the formula. |
| Agent that takes multi-step actions in your systems | Frontier, rarely | Only if 02-02's four criteria are all yes |

Rule: start one tier lower than you think. Move up only when your own test says so.

## The 60-minute bake-off

1. Pick one task you do weekly.
2. Collect 10 real inputs. Include 2 ugly ones.
3. Write the correct answer for each, by hand.
4. Run all 10 through 2-3 models. Same prompt.
5. Score: right / wrong / dangerous-wrong (e.g., invented a SKU that exists).
6. Winner = most right with zero dangerous-wrong. Tie → cheaper one.
7. Write down: model, date, score. Redo quarterly.

```
eval/<task>/
  001.input.txt   001.expected.txt
  ...
  results.csv     # date, model, prompt_hash, correct, dangerous, cost, latency
```

## Cost levers, in order

1. Smaller tier for the simple parts of a pipeline.
2. Prompt caching for repeated context (system prompt, catalog, rules): cached input is a fraction of full price.
3. Batch API for anything not urgent: typically ~50% off.
4. Lower "thinking" / "effort" for transforms; raise it for reasoning.
5. Local model only when data residency or volume forces it; count your own maintenance time as a cost.

License: CC0.
