# Build vs. buy: do you need your own AI model?

version: 1.0 · 2026-09-17 · MIT · from Distru's No Bullshit AI Course (08-moonshots/02). Built by Sebastian and the Distru team.

What it is: a decision checklist. Answer the questions in order. The first "stop" you hit is your answer. Most operators stop at question 2 or 3.
What to bring: one task you want the AI to do, described in one sentence, and ten real examples of it.
What comes out: one of five answers: better prompt, skill file, retrieval, small decision model, fine-tune.

## 0. Before anything

- [ ] Write the task in one sentence. "Sort incoming support emails into six buckets." "Turn buyer texts into order lines." If you cannot, stop. You do not have a task, you have a wish.
- [ ] Collect ten real examples with the right answer written next to each. Ugly ones included.
- [ ] Decide how you will know it worked. A number: "9 of 10 right", "under 2 wrong per 100".

## 1. Is the answer already in the model?

- [ ] Paste one example into a current frontier model with clear instructions. Does it get it right?
- [ ] Try all ten. Count.
- [ ] 8 or more right → **stop. Better prompt.** Write the instructions down and reuse them. Cost: an hour.

## 2. Does it fail because it does not know your rules?

- [ ] Look at the misses. Is the cause "it does not know our SOP / price sheet / state rule"?
- [ ] Paste the rule in. Rerun the misses.
- [ ] Fixed → **stop. Skill file.** Put the rules and the steps in a SKILL.md or a standing system prompt. Cost: an afternoon.

## 3. Does it fail because the facts change or there are too many to paste?

- [ ] Is the missing information a document set (SOPs, past tickets, COAs, price history) that changes weekly?
- [ ] Would the right page, pasted in, fix the miss?
- [ ] Yes → **stop. Retrieval.** Index the documents, fetch the relevant ones per question, let the model read them. Cost: days to set up, ongoing upkeep of the index. Fine-tuning does not add facts reliably (Ovadia et al. 2024, https://arxiv.org/abs/2312.05934); retrieval does.

## 4. Is the output a label, score or yes/no rather than text?

- [ ] Every answer is one of a fixed set (bucket, priority, which rep, approve/hold)?
- [ ] You have hundreds to thousands of past examples with the label already attached (closed tickets, past orders, past decisions)?
- [ ] Wrong answers are cheap to catch and reverse?
- [ ] All three yes → **stop. Small decision model.** A calibrated decision model or a classic classifier trained on your labelled history. Cheap to run, fast, testable. Needs a threshold and a monthly calibration check.

## 5. Only now: fine-tuning

- [ ] The failure is style, format or a consistent instruction-following defect, not missing facts.
- [ ] You have at least 50 to 100 clean examples for a first pass (OpenAI: minimum 10, start with 50; https://developers.openai.com/api/docs/guides/supervised-fine-tuning), and a held-out eval set the model never trains on.
- [ ] Volume is high enough that a shorter prompt saves real money.
- [ ] Someone owns re-running the eval every time the base model or your data changes.
- [ ] All four yes → fine-tune a small model. Otherwise go back to 1.

## Never

- Train "a cannabis LLM" from scratch. Nobody in this industry has the data or the budget, and a frontier model with your documents beats it.
- Fine-tune to teach the model your prices, your inventory or this month's state rule. Those change; weights do not.
- Skip the eval. Without a test set with known answers you cannot tell whether any of this made it better.
