# cannabis-eval-starter

version 1.0 · 2026-09-17 · MIT · from Distru's No Bullshit AI Course (08-moonshots/03). Built by Sebastian and the Distru team.

A private benchmark you can run in five minutes. Ten example tasks with fake data, a scorer, and a sample outputs file. Replace the tasks with 50 of your own and rerun it every time a model, prompt or vendor changes.

## Files

- `tasks.csv`: `task_id, category, input, expected, match`. `match` is `exact`, `numeric` or `contains`. The ten rows are examples with made-up numbers; the point is the shape.
- `outputs.sample.csv`: `task_id, model_output, confidence`. What a model run produced. Three are wrong on purpose (one of them overconfident) so you can see the report.
- `score.py`: scores outputs against tasks. Standard library only.

## Run

```
python3 score.py tasks.csv outputs.sample.csv
```

You get accuracy overall, per category, a list of misses, and, because the sample has a `confidence` column, a five-bin calibration table with expected calibration error. Drop the column and the table disappears.

## Make it yours

1. Copy `tasks.csv`. Replace every row with a real task from your week and the answer a careful person would give. Fifty is enough to start. Redact tags, license numbers and names.
2. Run each task through the model or tool you are testing with the same instructions every time. Save `task_id, model_output` and, if the tool gives one, `confidence`.
3. `python3 score.py my-tasks.csv my-outputs.csv`. Write the score, the model name and the date in a log.
4. Rerun on every model change. The number moving is the only signal that matters.

## Reading the calibration table

Each row is a confidence range. `mean conf` is what the model claimed; `accuracy` is what happened. If the 0.8-1.0 row says accuracy 60 percent, the model is overconfident there and your threshold for acting without a human should move up. Ten rows is too few to trust the table; it firms up around a few hundred.
