No Bullshit AINo Bullshit AIno-bullshit-aiCannabis Courseby Distru00 XP
Plain English

How to Know It Still Works Next Month

You cannot tell whether a prompt change helped by reading a few answers. You need the same questions, scored the same way, before and after. Thirty of them is enough.Build a small set from work you have already done, score it the same way every time, and you can answer whether a change helped instead of guessing.Private eval construction: transcript sourcing, rubric design, held-out split, and reading a score difference honestly.

10 min read

Questions this answersQuestions this answersfaq

can we just use a published AI benchmark?can we just use a published AI benchmark?why not a public benchmark?
They measure someone else's work. None of them contain your customers, your product names, your state's rules or the way your team writes. A model can score brilliantly on a public test and still be wrong about your catalogue, and that is the only question you actually have.They measure someone else's work. None of them contain your customers, your product names, your state's rules or the way your team writes. A model can score brilliantly on a public test and still be wrong about your catalogue, and that is the only question you actually have.Public benchmarks measure general capability on tasks that are not yours, and the widely published ones are old enough to have leaked into training data. Neither problem applies to thirty cases drawn from your own transcripts, which is why a small private set beats a large public one for deciding whether to ship a change.
how many test cases do we need?how many test cases do we need?how large should the set be?
Thirty is a real answer. Fewer than about twenty and a couple of lucky results move the score enough to fool you. The bigger risk is not size, it is building the set out of easy cases because they were quicker to write down.Thirty is a real answer. Fewer than about twenty and a couple of lucky results move the score enough to fool you. The bigger risk is not size, it is building the set out of easy cases because they were quicker to write down.Thirty to fifty for a first set. Below roughly twenty, ordinary variance swamps the effect you are trying to see. Composition matters more than count: include the failures that prompted the work, keep the ratio of hard to easy honest, and hold back a portion you do not look at while tuning.

Read the rest for freeSign in to read the restauth required past this point

A name and an email, once. The whole course is free, nothing is sold to you, and your place is kept so you can pick it up on your phone later.A name and an email, once. The course is free; the account exists so your progress and your level on the dial follow you between devices, and so you can be issued a certificate at the end.name + email. free. account carries completion state, data-level and certificate issuance.

Sign in or sign upSign in or sign upsign in

We store your name, your email, which lessons you open and which you finish, and the reading level you use. Distru can see that. Nothing is charged and nothing is sold on.Stored: name, email, which lessons you open and complete, and your level on the dial. Distru staff can see it. No payment, no third-party advertising trackers.stored: name, email, per-lesson view + completion, data-level. visible to Distru. no payment, no ad-tech, no third-party pixels.

Adapted from Microsoft's Generative AI for Beginners (Lesson 14: The generative AI application lifecycle), MIT License. Rewritten for cannabis operations; not endorsed by Microsoft.