AI Basics, Minus the Hype
Two kinds of AI, side by side. The kind that writes (ChatGPT, Claude) and the newer kind that only decides: pick one, rate it, yes or no. What each gets wrong, what you can safely paste in, and a first real task with each.Two kinds of model: LLMs that generate text, and System One decision models that return typed answers with calibrated probabilities. Failure modes, data-retention, and a first useful task for each.Tokens, context windows, sampling and failure modes for LLMs; state, typed questions, calibrated probabilities and confidence for System One decision models; data-retention policies; a zero-to-useful setup for an operator's first workflow with each.
Carry on withCarry on withnext:
The AI that writes (LLMs)Part A · LLMs: the AI that writesA · LLMs
- 017 min
What an AI Chatbot Actually IsWhat a language model (LLM) actually isLLMs: tokens, context windows, and why they make things upDoneDonedone
It is a very good guesser of the next word. That one fact explains almost everything it does well and badly.A language model guesses the next word-chunk (token), over and over. It only sees what is in the chat (its context window), which is why it drafts well and invents facts.Autoregressive next-token prediction, the context window as the only state, and why fluency is not accuracy.
- 027 min
Where AI Gets It Wrong (and How to Catch It)Where AI Gets It Wrong: the Failure Modes, and the Five-Second Check for EachFailure modes: hallucination, arithmetic, dates, stale knowledge, context lossDoneDonedone
AI will confidently invent a Metrc rule, a tag number, or a total. Here is the short list of things to never trust without checking, and the five-second check for each.AI will confidently invent a Metrc rule, a tag number or a total; engineers call this hallucination. Here is the short list of things never to trust without checking, the cheap check that catches each one, and why long chats make it worse.A taxonomy of LLM failure modes (sampling, uncalibrated confidence, tokenized arithmetic, no clock, training cutoff, context saturation) with the deterministic validator that catches each.
- 037 min
What You Can Paste Into AI (and What You Shouldn't)What You Can Paste Into AI: Data Retention, PII, and Credentials in Plain TermsData handling: retention tiers, PII redaction, and credentials behind a tool boundaryDoneDonedone
A plain rule set for what is fine to paste (a price list), what to black out first (a buyer's email) and what never goes in (a customer's ID, your Metrc password).Three piles for anything you paste: fine (a price list), redact first (a buyer's email, PII), never (a customer's ID, your Metrc password). Plus how retention and training policies differ by tier, so you can check yours in two minutes.Consumer vs. business vs. API data policies, a two-minute policy check, deterministic PII redaction with a reversible map, and credentials pushed behind an MCP/tool boundary.
- 048 min
Your First Real Task: 30 Minutes, One ProblemYour First Real Task: One Chore, One Project Workspace, 30 MinutesZero-to-useful: one project, one system prompt, one timed workflowDoneDonedone
Pick one annoying weekly chore, set up a workspace once, do the chore with AI three times with a stopwatch running, and keep it only if it is faster. That is the whole lesson.Pick one weekly chore, set up a project with standing instructions (a system prompt) and a few reference files once, run the chore three times against the clock, and keep it only if it beats doing it by hand.Set up a project with a persistent system prompt and attached reference context, run one repeatable workflow three times, time it end to end including review, and keep or drop on the numbers.
- 056 min
Where the AI Actually Runs (Same Brain, Different Keys)Where the AI Actually Runs: Chat, Desktop, Terminal, Server (Same Model, Bigger Permissions)Chat, desktop, terminal, server: one model, four permission setsDoneDonedone
The chat window, the desktop app, the terminal tool and the always-on server agent are the same AI with a bigger and bigger set of keys. Pick the smallest one that fits the job.Chat, desktop sandbox, terminal agent and self-hosted harness run the same model; what changes is which files it can reach and whether it needs you there. Four questions pick the smallest surface that fits the job.Surfaces differ by filesystem access and autonomy, not by model. A four-question decision tree, with Claude's Chat, Cowork, Code and a self-hosted harness as the worked example.
The AI that only decides (decision models)Part B · Decision models: pick, rate, yes/noB · System One decision models
- 067 min
The AI That Only DecidesThe AI that only decides: decision models (System One)System One decision models: typed questions in, calibrated answers outDoneDonedone
Not every AI writes. This kind reads what you give it and answers only with a pick, a rating or a yes/no, plus how sure it is. That is a different tool for different jobs.A decision model reads the same text an LLM would, but returns typed answers (Choice, Score, Noul) with probabilities instead of prose. What it is, what it cannot do, and what it costs.Jev (TypeSafe): state plus typed questions over POST /v1/systemone; choice/score/noul with probability distributions; no generation, text-only, $0.042 per Mtok input.
- 076 min
How Sure Is It? Reading the NumberHow sure is it? Probabilities, confidence and the three bandsConfidence-gated routing: probabilities, calibration, thresholds that scale with riskDoneDonedone
A decision model tells you how sure it is. That number decides whether the software acts on its own, asks someone, or hands the whole thing to a person.Probability per option, one confidence number per answer, and why calibrated matters. Three bands (act, confirm, escalate) with cut-offs that get stricter as the cost of a mistake goes up.Probability distributions vs. the derived confidence statistic, what calibration guarantees (in aggregate), three-band routing, per-action thresholds, and speculative fan-out.