No Bullshit AINo Bullshit AIno-bullshit-aiCannabis Courseby Distru00 XP
Plain English

Building It Yourself

For whoever writes the code. What a request actually is, how a tool loop works, how to stop it costing more than it should, and how to know it still works next month. Read this at level 1 if you want to understand what your tech person is doing.The developer track. One request end to end, the tool-call loop and where the write gate belongs, caching as an architectural choice rather than a setting, and an eval built from your own transcripts.Messages requests, the tool-use loop and parallel results, prompt-prefix design and measured token spend, and a private eval with rubric grading. Four lessons, written for the person with the API key.

5 lessons~47 min
  1. 01

    What Your Tech Person Is Actually BuildingOne request, end to end: what you send and what comes backMessages requests: system vs turns, stateless history, usage accounting, typed errors

    There is no magic in the wiring. Your tech person sends a block of text and gets a block of text back, and pays by the word in both directions. Knowing that is enough to follow the rest of this module.One call: the system prompt that never changes, the turns that do, and the usage numbers that come back with the answer. The service keeps no memory, so history is your job.Request anatomy and statelessness: system vs messages, history re-send, the usage block as the only honest cost signal, and which failures retry.

    • api
    • developer
    8 min
  2. 02

    How the AI Does Things Instead of Just TalkingTools and the loop: how the AI asks your code to do somethingTool use: schemas, the loop, parallel results in one turn, and the gate in code

    The AI cannot touch your systems. It can only ask your code to, and your code decides whether to. That gap is where every safety rail you have goes.The loop: you describe the tools, the model asks for one, your code runs it and returns the result, and round it goes until the model stops asking. The refusal point is yours.Tool definitions and strict schemas, the loop, parallel tool results in a single turn, error results, and why the gate lives in code and not in the system prompt.

    • api
    • developer
    • agents
    10 min
  3. 03

    Why the Bill Grew, and What to Do About ItDesigning for cost: what to measure before you change anythingPrompt-prefix design, measured spend, and the levers in the order that pays

    Bills grow for boring reasons, and almost never the reason people guess. Measure before you change anything, then take the free savings before the ones that cost you quality.Find where the tokens go, fix the stable part of your prompt so it can be cached, and work the levers in order: free wins before tradeoffs.Token profiling from usage logs, prefix stability and cache invalidation, and lever order: caching and hygiene before effort, batch or model changes.

    • api
    • developer
    • cost
    10 min
  4. 04

    How to Know It Still Works Next MonthA test set of your own: proving a change helpedPrivate eval sets: sourcing from transcripts, rubric grading, and change verification

    You cannot tell whether a prompt change helped by reading a few answers. You need the same questions, scored the same way, before and after. Thirty of them is enough.Build a small set from work you have already done, score it the same way every time, and you can answer whether a change helped instead of guessing.Private eval construction: transcript sourcing, rubric design, held-out split, and reading a score difference honestly.

    • api
    • developer
    • eval
    10 min
  5. 05

    When the AI Runs Out of RoomLong runs: what to do when the conversation stops fittingContext management: clearing tool results, compaction, external memory, and budgets

    Every job the AI does has a desk, and the desk has an edge. A long run fills it with paperwork from steps nobody cares about any more. You decide what gets cleared, or it fails at the worst moment.Everything a run reads stays in the conversation. Past a certain length that becomes the problem. Clear the old results, summarise them, or write them down outside, and choose deliberately.Context exhaustion and the three strategies: clearing tool results, compaction, and external memory, plus run budgets and stopping conditions.

    • api
    • developer
    • agents
    9 min