How It Actually Works
Why the tool behaves the way it does: how a machine learns from examples, why it sometimes memorises instead, how it turns words into numbers, and why the same question gets different answers.How machine learning actually works (training from examples, overfitting), how words become numbers (embeddings), how a transformer keeps track of a sentence (attention), why answers vary (sampling and temperature), and the different kinds of model.Symbolic AI to ML to LLMs, perceptrons and gradient descent, overfitting and generalisation, embeddings and cosine similarity, attention and transformers, sampling and temperature, the model taxonomy.
Carry on withCarry on withnext:
- 017 min
Rules, Then Learning, Then Guessing WordsRules, learning, and guessing words: three kinds of AI in your shopSymbolic AI, machine learning, LLMs: three approaches and where each still winsDoneDonedone
Computers have been made to seem smart in three ways: rules someone wrote down, patterns learned from examples, and guessing the next word. All three are already running in a dispensary, and each is right for different jobs.Seventy years of AI in three approaches: hand-written rules (expert systems), learning from labelled examples (machine learning), and next-word prediction at scale (LLMs). Where each lives in cannabis software today, and how to tell which one you are looking at.Top-down symbolic systems versus bottom-up learned models, the knowledge-acquisition bottleneck that ended the expert-system era, the path from perceptrons to transformers, and a working rule of thumb for when a rule beats a model.
- 028 min
How a Machine Learns From ExamplesHow a machine learns: one neuron, its weights, and fixing mistakesPerceptrons, the update rule, and why networks need layers and gradient descentDoneDonedone
Teach one tiny decider to guess which batches will sell, by showing it its mistakes one at a time. Every AI model you use, including the chat ones, learned the same way, just with billions of these deciders.A single artificial neuron (a perceptron) weighs a few facts and says yes or no. Training is showing it the examples it gets wrong and nudging its weights each time. You run it yourself, then see why real networks need many layers.The perceptron model and update rule, trained live on synthetic sell-through data; convergence on separable data and cycling on non-separable data; from there to multi-layer networks, differentiable losses, gradient descent and backpropagation.
- 037 min
When It Memorises Instead of LearningMemorising instead of learning: overfitting and the held-out testOverfitting, bias and variance, and why you always score on held-out dataDoneDonedone
A model can get so good at the examples it was shown that it gets worse at everything else. You will make that happen with a slider, and learn the one test that catches it: hide some answers and see if it can guess them.Overfitting is a model memorising the noise in its training data instead of the pattern. You will cause it on purpose, see why the error on the training data lies to you, and learn to check a model on data it never saw (a held-out set).Model capacity versus generalisation, the bias-variance trade-off made visible on a polynomial fit, held-out and validation sets, and what overfitting looks like in the vendor claims and fine-tunes you will be shown.
- 048 min
Words as NumbersWords as numbers: embeddings and similarityText representation: bag-of-words, embeddings, cosine similarityDoneDonedone
Computers only work with numbers, so every AI turns words into lists of numbers first. Done well, words that mean similar things get similar numbers. You will tap around a real map of cannabis words and see what the model thinks is close, including one it gets wrong.Why counting words is not enough, how an embedding turns a word or a paragraph into a list of numbers so that similar meanings land close together, how closeness is measured (cosine similarity), and what that makes possible: search by meaning.From bag-of-words and TF-IDF to learned dense embeddings (word2vec's distributional hypothesis, contextual sentence embeddings), cosine similarity, a real embedding map of 30 cannabis terms, and the failure mode of embedding bare keywords.
- 057 min
How It Keeps Track of a SentenceKeeping track of a sentence: attention and the transformerSequence models: RNNs, attention, transformers; encoder, decoder, encoder-decoderDoneDonedone
Word order changes meaning, and "it" has to point at something. Older AI read one word at a time and forgot the start of long messages. Today's AI lets every word look back at every other word at once. That one idea is what made chat AI possible, and it also explains its limits.Why word order breaks simple word-counting, how older models read one word at a time and lost track, how attention lets every word weigh every other word directly, and the three shapes of transformer model you will meet: readers, writers, and translators.From recurrent networks and their long-range memory problem to self-attention (queries, keys, values), transformer blocks, the quadratic cost of context, and encoder-only, decoder-only and encoder-decoder architectures with the jobs each fits.
- 066 min
Why the Same Question Gets Different AnswersSame question, different answers: odds, sampling and temperatureDecoding: next-token distributions, sampling, temperature, top-pDoneDonedone
The AI does not pick its next word. It works out the odds for every possible word, and then one gets drawn, a bit like a weighted raffle. That is why asking twice gets two answers. You will turn the randomness dial yourself and see what it can and cannot fix.A language model outputs a probability for every possible next word; software then samples one. Temperature controls how adventurous that draw is. Why answers vary, when to turn it down, and why turning it down makes answers consistent but not correct.Logits to probabilities via softmax, temperature scaling, greedy versus sampled decoding, top-k and top-p, what temperature 0 does and does not guarantee, and how to set decoding for extraction versus generation.
- 076 min
The Model ZooThe model zoo: kinds of AI model, sizes, and who runs themModel taxonomy: modality, size, access, and model versus serviceDoneDonedone
"AI model" covers very different things: ones that write, ones that only read, ones that only decide, ones that see pictures. They come big or small, and some you can run on your own computer. Three questions sort any of them, and you will use those three to read vendor pitches.Sort any model by three questions: what goes in and what comes out (text, numbers, images, decisions), how big it is (frontier or small), and who runs it (a vendor's service or open weights on your own machine). Plus the difference between a model and the product wrapped around it.Foundation models and their descendants by modality and output (generative, embedding, classifier, decision, multimodal), by scale (frontier versus small language models), by access (proprietary API versus open-weight), and why the model is not the service when it comes to data handling.