# Order normalizer

A buyer sends one spreadsheet for three stores, in their words ("BD 3.5", "2 cs", "gummies"). This turns it into draft order lines in your words, with a confidence per line and an "unknown" bucket for things you do not sell. From lesson 04-03 of the No Bullshit AI Course. Python 3, standard library, no AI call, no write to any system.

- `normalize.py`: the matcher. Thresholds (`MATCH`, `REVIEW`) and the buyer-shorthand table (`ALIASES`) are constants at the top. Edit those first.
- `catalog.csv`: a fake 15-product catalog. Replace with an export of yours: `sku, name, brand, category, size, unit, case_size`.
- `buyer-order.csv`: a fake three-store order with the usual mess. Replace with the buyer's sheet, reshaped to `store, item, qty`.

Run it:

    python3 normalize.py buyer-order.csv catalog.csv
    python3 normalize.py buyer-order.csv catalog.csv --out draft.csv
    python3 normalize.py --self-test

What the output means:

- ` ` (blank) `match`: confidence at or above `MATCH`. Put it on the draft.
- `?` `review`: two products fit, or the confidence is middling, or the quantity could not be read. The rep picks from the three candidates shown.
- `!` `unknown`: nothing in your catalog is close. Ask the buyer; do not guess.

How it scores: buyer text and catalog name are both normalised (aliases expanded, "3.5 g" to "3.5g"), then a word-overlap score and a string-similarity score are blended. A size that disagrees ("1g" vs "0.5g") is punished hard because it is a different product, not a typo. A close runner-up forces `review` even when the top score is high. Case quantities are multiplied by the catalog's `case_size`; if your catalog has no case size, the line goes to review instead of inventing a number.

Where a model helps: the `unknown` and `review` rows. Hand the buyer's text plus the three candidates to a model and ask it to pick or say "none"; or use a decision model as in lesson 04-06. The lesson shows both. Everything the matcher already resolved never touches a model.

Version 1.0, tested 2026-09-17 with Python 3.11 on the bundled files. CC0 1.0: copy it, change it, no attribution needed.
