# Product data cleanup

One script that makes a product CSV fit for a bulk import or a menu platform: SKUs, categories, units, sizes, prices, image URLs, and a description check against the forbidden patterns in `../menu-copy-rules.md`. From lesson 04-05 of the No Bullshit AI Course. Python 3, standard library, no AI call.

- `product-data-cleanup.py`: the script. `CATEGORY_MAP`, `UNIT_MAP` and `COPY_RULES` are constants at the top; edit those for your catalog and your state.
- `sample-products.csv`: twelve fake rows with the usual problems: a SKU with a trailing space, four spellings of "grams", a blank size, a `$` in a price, a duplicate SKU, and five descriptions that would not survive a compliance read.

Run it:

    python3 product-data-cleanup.py sample-products.csv --out clean.csv --image-base https://cdn.example.com/products/
    python3 product-data-cleanup.py sample-products.csv --check-copy
    python3 product-data-cleanup.py --self-test

What you get: a log with one line per change, a `clean.csv` with an `image_url` column (your file names joined to the base URL) and a `copy_flags` column, and a non-zero exit if a duplicate SKU is found. Rows with anything in `copy_flags` are not ready for a menu; the script never rewrites a description, it only points.

Image URLs: most menu and ERP bulk imports want a hosted URL per row, not a file. Upload the folder to wherever you host images, pass that folder's URL as `--image-base`, and pass the local folder as `--image-dir` so the script can tell you which rows point at a file that does not exist.

Version 1.0, tested 2026-09-17 with Python 3.11 on the bundled sample. CC0 1.0: copy it, change it, no attribution needed.
