# csv-head: look before you paste

Two small things for the moment right before you paste a spreadsheet export into an AI chat.

1. `csv-head.py`: prints the header row, row count, a few sample rows, blank cells per column, columns that look like they hold personal data, and a rough token estimate. Offline, standard library only, no network, no install.
2. `paste-this-first.md`: the prompt to send *before* your data, so the AI describes the columns back to you and asks for the formula-or-script route on anything numeric.

From Distru's No Bullshit AI Course, lesson 03-02 "Paste the spreadsheet". CC0 1.0: copy, edit, share, no attribution needed.

## Run it

    python3 csv-head.py packages.csv              # report
    python3 csv-head.py packages.csv -n 5         # show 5 sample rows instead of 3
    python3 csv-head.py packages.csv --paste      # header + 3 rows, ready to paste into a chat
    python3 csv-head.py --self-test               # checks itself on a built-in sample; prints ok/FAIL lines

Needs Python 3.8 or newer, which macOS and most Linux machines already have. On Windows, install Python from python.org and run the same commands in PowerShell.

## What the report tells you

- **data rows / columns.** Write these two numbers down. After any cleanup, the row count should match.
- **ragged rows.** Rows with the wrong number of columns. Usually a comma inside a product name. Fix before pasting; the AI will silently misread them.
- **token estimate.** A range, because tokenizers differ and this is a heuristic. A 500-row, 10-column package export with a Metrc tag on every row comes out around 25K to 29K tokens (the true count was 26.8K), well inside any current model's context window. Tags and numbers cost about one token per 2 characters, far more than prose. Past roughly 60K tokens the script tells you to upload as a file or ask for the script instead of the answer.
- **LOOKS PERSONAL.** Header names that usually hold customer or employee data. Drop the column or replace values with `Customer 1`, `Customer 2` before pasting. The check is on the header name only; it will miss a badly named column and will flag `card_count`. Read the header yourself too.
- **Metrc tags.** Columns where most values match the 24-character `1A…` tag shape. Keep them as text. Spreadsheets turn them into `1.23E+23`.

## How it was tested

2026-09-17, offline. `--self-test` passes on the built-in sample. The estimator's constants were fitted against `tiktoken` (`cl100k_base` and `o200k_base`) on three synthetic 500-row exports with realistic column shapes, then checked:

| file | characters | csv-head range | tiktoken cl100k | o200k |
|---|---|---|---|---|
| tag-heavy, 5 columns | 24,118 | 10.3K to 13.9K | 12,188 | 12,188 |
| prose-heavy, 3 columns | 49,063 | 10.9K to 14.7K | 12,290 | 12,164 |
| mixed package export, 10 columns | 54,473 | 21.3K to 28.8K | 26,830 | 26,653 |

Max error 7%. Anthropic and Google tokenizers were not checked; expect the same order of magnitude. The PII check is header-name only and was not tested beyond the column names listed in the script. Not tested on real customer exports; run it on yours and tell us what it got wrong.

## Change it

The regex `PII_HINTS` is where you add your own column names. `TAG_LIKE` matches Metrc tags; if your state's tags look different, edit it. The token estimate lives in `estimate_tokens()`; if you want the exact count, `pip install tiktoken` and encode the file yourself.
