# Runbook: <automation name>

One page per automation. Fill it in the day you switch the schedule on, not the day it breaks.
From Distru's No Bullshit AI Course, module 05, lesson 4. Version 1.0 (2026-09-17). CC0.

## What it does (one sentence)

> Every <when>, it <reads what> from <where> and <writes what> to <where>. Example: Every morning at 06:00 Pacific it reads today's READY_TO_SHIP orders from Distru and sets them to DELIVERING, then posts a line to #ops.

## Where it lives

| | |
|---|---|
| Runs in | n8n at `<url>` / a cron job on `<server>` / a script on `<laptop>` |
| Workflow or script name | |
| Schedule | `0 6 * * *` in `<time zone>` |
| Credentials it uses | Distru API key "<name>" (view orders, edit orders only), Slack "<name>" |
| Where the credentials are stored | n8n Credentials / 1Password vault "<name>" |
| Who owns it | <name>, backup <name> |
| Who reads the Slack channel | #ops, checked by <role> each morning |

## How to turn it off

1. Open the workflow. Toggle **Active** off. (Or: `crontab -e`, comment the line.)
2. Post in #ops: "Turned off <name> because <reason>. Doing it by hand until <date>."
3. The manual fallback is: <the by-hand steps, or a link to the SOP>.

Turning it off is always allowed. Nobody needs permission to stop an automation.

## What a good run looks like

- Slack line at ~06:01 starting with `Set N order(s) to DELIVERING`.
- N is roughly <normal range, e.g. 3 to 15>. Zero on a delivery day is suspicious.
- Execution list shows a green run under <n> seconds.

## Known failures and what to do

| you see | it means | do this |
|---|---|---|
| `FAILED to update ... 400 ... fulfilled` | the order is not ready (unfulfilled line, no customer, no transfer) | fix the order in Distru; the run does not retry, tomorrow's run will pick it up if still due |
| `401` | the API key was revoked or expired | make a new key with the same narrow permissions, update the credential, run once manually |
| `429` or `403 rateLimitExceeded` | too many calls too fast | nothing today; lower the per-run cap or widen the batch interval before tomorrow |
| No Slack line at all | the run did not happen, or Slack failed | open the executions list; if empty, check the schedule and the instance is up; if red, read the error |
| The same record twice | a write ran but the record of it was lost, then the write ran again | stop the workflow, clean up by hand, add or fix the find-before-create step |
| More than one page of results | more records matched than one run handles | run again, or add pagination |

## Rules this automation follows

- [ ] Every write is keyed on a stable id (order id, tag) and checks for an existing record first.
- [ ] Retries only on 429 and 5xx, never on 400. Retries wait between tries.
- [ ] A failed record is reported by name, in a place a human reads, the same day.
- [ ] A dry-run switch exists and is used after every change to the workflow.
- [ ] The workflow does at most <cap> writes per run.
- [ ] Nothing in it touches money, compliance or a customer without a person approving first.

## Change log

| date | who | what changed | dry run done? |
|---|---|---|---|
| | | | |

## Monthly review (first Monday, 15 minutes)

- [ ] Count runs, failures, and records touched last month. Write the three numbers here: ___ / ___ / ___
- [ ] Read one successful run end to end. Does the output still match what the sentence at the top says?
- [ ] Open the failure with the ugliest message. Does the table above cover it? If not, add a row.
- [ ] Did anyone turn it off last month? Why? Is the reason fixed?
- [ ] Are the credentials still narrow, and does the owner still work here?
- [ ] Is it still worth running? If it touched fewer than <n> records a month, consider deleting it.
