Runbook
The written procedure for operating a system or responding to a known failure, step by step.
Also written operations manual, procedure
A runbook documents how something is operated: the sequence to start a service, what to check during the batch cycle, and what to do when a particular job fails.
Its most valuable content is failure handling. For each known problem it states how to recognise it, what it means, what to do, and who to escalate to if the fix does not work. That turns an unfamiliar failure at three in the morning into a procedure to follow rather than a problem to solve from first principles.
It matters most where knowledge is concentrated in a few people. Writing down what those people know is one of the more effective things an organisation can do about the risk of them retiring, a real concern on systems whose original authors have long since left.
Good ones are specific and current, naming actual job names, dataset names and commands, and updated whenever a procedure changes. Stale ones are worse than none at all, because they are followed confidently and are wrong.
Related terms
- OperatorThe person who watches the running systems, responds to messages, and starts, stops and recovers work.
- Production supportKeeping live systems running: diagnosing failures, fixing what broke, and getting the service back.
- Batch schedulerSoftware that submits production jobs automatically in the right order, at the right time, based on what has already succeeded.
- Change controlThe process every change goes through before reaching production: reviewed, approved, scheduled and reversible.
- Batch windowThe period, usually overnight, in which batch work has to finish before the business needs the systems back.