Generation data group
A set of datasets holding successive versions of the same file, so today's copy, yesterday's and last week's all exist under one name.
Also written GDG, generation dataset
A generation data group, or GDG, is a way of keeping successive versions of a dataset without inventing new names for each one. The group has a base name, and each run creates a new generation under it.
You refer to generations relatively. Zero means the current one, minus one the previous, minus two the one before that, and plus one creates a new one. A daily job can therefore be written once: read generation zero, write generation plus one. Tomorrow, what was plus one has become zero, and the JCL does not change.
The group is defined with a limit, say the last ten generations, and older ones are dropped automatically as new ones arrive. That gives you a rolling history without any housekeeping.
This pattern is used constantly in batch. Daily extracts, transaction files, backups and interface files between systems are all typically GDGs, because it makes reruns and recovery straightforward: if today's run was wrong, yesterday's data is still there under a predictable name.
Related terms
- DatasetWhat a mainframe calls a file: a named collection of records with a defined structure, rather than a loose stream of bytes.
- Batch processingRunning work in bulk, without anyone watching, usually on a schedule. The overnight run that settles a day of transactions.
- JCLJob Control Language: the language used to tell the mainframe what programs to run, in what order, and with what files.
- DispositionThe DISP parameter, which says whether a dataset already exists, and what to do with it when the step ends normally or fails.
- IDCAMSThe utility used to create, list, copy and delete datasets, and the standard tool for anything involving VSAM.