Recover a failed run automatically
What this solves
When a job fails in the middle of the night, two things usually go wrong: the mess it leaves (a half-written file, a lock, a partial post) sits there until morning, and nobody finds out until a business process is already late. The run just stops, and recovery waits for a person who isn't watching.
What you get. The workflow cleans up and retries the recoverable failures on its own, and raises a clear alert the moment something needs a human, so the run keeps moving and an operator only steps in for the failures that genuinely need judgment.
Related topics
- Run events on job outcomes (Builder)
- Order jobs with dependencies (Builder)
- Respond to a failed job (Operator)
- Monitor the daily run (Operator)