Respond to a failed job
When a job fails, your goal is to find out why, fix or account for it, and get the run moving again. The job's own output is the fastest path to the cause.
A failed job stops the run. Without a clear path to the cause and the right action, it sits unresolved, or gets restarted before the real problem is fixed, so it just fails again.
Steps
To respond to a failed job, complete the following steps:
- In Processes, open the failed job to see its detail.
- Read the job output and logs to find the cause. In the Completion section (and on the Output tab), check the Exit Code and Termination description. The exit code shows as a badge, green for 0 and red for any non-zero value. On a legacy LSAM job, also check the Agent Message row below Termination — it carries the reason the agent itself gave, which is often the actual answer when Termination holds only the exit code. See Agent Message.
- Resolve the cause. Many failures are configuration or connection issues, so escalate to the Builder (job settings) or Administrator (connections, agents) as needed.
- Choose how to clear the job (see below).
The Exit Code is the code the command, script, or connector returned; it's a diagnostic detail, not the job's status. A failed job that reports no exit code shows a dash (—) rather than a
0, so "failed with code 0" reads differently from "failed, code unknown."
A job whose frequency configures failure retries does not read Failed while it still has attempts left. It waits in Wait start time, still showing the failed attempt's exit code and termination description, and its dependents keep waiting — so the run is not stalled and nothing downstream has reacted yet. The job's Summary tab shows Retry Count as used-of-maximum.
Let it take the attempts it has left if the cause might be transient. If it will not be — the failure is in the job's configuration, or a system it needs is down — Restart the job once that is resolved, or Cancel it, rather than waiting for the retries to run out. See Failure retries.
Your options on a failed job
| Action | Use it when |
|---|---|
| Restart | The cause is fixed and you want the job to run again. Reset Retry Count is selected by default, so the job runs with its full retry budget again; clear it to keep the count it had. |
| Mark Fixed | The work was completed or resolved outside the system, and the job should count as done. |
| Under Review | You need to flag the job for follow-up without resolving it yet. |
| Cancel | The job should not run, and downstream jobs should stop waiting on it. |
Restart and Mark Fixed are available from a Failed job; Mark Fixed is also available from a job Under Review. If an option isn't offered, the job isn't in a status that accepts it.
Before you restart
Confirm the underlying problem is actually resolved. Otherwise the job will just fail again. If the failure is in the job's configuration, the Builder needs to change it before a restart will succeed.
When to escalate
Escalate (with the job output, the job's status, and the workflow instance) when the failure isn't explained by the job's settings or a connection/agent issue, or when a job's status doesn't match what actually happened.
Related topics
- Operator response guide — full configuration reference and troubleshooting
- Job statuses and actions (Operator)
- Respond to a failed connector job (Operator)