Operator response guide
Guides an operator through the daily run. See How the daily run works and Job and workflow statuses.
Task walkthrough: Respond to a failed job. This page is the full configuration and troubleshooting reference.
Most of what an operator does during the daily run is read a status and decide whether it needs a response. This page maps each situation you are likely to meet to what it means and the action to take, then covers responding to a failed job, reading its exit code, and when to escalate.
Situation → response
| Situation | What it means | Action |
|---|---|---|
Job in a WAIT_* dependency status | Waiting on a predecessor job, threshold, resource, or expression | Confirm the blocker; if it should proceed now, Force Start. If it shouldn't run, Skip. |
Job WAIT_MACHINE | No agent/machine available — a specific-agent job holds here while its target agent is offline | Check agent/relay availability (Administrator); the job proceeds when the target is back online. A legacy agent showing UNKNOWN (amber) means its relay is stale — check the relay, not the agent. |
Job WAIT_START_TIME | Not yet at its start time | Normal; Force Start only if it must run early. |
Job in a running state (JOB_RUNNING, starting, LATE_TO_FINISH) | Actively running | To stop it, use Kill — Cancel applies only before a job starts running. To force its downstream outcome, Mark Finished OK or Mark Failed (a running container job accepts Kill only). |
Job ON_HOLD | Held intentionally | Release when ready. |
Job LATE_TO_START / LATE_TO_FINISH | Past a timing threshold | Investigate the delay; the job is still progressing. |
Job FAILED | The job failed | Read the job output/logs; Restart once fixed, Mark Fixed if resolved outside the system, Under Review to flag it, or Cancel. |
Job MISSED_START_TIME | Didn't start in its window | Decide whether to Skip or rerun via the next build. |
Job stuck in *_PENDING_TERM | Completion processing not settling | Transitional normally; if prolonged, escalate (see below). |
Workflow ON_HOLD / PARENT_HOLD | Held (or parent held) | Release when ready; releasing the parent releases children. |
Workflow COMPLETED, but a job in it has to run again | The run finished; you need more work out of that same instance rather than a fresh build | Hold, Release or Start the completed instance to reopen it, then Restart the job. Nothing runs on the reopen alone. A nested sub-schedule can't be reopened — act on its parent. |
Workflow STARTED_BY_USER with nothing running | An instance reopened with Start on which no job has been revived yet | Restart the job you reopened it for, or Close the instance to put it back to completed. |
Responding to a failed job
To respond to a failed job, complete the following steps:
- Open the job's detail and read its output/logs to find the cause.
- Resolve the cause (often a connector/config issue — route to the Builder or Administrator).
- Then choose: Restart (re-run), Mark Fixed (resolved outside the system), Under Review (flag for follow-up), or Cancel (abandon).
Reading the exit code
A completed job's detail now shows its real Exit Code (return code) and a Termination
description in the Completion section — and on the Output tab beside the run date. The Exit
Code renders as a badge: green for 0 (success), red for any non-zero value. This is a data
field only — it does not change the job's status classification (a job can be FAILED with the
exit code still shown).
- Previously the return code always displayed a dash even for a job that finished OK with code 0;
a real numeric code (including
0) is now persisted and shown for both universal-agent and legacy/LSAM jobs. - A failed job with no numeric code reported by the agent shows
—(not a spurious0) — distinguish "failed with code 0" from "failed, code unknown." - Mixed-version agents are supported: an older agent that omits the code leaves the field blank rather than breaking the report.
Which action is available
Actions are status-gated (see Job and workflow statuses). There are 11 job actions (realigned to the Classic matrix). Key points for triage:
- Cancel stops a job that hasn't started running; Kill stops one that is running. They don't overlap — offer Kill, not Cancel, for anything in a running status.
- Restart is available from any stopped state (
FAILED,MARKED_FAILED,CANCELLED,SKIPPED,INITIALIZATION_ERROR,UNDER_REVIEW,JOB_TO_BE_SKIPPED, and the finished states) — not onlyFAILED/CANCELLED. - Mark Fixed from
FAILED/MARKED_FAILED/INITIALIZATION_ERROR/UNDER_REVIEW; Release only fromON_HOLD; Force Start/Skip from waiting states. - Skip is deferred and in-sequence. In a live workflow, Skip marks the job
JOB_TO_BE_SKIPPEDrather than skipping it on the spot: it becomes terminalSKIPPED— and its dependents are released — only when it would otherwise have been its turn to start (its TIME/JOB/THRESHOLD/RESOURCE gates satisfied). Skipping a job that has already reached that point (FAILED,MISSED_START_TIME) resolves promptly. So after selecting Skip, expect the job to sit in "Marked to be skipped" until its turn comes — that is normal, not a stuck action. - Mark Finished OK / Mark Failed are broad manual overrides that force a job's outcome so dependents proceed/react; accepted from most non-terminal states. A running container job accepts Kill only (the mark-overrides are refused while it runs).
- Marking a running job doesn't stop it. Both overrides move the status the platform records and nothing else — the process carries on running on the agent until it finishes by itself. Kill is what stops the work. Mark a running job and then restart it and you can have two runs of the same program alive at once, so Kill first if the work has to stop too. See Marking a running job does not stop it.
If an action isn't offered, the job isn't in a status that accepts it.
Contact support when
- A job sits in a
*_PENDING_TERMstatus indefinitely. - A status-appropriate action is rejected by the system, or a job's status doesn't reflect what the agent actually reported.
Include the workflow instance, the job, its status, and the job output.