Skip to main content

Operator response guide

Guides an operator through the daily run. See How the daily run works and Job and workflow statuses.

Task walkthrough: Respond to a failed job. This page is the full configuration and troubleshooting reference.

Most of what an operator does during the daily run is read a status and decide whether it needs a response. This page maps each situation you are likely to meet to what it means and the action to take, then covers responding to a failed job, reading its exit code, and when to escalate.

Situation → response​

SituationWhat it meansAction
Job in a WAIT_* dependency statusWaiting on a predecessor job, threshold, resource, or expressionConfirm the blocker; if it should proceed now, Force Start. If it shouldn't run, Skip.
Job WAIT_MACHINENo agent/machine available — a specific-agent job holds here while its target agent is offlineCheck agent/relay availability (Administrator); the job proceeds when the target is back online. A legacy agent showing UNKNOWN (amber) means its relay is stale — check the relay, not the agent.
Job WAIT_START_TIMENot yet at its start timeNormal; Force Start only if it must run early.
Job in a running state (JOB_RUNNING, starting, LATE_TO_FINISH)Actively runningTo stop it, use Kill — Cancel applies only before a job starts running. To force its downstream outcome, Mark Finished OK or Mark Failed (a running container job accepts Kill only).
Job ON_HOLDHeld intentionallyRelease when ready.
Job LATE_TO_START / LATE_TO_FINISHPast a timing thresholdInvestigate the delay; the job is still progressing.
Job FAILEDThe job failedRead the job output/logs; Restart once fixed, Mark Fixed if resolved outside the system, Under Review to flag it, or Cancel.
Job MISSED_START_TIMEDidn't start in its windowDecide whether to Skip or rerun via the next build.
Job stuck in *_PENDING_TERMCompletion processing not settlingTransitional normally; if prolonged, escalate (see below).
Workflow ON_HOLD / PARENT_HOLDHeld (or parent held)Release when ready; releasing the parent releases children.
Workflow COMPLETED, but a job in it has to run againThe run finished; you need more work out of that same instance rather than a fresh buildHold, Release or Start the completed instance to reopen it, then Restart the job. Nothing runs on the reopen alone. A nested sub-schedule can't be reopened — act on its parent.
Workflow STARTED_BY_USER with nothing runningAn instance reopened with Start on which no job has been revived yetRestart the job you reopened it for, or Close the instance to put it back to completed.

Responding to a failed job​

To respond to a failed job, complete the following steps:

  1. Open the job's detail and read its output/logs to find the cause.
  2. Resolve the cause (often a connector/config issue — route to the Builder or Administrator).
  3. Then choose: Restart (re-run), Mark Fixed (resolved outside the system), Under Review (flag for follow-up), or Cancel (abandon).

Reading the exit code​

A completed job's detail now shows its real Exit Code (return code) and a Termination description in the Completion section — and on the Output tab beside the run date. The Exit Code renders as a badge: green for 0 (success), red for any non-zero value. This is a data field only — it does not change the job's status classification (a job can be FAILED with the exit code still shown).

  • Previously the return code always displayed a dash even for a job that finished OK with code 0; a real numeric code (including 0) is now persisted and shown for both universal-agent and legacy/LSAM jobs.
  • A failed job with no numeric code reported by the agent shows — (not a spurious 0) — distinguish "failed with code 0" from "failed, code unknown."
  • Mixed-version agents are supported: an older agent that omits the code leaves the field blank rather than breaking the report.

Which action is available​

Actions are status-gated (see Job and workflow statuses). There are 11 job actions (realigned to the Classic matrix). Key points for triage:

  • Cancel stops a job that hasn't started running; Kill stops one that is running. They don't overlap — offer Kill, not Cancel, for anything in a running status.
  • Restart is available from any stopped state (FAILED, MARKED_FAILED, CANCELLED, SKIPPED, INITIALIZATION_ERROR, UNDER_REVIEW, JOB_TO_BE_SKIPPED, and the finished states) — not only FAILED/CANCELLED.
  • Mark Fixed from FAILED/MARKED_FAILED/INITIALIZATION_ERROR/UNDER_REVIEW; Release only from ON_HOLD; Force Start/Skip from waiting states.
  • Skip is deferred and in-sequence. In a live workflow, Skip marks the job JOB_TO_BE_SKIPPED rather than skipping it on the spot: it becomes terminal SKIPPED — and its dependents are released — only when it would otherwise have been its turn to start (its TIME/JOB/THRESHOLD/RESOURCE gates satisfied). Skipping a job that has already reached that point (FAILED, MISSED_START_TIME) resolves promptly. So after selecting Skip, expect the job to sit in "Marked to be skipped" until its turn comes — that is normal, not a stuck action.
  • Mark Finished OK / Mark Failed are broad manual overrides that force a job's outcome so dependents proceed/react; accepted from most non-terminal states. A running container job accepts Kill only (the mark-overrides are refused while it runs).
  • Marking a running job doesn't stop it. Both overrides move the status the platform records and nothing else — the process carries on running on the agent until it finishes by itself. Kill is what stops the work. Mark a running job and then restart it and you can have two runs of the same program alive at once, so Kill first if the work has to stop too. See Marking a running job does not stop it.

If an action isn't offered, the job isn't in a status that accepts it.

Contact support when​

  • A job sits in a *_PENDING_TERM status indefinitely.
  • A status-appropriate action is rejected by the system, or a job's status doesn't reflect what the agent actually reported.

Include the workflow instance, the job, its status, and the job output.