Skip to main content

Job statuses and actions

A quick reference for what a job's status means and what you can do about it.

What this solves

A job's status is only useful if you know what it means and what you're allowed to do about it. Guess wrong, and you restart something that shouldn't be, or leave a run stalled.

What a status is telling you​

Status groupExamplesWhat it means
HeldOn HoldWon't run until released.
WaitingWaiting on a dependency, start time, resource, machine; Late to Start; Marked to be skippedEligible but waiting on something. Expected in a normal run — worth attention only once it shows Late to Start, or when nothing it's waiting on is going to arrive. "Marked to be skipped" means you skipped it and it will drop out when its turn comes.
RunningAttempting to start, Running, Late to FinishActively starting or running.
Finishing(the "pending" states)Done; completion is being processed. Transitional, not stuck.
FinishedFinished OK, FixedCompleted successfully.
FailedFailed, Initialization Error, Missed Start TimeNeeds attention.
Other terminalCancelled, SkippedEnded without running to success.
Under ReviewUnder ReviewA failed job you have flagged for review. Not final: dependents waiting for a failure are released, but the workflow does not complete until you resolve the job or Close the workflow.

Exceeded Max Run Time is a badge, not a status​

A job that has run longer than its Max Run Time carries an Exceeded Max Run Time badge beside its status. The status itself doesn't change: the job reads Running while it runs and then whatever it ends up with — including Finished OK.

Three things follow from that, and they matter during a run:

  • The job has not been stopped. The platform reports the overrun and lets the job finish. If you want it stopped, Kill it like any other running job.
  • The badge stays after the job completes. A job that overran and then finished OK is precisely the case the status can't show, so the marker outlives the run. Seeing it on a green job is not a contradiction.
  • There is no filter for it. The status filters work on statuses, and this isn't one. The Max Run Time Exceeded row on the job's Summary tab gives the time it was detected.
Good to know

Nothing reported an overrun in earlier builds — a job could run for hours past its limit and finish looking clean. If a long-running job suddenly starts carrying this badge, the job hasn't changed; what the platform tells you about it has.

What you can do, by situation​

The job is…You canTo…
On HoldRelease · Force Start · Skip · CancelLet it run, push it through, skip it, or stop it.
Waiting on a dependency/time/resource/machineForce Start · Skip · Hold · CancelRun it now, skip it, hold it, or stop it. A job waiting for its next failure-retry attempt is in this group too — see the note below.
RunningKillStop a job that's actively running.
FailedRestart · Mark Fixed · Under Review · Skip · CancelRe-run, accept as resolved, flag for review, skip it, or abandon.
Missed Start TimeForce Start · Skip · CancelRun it now, skip it, or abandon. On a job set to rerun, missing the deadline retires the rest of its cycle, so restarting it gives you a single run rather than the runs it had left.
Under ReviewMark Fixed · Restart · CancelResolve, re-run, or abandon.

Actions are status-specific: the run only offers an action a job's status allows. Cancel stops a job before it runs; Kill stops one that's already running. Restart is available once a job has stopped (Failed, Cancelled, Skipped, …); Release only from On Hold.

Mark Finished OK and Mark Failed are manual overrides available from most live states: use them to force a job's outcome (successful or failed) so the jobs waiting on it move on. (A container job, one whose body is another workflow, can only be Killed while it's running.)

Marking a running job doesn't stop it

Both overrides change the outcome the run records. Neither stops the work: the process carries on running on the machine until it finishes by itself. Kill is the action that stops it.

So if you mark a running job and then restart it, the same program can be running twice at once. Kill first when the work has to stop as well. When the original run does finish, its result is discarded rather than applied to the restarted job.

What the actions do​

ActionEffect
HoldPause the job so it won't run until released.
ReleaseTake a job off hold.
Force StartStart a waiting job now, regardless of what it's waiting on. Refused if the target agent is marked, or if the job's legacy agent group has nothing to dispatch to — the refusal names which.
SkipMark the job so it doesn't run. In a live workflow the skip is applied in sequence (when the job would have been its turn to run), and only then do its dependents stop waiting on it. Skipping a job that has already failed or missed its start time takes effect promptly.
RestartRun a stopped job again (failed, cancelled, skipped, or already finished). The dialog offers Restart on Hold, and Reset Retry Count — selected by default, so the job starts again with its full retry budget.
Mark FixedTreat a failed job as resolved (work done outside the system).
Under ReviewFlag a failed job for follow-up.
CancelStop a job that hasn't started running yet; dependents stop waiting on it.
KillStop a job that's actively running. (Cancel doesn't apply once a job is running.) On a legacy LSAM machine it takes up to one relay heartbeat to land — see the note below.
Mark Finished OKForce a job to a successful outcome so its dependents proceed. Does not stop a running job — see above.
Mark FailedForce a job to a failed outcome so its dependents react accordingly. Does not stop a running job — see above.
note
A job retrying after a failure reads Wait start time, not Failed

A job whose frequency configures failure retries does not go to Failed while it still has attempts left. It waits in Wait start time with the failed attempt's exit code and termination description still showing, and its dependents keep waiting — so nothing downstream reacts until the job is genuinely out of budget. The job's Summary tab shows Retry Count as used-of-maximum.

If you do not want the attempts it has left, Restart or Cancel the job rather than waiting for them. A restart clears the pending attempt, and returns the count to zero unless you clear Reset Retry Count. See Failure retries.

A cancel now stays cancelled. A job cancelled before it ever ran could previously be picked back up by the retry gate while it still had budget — so the cancel appeared to work and the job then ran anyway. A job that ran nothing is no longer eligible for a retry.

Killing a job on a legacy LSAM machine takes a moment

The kill is carried to the machine on the relay's next check-in, so allow around half a minute before the process actually stops. The job reads Running, to be terminated until then, and once the process stops the job reports Failed — there's no kill-specific outcome. A kill the machine can't act on yet is retried, not dropped, so you don't need to re-issue it.

Earlier builds never carried the request down to the machine at all: the job sat at Running, to be terminated and the process ran to completion. If you have a runbook step that works around that, it can go.

Holding a workflow holds everything nested beneath it

Hold on a workflow that contains container jobs holds every workflow below it, at every depth — each one reads Parent hold and none of their jobs starts. Releasing the one you held resumes the whole tree; you don't need to release each level.

Two things to expect. A workflow you held directly keeps reading On hold, not Parent hold, even while something above it is also held — so releasing the ancestor does not release it, and you have to release it yourself. And a held workflow still finishes work already running; a hold stops new jobs starting, it does not freeze the run mid-job.

Earlier builds only reached one level down, so holding the top of a three-level design stopped the middle and left the bottom running. If you have a runbook step that holds each level in turn, it can go.

A hold no longer costs you the jobs you held

While the hold is on, waiting jobs are not ended for missing their start deadline. Holding a workflow to investigate used to do exactly that: a job whose Latest Start passed during the hold went to Missed start time — terminal — and releasing the hold did not bring it back. So the hold you placed to buy time destroyed the run you were protecting. It doesn't now, at any depth.

The deadline does not move, though. It is measured from the schedule date, so a job whose Latest Start passed while you were holding is ended on the first pass after you release. The hold defers the decision rather than extending the window — if you need the job to run, have its Latest Start pushed out before you release, or Force Start it. Late to start is still raised under a hold, and a job already running is unaffected.

If an action says it's already in flight

Occasionally you'll act on a job at the same instant the platform is starting it, and the action reports that the operation is in flight. Nothing is wrong and nothing was half-applied — the job is momentarily busy. Wait a second and try again.

You're most likely to see this on Kill for a job that still reads as waiting to start. It's also why cancelling a whole workflow leaves a job alone if that job has just reached its own outcome: the cancel doesn't overwrite a result that already arrived.

Related topics