Agents and agent pools
Task walkthrough: Manage agents and agent pools. This page is the full configuration and troubleshooting reference.
An agent is the software that runs jobs on a machine. Agent pools group agents so jobs can be routed to whichever agent is available. The Universal Agent runs the built-in and downloadable connectors; legacy LSAM machines are reached through the relay.
Agent pool
| Field | Notes |
|---|---|
name | Unique; cannot be changed after creation. Up to 255 characters, lowercase letters, digits and - only — a pool name is an identifier, so it is one of the naming exceptions. |
description | Optional, editable. |
Pool views also show Queued Jobs and Running Jobs counts. Neither is a live measure of a Universal Agent pool's work:
- Queued Jobs counts only jobs whose work never reached the pool's queue — a job sitting on the queue waiting for an agent is not counted.
- Running Jobs reads 0, because the platform does not record how many jobs each Universal Agent is running.
To see what a pool is doing, look at the jobs themselves in the runtime views.
Agent
| Field | Values / notes |
|---|---|
name, hostname | Identity. |
os | windows, linux, darwin. |
architecture | x64, arm64. |
agentVersion | Installed agent version. |
agentType | UNIVERSAL, plus legacy WINDOWS / UNIX / IBM_i / SQL / MFT (displayed as OpCon MFT) / RPA (displayed as OpCon RPA) / EASE; other values (MCP, SAP, ZOS) exist in the type list. |
status | Observed connectivity: ONLINE, OFFLINE, UNKNOWN. UNKNOWN = relay-down (reachability can't be determined because the agent's relay went stale) — distinct from OFFLINE = agent-down. Shown amber, with a text label. Applies to legacy LSAM agents reached via a relay. BUSY and DRAINING are also in the status list, but nothing in this build sets either, so an agent never shows them. |
operatorStatus | Operator state — what you asserted, independent of what the platform observes: ACTIVE (the default), MARKED_OFFLINE, MARKED_DRAINING. See Taking an agent out of service. |
effectiveStatus | The single value the UI and reports show, derived from the two above. An operator mark always wins — there is no case in which observed status overrides one. |
updateStatus | UP_TO_DATE, UPDATE_AVAILABLE, UPDATING, UPDATE_FAILED. |
lastHeartbeat | Last check-in. A legacy agent shows Never until its relay first reports it. A Universal Agent never shows Never: registering it records the registration time as its last heartbeat and its status as ONLINE, before the agent program has ever run. If the program never starts, the agent reads OFFLINE once that time is about 90 seconds old. |
currentJobCount / maxConcurrentJobs | Load vs. its concurrency cap. For a Universal Agent, currentJobCount always reads 0, and the cap you edit here is saved but never sent to the agent. The agent runs up to the number set in its own MAX_CONCURRENT_JOBS environment variable — 5 if it is not set — whatever this field says. See Registering a Universal Agent. |
capabilities.runtimes | Detected runtimes: node, python, dotnet, java, powershell, docker. |
defaultEventEnvironmentId | Legacy agents only. The one environment that events this machine raises are applied to. Optional in the schema, but while it is unset every event that machine raises is refused — see Events raised by a legacy agent. Set it on the agent's register/edit dialog under Agents → Legacy Agents & Groups; clearing it back to unset stops the channel. |
ftEndpoint | Legacy agents only — null for a Universal Agent, and not offered for an OpCon MFT, OpCon RPA or EASE agent. The machine's file-transfer address, ports, role, and TLS switches. See File-transfer endpoint. |
hasApiToken | OpCon MFT and OpCon RPA agents only. Whether an API token is stored for the agent. The token itself is never returned by any read. See OpCon MFT Transfer job and OpCon RPA Task job. |
easeConfig | EASE agents only — null for every other type. The non-secret connector settings: customerId, easeScheduleName, easeUser, localEventUser, timeZone, localScheduleName and debug. An update replaces it wholesale, so send every field you want to keep. See EASE jobs. |
hasEasePassword / hasLocalEventToken | EASE agents only. Whether each of the two EASE secrets is stored. Neither value is ever returned by any read. |
The agent detail view also lists recent commands. Stub (not populated yet): CPU cores, memory,
free disk, and the installed-plugins section show —.
Registering a Universal Agent
New Agent on an agent pool opens an Install New Agent dialog with install commands for
Windows, Linux and macOS. Those commands do not work in this build: the agent program has no
configure or install command, and it does not register itself with the pool when it first
connects.
A Universal Agent is registered through the API, and the agent program then signs in with the credential that registration returns:
- Register the agent. An authenticated call to the agent registration endpoint,
POST /api/v1/agents/register(behind the environment's/agent-apipath prefix in the cloud), names the agent and the pool it joins and describes the machine —name,agentPoolId,hostname,os,architecture,agentVersion. Agent names are unique in the tenant. The response returns the new agent and aclientIdandclientSecret. The secret is returned only in that response — it is stored hashed and cannot be read back. - Write the credentials file on the agent machine. At startup the agent reads a JSON file with
exactly four keys —
agentId,agentPoolId,clientId,clientSecret— and refuses a file with any other key. The default location is~/.opcon-agent/credentials.jsonfor the account the agent runs as;CREDENTIALS_FILEchanges it. Restrict the file to that account — the agent logs a warning at startup when the file's mode is anything other than0600. - Point the agent at the platform and start it. Set
AGENT_SERVICE_URL, andAGENT_SERVICE_PATH_PREFIX(for example/agent-api) when the platform is reached through the cloud edge, then runopcon-agent start.
An agent with no credentials file stops at startup with Agent not registered. The program's only
other commands are opcon-agent version and opcon-agent status, which reports the agent and pool
IDs from the credentials file.
| Agent environment variable | Default | What it sets |
|---|---|---|
AGENT_SERVICE_URL | — | The platform's agent service. |
AGENT_SERVICE_PATH_PREFIX | empty | Path prefix for the cloud edge, for example /agent-api. |
CREDENTIALS_FILE | ~/.opcon-agent/credentials.json | Where the credentials file is read from. |
MAX_CONCURRENT_JOBS | 5 | How many jobs the agent runs at once. This, not the Max Concurrent Jobs value on the agent's row, is the cap a Universal Agent obeys. |
HEARTBEAT_INTERVAL_MS | 30000 | How often the agent checks in. |
How a Universal Agent takes and runs work
These apply to every job type a Universal Agent runs — Run Command, Run Script and Wait for File:
- Every job has a fixed one-hour limit. The platform sends every job with a timeout of 3,600 seconds, and there is no job setting that changes it. At the hour the agent stops the process (a graceful stop, then a forced stop five seconds later) and the job fails. A Wait for File job stops waiting at the hour.
- The agent takes a batch of work, then waits for all of it. Each time it polls, the agent takes as many jobs as it has free slots for, up to 10, and does not poll again until every job in that batch has finished. One long-running job therefore keeps the agent from picking up new work, even when it has free slots.
- Kill does not reach a Universal Agent. The job's status records the kill request, but the process on the agent keeps running until it ends on its own or reaches the one-hour limit. Kill stops work only on legacy agents, through their relay.
OpCon MFT agents
An OpCon MFT agent is registered under a relay, on the same page and with the same dialog as a
legacy LSAM machine, but it is not an LSAM. Three fields on the dialog differ: the host and port
are labelled Host and HTTPS Port (default 41100), there is no JORS Port because there
is no JORS listener, and there is an API Token. It also has no file-transfer endpoint, so it can
be neither end of a SMAFT transfer.
On the agent list an MFT row shows OpCon MFT as its type, its HTTPS port in the Port column, and the version the agent reports. Its actions menu carries one extra item, Reset API token.
Full detail — registering one, what the token is for, and what a reset invalidates — is on OpCon MFT Transfer job.
OpCon RPA agents
An OpCon RPA agent is registered the same way and is likewise not an LSAM. The dialog
differs in the same three places — Host and HTTPS Port (default 7047), no JORS Port,
and an API Token — and it too has no file-transfer endpoint, so it cannot take part in a
SMAFT transfer.
Two things differ from an OpCon MFT agent:
- The token must be a GUID, and it is what makes the agent usable at all: an RPA agent with no token stored reports DOWN, because the relay has nothing to authenticate with.
- There is no Reset API token action. Continuum cannot mint an RPA token — it is configured on the agent itself, and Continuum only stores a copy. The edit dialog offers Clear token to remove the stored copy instead.
On the agent list an RPA row shows OpCon RPA as its type, its HTTPS port in the Port column, and the version the agent reports. Its health is checked directly against the agent roughly every 10 seconds, and is deliberately not tied to whether the agent's own robot clients are connected.
Full detail — registering one and what the token is for — is on OpCon RPA Task job.
EASE agents
An EASE agent is the third kind registered under a relay without being an LSAM, and it is the one that is not a machine of yours: it is Jack Henry's hosted EASE OpCon, reached over its REST API. It has no JORS port and no file-transfer endpoint, so it cannot take part in a SMAFT transfer.
Two things make its dialog different from the other two:
- There is no API Token. An EASE agent authenticates with its own credentials — an EASE User and EASE Password — stored in an EASE Datacenter group on the form, alongside a Customer Id, an EASE Schedule Name and an optional Time Zone. A second group, Local (Continuum) events, holds the event user and token the SEQ property write-back uses. Both secrets are write-only.
- Its Default Event Environment means something narrower. An EASE agent has no MSGIN directory; the setting governs the one event it raises, the SEQ property write-back.
On the agent list an EASE row shows EASE as its type, its HTTPS port (default 443) in the
Port column, and the EASE OpCon's own REST API version. Its health is probed directly against
the EASE OpCon roughly every 10 seconds, and an agent with no Customer Id or no EASE Schedule
Name reports DOWN without a call being made. The probe sends no credentials, so a wrong
EASE Password leaves the agent UP and fails its jobs at start.
Full detail — registering one, every field and its limits — is on EASE jobs.
File-transfer endpoint (legacy agents)
A SMAFT File Transfer job connects one legacy machine directly to another. That is a different network path from the one the relay uses, and a different set of questions — which direction a machine may take part in, which port to dial, whether TLS is possible. Those answers live on the agent, as its file-transfer endpoint.
You configure it on an existing legacy agent: Agents → Legacy Agents & Groups, open the agent, then the File Transfer (SMAFT) section. It is not offered while registering a new agent — register the machine first, then set its endpoint.
The agent does not report these values and they are not read from your Classic configuration. An administrator sets them, which is also how Classic works.
| Field | Default when blank | Notes |
|---|---|---|
| FT IP Address | the agent's LSAM host | The address the other agent connects to for a transfer. Often different from the LSAM host, which only the relay uses. A hostname or IPv4 address only — no port, scheme, path, or spaces. |
| FT Port (non-TLS) | 3108 on UNIX, 3110 on Windows, none on SQL | 1–65535. |
| FT Port (TLS) | unset — there is no platform default | 1–65535. A transfer that asks for TLS cannot resolve a port until this is set. |
| FT Role | None - never a file-transfer participant | Source only, Destination only, Both source and destination, or None - never a file-transfer participant. Decides which machine pickers the agent appears in. File transfer is opt-in — see below. |
| Supports non-TLS transfers | on | Both ends must permit a mode for a transfer to use it. |
| Supports TLS transfers | off |
Full file-transfer support, which is what a push (Start Transfer On = Source) requires on
both machines, is part of the endpoint and defaults to enabled. It has no control in this
build's interface — it is reachable only through the API — so in practice every registered legacy
machine can take part in a push until one is changed deliberately.
Inherited or overridden
Every field is empty when nothing is stored, and the hint under it reads the effective value —
Inherited - currently 3108 (default for this platform) — rather than pre-filling it. That is
deliberate: a pre-filled form would freeze today's defaults in as explicit overrides the first time
you saved it.
Each hint states the value the way you would read it in the list beside it, not the way the
protocol carries it — so the role's hint reads
Inherited - currently None - never a file-transfer participant (default for this platform).
Once you store a value the hint reads Override - stored as … instead. Clearing the address, a
port, or the role empties the field and returns it to inheriting the default. The two TLS switches
cannot be returned to inherited from the interface — a switch has no "unset" position to show — so
once toggled they stay stored.
FT IP Address is reached agent-to-agent, not by the relay, so an address that works for the relay is not evidence it works for a transfer. If a transfer fails to connect while both machines report online, check this field before anything else.
FT Role inherits None when nothing is stored, so a machine only becomes eligible for a transfer once an administrator gives it a role. That includes agents registered before the file-transfer endpoint existed: no role was backfilled for them, so each one participates in nothing until you set one.
A job that names a machine with no role saves normally and fails at dispatch, naming the machine and telling you which role to set. If a machine you expect is missing from the Source Machine or Destination Machine list, its role is the first thing to check.
A machine whose role is None never appears in either picker and can never take part in a transfer. A Universal Agent has no endpoint at all and is never offered.
A SQL machine has no file-transfer capability either. It is the one legacy platform with no default non-TLS port, so it resolves to none rather than inheriting the UNIX 3108, and it never appears in a transfer's machine pickers. A port stored on one deliberately is still honoured.
Taking an agent out of service
An agent's status reports what the platform observes — whether the machine is reachable and
answering. That is not the same question as whether you want work to go there. Patching a host,
draining it before a reboot, or holding a misbehaving machine back are all decisions the platform
cannot infer from a heartbeat.
So an agent carries a second, separate value: its operator state. You set it, only you clear it, and a heartbeat never overwrites it. Marking an agent that later reconnects leaves it marked.
| Action | Operator state | What it does |
|---|---|---|
| Mark Offline | MARKED_OFFLINE | Treats the agent as out of service. No new work is dispatched to it and it is not selected as a pool or group member. |
| Mark Draining | MARKED_DRAINING | Stops new work reaching the agent. Jobs already running finish normally. Use before planned maintenance. |
| Clear Mark | ACTIVE | Returns the agent to service. Its status reverts to observed connectivity. |
Where to find them. The row menu on an agent list, the action menu on an agent's detail page, and the row menu on the Legacy Agents & Groups page — legacy relay-attached agents can be marked the same way. All three actions are always listed; the ones that would not change anything for the agent's current state are shown disabled rather than hidden.
A Reason is optional and, when supplied, is stored with the mark and shown in the status tooltip alongside who set it and when. If the agent is currently running jobs, the confirmation warns you and tells you how many — it warns, it does not block, because draining a host for maintenance is exactly the case where work is still on it.
What a mark changes
An agent is eligible for new work only when it is observed ONLINE and its operator state is
ACTIVE. Everything below follows from that one rule.
- Nothing new is dispatched. A marked agent is skipped when a pool or a legacy group picks a target, and the paths that hand work to an agent after selection are blocked too.
- Running jobs are left alone. A mark stops the next job; it never interrupts one already underway.
- A job targeting a marked agent holds. It waits in
WAIT_MACHINErather than failing, and proceeds once the mark is cleared. Unlike the observed-status gate, this one does not fail open on a stale relay feed — your intent is trusted whether or not the feed is fresh. - A fully drained legacy group defers rather than fails. When a group's jobs could have
dispatched but for the marks, they hold in
WAIT_MACHINEinstead of dyingINITIALIZATION_ERRORand releasing their dependents as failures. Marks are no longer the only case that holds — see When a legacy group has no dispatch target. - Force start is refused. Force-starting a job onto a marked agent, or onto a legacy group with no unmarked member available, is rejected with an error naming the agent and its mark reason. There is no override; clear the mark instead.
- Pool counts follow. A pool's online count means dispatchable, so a drained pool no longer reads "5 of 5 online" while accepting no work. A marked agent counts toward the pool's offline figure. Legacy groups report a dispatchable count alongside their observed online count.
A mark applies to work already being polled for within about five seconds, not instantly. Clearing one takes effect on the same window. A job dispatched in that gap runs normally.
How a marked agent is displayed
The status chip shows the mark, not the observed status, and composes the two only when they
disagree — a marked agent that is still reachable reads Marked Draining, while one that has since
gone down reads Offline (Marked Draining). Hovering shows the reason, who set the mark, and when.
Filtering keeps the two ideas apart, deliberately:
- The Operator state filter on an agent list selects
All,Active,Marked Draining, orMarked Offline. - The Status filter keeps its observed-only meaning. A
MARKED_DRAININGagent is usually still connected, so it does match aStatus is ONLINEfilter. If you want "agents that can actually take work," filter on operator state, not status.
Marks also reach the rest of the product: the Agents dashboard widget counts them as their own slices rather than folding them into Online, the Agents report can show and filter on them (see Reports), and they can raise a notification (see Notifications).
How quickly a stop is noticed
There are two ways an agent or relay can stop, and they are detected by completely different mechanisms — which is why one shows up almost immediately and the other takes a while.
A clean stop announces itself. A Universal Agent or a relay shut down normally reports itself offline as the first thing it does on the way down. That takes about a second: you can stop an agent and see it offline on the next refresh of the page.
An abrupt loss has to be inferred. If the process is killed, the host loses power, or the network drops, nothing announces anything. The platform notices only when heartbeats stop arriving, which takes:
| What was lost | Noticed after roughly | Result |
|---|---|---|
| A Universal Agent | 90 seconds — three missed heartbeats | The agent goes OFFLINE |
| A relay | 5 minutes | The relay goes offline and its legacy agents go UNKNOWN |
Five minutes is a deliberate debounce, not an oversight. A shorter window would turn a brief network
blip into five minutes of UNKNOWN legacy agents across the relay, and every queued legacy-group job
holding through it for nothing. The clean-stop signal means a real relay shutdown is still reflected
in seconds; the long window only applies when the relay vanishes without warning.
Either way, queued legacy-group work waits. The two paths leave the relay's legacy agents in different states, and the difference is now only how fast it is noticed:
- Clean relay stop → its legacy agents go
OFFLINE, about a second later. A legacy-group job holds inWAIT_MACHINEuntil the relay reports in again. - Abrupt relay loss → its legacy agents go
UNKNOWNafter about five minutes. A legacy-group job holds too, and resumes when the relay reconnects.
An abrupt loss used to be the worse of the two by far: the WAIT_MACHINE gate failed open on an
UNKNOWN feed, the job advanced, and a least-tasked group with no online member then failed at
dispatch — No ONLINE legacy-group member with a dispatch pool for job … — releasing its dependents
as failures. Stopping a relay cleanly is still worth doing, because a clean stop is reflected in
seconds rather than five minutes, but it is no longer the difference between work waiting and work
being destroyed.
Status lists refresh every 10 seconds while you have them open, so you do not have to reload the page to watch an agent come back.
After the platform itself restarts, it waits before judging anything stale — long enough to give every agent and relay a chance to check in. Nothing is marked offline just because the platform was the thing that was down.
How jobs are routed to agents
A job's agent assignment (set in the workflow) carries a kind (agent-pool or
legacy-group) and a mode:
| Assignment | Mode | Behavior |
|---|---|---|
| Specific agent | — | The primary agent. Secondary, tertiary and quaternary agents can be set and are saved, but dispatch uses only the primary — there is no failover. For a Universal Agent, see the note below. |
| Agent pool | least-tasked | One job placed on the pool's queue, run by whichever agent in the pool asks for work first. No agent is chosen by load. |
| Agent pool | run-all | The same as least-tasked: one job on the pool's queue, run by one agent. A Universal Agent pool job does not fan out to every agent. |
| Legacy agent group | least-tasked | An ONLINE, unmarked member of the group, picked at random — not by load. |
| Legacy agent group | run-on-all | Every member — fanned into one job instance per member, frozen at build. |
The platform does not pick a Universal Agent for a job. It puts the job on the queue of the job's pool, and the first agent in that pool to poll takes it. That is also true of a job assigned to a specific Universal Agent: it goes on the queue of that agent's pool, so any agent in the same pool can run it. To pin work to one machine, give that agent a pool of its own.
Legacy LSAM job types dispatch through a legacy agent group (or a specific legacy agent), not a Universal-Agent pool — the assignment selector only offers groups compatible with the job type, and a group holds agents of a single type. See Job configuration. Whichever route a job takes, the platform only places it on an agent whose type/OS matches the job type's requirements.
Group dispatch gating. A least-tasked group job advances only when the group has a member
dispatch could actually route to: one that is observed ONLINE, unmarked, and has a resolved
relay dispatch pool. The gate applies the same test dispatch does, so the two can never
disagree — which is what stopped a job passing the gate and then dying at dispatch. Otherwise the job
holds in WAIT_MACHINE, and this gate does not fail open on a stale or UNKNOWN relay feed
the way specific-agent gating does. The one thing it still fails open on is being unable to read the
group at all, so an unreachable agent-service cannot deadlock the scheduler.
A run-on-all group job resolves its members to a frozen snapshot at build and reuses the multi-instance loop (one instance per member); its children are pinned to a member and take the specific-agent gating above.
When a legacy group has no dispatch target
A least-tasked group job is released from WAIT_MACHINE only when the group has a member dispatch
can route to, so in every one of these cases the job holds and starts on its own once the group
has one:
| The group | What happens |
|---|---|
| Has no members | Holds — a member may still be added |
Has members, all observed UNKNOWN (its relay was lost abruptly) | Holds — status will come back when the relay does |
Has an ONLINE, unmarked member whose relay dispatch pool is not resolved yet | Holds — the pool resolves on its own |
| Is blocked only because its members are marked | Holds until a mark is cleared |
Has members that are all observed OFFLINE, with no mark | Holds until a member comes back online |
A job holding this way is tracked as a long hold like any other, so it shows up as waiting rather than disappearing.
An all-OFFLINE group fails a job INITIALIZATION_ERROR — No ONLINE legacy-group member with a
dispatch pool for job … — in only two situations: the job was force-started, which skips the
WAIT_MACHINE gate, or the group's last usable member went offline or was marked in the moment
between the gate releasing the job and dispatch.
Force start is refused for the first four cases. Force start bypasses WAIT_MACHINE gating, so
permitting it left the forced job deferring silently at the dispatcher with nothing to say why. It is
now rejected up front with the reason:
| The group | The refusal says |
|---|---|
| No members | the target legacy agent group has no members. Add a member to proceed. |
All UNKNOWN | every member of the target legacy agent group has UNKNOWN status (relay unreachable). Wait for the relay to reconnect. |
| No member with a dispatch pool | no ONLINE member of the target legacy agent group has a dispatch pool. |
| Blocked only by marks | no unmarked member of the target legacy agent group is available. Clear a mark to proceed. |
Force-starting onto an all-OFFLINE group is allowed, and fails the job at dispatch with
INITIALIZATION_ERROR.
- Heartbeat "Never" on a legacy agent means its relay hasn't reported it yet — a common first thing to check for "no agent available." A Universal Agent shows its registration time instead, so a recent heartbeat on a new Universal Agent is not proof the program is running.
- To take an agent out of service, use the operator mark (Mark Draining or Mark Offline).
The observed
DRAININGandBUSYstatuses are never set in this build, and there is no observed disabled state —DISABLEDwas removed from the status list because nothing ever set it. UNKNOWNvsOFFLINEfor legacy agents:UNKNOWN(amber) means the relay is stale, so a legacy LSAM agent's reachability can't be confirmed — not that the agent is down.OFFLINEmeans the agent itself is down. When a relay recovers, its agents leaveUNKNOWNon the next heartbeat. How fast a stop is noticed depends on how it stopped — see How quickly a stop is noticed.- Legacy agents and groups are managed in-product on the Agents page, on its Legacy Agents & Groups tab (a sibling of the Agent Pools tab; the two former navigation entries were merged into a single Agents item). Register/edit/delete a legacy agent there; the relay hot-reconciles its LSAM connections within one heartbeat — no relay restart needed. Agent names are unique per relay, and a legacy agent's name must be 128 characters or fewer: the platform accepts up to 255, but the relay refuses to connect to an agent with a longer name, so it never comes online.
- Legacy agent groups are type-homogeneous (all members share one legacy type) and may span relays. A group is an operational dispatch target for a legacy job (least-tasked or run-on-all; see the routing table). Deleting a group soft-deletes the group only — never its member agents or their queues.
- CPU/memory/disk and installed-plugins on the agent detail are not implemented yet — the
—is expected, not missing data.
Troubleshooting
| Symptom | Likely cause | Resolution |
|---|---|---|
Job can't be assigned / WAIT_MACHINE | No matching agent ONLINE (type/OS mismatch, all offline, or every candidate marked) | Check agent status/type vs the job's requirements (Administrator). |
Specific-agent job sits in WAIT_MACHINE | The job targets one specific agent that is OFFLINE on a fresh status feed — it holds rather than fail, and proceeds when that agent is ONLINE again | Bring the target agent online, or reassign the job (Administrator/Builder). |
| A job assigned to one Universal Agent ran on a different one | The job went on the queue of that agent's pool, and another agent in the pool took it first | Give the agent a pool of its own if the work must run there (Administrator/Builder). |
| A run-all pool job ran on only one agent | A Universal Agent pool job is one job on the pool's queue, whichever mode is set | Add one job per agent, each assigned to an agent in its own pool (Builder). |
| A new Universal Agent never comes online | It was never registered through the API, or its credentials file is missing or has extra keys — the Install New Agent commands do not register it | Register it and write its credentials file — see Registering a Universal Agent (Administrator). |
| Specific-agent job advances despite an offline-looking target | By design: when the status feed is stale or UNKNOWN, the gate fails open so the scheduler never deadlocks on untrusted status. | Confirm the target agent/relay is actually reachable; investigate the stale feed (Administrator). |
Least-tasked group job sits in WAIT_MACHINE | The group has no member dispatch could route to — none ONLINE and unmarked with a resolved dispatch pool. It holds and proceeds when one appears. An all-OFFLINE, all-UNKNOWN (relay-lost) or empty group holds here too, rather than failing | Bring a group member online, or check the relay — see When a legacy group has no dispatch target (Administrator). |
Least-tasked group job fails INITIALIZATION_ERROR | It was force-started onto a group with no member online, or the group's last usable member dropped in the moment between the gate releasing the job and dispatch | Bring a member online and add the job again (Operator/Administrator). |
| Agent shows OFFLINE / heartbeat stale | Connectivity or the agent process is down | Check the agent host and network (Administrator). |
| Agent looks healthy but takes no work | It's marked — the chip reads Marked Offline or Marked Draining. A mark survives a reconnect, so one set during past maintenance is still in force. | Hover the chip for the reason and who set it, then Clear Mark (Administrator). |
| Force start rejected, naming the agent | The target agent is marked, or every usable member of the target legacy group is | Clear the mark, or retarget the job. There is no force-start override (Operator/Administrator). |
| Force start rejected, naming the legacy group | The group is empty, entirely UNKNOWN, or has no ONLINE member with a dispatch pool. The refusal names which | Fix what it names — add a member, wait for the relay — then force start again (Operator/Administrator). |
| Pool reads fewer online than you expect | Pool online counts only agents that can actually take work, so marked members are excluded | Check the pool's agents for marks (Administrator). |
Legacy agent shows UNKNOWN (amber) | Its relay is stale/down — agent reachability can't be determined (not the same as the agent being down) | Check the relay; on recovery, agents leave UNKNOWN within a heartbeat (Administrator). |
| Jobs pile up on one agent | The agent's concurrency cap is reached. On a Universal Agent the cap is its MAX_CONCURRENT_JOBS setting, not the value on its row | On a Universal Agent, raise MAX_CONCURRENT_JOBS in the agent's environment and restart it, or add agents to the pool (Administrator). |
| A Universal Agent with free slots takes no new work | It is still running a job from its last batch, and it does not poll again until the whole batch has finished | Expected in this build. Keep long-running jobs on agents of their own (Administrator / Builder). |
| A Universal Agent job fails after exactly one hour | Every Universal Agent job has a fixed one-hour limit | Split the work so each job finishes within the hour (Builder). |
| Kill on a Universal Agent job doesn't stop it | Kill reaches legacy agents only | The job ends when its process ends, or at the one-hour limit. Stop the process on the agent machine if it must end sooner (Administrator). |