Manage agents and agent pools
Agents run your jobs; agent pools group agents so work can be routed to whichever one is available. As an administrator you keep agents healthy and organized so workflows always have somewhere to run.
If workflows target a single named machine, one agent going down or filling up stalls the work. No one has a clear view of which agents are healthy or overloaded.
Agent pools
Group related agents into a pool, then let workflows target the pool instead of a single machine. Create a pool with a name (the name can't be changed later) and a description. The pool view also shows queued and running job counts, but don't rely on them for a Universal Agent pool: Queued Jobs counts only work that never reached the pool's queue, and Running Jobs always reads 0. Check the jobs themselves in the runtime views instead.
A pool doesn't pick an agent for a job. The job goes on the pool's queue and the first agent to ask for work takes it — see How workflows pick an agent.
Add a Universal Agent
Don't use the commands in the Install New Agent dialog that New Agent opens on a pool. They
don't work in this build: the agent program has no configure or install command, and it doesn't
register itself when it first connects.
A Universal Agent is registered through the API, which returns the credential the agent signs in with. You then put that credential in a file on the agent machine, point the agent at the platform and start it. The steps, the file format and the agent's settings are in Registering a Universal Agent.
Read an agent's state
Each agent shows the essentials you need at a glance:
- Status: Online, Offline, or Unknown. Unknown (shown amber) applies to a legacy agent whose relay can't be reached, so its state can't be confirmed. This is different from Offline, which means the agent itself is down. If someone has deliberately taken the agent out of service, the status reads Marked Offline or Marked Draining instead — hover it for the reason, who set it, and when. (Busy and Draining are in the status list, but nothing sets them in this build.)
- Last heartbeat: when it last checked in. On a legacy agent, "Never" means its relay hasn't reported it yet. A Universal Agent shows its registration time until the program first checks in, so a recent time on a new agent doesn't prove it's running.
- Type and OS: e.g. a Universal Agent on Windows or Linux.
- Load: current jobs vs. its maximum concurrent jobs. On a Universal Agent the current count
always reads 0, and the maximum you edit here isn't sent to the agent — it runs up to its own
MAX_CONCURRENT_JOBSsetting (5 unless changed). Change that on the agent machine instead. - Update status: whether the agent version is up to date or an update is available.
- Runtimes: detected tools (Node, Python, .NET, Java, PowerShell, Docker), useful for the Run Script job.
Take an agent out of service for maintenance
Patching a host, rebooting it, or holding a misbehaving machine back are all decisions the platform can't infer from a heartbeat — so you tell it directly. From an agent's row menu, its detail page, or the Legacy Agents & Groups page:
- Choose Mark Draining to stop new work reaching the agent while jobs already on it finish, or Mark Offline to treat it as out of service entirely.
- Add a Reason — optional, but it's what the next person sees when they hover the status. If jobs are currently running, you're told how many. That's a warning, not a block.
- When the work is done, choose Clear Mark. The agent's status goes back to reporting its actual connectivity.
A mark is yours, not the platform's: reconnecting doesn't clear it, and a heartbeat never overwrites it. While a mark is in place:
- No new jobs are sent to the agent, and it isn't picked as a pool or group member.
- A job that targets it waits rather than failing, and runs once you clear the mark.
- Force start is refused for a job aimed at a marked agent, naming the agent and your reason. Clear the mark instead — there's no override. Force start is also refused for a job aimed at a legacy agent group that has nothing to dispatch to at all — an empty group, one whose members are all Unknown, or one with no online member holding a relay dispatch pool — and the refusal names which of those it is.
- Pool and dashboard counts stop calling it available, so a fully drained pool no longer reads as fully online.
Marks take effect within a few seconds rather than instantly, so a job dispatched as you set one still runs. To find every agent currently marked, use the Operator state filter on an agent list or the Agents report. Filtering on Status won't do it — a marked agent that's still connected is genuinely still Online.
You can also be told when someone marks an agent: agent notification groups offer Marked Offline, Marked Draining, and Mark Cleared triggers (Manage notifications).
How workflows pick an agent
When a Builder designs a job, they choose a specific agent or an agent pool. A legacy LSAM job can instead target a legacy agent group. A job is only sent to an agent whose type and OS match the work. A job pinned to a specific agent (or to a group with no member online) waits rather than failing, so check the target if such a job sits waiting.
What actually happens to the job:
- Agent pool, whichever mode is chosen: the job goes on the pool's queue once, and the first Universal Agent in the pool to ask for work runs it. Nothing picks the least-busy agent, and run on all does not run the job on every agent.
- Specific Universal Agent: the job goes on the queue of that agent's pool, so any agent in the same pool can take it. If the work must run on one machine, give that agent a pool of its own.
- Backup agents (secondary, tertiary, quaternary) are saved with the job but never used — there is no failover to them.
- Legacy agent group: least-tasked picks an online member at random; run on all makes one job per member.
A few things about how a Universal Agent runs work are worth knowing before they surprise you: every job stops at one hour, the agent doesn't take new work until the whole batch it last picked up has finished, and Kill doesn't reach it. See How a Universal Agent takes and runs work.
A legacy agent group holds one agent type, so adding an agent of any other type to it is refused — including an OpCon MFT agent to an OpCon RPA or EASE group, and any other pairing of the three. The REST-reached types are as distinct from one another as any of them is from a Windows LSAM.
Manage legacy agents and groups
Legacy LSAM agents are registered on the Agents page, on its Legacy Agents & Groups tab (alongside the Agent Pools tab): add, edit, or remove a legacy agent there. OpCon MFT, OpCon RPA and EASE agents are registered on the same tab with the same dialog, even though none of them is an LSAM — each is reached over its own HTTPS REST API rather than the LSAM wire, and each needs a relay that has reported it can serve that type. Changes take effect automatically: the connecting relay picks them up within about a heartbeat, with no restart. Give each legacy agent a name that's unique for its relay and 128 characters or fewer — the platform accepts a longer name, but the relay won't connect to that agent, so it never comes online.
Set each legacy agent's default event environment
A legacy agent's register/edit dialog carries a Default event environment. It has nothing to do with where the agent's jobs run — it is the one environment that events the machine itself raises are applied to, when a job or process on it writes an event file into the agent's MSGIN directory.
While it is unset, every event that machine raises is refused, and the machine deletes the event file before the platform sees it — so the event is lost, and a refusal at this stage doesn't appear in the event log either. Clearing the field back to unset has the same effect, which is why the dialog says so at the field.
It is per machine: one machine sends everything to one environment, and there is no per-event override. A machine shared across environments still needs one chosen. See Events raised by a legacy agent.
Give a legacy agent a file-transfer endpoint
A SMAFT File Transfer job connects one legacy machine directly to another, so every machine that takes part needs a file-transfer endpoint. Open the agent on the Legacy Agents & Groups tab and find the File Transfer (SMAFT) section. It appears on an existing agent only — register the machine first, then set its endpoint.
- Set FT IP Address to the address the other agent uses to reach this machine. Blank inherits the agent's LSAM host.
- Set FT Port (non-TLS) if the machine does not use the platform default — 3108 on UNIX, 3110 on Windows. Set FT Port (TLS) only if you intend to use TLS; there is no default for it.
- Set FT Role to decide which transfers this machine may take part in: source only, destination
only, both, or none. Blank inherits
None, so this step is not optional — a machine with no role set appears in neither picker and takes part in no transfer. - Leave the two Supports … transfers switches alone unless the machine differs from the defaults — non-TLS on, TLS off.
Leave a field blank to inherit its default. The hint under each field tells you which you are
looking at: Inherited - currently 3108 (default for this platform) or Override - stored as ….
Clearing the address, either port, or the role returns it to inheriting. The two switches cannot be
returned to inherited once toggled.
The relay reaches an agent on one address; for a transfer, the other agent reaches it directly, which is often a different route entirely. An address that works for the relay is no evidence it works for a transfer — so when a transfer fails to connect while both machines report online, check this field first.
No role was backfilled onto agents that were registered before the file-transfer endpoint existed, so each of them inherits None and takes part in nothing until you set a role. A job that already names such a machine still saves, then fails at dispatch naming the machine and the role to set.
If a machine is missing from a transfer's Source Machine or Destination Machine list, check its role before anything else.
For the full field reference, the defaults, and what a role of None excludes, see File-transfer endpoint.
On the same page you can gather legacy agents into a legacy agent group. Every member of a group must be the same legacy type (for example, all UNIX), and a group can span more than one relay. A Builder can then point a legacy job at the group instead of a single machine, to run on one online member (picked at random, not by load) or on every member at once. Removing a group leaves its member agents and their work untouched.
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
| A job is waiting on a machine | No matching agent is online (wrong type/OS, or all offline/draining). A job pinned to a specific agent waits for that agent. | Bring a matching agent online; check the job's required agent type, or reassign the job. |
| An agent shows offline / stale heartbeat | The agent host or its connection is down | Check the agent machine and network. |
| A legacy agent shows Unknown (amber) | Its relay can't be reached, so the agent's state is uncertain, not necessarily down | Check the relay; the agent leaves Unknown on its own once the relay reconnects. |
| A legacy job targeting a group waits | The group has no member the platform can dispatch to, and the job holds rather than failing: no members yet, every member Offline, members all Unknown because the relay was lost, an online member whose relay dispatch pool hasn't resolved, or every usable member marked | Bring a group member online, add one, clear a mark, or check the relay if members show Unknown. |
| A legacy job targeting a group fails to initialize | It was force-started onto a group with no member online, or the last usable member dropped just as the job was released | Bring a member online, then restart or re-add the job. |
| An agent looks healthy but takes no work | It's marked. A mark survives a reconnect, so one set during earlier maintenance is still in force. | Hover the status for the reason, then Clear Mark. |
| Work piles up on one agent | Its concurrency cap is reached — on a Universal Agent, its own MAX_CONCURRENT_JOBS setting | Raise MAX_CONCURRENT_JOBS on the agent machine and restart the agent, or add agents to the pool. |
| A new Universal Agent never comes online | It was installed with the Install New Agent commands, which don't register it | Register it through the API and give it its credentials file — see Registering a Universal Agent. |
| A job for one Universal Agent ran on another | It went on the queue of that agent's pool, and another agent in the pool took it | Give the agent a pool of its own. |
Some agent details (CPU, memory, disk, installed plugins) aren't populated yet and show "—".
Related topics
- Agents and agent pools — full configuration reference and troubleshooting
- Prepare agents for connectors (Administrator)
- Manage plugins (Administrator)