Relays
Task walkthrough: Set up and manage relays. This page is the full configuration and troubleshooting reference.
A relay is the on-prem software that lets legacy LSAM machines run jobs dispatched from OpCon
Continuum in the cloud. It connects outbound only to the OpCon agent-service (no inbound firewall
changes), authenticates with a downloaded credentials.json, and bridges dispatched work to the LSAM
machines behind it. Legacy agents are reached through a relay; a relay may front more than one
legacy agent and legacy agent groups may span relays.
The relay is a pure credentials consumer — its only inputs are a credentials.json file and a few
connection environment variables. Relay support covers registration and credential download,
credential rotation, artifact distribution, installation, and a /health readiness endpoint.
Where the relay sits
The two things this picture is here to settle:
- Nothing connects inward. The relay dials out to the agent service and polls for work, so running one needs no inbound firewall change.
- A Universal Agent does not need a relay. It talks to the platform directly. The relay exists only to reach machines that cannot. If a Universal Agent job fails, the relay is not involved — and if a legacy job can't be placed, the relay is the first thing to check.
Most machines behind a relay are legacy LSAM machines, reached over the LSAM wire protocol. Three kinds are not: an OpCon MFT agent, an OpCon RPA agent and an EASE agent are each reached over their own HTTPS REST API. All three still need a relay, and all three still appear under one on Legacy Agents & Groups, but none of them is an LSAM — see Serving OpCon MFT agents, Serving OpCon RPA agents and Serving EASE agents.
An EASE agent is the odd one out even among those three: it is not a machine on your network at all, but Jack Henry's hosted EASE OpCon, which the relay reaches out to over the internet.
Download the relay without registering one
Agents → Legacy Agents & Groups → Download Relay gets you the latest published relay build on its own. It is offered in the page's toolbar and again in the empty state when no relay exists yet — you generally want the software before you have anything to register.
The dialog shows the latest release's version and publish date, and a download button per platform. Downloading creates no relay and issues no credential.
The artifact used to be reachable only from the Register Relay success screen, so getting the binary meant registering a relay you may not want — creating a record and burning a one-time client secret. Worse, reopening that dialog resets its result, so a download that failed meant registering again. Neither is necessary now.
Two things to know:
- Only the platforms the release actually publishes are offered. A button appears for a given operating system and architecture only when the release carries that artifact, so you cannot be shown a download that would then fail.
- The artifact alone is not runnable. The installer does not set the connection environment, and the relay needs credentials. Use Register Relay for both — see below.
The Download Relay action is disabled rather than hidden when there is nothing to download, and its tooltip says which of four situations you are in:
| Tooltip | Meaning |
|---|---|
| Checking for the latest relay release… | The lookup is still in flight |
| No relay release is available. | Nothing is published for this environment |
| No relay build is available for a supported platform. | A release exists but publishes nothing for a platform this page offers |
| Could not check for the latest relay release. | The lookup itself failed — this is a fault to chase, not an absent release |
Register a relay
In the UI: Agents → Legacy Agents & Groups → Register Relay (backend POST /api/v1/relays/register,
tenant-scoped). On the success screen the dialog provides everything needed to stand up the relay:
| On the register dialog | Detail |
|---|---|
| Download credentials.json | Writes the strict { relayId, clientId, clientSecret } file the relay requires at startup. This is the only way to get the credential — see below. |
| Relay download links | The latest published release, served from the relay-releases manifest. The same per-platform buttons as Download Relay, so the two surfaces always offer the same builds. |
| Connection details | The runtime-derived AGENT_SERVICE_URL and the /agent-api path prefix for the environment, each with a copy button. |
| Install the relay | Per-platform Linux and Windows command sequences, pre-filled with this environment's URL and the release's version, and a pointer to the README.md inside the extracted package. |
The Client ID and Client Secret are never displayed and cannot be copied to the
clipboard. They exist only inside the credentials.json you download, and the secret cannot be
retrieved later. Download the file before closing the dialog; if you lose it,
regenerate the credentials rather than re-registering the relay.
credentials.json must match the exact three-key shape — a strict schema is enforced at boot, so
an extra key such as the agent-only agentPoolId causes a parse failure. The relay package ships a
shape reference at config/credentials.json.example; it is not a working credential.
Connection environment variables
| Variable | Required | Default | Purpose |
|---|---|---|---|
AGENT_SERVICE_URL | ✅ | — | Base URL of the agent-service for the environment (from the register dialog). |
AGENT_SERVICE_PATH_PREFIX | ⚠️ cloud only | "" | Service path prefix for ALB/CloudFront routing, e.g. /agent-api. Omit for local/dedicated-origin. |
CREDENTIALS_FILE | ✅ | ~/.opcon-relay/credentials.json | Absolute path to credentials.json (installers pre-set this). |
RELAY_STATE_DIR | ❌ | the directory holding CREDENTIALS_FILE | Where the relay keeps the state it needs to survive a restart. The Linux installer sets /var/lib/opcon-relay. |
WORK_POLLER_WAIT_SECONDS | ❌ | 20 (both installers pass 10) | Edge long-poll wait (seconds, max 20). The relay's own default is 20; the Linux and Windows installers set it to 10 unless you pass another value. Keep ≤10 on-prem behind CloudFront/envoy (~15s upstream timeout) to avoid 504s — so a relay started without an installer needs it set explicitly. |
HEALTH_PORT | ❌ | 8080 | Port for the HTTP /health readiness endpoint. |
LOG_LEVEL | ❌ | info | trace | debug | info | warn | error | fatal. |
The relay reads its configuration entirely from environment variables. There is no
relay.yaml— the old YAML template has been removed.config/relay.env.examplein the relay package is the authoritative reference for every variable.
A wrong
AGENT_SERVICE_URL/ missingAGENT_SERVICE_PATH_PREFIXtypically surfaces as an opaque403/404from the cloud edge. Use the exact values from the register dialog.
Install
There are two install paths — Linux and Windows, and both are native packages. Every artifact is
self-contained: the packages and the LSAM connector bundle their own runtimes, so no host Node
or .NET runtime is required on either platform. A later release slimmed the Windows artifact and
made the connector self-contained to start, closing out the host-runtime requirement. Full
per-platform steps are in the README.md at the root of the relay package.
The native installers take the connection details as arguments and set the environment for you;
only credentials.json is a file. --agent-service-url (Linux) / -AgentServiceUrl
(Windows) is required — the installer aborts without it.
Run the installer from inside the extracted package, not from the folder you downloaded into.
Both platforms' commands on the register dialog start by changing into the download folder and
extracting, because an elevated PowerShell opens in System32 and a Linux shell in your home
directory — so the first command would otherwise not find the archive.
- Linux:
sudo ./install/install.sh --agent-service-url <url>from the extractedopcon-relay-<version>/folder → app under/opt/opcon-relay, config/etc/opcon, logs/var/log/opcon-relay,opcon-relaysystemd unit. Other options (with their defaults):--agent-service-path-prefix(/agent-api; pass''for a direct origin),--work-poller-wait-seconds(10),--install-dir,--config-dir,--log-dir,--user/--group(opcon). The installer generates/etc/opcon/relay.envfrom its arguments and systemd loads it viaEnvironmentFile— edit that file and restart to change values. Logs viajournalctl -u opcon-relay. The unit also declares a state directory,/var/lib/opcon-relay, which systemd creates and hands to the service account;/etc/opconstays read-only, as doescredentials.jsoninside it. - Windows:
powershell -ExecutionPolicy Bypass -File .\install\install.ps1 -AgentServiceUrl <url>from an Administrator PowerShell, in the extractedopcon-relay-<version>\folder. The-ExecutionPolicy Bypassis needed because a downloaded script is blocked by the default policy. →C:\Program Files\OpCon\Relay, configC:\ProgramData\OpCon\Relay(ACL: Administrators + SYSTEM), logsC:\ProgramData\OpCon\Relay\logsandC:\ProgramData\OpCon\Relay\dotnet; serviceOpConRelay. Other parameters:-AgentServicePathPrefix,-WorkPollerWaitSeconds,-InstallDir,-ConfigDir,-LogDir,-ConnectorLogDir,-LogRetentionDays(30, 1–3650),-ServiceName. The service is managed by the bundlednssm.exe(copied into the install dir and registered), which sets the connection settings as service environment variables and captures stdout/stderr torelay-stdout.log/relay-stderr.log. NSSM is the only supported Windows service mechanism — the old-UseNSSMswitch and thesc.exenative-service path have been removed, because the relay cannot run reliably as a bare native Windows service (it fails with error 1053). See Logs, rotation and retention for what the two log directories hold and the rules the installer applies to them.
Two things about the installer that are easy to get wrong:
- It does not start the service, and it places a placeholder
credentials.json. Copy your downloadedcredentials.jsonover the placeholder first — copying onto the existing file keeps the owner and0600mode (Linux) or the restricted ACL (Windows) the installer applied — thensudo systemctl enable --now opcon-relayorStart-Service OpConRelay. A relay that starts but cannot authenticate immediately after an install is almost always a placeholder that was never replaced. - On Windows, do not set the connection values as machine environment variables. The service's own environment, set by the installer, takes precedence over them.
Re-running the installer to upgrade or change a setting rebuilds relay.env (Linux) or the
service's settings (Windows) from its arguments. Any option you passed the first time and omit the
second reverts to its default, and hand edits to relay.env are lost. Re-pass every option you
used — the URL, and also --agent-service-path-prefix, --work-poller-wait-seconds,
--install-dir and their Windows equivalents. An existing credentials.json is kept.
Every downloaded package also contains a README.md with numbered per-platform steps and the full
parameter tables. It is the authoritative install guide for the release you downloaded.
Logs, rotation and retention
A relay writes two sets of logs, because it runs two processes: the relay itself, and the LSAM Connector that speaks the legacy protocol for it.
| Where | Rotation | Retention | |
|---|---|---|---|
| Relay (Linux) | journalctl -u opcon-relay | systemd's journal | systemd's own journal settings |
| Relay (Windows) | relay-stdout.log and relay-stderr.log in -LogDir (C:\ProgramData\OpCon\Relay\logs) | NSSM rotates at 10 MB while the service is running, and the retention task adds a daily boundary | Rotated relay-stdout-<timestamp>.log / relay-stderr-<timestamp>.log files are deleted after -LogRetentionDays (default 30) |
| LSAM Connector (Linux) | lsam-connector-<date>.log in --log-dir (/var/log/opcon-relay) | A new file per day | The newest 10 files are kept; the connector deletes the rest itself |
| LSAM Connector (Windows) | lsam-connector-<date>.log in -ConnectorLogDir (C:\ProgramData\OpCon\Relay\dotnet) | A new file per day | The newest 10 files are kept; the connector deletes the rest itself |
The LSAM Connector log now redacts what a legacy job carries as command data. Until it did, all of the following were written to it in clear text, at levels that are on in every default install:
- the command line, parameters and prerun command of a Windows or UNIX job;
- the values of its environment variables (each variable's name is still shown);
- an Embedded Script's arguments and its whole body — the body was previously abbreviated only when it ran past 1,024 characters, so a short script was logged verbatim;
- an IBM i job's prerun and call-script commands;
- a SQL job's script statements, its environment-variable values, and its other options — which
is where a
sqlcmd -Ppassword sits; - the agent's own echo of the command line back on every job-status message.
So a password typed onto a command line, into an environment variable or into a script body was readable by anyone who could read that file on the relay host. Treat connector logs written by an earlier relay as sensitive: review them, and rotate any credential that was passed this way.
Redaction rewrites the log and nothing else. The agent still receives the real command line and environment, so no job behaves differently. Relatedly, a failure to send a frame no longer logs the frame's contents at all — only its length and message type.
On Windows, the deleting is done by a scheduled task the installer registers,
\OpCon\OpConRelay-LogRetention — daily at 03:00, as SYSTEM. It asks NSSM to rotate, then
deletes rotated relay logs past the retention period; the active relay-stdout.log and
relay-stderr.log are never deleted. Re-running the installer re-registers the task with whatever
-LogRetentionDays you pass, and uninstall.ps1 removes it. If the relay's log directory has been
turned into a junction or handed to a non-administrator, the task refuses to prune rather than
deleting through it, and the refusal shows up as a non-zero Last Run Result on the task.
-LogDir and -ConnectorLogDir must be folders that hold nothing but relay logsSYSTEM creates and deletes files in both, so the installer will not accept a folder somebody else controls. It aborts on a drive root, on a junction or symbolic link, and on an existing folder containing anything that is not a relay log file — a sub-folder, the config directory, a shared log folder. Point each at its own empty folder. The installer then restricts both to Administrators and SYSTEM, with Users: read.
The same care shapes uninstall: it deletes only the relay's own log files and then the folder if
nothing is left, never a recursive delete. A folder holding anything else is kept, and says so.
-KeepLogs keeps both folders untouched.
Earlier installs wrote the LSAM Connector's log under C:\Program Files\OpCon\Relay\dotnet\logs\.
The installer now points the connector at C:\ProgramData\OpCon\Relay\dotnet instead — writing
application logs under Program Files needs privileges the service should not want. The old files
are left where they are and no longer written to; delete them when you are satisfied you don't
need them.
Readiness / health
The relay runs a minimal zero-dependency /health server: 200 {"status":"ok"} when
ready, 503 {"status":"starting"} during startup, 404 otherwise. Readiness follows the relay's own
startup (ready after the core loops start, not-ready on connector max-restarts, closed on shutdown).
It is for local and orchestrator checks; nothing outside the host needs to reach it.
Rotate credentials
If a clientSecret is lost, rotate rather than delete + re-register. UI: the
Regenerate credentials dialog on the relay; backend POST /api/v1/relays/:id/credentials/rotate
(tenant-scoped, authenticated). It generates a new clientId/clientSecret, rotates the active
credential in place, and offers the new pair as a credentials.json download once; it
records who performed the rotation and when.
As on the register dialog, the new Client ID and Client Secret are never displayed — the downloaded file is the only copy. If you lose it, regenerate again.
Rotating does not restart anything. The relay keeps using its old file until you replace it, so finish the job on the relay host:
| Platform | Replace | Then |
|---|---|---|
| Linux | /etc/opcon/credentials.json | sudo systemctl restart opcon-relay |
| Windows | C:\ProgramData\OpCon\Relay\credentials.json | Restart-Service OpConRelay |
After rotation, old credentials fail OAuth: refresh-token validation verifies the token's
clientId still resolves to an active credential, so a pre-rotation refresh token cannot keep
minting ~24h access tokens. Cross-tenant/missing relay → 404.
Relay ↔ agent-service runtime
- The relay reports in on
POST /api/v1/relays/:id/heartbeatand fetches its members fromGET /api/v1/relays/self/agents(relay-authenticated). - Legacy agent changes made in-product hot-reconcile to the relay within about one heartbeat — no relay restart. See LSAM legacy connectors.
- An operator's Kill on a job running behind the relay also rides the heartbeat: the platform offers the pending kill on the relay's next check-in, the relay stops the process on the machine, and it acknowledges the dispatch on the beat after that. A kill that isn't acted on is re-offered every heartbeat until it is. See Killing a job on a legacy (LSAM) agent. This path exists only through a relay: a Kill never reaches a Universal Agent.
- A legacy agent's name must be 128 characters or fewer. The platform accepts names up to 255 characters, but the relay refuses to open a connection for a longer one, so that agent never comes online.
- A host name is looked up once. When the relay first connects to a legacy agent it resolves the agent's LSAM host name to an address and keeps using that address, including when it reconnects. If the address behind the name changes, restart the relay — or change the agent's LSAM host — so it is looked up again.
- When a relay goes stale, its legacy agents show
UNKNOWN(amber) — relay-down, distinct fromOFFLINE(agent-down). Specific-agentWAIT_MACHINEgating fails open on a stale/UNKNOWNfeed so the scheduler can't deadlock; legacy-group gating does not — it holds the job until the group has a member dispatch can route to. See Agents and agent pools and Operator response guide.
Stopping a relay
Stop the relay cleanly whenever you have the choice. The two ways a relay can stop are detected differently and leave its legacy agents in different states:
| How it stopped | Noticed after | Its legacy agents become | Queued legacy-group work |
|---|---|---|---|
| Cleanly — the service is stopped | About a second. The relay reports itself offline as the first step of its shutdown | OFFLINE | Holds in WAIT_MACHINE until the relay reports in again |
| Abruptly — killed, power lost, network dropped | About 5 minutes of missed heartbeats | UNKNOWN | Holds too, and resumes when the relay reconnects |
Queued legacy-group work now waits either way. An abrupt loss used to be far worse than a clean
stop: the WAIT_MACHINE gate failed open on an UNKNOWN feed, so the job advanced and then
failed at dispatch with no online group member. An all-UNKNOWN group now holds instead — see
When a legacy group has no dispatch target.
Stopping cleanly is still the better choice, for a smaller reason: a clean stop is reflected in
seconds rather than five minutes, so the status you see after a planned stop is the status. The
five-minute window on an abrupt loss stays deliberate — a shorter one would turn a brief network blip
into five minutes of UNKNOWN agents and queued jobs holding through it for nothing.
A relay's shutdown reports itself offline before it drains its work poller, so the status you see is not waiting on in-flight work. If the report cannot be delivered, shutdown continues anyway and the relay is detected by the ordinary five-minute window.
A clean stop now raises an Agent Offline notification for each of the relay's legacy agents, not
only a Relay Offline for the relay. It used to raise the relay's event alone, so a planned relay
restart sent a run of Online notifications on the way back up with no Offline to match them —
which reads like a run of unexplained outages. The Offline events are raised only for agents the stop
actually moved, so an agent already OFFLINE does not produce a second one.
This applies to a clean stop. An abrupt loss moves the agents to UNKNOWN rather than OFFLINE
and raises no per-agent event, which is deliberate: relay-down is not the same claim as
agent-down, and the agents' real state is not known. See
Agent (connectivity) triggers.
When a legacy agent stops answering without disconnecting
A legacy agent normally goes away visibly: the TCP connection closes and the connector notices at once. The hard case is an agent that stays connected and stops answering — the host was paused or frozen, or a NAT port-forward in front of it keeps the socket open after the machine behind it is gone. Nothing announces anything, and the connector's send is waiting on an acknowledgement that is never coming.
That wait used to block the whole machine. Every queued job start for that agent queues behind
the outstanding acknowledgement, and so does the connector's own heartbeat to it — so jobs already
running on it stayed RUNNING with no completion, nothing new could be sent to it, and the machine
went on reading as reachable. The only thing that eventually broke the deadlock was the send retry
loop giving up, which takes 31 attempts at the response timeout — about two and a half hours
at the 300-second default.
There is now a liveness deadline on the acknowledgement itself. If one has been outstanding
longer than LsamDefaults:AckLivenessTimeoutSeconds — 90 seconds by default — the connector
declares the agent dead on its next machine-status tick: the agent goes down, the socket is released,
and it reconnects. That is the same give-up path the retry loop always took, reached in a minute and
a half instead of two and a half hours. Setting the value to 0 or less turns the check off. The
send retry loop, the response timeout and the kill timeout are all unchanged.
The deadline is measured from when the acknowledgement first went outstanding, not from the last send, so re-sending the same frame does not push it back.
The relay re-seeds each legacy agent's state from the connector on every heartbeat, and the field it read reported the connection as up whenever it was configured up — a value that is only ever true. So a correct "agent is down" was overwritten as reachable once its event aged out, roughly three minutes later, and the agent read as available while nothing could be sent to it.
That field now reports up only when the connection is marked up, the socket really is connected, and the agent has confirmed the legacy handshake on it. Anything else is down, and a machine the connector has no connection for at all reads as unknown rather than as either.
Surviving a relay restart
A relay writes what it needs to pick up where it left off into its state directory —
RELAY_STATE_DIR, or the directory holding credentials.json when that is unset. Two files live
there, relay-state.json and the LSAM Connector's connector-state.json, and they have to share one
directory. The relay's own file is kept current as it works and is readable only by the account the
relay runs as.
What that state buys, when a relay restarts with work in flight:
- Legacy jobs that were running are picked up again rather than left untracked, so their completions still reach the platform.
- In-flight OpCon MFT, OpCon RPA and EASE runs are recovered and continue to be followed. An EASE run that was in flight is never re-added to the EASE schedule — see A relay restart resumes the run.
- Job identifiers do not collide. The relay resumes its internal job counter past where it left off instead of restarting from zero and reusing a number a job still running was given.
Neither file holds a secret. Both carry only what identifies the work — job handles and identifiers, machine names, statuses and timestamps — never a credential, an API token, a property value, a command line or job output.
Earlier Linux relays could not write either file: both went to /etc/opcon, which the installer's
systemd unit mounts read-only. Each save failed with a warning in the log and nothing else, so
everything above silently did not happen on a native Linux relay — a restart lost track of the jobs
that were running, and MFT and RPA runs in flight were not recovered. Docker and Windows relays were
never affected. Re-run the installer to pick up the state directory; no configuration of your own is
needed.
The same read-only sandbox stopped the LSAM Connector writing its own log file on Linux, which is
why that row appears in the log table only now. Its output was
already reaching journalctl -u opcon-relay; the file is what was missing.
Serving OpCon MFT agents
A relay reports on each check-in whether it can serve OpCon MFT agents. Two things follow from that report:
- Registering an OpCon MFT agent under a relay that does not support them is refused, with Upgrade the relay to register MFT agents. Resetting an existing MFT agent's API token under such a relay is refused the same way.
- A relay that does not support OpCon MFT is never shown MFT agents at all, so an already- registered MFT agent behind a downgraded relay simply gets no work rather than failing oddly.
If you are adding your first OpCon MFT agent, upgrade the relay first. Everything else about the relay — registration, credentials, install, health — is unchanged: an MFT agent rides the relay you already have.
Serving OpCon RPA agents
A relay reports on each check-in whether it can serve OpCon RPA agents, on a separate flag from its OpCon MFT one. A relay that can serve MFT agents is not thereby able to serve RPA agents — each capability is reported and gated on its own. Three things follow from that report:
- Registering an OpCon RPA agent under a relay that does not support them is refused, with This relay must be upgraded to run OpCon RPA agents. Editing an existing RPA agent under such a relay is refused the same way, and so is looking up its tasks from the job editor.
- A relay that does not support OpCon RPA is never shown RPA agents at all, so an already- registered RPA agent behind a downgraded relay simply gets no work rather than failing oddly.
- An OpCon RPA job is refused rather than queued when the relay serving its agent pool has not reported RPA support. The message names the pool: the relay serving agent pool '…' has not reported RPA support.
If you are adding your first OpCon RPA agent, upgrade the relay first. Everything else about the relay — registration, credentials, install, health — is unchanged: an RPA agent rides the relay you already have.
Serving EASE agents
A relay reports, on each check-in, whether it can serve EASE agents — a third flag, independent of its OpCon MFT and OpCon RPA ones. The same three consequences follow:
- Registering an EASE agent under a relay that does not support them is refused, with This relay must be upgraded to run EASE agents. Editing an existing EASE agent under such a relay is refused the same way.
- A relay that does not support EASE is never shown EASE agents at all, so an already-registered EASE agent behind a downgraded relay simply gets no work rather than failing oddly.
- An EASE job is refused rather than queued when the relay serving its agent pool has not reported EASE support. The message names the pool: the relay serving agent pool '…' has not reported EASE support. Upgrade the relay to run EASE jobs.
If you are adding your first EASE agent, upgrade the relay first. Everything else about the relay is unchanged: an EASE agent rides the relay you already have. See EASE jobs.
Events raised by the machines behind a relay
Traffic through a relay is not one-way. A job or process on a legacy LSAM machine can raise a Continuum event by writing a file of event syntax into the agent's MSGIN directory; the agent sends each line up the connection it already holds, and Continuum applies it. No inbound port is needed; each line carries a service-account credential, which Continuum verifies and checks against that account's roles.
Each legacy agent registration carries a Default event environment for this, and it is not optional in practice: until it is set, every event that machine raises is refused. See Events raised by a legacy agent for the line format, which event families are applied, and where a rejected line goes.
Security notes
- Keep
credentials.jsonat 0600 (Linux) or a restricted ACL (Windows); installers set this, and copying your downloaded file onto the placeholder preserves it — replacing the file another way may not. - Outbound only — no inbound internet ports.
8080(/health) is for local/orchestrator checks, not external exposure. - Treat
clientSecretlike a password; never commit or share it.
Keeping a relay current
A relay is installed on your side of the boundary and is not upgraded for you, so it can be older than the platform it talks to. Several current behaviours depend on the relay carrying them, and each fails in a way that points at the wrong thing:
| What needs a current relay | What an older relay does instead |
|---|---|
| Exit criteria using Exit Criteria Result, a Range condition, or more than five conditions | Refuses the dispatch, so the job never starts. Loud, and the safe direction — but the job does not run until the relay is upgraded. |
| Retrieving output files from a UNIX or IBM i agent | Names the job the way only a Windows agent expects. On UNIX the agent answers with its own global log files and the request reports success, so the wrong output is stored as the job's. On IBM i it returns nothing. IBM i took a second relay change after that, to read the OS/400 job number the agent reports — a relay current for UNIX may still return nothing for IBM i. |
| Submitting an IBM i batch job | The name the relay sends contained a character an IBM i job name cannot hold, so the agent refused the submission and the job failed immediately with Initialization error. |
| Sending a legacy job its workflow's schedule date | Sends the relay's own UTC date instead. West of UTC an evening dispatch therefore carries tomorrow's date, so a File Arrival window watches the wrong day and a path built from %SMA_MSLSAM_SCHEDULE_DATE% names the wrong day's file. Quiet, and wrong only some of the time, which is what makes it worth checking for. |
| Detecting a legacy agent that stops answering without disconnecting | Waits for the send retry loop to give up — about two and a half hours at the default response timeout — during which the agent's running jobs sit in RUNNING with no completion and it still reads as reachable. |
None of these can be detected from the cloud side today: nothing distinguishes an out-of-date relay from a job with no output. When legacy behaviour looks wrong in a way only some platforms show, check the relay build first. Both Download Relay and the register dialog's download links always serve the latest published release, and Download Relay gets you it without registering anything.
A legacy job's output files are named by the agent from the job name the relay sent at dispatch, and the request that lists them is built independently by the platform at read time. The two must agree exactly, so the character rule they both apply changed on both sides at once. Upgrade the relay when you take the platform release, not separately — a skew in either direction stops new jobs' output from being found, and there is no safe side to be on.
Two consequences for jobs that ran before the upgrade:
- Their output files keep the old names on the agent's disk, permanently. Output that was already viewed stays available, because it is held after the first read. Output that was never viewed comes back empty, and no upgrade order changes that.
- Refresh from agent on such a job clears what is held and re-reads, so it turns an old cached listing into an empty one. That is deliberate: a listing the platform cannot vouch for is dropped rather than kept, which is what removed the wrong-output exposure below.
Running jobs is unaffected either way — this is about retrieving output, not dispatching work.
An output listing is now checked against the job that asked for it
The platform discards returned files whose names do not belong to the requesting job. Before this, a request that named the job in a form the agent did not recognise still got an answer: a UNIX agent appends its own global log and error files to a listing that matched nothing, and that came back as a successful, non-empty result and was stored and served as the job's output. An operator could read another job's output believing it was their own.
The check applies where the request carries a job name — UNIX and IBM i. Windows listings are not filtered: the request carries no short name there and the file names have no such prefix. Files that do not match are dropped rather than the whole request failing, so a job that legitimately produced no output of its own still reports an honest empty listing.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
A UNIX or IBM i job's output files are empty, or hold the agent's own logfile/errfile | The relay predates per-platform output requests | Upgrade the relay, then use Refresh from agent on any job whose wrong list was already stored. See Keeping a relay current. |
| A UNIX or IBM i job that ran before a relay upgrade now lists no output files | Its files were written under the old naming rule and no longer match what is asked for | Expected, and not repairable — the names on the agent's disk do not change. Output viewed before the upgrade is still held. Re-run the job to get output under the current rule. |
| An IBM i job fails immediately with Initialization error and the agent reports a submission failure | The relay predates the IBM i job-name character fix | Upgrade the relay. IBM i enablement is still in progress — see IBM i (AS/400) Batch Job. |
| A legacy job never starts and its dispatch is refused on exit criteria | The relay predates Exit Criteria Result, range conditions, or tables over five rows | Upgrade the relay, or have the table rewritten as up to five plain conditions with Exit Criteria Result Fail. |
| Boot error: "Unrecognized key" / "relayId is required" | Edited/old credentials.json with wrong keys | Re-download from the UI; the file must be exactly { relayId, clientId, clientSecret }. |
Opaque 403/404 from the cloud on startup | Wrong AGENT_SERVICE_URL / missing AGENT_SERVICE_PATH_PREFIX | Use the exact connection values from the register dialog. |
Install aborts: --agent-service-url is required | Installer run without the connection URL | Pass --agent-service-url (Linux) / -AgentServiceUrl (Windows) with the value from the register dialog. |
| Windows install fails: "Bundled nssm.exe not found" | Incomplete/partial relay artifact | Re-download the relay artifact; the package must ship nssm.exe alongside install.ps1. |
| Windows install fails: "Bundled prune-relay-logs.ps1 not found" | Incomplete/partial relay artifact | Re-download the relay artifact; the package must ship prune-relay-logs.ps1 alongside install.ps1. |
| Windows install aborts: the log folder "must be a dedicated relay log folder" | -LogDir or -ConnectorLogDir points at a drive root, a junction, or a folder holding something else | Point it at its own empty folder — see Logs, rotation and retention. |
| Rotated relay logs keep accumulating past the retention period | The \OpCon\OpConRelay-LogRetention task is missing, disabled, or refusing to prune | Check the task's Last Run Result. A non-zero result means it refused the folder (a junction, or an owner that is not Administrators/SYSTEM) or could not delete a file. Re-running the installer re-registers the task. |
| Old credentials still work after a suspected leak | Credentials were never rotated | Rotate credentials (invalidates the old pair, incl. refresh tokens). |
Legacy agents show UNKNOWN | The relay is stale/down | Check the relay (health, logs, connectivity) — not the individual agents. |