Skip to content

RETRY

Runs the ops nested underneath it and, when any of them fails, runs the whole block again from the top. Reach for it around a step that fails for reasons outside the rule: a file still locked by another process, an agent that has not finished booting, an endpoint that answers with an error for a few seconds after a restart.

It repeats and nothing else. If you need to react to the failure rather than repeat it, use TRY statement with CATCH statement. When the attempts run out, RETRY rethrows the last error and the job stops there.

run the nested ops error? no on to the next op yes attempts left? yes: wait Pause (s) job fails, last error

Fields

Attempts

How many times the block may run in total, not how many extra tries follow the first one. 5 means the block runs at most five times. 1, the default, disables retrying and leaves you with a plain group of ops.

The field will not save blank and accepts nothing but digits and a sign.

0 means retry forever. There is no iteration cap behind it, so a block that always fails with 0 attempts keeps the job running until someone cancels it.

A negative number gets past the field and behaves like neither a count nor 0. The block runs zero times and the job fails on the spot with an error that carries no message, which reads in the log like anything but a retry problem.

Pause (s)

Seconds to wait after a failed attempt before the next one starts. Default 0, which retries with no delay at all. Whole seconds only.

The wait also happens after the final failed attempt, when there is nothing left to retry. A block set to 5 attempts and a 2 second pause therefore spends 10 seconds asleep before the job fails, not 8.

What counts as a failure

Any error raised by a nested op, including one you raise yourself with FAIL. The error is caught by the RETRY wrapper, not by the job, so the job keeps running while the attempts last.

A nested op set to Ignore Errors swallows its own failure before RETRY can see it. The pass then looks successful and the block never repeats. Trap Errors does something different again: it pauses the job and waits for an operator instead of retrying, and what happens next depends on the action they pick. Skip returns as though the op had worked, so the block does not repeat. Abort raises the error, which RETRY catches and answers with another attempt. Leave the ops you want retried on Throw Errors.

Retrying re-runs the block from the first nested op every time. Nothing the earlier ops did on the failed pass is undone first, so anything in the block that is not safe to repeat, a counter you increment or a file you append to, gets done once per attempt.

Timeout on the Options tab

A value in Timeout arms a clock once, before the first attempt starts, and it is never rearmed. Whichever attempt is running when it expires is aborted, that abort is caught like any other error, and every attempt after it runs with no clock at all. So Timeout bounds neither the block as a whole nor each attempt individually. Nor is the clock cancelled when the block finishes early: a timeout longer than the block itself is still ticking, and it can fire during a later op.

Error Handling on this same op is independent of the attempt count. Set to Trap Errors, it takes over only after the attempts are used up, pausing the job and waiting for an operator decision instead of failing it.

What you see in the job log

No line announces an attempt. There is no "attempt 2 of 5" marker, and the error that ended a failed pass is not logged anywhere: it is held only until the next attempt replaces it. What you see is the nested ops running again and whatever they log for themselves repeated once per pass, with nothing saying why. The step entry each op writes is at debug level, so raise the job log detail when you need to follow which pass you are looking at.

Only the last error survives to stop the job. A block that fails first on a timeout and then on a missing file reports the missing file, and the timeout leaves no trace.

In a rollback pass

The block behaves the same way, with the same attempt count, for whichever nested ops are set to run in rollback. Clear Run Rollback on this op to keep the whole block out of the rollback pass.

Combining with other ops

Wrap Run a Remote Script or Ship File Remotely when the target node comes up slowly. One or two extra attempts a few seconds apart covers most of it.

Put TRY statement and CATCH statement around the RETRY block, not inside it, when you want a fallback after every attempt has failed. A catch inside the block absorbs the error and the block stops repeating.

Sleep for a number of seconds is the plain wait when you want a delay without a failure driving it.

To poll for a condition instead of retrying a failure, use WHILE condition with a stash variable that a nested op refreshes on each pass.

Examples

Give a remote service five chances to answer, one second apart.

Attempts   5
Pause (s)  1
  -> Run a Remote Script    health check on bar_node

Retry a file copy four times with a longer gap, for a share that gets remounted.

Attempts   4
Pause (s)  15
  -> Ship File Remotely     myapp.tar to bar_node