Skip to content

WAIT for children

Blocks the rule until every op that was started in the background has finished, then collects what those ops produced and decides what to do about the ones that failed. It is the join point for parallel work.

An op runs in the background when you set Parallel Mode on its Options tab to [FORK] Fork and Wait. That op is handed to a separate process, the rule carries on to the next op immediately, and this op is where the rule catches up. Ops set to [NOHUP] Fork and Leave are not tracked, are not waited for, and never appear in the results.

op A [FORK] op B [FORK] op C [NOHUP] never joined WAIT for children exported keys land in the Return Key a child failed: fail / warn / silent

Fields

Children Stash Keys to Export

A one-column grid of stash variable names. Each background op works on its own private copy of the stash, so anything it writes is thrown away when it ends. The names you list here are the only values that make the trip back.

Add appends a row pre-filled with variable_name_here. One click on a row starts editing it, and Delete removes the row you have selected. A row cannot be left blank, so an unwanted row has to go through Delete rather than being emptied. Leaving the placeholder text in place exports a variable that does not exist, which costs you nothing but gives you nothing either.

Write bare names. This field does not expand ${var} placeholders, and it does not follow dotted paths, so results.total looks for a variable literally called results.total rather than walking into results. Export the top-level variable and reach into it afterwards.

A name that the background op never set comes back with no value and no warning.

Leave the grid empty and the wait still happens. You get the join and the error handling, only no data.

Errors

What to do when one of the background ops ended with an error.

Errors Effect
fail the job stops here, with a summary naming the process ids that failed and the ones that succeeded, followed by each failure's error text
warn that same summary is recorded as a warning and the rule continues to the next op
silent no summary, and the rule continues

fail is the default and the value you want unless the background work is genuinely optional.

silent is quieter than it sounds. The error text of every failed child is written to the log as an error line while the results are collected, whatever this field says. What the field controls is the summary and the decision that follows it. Under warn and silent the job runs to its normal end even though the step is marked at error level in the job summary.

Where the results go

Put a name in Return Key on the op's Options tab and the collected results land in that stash variable as a list, one entry per background op, in the order they finished rather than the order you placed them in the rule.

Each entry carries err, holding the error text when that op failed, ret, holding whatever the op evaluated to last and rarely worth reading, and one value per name from Children Stash Keys to Export. Read them by position, for example bar_results[0].test_result.

Because the order is the finishing order, position zero is whichever child was quickest, not the first one in the tree. When you need to tell the entries apart, have each child write a name of its own into an exported variable and match on that instead of counting positions.

Leave Return Key blank and the results are collected and discarded. The error handling still applies.

A child that is killed outright, by the machine running out of memory for instance, never gets to send its results back. It is still waited for, but it contributes no entry to the list and counts as neither a failure nor a success, so the list can come back shorter than the number of ops you forked with no error to explain it.

Things that catch people out

If no op forked before this one, the op does nothing and logs nothing. A WAIT placed above the ops it was meant to join is a silent no-op, not an error.

One WAIT consumes the children. A second WAIT further down finds an empty set and returns immediately, so pair each batch of background ops with exactly one WAIT.

Forget the WAIT entirely and the rule still blocks at its very end until the background ops finish, because nothing is allowed to outlive the rule except [NOHUP] ops. What you lose is the exported values and the failure handling: errors from those ops never reach the job's outcome.

The wait polls, it does not sleep until woken. It checks the children, then waits half a second, then checks again, so this op costs at least half a second of wall clock whenever anything was forked, however fast the children were.

The job log shows a line as each op forks, naming it and its process id, then a line here listing the process ids being waited on. Failures are reported per child. When every child succeeded, a closing line says the wait is done and lists an empty set of process ids, because by then there are none left to name. None of that text is configurable.

Combining with other ops

Set Parallel Mode to [FORK] Fork and Wait on the ops you want overlapped, then place this op after the last of them. CALL rule is a common thing to fork, since it gives each branch a whole rule of its own.

Feed the results into IF var condition THEN to branch on what came back, or into FOREACH CI when the children exported resource identifiers.

Use Semaphore Key on the forked ops when several of them touch the same resource and must not overlap after all.

Examples

Two test suites run at the same time, each writing its verdict into the stash, joined here with both verdicts brought back and a failure in either one stopping the job.

Children Stash Keys to Export
  test_result
  test_error
Errors    fail
Return Key (Options tab)    bar_results

Fire off a cache warm-up alongside the deployment, note it in the log when it fails, and carry on regardless.

Children Stash Keys to Export
  (empty)
Errors    warn