WAIT for children
Blocks the rule until every op that was started in the background has finished, then collects what those ops produced and decides what to do about the ones that failed. It is the join point for parallel work.
An op runs in the background when you set Parallel Mode on its Options tab to
[FORK] Fork and Wait. That op is handed to a separate process, the rule carries on to the next op
immediately, and this op is where the rule catches up. Ops set to [NOHUP] Fork and Leave are not
tracked, are not waited for, and never appear in the results.
Fields¶
Children Stash Keys to Export¶
A one-column grid of stash variable names. Each background op works on its own private copy of the stash, so anything it writes is thrown away when it ends. The names you list here are the only values that make the trip back.
Add appends a row pre-filled with variable_name_here. One click on a row starts editing it, and
Delete removes the row you have selected. A row cannot be left blank, so an unwanted row has to go
through Delete rather than being emptied. Leaving the placeholder text in place exports a variable
that does not exist, which costs you nothing but gives you nothing either.
Write bare names. This field does not expand ${var} placeholders, and it does not follow dotted
paths, so results.total looks for a variable literally called results.total rather than walking
into results. Export the top-level variable and reach into it afterwards.
A name that the background op never set comes back with no value and no warning.
Leave the grid empty and the wait still happens. You get the join and the error handling, only no data.
Errors¶
What to do when one of the background ops ended with an error.
| Errors | Effect |
|---|---|
fail |
the job stops here, with a summary naming the process ids that failed and the ones that succeeded, followed by each failure's error text |
warn |
that same summary is recorded as a warning and the rule continues to the next op |
silent |
no summary, and the rule continues |
fail is the default and the value you want unless the background work is genuinely optional.
silent is quieter than it sounds. The error text of every failed child is written to the log as an
error line while the results are collected, whatever this field says. What the field controls is the
summary and the decision that follows it. Under warn and silent the job runs to its normal end
even though the step is marked at error level in the job summary.
Where the results go¶
Put a name in Return Key on the op's Options tab and the collected results land in that stash
variable as a list, one entry per background op, in the order they finished rather than the order
you placed them in the rule.
Each entry carries err, holding the error text when that op failed, ret, holding whatever the op
evaluated to last and rarely worth reading, and one value per name from
Children Stash Keys to Export. Read them by position, for example bar_results[0].test_result.
Because the order is the finishing order, position zero is whichever child was quickest, not the first one in the tree. When you need to tell the entries apart, have each child write a name of its own into an exported variable and match on that instead of counting positions.
Leave Return Key blank and the results are collected and discarded. The error handling still
applies.
A child that is killed outright, by the machine running out of memory for instance, never gets to send its results back. It is still waited for, but it contributes no entry to the list and counts as neither a failure nor a success, so the list can come back shorter than the number of ops you forked with no error to explain it.
Things that catch people out¶
If no op forked before this one, the op does nothing and logs nothing. A WAIT placed above the ops it was meant to join is a silent no-op, not an error.
One WAIT consumes the children. A second WAIT further down finds an empty set and returns immediately, so pair each batch of background ops with exactly one WAIT.
Forget the WAIT entirely and the rule still blocks at its very end until the background ops finish,
because nothing is allowed to outlive the rule except [NOHUP] ops. What you lose is the exported
values and the failure handling: errors from those ops never reach the job's outcome.
The wait polls, it does not sleep until woken. It checks the children, then waits half a second, then checks again, so this op costs at least half a second of wall clock whenever anything was forked, however fast the children were.
The job log shows a line as each op forks, naming it and its process id, then a line here listing the process ids being waited on. Failures are reported per child. When every child succeeded, a closing line says the wait is done and lists an empty set of process ids, because by then there are none left to name. None of that text is configurable.
Combining with other ops¶
Set Parallel Mode to [FORK] Fork and Wait on the ops you want overlapped, then place this op
after the last of them. CALL rule is a common thing to fork, since it
gives each branch a whole rule of its own.
Feed the results into IF var condition THEN to branch on what came back, or into FOREACH CI when the children exported resource identifiers.
Use Semaphore Key on the forked ops when several of them touch the same resource and must not overlap after all.
Examples¶
Two test suites run at the same time, each writing its verdict into the stash, joined here with both verdicts brought back and a failure in either one stopping the job.
Children Stash Keys to Export
test_result
test_error
Errors fail
Return Key (Options tab) bar_results
Fire off a cache warm-up alongside the deployment, note it in the log when it fails, and carry on regardless.
Children Stash Keys to Export
(empty)
Errors warn