Skip to content

Pause a Job

Stops the job in the middle of a step and waits for a person to release it from the job monitor. The job stays alive the whole time, holding its process and its working directory, so everything the earlier ops set up is still there when it resumes.

Reach for it when a human has to do something outside Clarive before the job can go on: catalogue a file, open a firewall, confirm a window. When the decision itself needs recording against named approvers, use Request Approval. When you only need a fixed delay with nobody involved, use Sleep for a number of seconds.

op reached job set to paused Reason logged re-check the status every 5 seconds still paused resumed from the monitor next op runs cancelled or set to error the op fails 1 day with no answer fails unless allowed

Fields

Reason

The one-line message that tells whoever is looking at the monitor why the job stopped. It goes into the job log as Paused. Reason: followed by your text, marked as a milestone so the line renders in bold.

Leave it empty and the message reads unspecified, which helps nobody. Write a sentence that says what has to happen before the job is released.

The field expands ${} placeholders and the text is rendered as HTML, so a link is the most useful thing you can put here: point it at the screen where the work has to be done, and the person releasing the job has one click rather than a hunt. Text between asterisks comes out bold and text between backticks comes out as code, so escape those characters if you meant them literally.

The box is a plain multi-line editor and there is no length check on it, but two limits apply on the way out. A message longer than 2000 characters is cut down to 2000 and ends in (continue...), and the whole untruncated text is appended to the Details attachment under a line of equals signs. Of what survives, the log grid renders the first 512 characters and puts an ellipsis after them. Keep the message to a sentence and put the bulk in Details.

Anything that looks like a password is masked before the line is stored.

Don't Fail On Timeout

Off by default. It decides what happens when nobody releases the job within one day.

Off, the job fails at the deadline with an error line saying the pause timed out, and the step is treated as a failure like any other, rollback included. On, the deadline is logged as an error all the same, then the job resumes on its own and the next op runs as if somebody had released it.

Tick it for optional checkpoints, where a stalled job should carry on rather than sit there. Leave it clear for a checkpoint that is actually a gate.

The one-day limit is fixed and is not on this form. Neither is the five-second polling interval, which is why a pause released from the monitor takes a few seconds to take effect and why the shortest possible pause is about five seconds.

Details

The long form: the block of text attached to the log line, reached from an icon at the end of that row rather than shown inline. It is the place for the list of files, the command output or the table of items that explains the reason.

Expands ${} placeholders, which is what makes it useful: put the variable holding the offending items straight into the field and the reader sees them without hunting through the log. Leave it empty and the log line carries no attachment at all, unless an over-long Reason spilled into it.

Up to 250,000 characters the attachment opens in a panel inside the log. Past that the icon changes to one that opens the content in a separate browser tab, with a download link beside it. Passwords are masked here too, and the content is compressed before it is stored.

When the op refuses to pause

A warning line saying the job is pausing is written before anything else is decided, so it appears even on the runs described below where no pause happens.

A job cannot be paused during its INIT or CHECK steps. Reaching this op in either of them sets the job to error and fails the op, which is a very different outcome from pausing. Keep it in PRE, RUN or POST.

While paused, two things end the pause unhappily. Cancelling the job from the monitor fails the op with a message saying the job was cancelled while paused. Something else pushing the job into error does the same. In both cases the pause does not resume the step, it ends it.

When the pause does end normally, the job returns to whatever status it had before it stopped and continues with the very next op in the same step. Nothing is re-run, and nothing is skipped.

Logging and notifications

The Reason goes in as a milestone with Details attached, and the timeout writes an error line when it arrives. Neither can be turned off from this form. Releasing the job from the monitor adds a line naming the person who did it.

Pausing raises a job-paused event scoped to the job's projects and environment, and releasing raises a matching job-unpaused event, so notification rules subscribed to either can mail whoever needs to act. Nobody is mailed automatically by this op on its own.

The job shows as paused in the monitor for the entire wait, and counts as an active job, so any attempt to delete the topics it carries is refused while it sits there. The process behind it has to stay alive for the whole pause: if it dies, the housekeeping daemon notices the missing process and marks the job killed.

Rollback

This op runs again during a rollback pass by default. A job that failed after pausing once will pause a second time on the way back, with the same reason, and wait for another release. That is rarely what anyone wants. Clear Run Rollback on the op's Options tab (see the rule palette), or wrap the op in IF ROLLBACK if the rollback needs a pause of its own with a different message.

Combining with other ops

Put it behind IF var condition THEN so the job only stops when there is something to stop for. An unconditional pause in a shared rule is how a nightly pipeline ends up waiting all night.

Collect what the reader needs first, then pause: whatever an earlier op left in the stash can be dropped into Details with a placeholder.

Request Approval is the alternative when the wait needs a named set of approvers and an approve or reject decision on the record. This op accepts a release from anyone who can act on the job in the monitor.

Examples

Stop the job when uncatalogued objects turned up, and send the reader straight to the screen where they can fix it.

Reason    New objects are not catalogued yet. Catalogue them
          <a href="${catalog_url}" target="_blank">here</a>, then resume this job.
Don't Fail On Timeout   [ ]
Details   ${non_catalogued_objects}

An optional review checkpoint that gives the reviewer a day and then carries on by itself.

Reason    Optional review window. The job continues on its own tomorrow if nobody looks.
Don't Fail On Timeout   [x]
Details   ${changed_files}