Skip to content

Windows service

Starts, stops or restarts a Windows service on one or more servers, or reads its current state without touching it. Each server you pick is handled in turn, and the state the service ended in is written back into the stash keyed by server name.

Reach for this when the service state itself is what you care about. To run an arbitrary command on the same box, use Run a Remote Script. To find out which services exist before you act on one, run List Windows Services first.

one server at a time, in the order you picked them server active? yes no SKIPPED run the action Wait until status ticked poll every 2s unticked read the status once record next server

Every field on this form expands ${var} placeholders against the stash before the op runs, the number field included.

Fields

Server

The Windows servers to act on. The picker lists the resources that can take an agent connection, plus the variables that resolve to one, so you can select boxes by hand or feed in a variable that a previous op filled.

More than one server can be selected. They are processed one after another in the order shown, not in parallel, so a slow wait on the first server delays the rest. A variable holding a comma-separated list is split on the commas and treated the same way.

The field is required. A value that does not resolve to a real resource ends the step with a resource error, and the Errors setting does not soften that. The same is true of a server that cannot be reached: connection problems always fail the step.

A server marked inactive is not contacted. A warning goes into the job log, the server is recorded with the status SKIPPED, and the run moves on to the next one.

The connection uses whatever the server resource itself is configured with, including its user and its agent timeout. This op has no user field of its own.

Only the Clax and Worker connection methods hand back the command result in the shape this op reads, so Connect by on the server resource has to be set to one of those. With SSH, FTP or Balix the step ends with an error the first time it reads the service state. SSH is the one that catches people out: it is what a resource falls back to when nobody ever picked a method, and it is tried ahead of Clax on a resource that has both enabled.

Service

The short service name, the one you would pass to net start, not the display name shown in the Services console. wuauserv works, Windows Update does not, because the value goes into the command unquoted and the space splits it.

The field is required. Nothing here checks that the service exists before acting on it.

Action

What to do. Defaults to start when left unset.

Action What happens
start starts the service, unless it is already running
restart stops it, waits until it reports stopped, then starts it
stop stops the service, unless it is already stopped
status reads the state and changes nothing

start and stop check the current state first and do nothing when the service is already in the target state. That is a silent no-op at the default log level: no command runs and nothing is written to the job log, but the status still gets recorded.

restart always waits for the stopped state between the two halves, whether or not Wait until status is ticked. It uses the same Wait timeout value. With the default timeout of 0 that wait has no upper bound, so a service that refuses to stop leaves the job sitting there indefinitely. Set a timeout when you choose restart.

status ignores Errors, Wait until status and Wait timeout entirely. It reads once and records what it found.

Errors

How a failed start or stop is treated. Defaults to fail.

Errors Effect
fail the step fails and the job stops
warn the failure is recorded at debug level and the run continues
silent the failure is swallowed and the run continues

warn does not put a warning in the job log. The message comes out at debug level, so unless the job is running with debug logging turned on, warn and silent look identical from the outside. In both cases the recorded status for the failed server is whatever the service actually reports, which is usually the state it was already in, and the run carries on to the next server. The exception is a ticked Wait until status with no timeout: the wait that follows the failed command polls for a state that is never coming, and no later server is ever reached.

This setting covers only the start and stop commands. Resource lookup failures, connection failures and agent timeouts fail the step no matter what you choose here.

A command counts as successful when it returns a zero exit code, or when it returns a non-zero code but still prints a "was ... successfully" message. That second case exists because some Windows builds report success with a non-zero code.

Wait until status

Off by default. When ticked, the op keeps polling after the action until the service reaches the state that action implies: running for start and restart, stopped for stop. Polling happens every 2 seconds.

With Wait timeout left at its default of 0 that polling has no upper bound. A service that failed to start, or one that is slower than you expected, holds the job in the loop with nothing above debug level in the log, and the servers after it in the list are never touched. Give the wait a timeout whenever you tick this box.

When the wait gives up because Wait timeout expired, it writes a warning to the job log and carries on. It does not fail the step, and Errors has no say in it. The status recorded for that server is the state the service was in at the moment the wait gave up, so read the result if the difference matters to you.

Leaving it off does not skip the status read. The op still queries the service once after the action and records whatever it sees at that instant, which for a slow service is often START_PENDING rather than RUNNING.

Wait timeout (seconds, 0=infinite)

A blank field behaves the same as 0. Any value above zero is only checked between polls, so the wait ends on the first 2-second boundary past the limit rather than exactly on it, and a 1-second timeout still costs you 2 seconds.

The timeout bounds the polling loop only. Each individual command sent to the server is bounded separately by the agent timeout configured on the server resource, and that one does fail the step when it trips.

Status values

The status recorded for each server is one of these, in English, whatever display language the Windows box runs.

Status Meaning
RUNNING the service is running
STOPPED the service is stopped
START_PENDING the service is starting
STOP_PENDING the service is stopping
PAUSE_PENDING the service is pausing
CONTINUE_PENDING the service is resuming
PAUSED the service is paused
SKIPPED the server was inactive and never contacted

STOPPED doubles as the fallback. When the state cannot be read at all, including the case where the service name does not exist on that box, the op reports STOPPED.

A typo in Service therefore passes for success under both status and stop. stop reads the state, gets STOPPED, concludes the service is already stopped, sends no command and records STOPPED, all without a word above debug level. start and restart do expose the mistake, because the start command runs against the bad name and fails. Confirm the name with List Windows Services if the result surprises you.

What lands in the stash

The op writes windows_service_status, a map of server name to status:

windows_service_status:
  bar_host:  RUNNING
  baz_host:  SKIPPED

The key is fixed. A second Windows service op later in the rule overwrites it, so read the value before the next one runs, or copy it out with SET VAR.

Filling in Return Key on the op stores the same map under your own name as well, which is the safer option when the rule touches several services.

Two selected servers that share a name collapse into one entry.

Rollback

This op never marks the job as needing a rollback on its own, and it registers no undo. Stopping a service here and failing two ops later leaves the service stopped.

Rollback Needed Always does work. That marker is written for you before the op runs, so the job gets a rollback pass and the op runs again on the way back with the same configuration, which for stop means a second stop.

Rollback Needed Before and Rollback Needed After are modes the op itself has to act on, and this one does not, so neither ever marks anything. They are not inert, though. Picking either fills Needs Rollback Key with the op's own name, and from then on the op is conditional on the rollback pass: it runs on the way back only if that key was marked by something. Left at the auto-filled name nothing marks it, so the effect of choosing one of these two modes is that the op stops running during rollback. Point Needs Rollback Key at the key of an op that does mark it when you want the two tied together.

For real compensation, put a second Windows service op with the opposite action in the rollback branch of your rule.

Combining with other ops

Wrap the op in IF var condition THEN to act only on the environments that should be touched, and read windows_service_status afterwards with the same op to branch on the outcome.

FOREACH CI gives you per-server control the built-in loop does not, which is what you want when each server needs a different service name.

Pair a stop with Run a Remote Script for the file copy in between, then a start to finish. Add Sleep for a number of seconds between the stop and the copy when the service holds file locks after it reports stopped.

When a restart is flaky on a particular box, put the op inside RETRY with Errors left at fail, so a failed attempt is retried rather than logged and ignored.

Use FAIL after a status check to stop the job on a state you refuse to proceed from.

Examples

Restart a service on two hosts and refuse to wait more than a minute at either end.

Server                              bar_host, baz_host
Service                             foosvc
Action                              restart
Errors                              fail
Wait until status                   ticked
Wait timeout (seconds, 0=infinite)  60

Stop a service before a deployment, tolerating hosts where it is not installed. silent covers the stop command and nothing else: an entry in ${target_servers} that does not resolve to a resource, or a host that refuses the connection, still fails the step. Hosts without the service record STOPPED with no command sent, so read the recorded statuses afterwards if that matters.

Server                              ${target_servers}
Service                             bazagent
Action                              stop
Errors                              silent
Wait until status                   ticked
Wait timeout (seconds, 0=infinite)  30

Read the state across a group of servers without changing anything, then branch on it.

Server                              ${app_servers}
Service                             fooworker
Action                              status
Errors                              fail
Wait until status                   unticked
Wait timeout (seconds, 0=infinite)  0

followed by:

IF var condition THEN
  When: Any
    windows_service_status.bar_host   NOT EQUALS   RUNNING