Windows service
Starts, stops or restarts a Windows service on one or more servers, or reads its current state without touching it. Each server you pick is handled in turn, and the state the service ended in is written back into the stash keyed by server name.
Reach for this when the service state itself is what you care about. To run an arbitrary command on the same box, use Run a Remote Script. To find out which services exist before you act on one, run List Windows Services first.
Every field on this form expands ${var} placeholders against the stash before the op runs, the
number field included.
Fields¶
Server¶
The Windows servers to act on. The picker lists the resources that can take an agent connection, plus the variables that resolve to one, so you can select boxes by hand or feed in a variable that a previous op filled.
More than one server can be selected. They are processed one after another in the order shown, not in parallel, so a slow wait on the first server delays the rest. A variable holding a comma-separated list is split on the commas and treated the same way.
The field is required. A value that does not resolve to a real resource ends the step with a
resource error, and the Errors setting does not soften that. The same is true of a server that
cannot be reached: connection problems always fail the step.
A server marked inactive is not contacted. A warning goes into the job log, the server is recorded
with the status SKIPPED, and the run moves on to the next one.
The connection uses whatever the server resource itself is configured with, including its user and its agent timeout. This op has no user field of its own.
Only the Clax and Worker connection methods hand back the command result in the shape this op reads,
so Connect by on the server resource has to be set to one of those. With SSH, FTP or Balix the step
ends with an error the first time it reads the service state. SSH is the one that catches people
out: it is what a resource falls back to when nobody ever picked a method, and it is tried ahead of
Clax on a resource that has both enabled.
Service¶
The short service name, the one you would pass to net start, not the display name shown in the
Services console. wuauserv works, Windows Update does not, because the value goes into the
command unquoted and the space splits it.
The field is required. Nothing here checks that the service exists before acting on it.
Action¶
What to do. Defaults to start when left unset.
| Action | What happens |
|---|---|
start |
starts the service, unless it is already running |
restart |
stops it, waits until it reports stopped, then starts it |
stop |
stops the service, unless it is already stopped |
status |
reads the state and changes nothing |
start and stop check the current state first and do nothing when the service is already in the
target state. That is a silent no-op at the default log level: no command runs and nothing is
written to the job log, but the status still gets recorded.
restart always waits for the stopped state between the two halves, whether or not Wait until
status is ticked. It uses the same Wait timeout value. With the default timeout of 0 that wait
has no upper bound, so a service that refuses to stop leaves the job sitting there indefinitely. Set
a timeout when you choose restart.
status ignores Errors, Wait until status and Wait timeout entirely. It reads once and
records what it found.
Errors¶
How a failed start or stop is treated. Defaults to fail.
| Errors | Effect |
|---|---|
fail |
the step fails and the job stops |
warn |
the failure is recorded at debug level and the run continues |
silent |
the failure is swallowed and the run continues |
warn does not put a warning in the job log. The message comes out at debug level, so unless the
job is running with debug logging turned on, warn and silent look identical from the outside. In
both cases the recorded status for the failed server is whatever the service actually reports, which
is usually the state it was already in, and the run carries on to the next server. The exception is a
ticked Wait until status with no timeout: the wait that follows the failed command polls for a
state that is never coming, and no later server is ever reached.
This setting covers only the start and stop commands. Resource lookup failures, connection failures and agent timeouts fail the step no matter what you choose here.
A command counts as successful when it returns a zero exit code, or when it returns a non-zero code but still prints a "was ... successfully" message. That second case exists because some Windows builds report success with a non-zero code.
Wait until status¶
Off by default. When ticked, the op keeps polling after the action until the service reaches the
state that action implies: running for start and restart, stopped for stop. Polling happens
every 2 seconds.
With Wait timeout left at its default of 0 that polling has no upper bound. A service that failed
to start, or one that is slower than you expected, holds the job in the loop with nothing above debug
level in the log, and the servers after it in the list are never touched. Give the wait a timeout
whenever you tick this box.
When the wait gives up because Wait timeout expired, it writes a warning to the job log and
carries on. It does not fail the step, and Errors has no say in it. The status recorded for that
server is the state the service was in at the moment the wait gave up, so read the result if the
difference matters to you.
Leaving it off does not skip the status read. The op still queries the service once after the action
and records whatever it sees at that instant, which for a slow service is often START_PENDING
rather than RUNNING.
Wait timeout (seconds, 0=infinite)¶
A blank field behaves the same as 0. Any value above zero is only checked between polls, so the
wait ends on the first 2-second boundary past the limit rather than exactly on it, and a 1-second
timeout still costs you 2 seconds.
The timeout bounds the polling loop only. Each individual command sent to the server is bounded separately by the agent timeout configured on the server resource, and that one does fail the step when it trips.
Status values¶
The status recorded for each server is one of these, in English, whatever display language the Windows box runs.
| Status | Meaning |
|---|---|
RUNNING |
the service is running |
STOPPED |
the service is stopped |
START_PENDING |
the service is starting |
STOP_PENDING |
the service is stopping |
PAUSE_PENDING |
the service is pausing |
CONTINUE_PENDING |
the service is resuming |
PAUSED |
the service is paused |
SKIPPED |
the server was inactive and never contacted |
STOPPED doubles as the fallback. When the state cannot be read at all, including the case where
the service name does not exist on that box, the op reports STOPPED.
A typo in Service therefore passes for success under both status and stop. stop reads the
state, gets STOPPED, concludes the service is already stopped, sends no command and records
STOPPED, all without a word above debug level. start and restart do expose the mistake, because
the start command runs against the bad name and fails. Confirm the name with
List Windows Services if the result surprises you.
What lands in the stash¶
The op writes windows_service_status, a map of server name to status:
windows_service_status:
bar_host: RUNNING
baz_host: SKIPPED
The key is fixed. A second Windows service op later in the rule overwrites it, so read the value
before the next one runs, or copy it out with SET VAR.
Filling in Return Key on the op stores the same map under your own name as well, which is the
safer option when the rule touches several services.
Two selected servers that share a name collapse into one entry.
Rollback¶
This op never marks the job as needing a rollback on its own, and it registers no undo. Stopping a service here and failing two ops later leaves the service stopped.
Rollback Needed Always does work. That marker is written for you before the op runs, so the job
gets a rollback pass and the op runs again on the way back with the same configuration, which for
stop means a second stop.
Rollback Needed Before and Rollback Needed After are modes the op itself has to act on, and this
one does not, so neither ever marks anything. They are not inert, though. Picking either fills
Needs Rollback Key with the op's own name, and from then on the op is conditional on the rollback
pass: it runs on the way back only if that key was marked by something. Left at the auto-filled name
nothing marks it, so the effect of choosing one of these two modes is that the op stops running
during rollback. Point Needs Rollback Key at the key of an op that does mark it when you want the
two tied together.
For real compensation, put a second Windows service op with the opposite action in the rollback
branch of your rule.
Combining with other ops¶
Wrap the op in IF var condition THEN to act only on the
environments that should be touched, and read windows_service_status afterwards with the same op
to branch on the outcome.
FOREACH CI gives you per-server control the built-in loop does not, which is what you want when each server needs a different service name.
Pair a stop with Run a Remote Script for the file copy in between, then
a start to finish. Add Sleep for a number of seconds between the stop and the
copy when the service holds file locks after it reports stopped.
When a restart is flaky on a particular box, put the op inside RETRY with
Errors left at fail, so a failed attempt is retried rather than logged and ignored.
Use FAIL after a status check to stop the job on a state you refuse to
proceed from.
Examples¶
Restart a service on two hosts and refuse to wait more than a minute at either end.
Server bar_host, baz_host
Service foosvc
Action restart
Errors fail
Wait until status ticked
Wait timeout (seconds, 0=infinite) 60
Stop a service before a deployment, tolerating hosts where it is not installed. silent covers the
stop command and nothing else: an entry in ${target_servers} that does not resolve to a resource,
or a host that refuses the connection, still fails the step. Hosts without the service record
STOPPED with no command sent, so read the recorded statuses afterwards if that matters.
Server ${target_servers}
Service bazagent
Action stop
Errors silent
Wait until status ticked
Wait timeout (seconds, 0=infinite) 30
Read the state across a group of servers without changing anything, then branch on it.
Server ${app_servers}
Service fooworker
Action status
Errors fail
Wait until status unticked
Wait timeout (seconds, 0=infinite) 0
followed by:
IF var condition THEN
When: Any
windows_service_status.bar_host NOT EQUALS RUNNING