Save my stats
Takes a memory and CPU reading of the Clarive process that is running the rule and files it under a label you choose. The readings show up in the System Statistics screen, where the label is one of the filters, so you can plot your own measurement points next to the ones Clarive takes on its own.
Drop one op before a heavy stretch of a rule and another after it, with different labels, and the two lines on the chart tell you what that stretch costs. It is a measuring instrument, not something a production rule needs.
Fields¶
Statistics Key¶
The label the sample is filed under. A new op arrives with rule.stats already in the box, and the
form will not accept an empty value.
Any text works. There is no list of allowed labels, no validation and no length limit that you will
meet in practice. Whatever you type becomes an entry in the Key filter on the
System Statistics screen, and each distinct label is drawn as its own
line on the charts. Give each measurement point in a rule its own label. Two ops sharing a label and
landing in the same second are averaged into a single point on the chart, so a before and an after
reading under one label cancel each other out.
The field expands ${var} placeholders, so ${rule_name}.before is legal. Be careful with labels
that change on every run: each new value adds another entry to the filter and another line to the
chart, and there is no way to merge them afterwards. Keep the varying part out of the label and use a
fixed one such as myrule.before.
Hand-edit a rule to leave the field empty and the sample is filed under custom.stats instead. The
form cannot produce that.
What a sample contains¶
One row per run of the op, holding:
| Recorded | Meaning |
|---|---|
| memory | total for the process and every process it has started, in kB, MB and GB |
| CPU | percentage and accumulated CPU seconds, for that one process |
| timestamp | to the millisecond |
| host and instance | which machine and which Clarive instance produced it, so several nodes stay apart |
| process name and pid | which process was measured |
| since start | seconds since Clarive started up in that process |
| since last | seconds since the previous sample that same process saved |
Memory and CPU do not cover the same ground. Memory walks down into everything the process has started, so a rule that shells out to a build tool shows the build tool's footprint. CPU is read off the single process, so that same build tool contributes nothing to the CPU figure.
"Since last" is measured against the previous sample from the same process, whichever rule saved it, so a second rule running in the same worker resets the clock on yours. On the first sample a process ever takes, it is counted from Clarive starting up, which makes it equal to "since start" rather than zero.
The sample covers the Clarive process itself. Work a rule pushes out to a remote server or into the database is not in the numbers.
Taking a sample means reading the machine's full process table, twice. That is cheap next to a deployment and expensive inside a loop that runs hundreds of times, so place the op at a few deliberate points.
What comes back¶
With a Return Key set, you get a small map with status, always success, and key, the label
that was actually used. status says the op ran, not that anything useful was measured, so there is
no point testing it.
Setting Return Key to = merges both entries into the top of the stash, which drops anything the
rule already kept under key or status. Use an ordinary name or leave it blank.
A line naming the label is written to the log. The text is fixed and no field changes it. Inside a job it is written at debug level, so it only shows up when the job log is set to include debug output.
Retention¶
Samples are cleaned up by the purge daemon, which removes anything older than the
Number of days to keep process statistics setting, 30 days out of the box. See
Purge Daemon Configuration if you need a longer window for a
measurement campaign.
The System Statistics screen reads at most 2000 samples per query and 500 by default, so a label written in a tight loop fills the window and hides everything else.
Failure and rollback passes¶
The op has no input it can reject beyond an empty label the form will not let you save, so its
Error Handling setting rarely has anything to catch. When the host's process table cannot be read,
the sample is still written and still reports success, with the memory and CPU figures at zero and
the pid recorded as one the reading could not cover. Flat zeroes on the chart mean that, not an idle
server.
The op runs again on a rollback pass unless you untick Run Rollback, writing
a second sample under the same label. That is what you want when you are measuring a rollback, and
misleading when you are not.
Combining with other ops¶
Put one op either side of the section you are measuring, with distinct labels, then compare the two lines on the chart.
Statistics Key myrule.before
... the ops you want to measure ...
Statistics Key myrule.after
For rules that take a different path depending on the data, sample inside the branch rather than around it, so a run that skipped the work does not flatten the average. Put the op inside IF var condition THEN next to the work itself.
To time a stretch of a rule rather than measure it, Get Date with %s
gives you seconds since the epoch at two points and leaves the arithmetic to
Server CODE. That is lighter than a full sample.
For the cost of a rule broken down by op instead of by process, look at Rule Profiling, which needs no op at all.
Examples¶
Measure a whole rule under one label, with the reading taken at the end.
Statistics Key myrule.stats
Bracket the expensive part of a rule.
Statistics Key myrule.checkout.before
... checkout and build ops ...
Statistics Key myrule.checkout.after
Keep the samples from a scheduled rule apart from everything else, without letting the label vary per run.
Statistics Key nightly.foo_job