LoadWave
Reference

Metrics

Every metric LoadWave records, the labels it attaches, how failures are captured, and how to emit your own.

What LoadWave measures, and where the numbers can mislead you. How measurements from many machines become one number is a separate topic — see Aggregation.

The built-in metrics

The HTTP client records these automatically. Durations are in milliseconds; rates are fractions of one.

MetricKindWhat it is
http_reqscounterRequests issued.
http_req_durationtrendThe whole request, first byte written to last byte read.
http_req_waitingtrendTime to first byte — the server's own think time, with connection setup and body transfer excluded.
http_req_connectingtrendTCP setup. Recorded only on a fresh connection.
http_req_tls_handshakingtrendTLS handshake. Fresh connections only.
http_req_failedrateShare of requests judged unsuccessful.
http_req_bytes_in / _outcounterBytes read and written.
iterationscounterCompleted scenario iterations.
iteration_durationtrendOne iteration, excluding think time.
iteration_failedrateShare of iterations that returned an error.
vusgaugeVirtual users currently executing.
checksrateShare of checks that passed.
errorscounterErrors reported by scenarios.

The four metric kinds

KindAggregated byExample
counteraddinghttp_reqs
gaugeadding across nodes at one instant, never across instantsvus
trendmerging HDR histograms, so percentiles stay correcthttp_req_duration
rateadding numerator and denominator separatelyhttp_req_failed

http_req_duration versus http_req_waiting

duration is what a user experiences. waiting is what the server is responsible for. When duration rises but waiting does not, the server is fine and you are looking at connection churn, TLS, or a saturated generator. That distinction is usually the first thing worth checking.

Connection metrics are only recorded on fresh connections

A reused connection has no setup cost, and recording a zero for it would drag every percentile toward zero. So http_req_connecting and http_req_tls_handshaking count only the requests that actually opened a connection — their count is deliberately lower than http_reqs.

What counts as a failure

By default, a transport error or a status of 400 or above. A step's expect list overrides that, so a scenario probing for a 404 is not charged a failure for finding one.

An HTTP 500 is not a transport error. At load-test altitude it is a successful exchange with a bad status — the scenario gets the response and can decide what it means.

Labels

Every observation carries labels. The dashboard groups and filters on them.

LabelValue
scenarioThe scenario that produced it.
nameThe call site's stable identity.
methodHTTP method.
statusStatus code as a string, or 0 if no response arrived.
errorA bounded classification of a transport failure.
checkThe name given to a check.
expectedWhether the scenario anticipated the failure.

Cardinality is the thing to be careful about

Every distinct label combination is a separate time series held in memory for the length of the run. A label with unbounded values — a user id, a timestamp, a raw error string — creates a series per value and will exhaust memory.

Two defences are built in:

Request names collapse automatically. DeriveRequestName replaces path segments that look like identifiers with *, so /users/1 and /users/2 share one series instead of two million. Numeric segments, UUIDs and long hex strings are collapsed. Set name explicitly whenever the heuristic would still leave you with high cardinality.

Transport errors are classified, not quoted. The error label is drawn from a fixed vocabulary — timeout, connection_refused, connection_reset, dns, tls, eof, too_many_redirects, canceled, unknown — because raw error text embeds addresses and ports and would be unbounded.

Beyond that, each node caps itself at 5,000 series and the coordinator caps itself again. Past the cap, observations for new series are dropped and counted. A bounded, visibly lossy run beats an out-of-memory kill, and the report says so explicitly:

warning: 1,284 samples were dropped by nodes that hit their series cap;
       a high-cardinality tag is the usual cause

If you see that, find the runaway tag.

Why a request failed

Metrics can say that 3% of requests failed. They cannot say that they were all 502 payment declined on one endpoint, and that is usually the question.

So alongside the counters, LoadWave keeps a small table of failure kinds, aggregated by request name, method, status and transport error class — every one of those already bounded. Each row carries a short excerpt of what the server actually said: the response body, or the transport error's text, collapsed to a single line and clipped to 240 characters.

The excerpt is captured the first time each kind is seen and not again. A run where everything fails would otherwise spend its time copying error strings, and the second occurrence's body almost never says anything the first did not. The count keeps rising regardless.

Distinct kinds are capped, per node and again on the coordinator. Past the cap new kinds are dropped and counted, and the dashboard and report both say so rather than presenting a partial list as complete.

This appears as the Failed requests panel in the dashboard and a section of the HTML report.

Custom metrics

Scenarios can emit their own. They are ordinary metrics: they appear in the dashboard with percentiles and can carry thresholds.

const cartValue = "cart_value"

func checkout(ctx context.Context, vu *loadwave.VU) error {
    // ...
    vu.Metrics().Trend(cartValue, vu.Labels(), receipt.Total)
    return nil
}
vu.Metrics().Count("orders_placed", vu.Labels(), 1)
vu.Metrics().Rate("payment_succeeded", vu.Labels(), ok)
vu.Metrics().Gauge("queue_depth", vu.Labels(), float64(depth))

Trend values must be in the same 0.1ms-to-60s window as the built-ins, scaled into whatever unit makes sense for your metric.

Add labels with vu.Tag, which applies to everything the VU emits for the rest of the iteration:

vu.Tag("flow", "purchase")

Keep tag values to a small fixed set. A tag per customer is a series per customer.

Checks

A check records a named assertion and returns its result, so it composes into control flow:

if !vu.Check("logged in", resp.StatusCode == http.StatusOK) {
    return fmt.Errorf("login failed: %d", resp.StatusCode)
}

A failing check does not by itself fail the iteration — return an error to do that. Checks measure how often something held; errors say the iteration did not accomplish what it set out to.

Checkf adds a message logged on failure. The message is not used as a label, so it can safely contain specific values.

Use the totals, not the per-series numbers

The API exposes both series — one entry per label combination — and totals, one correctly merged aggregate per metric. Always use totals for a whole-run figure. Folding series yourself produces plausible-looking numbers that disagree with the thresholds, because means need re-weighting by count and percentiles need the distributions.

The same applies to endpoints: those percentiles are recomputed from each endpoint's merged distribution across all of its status codes. Taking the maximum of the per-status percentiles — the obvious shortcut — reports the tail of whichever status happened to be slowest, and the two diverge most exactly when an endpoint starts failing.

See also

On this page