Guides

Reading a flake report and the --jobs hint

vero run --repeat N plan.yaml runs the plan N times and reports, per task, how often it passed. A task whose outcome varied is marked FLAKY, with the assertions that varied and every value their paths took.

vero run --repeat N --compare-jobs --jobs J plan.yaml runs the N repeats at --jobs 1 and then N more at --jobs J, and prints both pass rates side by side.

Why --jobs 1 is the control

At --jobs 1 tasks run one at a time in plan order, so against a deterministic server a run is deterministic. A task that passes every run there and flakes at --jobs J is racing another task for some state: a rate-limit budget, a row, a balance. That pattern has one usual cause, a missing holds, and vero prints a hint:

hint: rate-limit-trips-on-fourth passed 20/20 at --jobs 1 and 5/20 at --jobs 8.
      likely shares state with pagination (overlapped in 15 of 15 failing runs, 3 passing), ...
      consider holds: [127.0.0.1:18080/rate-limit-trips-on-fourth] on rate-limit-trips-on-fourth and on pagination
      (the resource name is a placeholder: name it for the state they share)
      This is a heuristic: vero sees when tasks ran and which hosts they called, not the server's state.

A task that is flaky at --jobs 1 as well gets a different hint: nothing ran beside it, so no holds can fix it, and the cause is the server's state between runs or the plan's lifecycle.

A ws step, or a grpc stream with expect, collects messages for expect.timeout. When every failure of a flaky task was that window running out before expect.count messages came, the hint says so instead of suggesting holds, at either job count:

hint: ws-c01 passed 30/30 at --jobs 1 and 25/30 at --jobs 40.
      every failure was step 1's collection window: fewer than 3 messages within expect.timeout 10ms.
      The latest message that made it in at --jobs 40 arrived at 9.8ms; vero sees nothing after the window closes.
      The server's messages come later than the window allows: that is its pace, not state another task
      shares, and holds will not change it. Raise expect.timeout above what the server needs, or keep it
      and read this as a latency finding.

Under load a server's messages arrive later. G05 measured a WebSocket echo through nginx at a median of 0.5 ms for its last message at --jobs 1 and 8 ms at --jobs 40, with a worst case of 216 ms (real ws and gRPC servers). A window sized from a quiet run fails under load for that reason alone.

The hint is a heuristic

"Likely shares state with" is an inference, not a finding. vero does not see the server's state. It sees, for each run at the higher job count, when each task started and ended (the task.start and task.end events) and which hosts it sent requests to (hosts on task.end). A candidate is a task that:

  • ran at the same time as the flaky task in a run where the flaky task failed, and
  • sent requests to a host the flaky task also called.

Candidates are ranked by how many failing runs they overlapped, then by how often an overlap came with a failure rather than a pass. At --jobs 8 nearly everything overlaps something, so an innocent task can rank high; the ranking and the counts are printed so a reader can judge. The suggested resource name reuses a resource the flaky task already holds, if any; otherwise it is a placeholder built from the host and the task name.

The hint does not derive resources from request contents (API keys or paths): it cannot infer which requests share state and never reads those values.

This page is docs/content/flake-hints.md in the repository.

verodocs