Guides

Running in CI

A CI job needs three things from a test runner: one number that says whether to stop the pipeline, files the CI system can show, and no surprises between the laptop and the runner. vero gives the first through its exit code, the second through four export flags, and the third by reading everything it needs from the environment and the command line, never from a file the plan points at. This page says which flags to set and why, and ends with a job skeleton.

The exit code is the gate

Read the exit code, not the JUnit file. Every task ends in one of seven states, and the exit code is computed from them in one place, by precedence:

Code Meaning
64 the plan is invalid and nothing was sent
3 the run did not finish: a task was blocked or not run, and nothing is known about it
2 a task errored or timed out
1 an assertion failed
4 every task passed but a teardown failed
0 every task passed, or was skipped by its own guard

When several apply, the one higher in the table wins. A run that did not finish outranks one that failed, because an incomplete run may hide more failures than the one it reported; a job that treats 3 like 1 learns nothing from a plan whose second half never ran. States and exit codes has how each state is reached.

JUnit, TAP and the HTML report never change the code. They are written after the code is known, and a file vero cannot write raises the code to at least 2, so a missing artefact directory fails the job rather than passing it quietly:

$ vero run --junit /nonexistent/dir/j.xml examples/rate-limit.yaml
vero run: open /nonexistent/dir/j.xml: no such file or directory
4 passed  0 failed  0 errored  0 timed out  0 skipped  0 blocked  0 not run   (4 tasks, 12 requests, 0.0s)
$ echo $?
2

Flags for a job

--strict turns the two lint warnings into load errors. An all over a collection nothing in the step sizes, or a unique plan that never uses run.id, then exits 64 before any socket opens. On a laptop the warning is enough; in CI a warning nobody reads is a pass.

--strict-paths is off by default and, in the words of its help text, meant for CI. It fails a step when a path an assertion reads is absent from the result and not listed in the step's optional:. The hole it closes is a typo that holds: body.eror == nil is true on every response, because body.eror is nil on all of them. Without the flag the plan below passes; with it:

FAIL  health › step 1 › assert #2                                 typo.yaml:10

  GET http://127.0.0.1:18080/healthz                              HTTP/1.1 200 OK  1.3ms
    ...
    < {"status":"ok"}

  assert  body.eror == nil    failed
          --strict-paths: body.eror is absent from the result and not declared in optional:
          typo.yaml:10

An absence the plan expects goes in optional:, optional: [body.debug], and stays legal. The rules for which paths are checked are on Assertions.

--jobs defaults to 4 on this machine (GOMAXPROCS). --jobs 1 is deterministic against a deterministic server and is the right setting for a first CI run of a new plan; raise it once the plan passes --repeat at the higher count (below). --timeout caps the whole run, and --finally-timeout, 30 s by default, caps the teardown after it; a job with its own wall-clock limit should set the run timeout below that limit, so the finally tasks still get to run and the summary still gets printed.

--include-tag slow runs the tasks tagged slow, which are skipped by default. Put them in a nightly job and keep the per-commit job without the flag. A task downstream of an excluded one is blocked, with a reason that names the tag, and the run exits 3; give it the same tag.

--only task runs one task and what it needs, for a focused job. The summary says what was left out and adds a warning that a partial run cannot check holds against the tasks it skipped:

only  --only pagination: 2 of 4 tasks selected; left out: rate-limit-trips-on-fourth, filtering
only  a selected run does not check holds against the tasks it left out; a pass here can still flake in the whole plan
2 passed  0 failed  0 errored  0 timed out  0 skipped  0 blocked  0 not run   (2 tasks, 5 requests, 0.0s)

Colour is used only when stdout is a terminal, so a CI log gets none; NO_COLOR=1 turns it off where a runner allocates a pseudo-terminal.

Artefacts

The four export flags can all be set on one run:

vero run --strict-paths --junit junit.xml --tap results.tap --events events.ndjson --html report.html examples/rate-limit.yaml

That run against the fixture server printed one summary line and exited 0:

run 01M3ZSJ00AQTR1D5GZ4DD8712Y  rate-limit.yaml  --jobs 4  lifecycle reset

4 passed  0 failed  0 errored  0 timed out  0 skipped  0 blocked  0 not run   (4 tasks, 12 requests, 0.0s)

The event log is the record. --events events.ndjson writes one JSON object per line as the run happens, one write per line, so a run the CI system kills leaves a file whose every complete line parses. That run's log has 53 lines: run.start, 4 task.start, 9 step.start, 12 request, 13 assert, 9 step.end, 4 task.end, run.end. It carries all seven states, the bodies up to --events-body-cap (65536 bytes each), and every flag the run was given. --events - writes it to stdout and moves the human report to stderr, for a job that pipes the log elsewhere. The format is vero.events/v1.

JUnit shows green for things that are not. JUnit XML has passed, failed, errored and skipped, and most CI systems render <skipped> as green. vero maps skipped, blocked and not_run all to <skipped>, each with a message naming the cause, because a labelled wrong state beats a silent one. A job that reads only the JUnit file sees this run as all green:

<testsuite name="nightly" tests="3" failures="0" errors="0" skipped="2" time="0.001" timestamp="2026-10-03T02:30:12" id="01M3ZSJ02KDXS5XASYPTHVF1PS" file="blocked.yaml">
  <properties>
    <property name="vero.runId" value="01M3ZSJ02KDXS5XASYPTHVF1PS"></property>
    <property name="vero.summary" value="1 passed  0 failed  0 errored  0 timed out  1 skipped  1 blocked  0 not run   (3 tasks, 1 requests, 0.0s)"></property>
  </properties>
  <testcase name="health" classname="nightly.tasks" time="0.001"></testcase>
  <testcase name="optional-feature" classname="nightly.tasks" time="0.000">
    <skipped message="when guard false: env.feature == &#34;on&#34;"></skipped>
  </testcase>
  <testcase name="uses-it" classname="nightly.tasks" time="0.000">
    <skipped message="blocked by optional-feature"></skipped>
  </testcase>
</testsuite>

The run exited 3. The vero.summary property carries the summary line, so the counts are in the file for anyone who looks, and after --only a vero.only property says what was selected, so a partial run is never mistaken for the whole plan. One <testsuite> per run, <testcase> per task, classname ending in .tasks or .finally.

TAP version 14 has the same labels after # SKIP and a YAML block under each not ok:

TAP version 14
1..4
ok 1 - login
ok 2 - rate-limit-trips-on-fourth
ok 3 - pagination
ok 4 - filtering

The HTML report is one file, no script, nothing loaded from elsewhere, with a dark mode. It opens with the exit code and its meaning, then lists the tasks with failed first, then errored, timed out, blocked, not run, skipped and passed last, the failures expanded with their exchange and values. It is rendered from the event log, so vero report --html report.html events.ndjson produces the same page later from a saved log: the file it wrote for the run above is byte-identical to the one --html wrote during it. Keep the log and the page can be regenerated; keep only the page and it cannot.

Snapshots

A matches snapshot "file" assertion compares against a JSON file beside the plan. In CI the file must already be committed: a missing one fails the step with no snapshot at <file>; run with --update-snapshots to write it. --update-snapshots writes every snapshot file instead of comparing, fails each assertion it wrote, and exits 1, so it is never the green run. Run it on a laptop, read the diff, commit the files, and let CI compare:

$ vero run --update-snapshots snap/plan.yaml
  assert  body matches snapshot "health.json"    failed
          snapshot health.json written, not compared
0 passed  1 failed  ...   (1 tasks, 1 requests, 0.0s)
$ vero run snap/plan.yaml
1 passed  0 failed  0 errored  0 timed out  0 skipped  0 blocked  0 not run   (1 tasks, 1 requests, 0.0s)

A value that holds a secret is refused rather than written.

Flake gating

--repeat N runs the plan N times against the same server and prints a pass count per task; the exit code is the worst of the N runs, so one failure in twenty fails the job. --repeat N --compare-jobs runs N times at --jobs 1 and N at --jobs, prints both pass rates and, when a task passes serially and flakes in parallel, names the task it overlapped with and the holds that would separate them. Run the comparison when a plan is new or when a task starts flaking; run the plain repeat as a nightly guard. Reading a flake report and the --jobs hint explains the output line by line.

The environment is part of the plan

vero plan in CI needs the same environment as vero run, because the plan resolves against it at load. ${VAR} and ${VAR:-default} read it; a secretRef reads it and refuses a value under 4 bytes; a script step's runtime must be on PATH, or the plan does not load; and vero plan opens every Postgres connection a sql step reads through, to warn when the role could write. A vero plan --strict step in the job therefore catches a missing variable or runtime before the run starts, with exit 64 and the name of what is missing.

Secrets come from the CI system's secret store into environment variables, and the plan names the variable with secretRef. Nothing secret goes in the plan file, and every output vero writes is scrubbed of the values it knows are secret. Secrets has the rules.

TLS material comes from flags, never from the plan, because certificates belong to the environment a run targets. --ca-cert file adds a private CA on top of the system roots, so public hosts still verify; --client-cert and --client-key go together, and a key that group or others can read draws vero run: warning: --client-key <file> is readable by group or others (mode 0644); chmod 600 it. A CI secret store that writes the key with loose permissions gets that warning on every run until the job tightens them.

A job skeleton

Shell only, no vendor assumed. Each line's job is in the comment above it.

set -eu

# One static binary; build it, or fetch the one the release job built.
CGO_ENABLED=0 go build -o bin/ ./cmd/vero

# The target. Start it here, or export the URL of one that is already up.
export VERO_BASE_URL=http://127.0.0.1:18080
export ORDERS_API_KEY="$CI_ORDERS_API_KEY"

# Load the plan with the job's real environment. A missing variable, runtime or role is exit 64
# here, before anything is sent, and a lint is an error.
bin/vero plan --strict plan.yaml

# Run with strict paths, and keep every artefact. --events is the record; the others are views.
# Do not exit here: the artefacts upload whatever the code was.
set +e
bin/vero run --strict-paths \
  --junit artefacts/junit.xml --tap artefacts/results.tap \
  --events artefacts/events.ndjson --html artefacts/report.html \
  plan.yaml
code=$?
set -e

# Upload artefacts/ with the CI system's own mechanism here.

# The exit code is the verdict. 64, 3, 2, 1, 4 and 0 each mean one thing.
exit $code

The set +e around the run is deliberate: the artefacts have to be uploaded on a failure, which is when they matter, so the script keeps the code and exits with it after the upload. A nightly variant adds --include-tag slow and --repeat 10, and a focused job adds --only task. Diagnosing a failing plan is where to go when the code is not 0.

This page is docs/content/guides/ci.md in the repository.

verodocs