Guides
Running in CI
A CI job needs three things from a test runner: one number that says whether to stop the pipeline, files the CI system can show, and no surprises between the laptop and the runner. vero gives the first through its exit code, the second through four export flags, and the third by reading everything it needs from the environment and the command line, never from a file the plan points at. This page says which flags to set and why, and ends with a job skeleton.
The exit code is the gate
Read the exit code, not the JUnit file. Every task ends in one of seven states, and the exit code is computed from them in one place, by precedence:
| Code | Meaning |
|---|---|
| 64 | the plan is invalid and nothing was sent |
| 3 | the run did not finish: a task was blocked or not run, and nothing is known about it |
| 2 | a task errored or timed out |
| 1 | an assertion failed |
| 4 | every task passed but a teardown failed |
| 0 | every task passed, or was skipped by its own guard |
When several apply, the one higher in the table wins. A run that did not finish outranks one that failed, because an incomplete run may hide more failures than the one it reported; a job that treats 3 like 1 learns nothing from a plan whose second half never ran. States and exit codes has how each state is reached.
JUnit, TAP and the HTML report never change the code. They are written after the code is known, and a file vero cannot write raises the code to at least 2, so a missing artefact directory fails the job rather than passing it quietly:
$ vero run --junit /nonexistent/dir/j.xml examples/rate-limit.yaml
vero run: open /nonexistent/dir/j.xml: no such file or directory
4 passed 0 failed 0 errored 0 timed out 0 skipped 0 blocked 0 not run (4 tasks, 12 requests, 0.0s)
$ echo $?
2
Flags for a job
--strict turns the two lint warnings into load errors. An all over a collection nothing in
the step sizes, or a unique plan that never uses run.id, then exits 64 before any socket
opens. On a laptop the warning is enough; in CI a warning nobody reads is a pass.
--strict-paths is off by default and, in the words of its help text, meant for CI. It fails a
step when a path an assertion reads is absent from the result and not listed in the step's
optional:. The hole it closes is a typo that holds: body.eror == nil is true on every
response, because body.eror is nil on all of them. Without the flag the plan below passes;
with it:
FAIL health › step 1 › assert #2 typo.yaml:10
GET http://127.0.0.1:18080/healthz HTTP/1.1 200 OK 1.3ms
...
< {"status":"ok"}
assert body.eror == nil failed
--strict-paths: body.eror is absent from the result and not declared in optional:
typo.yaml:10
An absence the plan expects goes in optional:, optional: [body.debug], and stays legal. The
rules for which paths are checked are on Assertions.
--jobs defaults to 4 on this machine (GOMAXPROCS). --jobs 1 is deterministic against a
deterministic server and is the right setting for a first CI run of a new plan; raise it once the
plan passes --repeat at the higher count (below). --timeout caps the whole run, and
--finally-timeout, 30 s by default, caps the teardown after it; a job with its own wall-clock
limit should set the run timeout below that limit, so the finally tasks still get to run and
the summary still gets printed.
--include-tag slow runs the tasks tagged slow, which are skipped by default. Put them in a
nightly job and keep the per-commit job without the flag. A task downstream of an excluded one
is blocked, with a reason that names the tag, and the run exits 3; give it the same tag.
--only task runs one task and what it needs, for a focused job. The summary says what was left
out and adds a warning that a partial run cannot check holds against the tasks it skipped:
only --only pagination: 2 of 4 tasks selected; left out: rate-limit-trips-on-fourth, filtering
only a selected run does not check holds against the tasks it left out; a pass here can still flake in the whole plan
2 passed 0 failed 0 errored 0 timed out 0 skipped 0 blocked 0 not run (2 tasks, 5 requests, 0.0s)
Colour is used only when stdout is a terminal, so a CI log gets none; NO_COLOR=1 turns it off
where a runner allocates a pseudo-terminal.
Artefacts
The four export flags can all be set on one run:
vero run --strict-paths --junit junit.xml --tap results.tap --events events.ndjson --html report.html examples/rate-limit.yaml
That run against the fixture server printed one summary line and exited 0:
run 01M3ZSJ00AQTR1D5GZ4DD8712Y rate-limit.yaml --jobs 4 lifecycle reset
4 passed 0 failed 0 errored 0 timed out 0 skipped 0 blocked 0 not run (4 tasks, 12 requests, 0.0s)
The event log is the record. --events events.ndjson writes one JSON object per line as the
run happens, one write per line, so a run the CI system kills leaves a file whose every complete
line parses. That run's log has 53 lines: run.start, 4 task.start, 9 step.start, 12
request, 13 assert, 9 step.end, 4 task.end, run.end. It carries all seven states, the
bodies up to --events-body-cap (65536 bytes each), and every flag the run was given.
--events - writes it to stdout and moves the human report to stderr, for a job that pipes the
log elsewhere. The format is vero.events/v1.
JUnit shows green for things that are not. JUnit XML has passed, failed, errored and
skipped, and most CI systems render <skipped> as green. vero maps skipped, blocked and
not_run all to <skipped>, each with a message naming the cause, because a labelled wrong
state beats a silent one. A job that reads only the JUnit file sees this run as all green:
<testsuite name="nightly" tests="3" failures="0" errors="0" skipped="2" time="0.001" timestamp="2026-10-03T02:30:12" id="01M3ZSJ02KDXS5XASYPTHVF1PS" file="blocked.yaml">
<properties>
<property name="vero.runId" value="01M3ZSJ02KDXS5XASYPTHVF1PS"></property>
<property name="vero.summary" value="1 passed 0 failed 0 errored 0 timed out 1 skipped 1 blocked 0 not run (3 tasks, 1 requests, 0.0s)"></property>
</properties>
<testcase name="health" classname="nightly.tasks" time="0.001"></testcase>
<testcase name="optional-feature" classname="nightly.tasks" time="0.000">
<skipped message="when guard false: env.feature == "on""></skipped>
</testcase>
<testcase name="uses-it" classname="nightly.tasks" time="0.000">
<skipped message="blocked by optional-feature"></skipped>
</testcase>
</testsuite>
The run exited 3. The vero.summary property carries the summary line, so the counts are in the
file for anyone who looks, and after --only a vero.only property says what was selected, so
a partial run is never mistaken for the whole plan. One <testsuite> per run, <testcase> per
task, classname ending in .tasks or .finally.
TAP version 14 has the same labels after # SKIP and a YAML block under each not ok:
TAP version 14
1..4
ok 1 - login
ok 2 - rate-limit-trips-on-fourth
ok 3 - pagination
ok 4 - filtering
The HTML report is one file, no script, nothing loaded from elsewhere, with a dark mode. It
opens with the exit code and its meaning, then lists the tasks with failed first, then errored,
timed out, blocked, not run, skipped and passed last, the failures expanded with their exchange
and values. It is rendered from the event log, so vero report --html report.html events.ndjson produces the same page later from a saved log: the file it wrote for the run
above is byte-identical to the one --html wrote during it. Keep the log and the page can be
regenerated; keep only the page and it cannot.
Snapshots
A matches snapshot "file" assertion compares against a JSON file beside the plan. In CI the
file must already be committed: a missing one fails the step with no snapshot at <file>; run with --update-snapshots to write it. --update-snapshots writes every snapshot file instead of
comparing, fails each assertion it wrote, and exits 1, so it is never the green run. Run it on a
laptop, read the diff, commit the files, and let CI compare:
$ vero run --update-snapshots snap/plan.yaml
assert body matches snapshot "health.json" failed
snapshot health.json written, not compared
0 passed 1 failed ... (1 tasks, 1 requests, 0.0s)
$ vero run snap/plan.yaml
1 passed 0 failed 0 errored 0 timed out 0 skipped 0 blocked 0 not run (1 tasks, 1 requests, 0.0s)
A value that holds a secret is refused rather than written.
Flake gating
--repeat N runs the plan N times against the same server and prints a pass count per task; the
exit code is the worst of the N runs, so one failure in twenty fails the job. --repeat N --compare-jobs runs N times at --jobs 1 and N at --jobs, prints both pass rates and, when a
task passes serially and flakes in parallel, names the task it overlapped with and the holds
that would separate them. Run the comparison when a plan is new or when a task starts flaking;
run the plain repeat as a nightly guard. Reading a flake report and the --jobs hint
explains the output line by line.
The environment is part of the plan
vero plan in CI needs the same environment as vero run, because the plan resolves against it
at load. ${VAR} and ${VAR:-default} read it; a secretRef reads it and refuses a value under
4 bytes; a script step's runtime must be on PATH, or the plan does not load; and vero plan
opens every Postgres connection a sql step reads through, to warn when the role could write.
A vero plan --strict step in the job therefore catches a missing variable or runtime before
the run starts, with exit 64 and the name of what is missing.
Secrets come from the CI system's secret store into environment variables, and the plan names
the variable with secretRef. Nothing secret goes in the plan file, and every output vero writes
is scrubbed of the values it knows are secret. Secrets has the rules.
TLS material comes from flags, never from the plan, because certificates belong to the
environment a run targets. --ca-cert file adds a private CA on top of the system roots, so
public hosts still verify; --client-cert and --client-key go together, and a key that group
or others can read draws vero run: warning: --client-key <file> is readable by group or others (mode 0644); chmod 600 it. A CI secret store that writes the key with loose permissions gets
that warning on every run until the job tightens them.
A job skeleton
Shell only, no vendor assumed. Each line's job is in the comment above it.
set -eu
# One static binary; build it, or fetch the one the release job built.
CGO_ENABLED=0 go build -o bin/ ./cmd/vero
# The target. Start it here, or export the URL of one that is already up.
export VERO_BASE_URL=http://127.0.0.1:18080
export ORDERS_API_KEY="$CI_ORDERS_API_KEY"
# Load the plan with the job's real environment. A missing variable, runtime or role is exit 64
# here, before anything is sent, and a lint is an error.
bin/vero plan --strict plan.yaml
# Run with strict paths, and keep every artefact. --events is the record; the others are views.
# Do not exit here: the artefacts upload whatever the code was.
set +e
bin/vero run --strict-paths \
--junit artefacts/junit.xml --tap artefacts/results.tap \
--events artefacts/events.ndjson --html artefacts/report.html \
plan.yaml
code=$?
set -e
# Upload artefacts/ with the CI system's own mechanism here.
# The exit code is the verdict. 64, 3, 2, 1, 4 and 0 each mean one thing.
exit $code
The set +e around the run is deliberate: the artefacts have to be uploaded on a failure, which
is when they matter, so the script keeps the code and exits with it after the upload. A nightly
variant adds --include-tag slow and --repeat 10, and a focused job adds --only task.
Diagnosing a failing plan is where to go when the code is not 0.
This page is docs/content/guides/ci.md in the repository.