Guides

Diagnosing a failing plan

A plan goes wrong in one of three places: it does not load, it loads and a task does not pass, or it passes and did not check what you meant. vero puts a different kind of evidence in front of you for each, and this guide reads them in that order. The flags live on the command line page; the states a task can end in, and what each does to the exit code, are on States and exit codes.

A plan that does not load

A load error names the file, the line, the place in the plan and what is wrong, one per line, and the last line counts them. Nothing was sent: the exit code is 64.

$ vero plan load.yaml
load.yaml:13: tasks.create-order.steps[0].http.headers.X-Trace: header X-Trace is an int (0755 decodes as 493); quote it: "0755"
load.yaml:15: tasks.create-order.steps[0].asert: field asert not found in a step (did you mean assert?)
load.yaml:13: tasks.create-order.steps[0].http.headers.Authorization: {{ tasks.login.results.token }} names task login, which does not exist
load.yaml: 3 load error(s); nothing ran

The path reads like the YAML: tasks.create-order.steps[0].http.headers.X-Trace is the X-Trace header of the first step of the task named create-order. Tasks are named, not numbered, so a path survives reordering the plan. The loader runs every stage even after one has found errors, so one load reports the typo, the unquoted number and the missing task together rather than one per attempt.

The refusals you will meet most:

Looks like Message
a bare number where a string belongs header X-Trace is an int (0755 decodes as 493); quote it: "0755"
a misspelled key field asert not found in a step (did you mean assert?)
the same key twice duplicate key "timeout" at line 12, first defined at line 10
a template that names a task the plan lacks {{ tasks.login.results.token }} names task login, which does not exist
a result the task does not declare {{ tasks.create-order.results.orderld }} names result orderld, which task create-order does not declare (it has: orderId)
an assertion reading a task it does not run after "tasks.create-order.results.orderId != nil" reads a result of task create-order, but check does not run after create-order: the graph has no path from create-order to check. Add needs: [create-order] to check, or template the value so the edge is derived

A duplicate key stops that task from decoding, so it is followed by task check has no steps; fix the duplicate and the second error goes with it. A cycle is reported as a whole, with the tasks that could still be scheduled and the ones stuck in the loop:

cycle.yaml: the task graph has a cycle and nothing will run
  scheduled levels: [[d]]
  unreachable (in a cycle): [a b c]

Templates and Assertions list every message of their own. The YAML rules behind the first three rows are on Plan format.

Seeing what vero will do

Most surprises in a run are edges: a task ran before the one it needed, or two tasks that spend the same budget ran at once. vero plan --graph prints every edge with the place in the plan that created it, and every holds claim:

$ vero plan --graph graph.yaml
task login
task list-orders
  needs login derived:template at tasks.list-orders.steps[0].http.headers.Authorization
  holds catalog shared
task reindex
  needs login explicit at tasks.reindex.needs[0]
  holds catalog exclusive
task feature-check
  needs login derived:when at tasks.feature-check.when[1]
  holds catalog shared

resource catalog shared
  held by list-orders, reindex (exclusive), feature-check

Three kinds of edge show up. explicit is a needs: entry. derived:template came from a {{ tasks.login.results.token }} in a string, and derived:when from a guard that reads tasks.login. If an edge you expected is missing, the string that should create it does not reference the task, and the fix is to template the value or add needs. If a task does not hold what it should, the resource line says who does; reindex (exclusive) marks a claim that overrides the resource's default. --only <task> cuts the graph to what vero run --only would run. --explain <task> answers the question for one task:

$ vero plan --explain feature-check graph.yaml
task feature-check
  needs
    login derived:when at tasks.feature-check.when[1]
  upstream, transitively: login
  downstream: none
  holds
    catalog shared, also held by list-orders, reindex (exclusive)
  when
    env.feature == "on"   (reads none)
    tasks.login.results.token != nil   (reads login)
  steps
    1. http GET {{ env.baseUrl }}/healthz: 1 request
  assertion paths (checked by --strict-paths): none

The last line is what --strict-paths will look for; a path you expected to see there is one no assertion reads. How a run works explains what the scheduler does with these edges and claims.

Warnings vero prints before a run

Two lints and one piece of advice come out of vero plan and vero run, on stderr, before anything is sent. They are not errors; the plan loads and runs.

$ vero plan lint.yaml
warning: lint.yaml:12: tasks.orders.steps[0].assert[1]: "body.items all (it.status != \"cancelled\")" holds on an empty body.items; add `body.items count N` or a len() assertion on it in the same step (chapter 5.5)
warning: lint.yaml:5: lifecycle.strategy: strategy unique, but nothing in the plan references run.id, so nothing it creates is namespaced by the run
advice: lint.yaml:13: tasks.orders.steps[0].assert[2]: this assertion is 39 nodes with 2 closure(s) nested 2 deep (advice past 40 nodes, 2 closures or nesting 1); one that size is a program, and a script step may be the honest place for it (chapter 10.3)
ok: 1 tasks, 0 finally tasks; the plan sends 1 HTTP request

The first lint fires on an all or none over a collection whose size nothing in the same step pins: all over an empty list is true, so a server that returned no items would pass. It goes quiet when the step also asserts body.items count N, len(body.items) == N, or anything with len, count or filter over the same collection. The second fires on a unique plan that never mentions run.id, because then nothing the plan creates is namespaced by the run. The advice fires on an assertion past 40 nodes, past 2 closures, or with closures nested inside closures; it is printed and never refused.

--strict turns the two lints into load errors. Each line is printed as error (--strict):, the run ends with lint.yaml: 2 lint warning(s) are errors under --strict; nothing ran, and the exit code is 64. Advice stays advice under --strict.

Reading a failure block

Every task that did not pass gets a block. Its heading says which of three things to look at, and the rest of the block is the evidence for it:

$ vero run fail.yaml

FAIL  wrong-status › step 1 › assert #1                           fail.yaml:11

  GET http://127.0.0.1:18080/fixtures/status?code=503             HTTP/1.1 503 Service Unavailable  0.9ms
    > accept-encoding: gzip
    > user-agent: vero
    < content-length: 17
    < content-type: application/json
    < date: Sat, 03 Oct 2026 02:31:19 GMT
    < {"status":"503"}

  assert  status == 200    failed
          status = 503
          fail.yaml:11
  assert  duration.tls < 5ms    n/a
          n/a: it compares a timing phase that did not happen on this request, so it is not held
          duration.tls = n/a
          fail.yaml:13

  timing  dns n/a  connect 0.5ms  tls n/a  ttfb 0.3ms  total 0.9ms  (new connection)

ERROR  no-server › step 1                                         fail.yaml

  Get "http://127.0.0.1:1/orders": dial tcp 127.0.0.1:1: connect: connection refused

  GET http://127.0.0.1:1/orders                                   no response  0.1ms
    > accept-encoding: gzip
    > user-agent: vero


  timing  dns n/a  connect 0.1ms  tls n/a  ttfb n/a  total 0.1ms  (new connection)

TIMEOUT  too-slow › step 1                                        fail.yaml

  deadline exceeded waiting for first byte

  GET http://127.0.0.1:18080/fixtures/slow?ms=3000                no response  400.8ms
    > accept-encoding: gzip
    > user-agent: vero


  timing  dns n/a  connect 0.2ms  tls n/a  ttfb n/a  total 400.8ms  (new connection)  deadline hit waiting for first byte

0 passed  1 failed  1 errored  1 timed out  0 skipped  0 blocked  0 not run   (3 tasks, 3 requests, 0.4s)

FAIL means the server answered and an assertion did not hold. Read the assert lines: each one that did not hold is followed by every path it read and the value it found, status = 503, and then the plan line to go to. An assertion marked n/a compared a timing phase this request did not have, here tls on a plain HTTP connection, and the note says so; n/a fails the task rather than passing it. The heading points at the first assertion that did not hold.

ERROR means the step did not get an answer to judge: a refused connection, a name that did not resolve, a certificate the run does not trust, a body that is not the JSON it claims to be. The reason is the second line of the block. Look at the server, the network or the URL, not at the assertions; they were not evaluated.

TIMEOUT means a deadline fired first. The reason says which phase it fell in, and the timing line repeats it as deadline hit waiting for first byte. The phases are dns, connect, tls, ttfb and total; a phase that did not happen on this request is n/a, and (reused connection) tells you why dns, connect and tls can be absent. The reason also names the deadline when it was not the step's own: run timeout 300ms fired while it was running: ..., or cancelled (interrupt).

The exchange between the heading and the assertions is the request as sent, the response as received, headers lower-cased and sorted, and the body cut at 2 KiB. Each step kind prints its own shape there; the step pages, such as the http step, show theirs.

A path that is not there

body.eror == nil holds on every response, because a key the body lacks reads as nil. That is the hole --strict-paths closes. Under it a step fails when an assertion reads a path the result does not have, and the detail says which:

$ vero run --strict-paths strict.yaml

FAIL  big-id › step 1 › assert #2                                 strict.yaml:10

  GET http://127.0.0.1:18080/fixtures/big-id                      HTTP/1.1 200 OK  1.0ms
    ...
    < {"id":10000000000000001,"total":10.10}

  assert  body.eror == nil    failed
          --strict-paths: body.eror is absent from the result and not declared in optional:
          strict.yaml:10
  assert  missing body.debug    failed
          --strict-paths: body.debug is absent from the result and not declared in optional:
          strict.yaml:10

The same plan passes without the flag. The first line is a typo to fix. The second is an absence the plan means: missing body.debug asserts that a debug field is not leaking, and under --strict-paths the operand of missing is checked like any other path. Declare it in the step's optional: list, optional: [body.debug], and the step passes with the flag on. exists body.x is never checked this way; it is the question, not a read. The flag is off by default and meant for CI, where a plan that quietly asserts nothing is the expensive kind of green.

Snapshots, schemas and OpenAPI

Three assertions compare a value against a file beside the plan. Each explains a mismatch path by path. Assertions has the full rules; this is what the output looks like.

A snapshot that does not exist yet fails and says how to write it. --update-snapshots writes every snapshot file the plan names, fails each assertion it wrote, and exits 1, so the writing run is never the green one. The next run compares:

  assert  body matches snapshot "snapshots/big-id.json"    failed
          body = {"id":10000000000000001,"total":10.10}
          no snapshot at snapshots/big-id.json; run with --update-snapshots to write it

$ vero run --update-snapshots snap.yaml
  ...
          snapshot snapshots/big-id.json written, not compared

$ vero run snap.yaml
1 passed  0 failed  0 errored  0 timed out  0 skipped  0 blocked  0 not run   (1 tasks, 1 requests, 0.0s)

When the response drifts, every differing path is listed with both values. Numbers compare exactly, so 10.250 and 10.10 are a difference and so is a 17-digit id off by one:

  assert  body matches snapshot "snapshots/big-id.json"    failed
          body = {"id":10000000000000001,"total":10.10}
          snapshot snapshots/big-id.json: body.id: expected 10000000000000002, actual 10000000000000001
          snapshot snapshots/big-id.json: body.total: expected 10.250, actual 10.10

A field that changes on every run, a timestamp or a generated id, goes in ignoring: body matches snapshot "snapshots/big-id.json" ignoring ["total"] reports only body.id. An ignoring entry covers the path and everything under it, and items[*].createdAt matches every element.

A schema mismatch prints one line per violation, with the JSON pointer into the body and the keyword that failed, up to 20 of them:

  assert  body matches schema "order.schema.json"    failed
          body = {"id":10000000000000001,"total":10.10}
          schema order.schema.json: /  required: missing property 'currency'
          schema order.schema.json: /id  maximum: maximum: got 10000000000000001, want 1000000
          schema order.schema.json: /total  type: got number, want string

response matches openapi "openapi.yaml" checks the whole exchange against the document: the operation by method and path, the response by status, the required headers and the body against the media type's schema. Each miss is its own line, and a request the document does not describe at all is one line:

  assert  response matches openapi "openapi.yaml"    failed
          request = {GET http://127.0.0.1:18080/fixtures/big-id}
          status = 200
          headers = headers(content-length, content-type, date)
          body = {"id":10000000000000001,"total":10.10}
          openapi openapi.yaml: GET /fixtures/big-id 200 requires header x-request-id, and the response has none
          openapi openapi.yaml: GET /fixtures/big-id 200: /  required: missing property 'currency'
          openapi openapi.yaml: GET /fixtures/big-id 200: /total  type: got number, want string

  assert  response matches openapi "openapi.yaml"    failed
          ...
          openapi openapi.yaml: no operation for GET /healthz

Passing alone, failing together

A task that passes at --jobs 1 and fails some of the time above it is almost always sharing a budget, a row or a balance with another task without a holds claim on it. Run vero run --repeat 20 --compare-jobs and read the hint as Reading a flake report and the --jobs hint describes; what holds does to the schedule is on How a run works.

The event log as evidence

Everything the terminal shows about an assertion is also in the event log, as data. Each assert event carries the expression, its outcome, where it is in the plan, and for one that did not hold the same values and detail lines the failure block printed:

{"schema":"vero.events/v1","ts":"2026-10-03T02:31:20.547Z","runId":"01M3ZSM2Q0CJQKG97Y7CH67EXF","type":"assert","task":"wrong-status","step":1,"index":1,"expr":"status == 200","location":"fail.yaml:11","outcome":"failed","values":[{"path":"status","value":"503"}],"detail":["status = 503"]}

--events - writes the log to stdout and moves the human report to stderr, so a run can be piped into jq or saved as a build artefact. A saved log renders into the HTML report later with vero report --html out.html events.ndjson, with the same failures in full. The fields are on vero.events/v1. A value a plan marked secret is already scrubbed in the log, so what it shows is what the run showed; Secrets has the rules.

This page is docs/content/guides/diagnosing.md in the repository.

verodocs