Evidence

The ws and grpc steps against servers vero did not write

Historical experiment (2026-10-01). Its machines, installed tools, and deployment topology describe that run, not vero prerequisites. For current setup see Getting started and Running the examples.

G05, run on 2026-10-01 between 02:38Z and 02:50Z on host worker, against k3s. Until this run, the ws step (B11) and the grpc step (B12, B13) had only met fixtures vero wrote itself, in the same process, over loopback. Here they talk to a gRPC server and a WebSocket server from other people, through nginx with TLS, the way they would reach a service in a cluster.

Result. Every call shape works as written, and 300 runs at --jobs 1 and at --jobs 8 were 300/300 for every task. The servers were never the problem. The real setup found four places where vero's report was wrong or unhelpful, and each is fixed in its own commit:

Found vero said Fixed in
a deadline that grpcbin's copy of grpc-timeout answered first failed, code = "DEADLINE_EXCEEDED", 4 of 300 runs fb78215 B12: now timed_out, 0 of 300
nginx's proxy_read_timeout dropping an idle socket errored, "failed to read frame header: EOF" 100d8fe B11: closed.code == 1006, a failure that shows what came
a collection window too short for the server under load "consider holds" on 34 of 40 tasks, "look at the server's state" on 6 d05b236 E02: names the window, 0 holds hints
ws and grpc tasks in the overlap heuristic "no other task overlapped it on the same host" about 80 tasks on one proxy d05b236 E02: ws and grpc hosts count

B11's risk section predicted that a timeout-bounded collection would look like flake under load. It does, but only once the window gets near the server's pace under load. At 1 s it never came close (worst case 216 ms). At 10 ms it flaked at --jobs 40 only, which is exactly the pattern the --jobs hint reads as a missing holds. So the risk showed up in the advice, not in the pass rates, and the advice is what got fixed.

Historical setup

Manifests in deploy/experiments/real-streams/, applied with kubectl apply -k into namespace testing-platform, after make -C deploy/experiments/real-streams tls had made a CA and a server certificate for 10.233.1.1 (the node, $POD) and loaded them as a Secret. The certificate files stay in the git-ignored tls/.

Piece Image, as pulled
gRPC server docker.io/moul/grpcbin@sha256:bd8f2ffdd02d0849fad2d1c754eff4402c867e7a3e0552b8992f4590f5687d20
WebSocket server docker.io/jmalloc/echo-server@sha256:86f2c45aa7e7ebe1be30b21f8cfff25a7ed6e3b059751822d4b35bf244a688d5
proxy docker.io/library/nginx@sha256:65645c7bb6a0661892a8b03b89d0743208a18dd2f3f17a54ef4b76fb8e2f2a10 (1.27.5, the same image as G02)

nginx listens with TLS and http2 on on 8443 (NodePort 30562): grpc_pass to grpcbin for /hello.HelloService/ and /grpcbin.GRPCBin/, proxy_pass with the upgrade headers to echo-server for /ws and for /ws-idle, which sets proxy_read_timeout 1s. Port 8444 (NodePort 30563) sends every gRPC call to 127.0.0.1:9, where nothing listens. Requests went from worker to 10.233.1.1:30562, over the veth gateway, by NodePort, with no port-forward. vero trusted the CA with --ca-cert deploy/experiments/real-streams/tls/ca.pem. The protos are upstream's hello.proto and grpcbin.proto from github.com/moul/pb at bca18df4138cc423a9f8513f9c7d5f71f1ea4b35, in testdata/experiments/real-streams/ with the plans.

vero built from main at each step, go1.26.7. Node port 30560, vero's usual Postgres port, now belongs to Penpot again, and the testing-platform namespace did not exist when this started. The run created the namespace and deleted it afterwards. This cleanup was specific to a namespace created by this experiment; deleting an existing shared namespace is unsafe.

The four call shapes and a ws sequence

streams.yaml has five independent tasks: SayHello (unary), LotsOfReplies (server stream), LotsOfGreetings (client stream, three sends), BidiHello (three sends) and a ws step that sends two messages to echo-server with expect: { count: 3, timeout: 1s }.

$ bin/vero run --ca-cert deploy/experiments/real-streams/tls/ca.pem --repeat 300 --compare-jobs --jobs 8 \
    testdata/experiments/real-streams/streams.yaml          # 02:38:42Z to 02:38:56Z, exit 0
300 runs at --jobs 1 and 300 at --jobs 8, lifecycle isolated:
  task           --jobs 1    --jobs 8
  unary          300/300     300/300
  server-stream  300/300     300/300
  client-stream  300/300     300/300
  bidi           300/300     300/300
  ws-sequence    300/300     300/300

What the real servers do that the fixtures did not:

  • echo-server greets first. Its first message is Request served by echo-674bb44fc8-7mv56, before any echo. vero records it as messages[0], in arrival order, so the plan asserts on it and reads the echoes at [1] and [2]. A plan written against a fixture that only echoes would fail here on its first index. The step reports the truth, and the plan was the thing to fix.
  • grpcbin's LotsOfReplies sends ten replies and the stream's headers come from nginx (server: nginx/1.27.5), so headers is what the client saw at the proxy, not the upstream's.
  • LotsOfGreetings answers once after the half-close (hello a, b, c), and BidiHello answers each send in order. Both worked with vero sending and reading at once (B13).

Load, and where the collection window stands

To give B11's risk a fair chance, streams-load.yaml expands the ws sequence and the bidi call into 40 copies each with a matrix (A07): 80 sockets and streams through one nginx at --jobs 40.

$ bin/vero run ... --repeat 100 --compare-jobs --jobs 40 testdata/experiments/real-streams/streams-load.yaml
                                                             # 02:39:16Z to 02:40:27Z, exit 0
all 80 tasks: 100/100 at --jobs 1 and 100/100 at --jobs 40

When the last message arrived, from the event logs (received[-1].atMs; for ws, time since the upgrade, and for grpc, since the call started):

Plan, jobs Step n p50 ms p99 ms max ms
streams, 1 ws, 3 messages 300 0.39 4.40 48.88
streams, 8 ws, 3 messages 300 2.00 19.87 34.66
streams, 1 grpc LotsOfReplies, 10 300 3.59 12.00 14.76
streams, 8 grpc LotsOfReplies, 10 300 12.54 51.88 69.82
streams, 1 grpc BidiHello, 3 300 3.65 12.50 27.49
streams, 8 grpc BidiHello, 3 300 12.53 52.13 118.36
load, 1 ws 4000 0.49 6.03 37.84
load, 40 ws 4000 8.25 88.64 215.73
load, 1 grpc BidiHello, 2 4000 3.98 22.28 100.98
load, 40 grpc BidiHello, 2 4000 59.86 210.60 248.92

Load moves the ws median from 0.5 ms to 8 ms, 17 times, and the worst case to 216 ms. A 1 s window is never threatened. streams-tight.yaml sets the window to 100 ms, near that p99. It held too, at 100/100 for all 80 tasks at both job counts (02:41:01Z to 02:42:11Z), because the worst arrival that time was 78.65 ms. The worst case moves from run to run.

At 10 ms, between the --jobs 1 p99 and the --jobs 40 p50 (the same plan with the one value changed, run 30 times, not committed):

ws passes at --jobs 1 ws passes at --jobs 40 bidi hints
before d05b236 1193/1200 1022/1200 1200/1200 both 34 "consider holds: [target/ws-cNN]", 6 "look at the server's state between runs"
after d05b236 1200/1200 1049/1200 1200/1200 both 40 "every failure was step 1's collection window", 0 holds

Before the fix, vero told the user to add a holds to 34 tasks that share nothing, under a line claiming "no other task overlapped it on the same host". It said that about 80 tasks calling one proxy, because the overlap heuristic only counted HTTP exchanges as hosts. After it:

hint: ws-c01 passed 30/30 at --jobs 1 and 25/30 at --jobs 40.
      every failure was step 1's collection window: fewer than 3 messages within expect.timeout 10ms.
      The latest message that made it in at --jobs 40 arrived at 9.8ms; vero sees nothing after the window closes.
      The server's messages come later than the window allows: that is its pace, not state another task
      shares, and holds will not change it. Raise expect.timeout above what the server needs, or keep it
      and read this as a latency finding.

Did B11's flake risk show up? Yes, and it showed up as wrong advice, not as noise in the pass rates. A window sized with room for load (1 s against a 216 ms worst case) never flaked in 8 600 collections. A window sized from a quiet run flakes only above --jobs 1, and that pattern is the same one a missing holds leaves. The pass rates alone cannot tell the two apart. The failure's shape can (fewer than count, and nothing after the window), and the hint now reads it.

The probes

probes.yaml asserts what vero reports now.

A gRPC call through nginx to an upstream that is down. nginx answers HTTP 502 with an HTML page, and grpc-go turns that into UNAVAILABLE with the message unexpected HTTP status code received from server: 502 (Bad Gateway); transport: received unexpected content-type "text/html" and the page's text. vero reports code == "UNAVAILABLE", an assertable code, with the message shown. Right: the gRPC spec maps HTTP 502 to UNAVAILABLE, and the proxy did answer, so it isn't B12's "no server answered" (errored), which stays for a refused dial. 300 of 300 runs said the same.

A deadline that nginx passes on. grpcbin.GRPCBin/NoResponseUnary was meant to hang. It doesn't: grpcbin fails it at once with INTERNAL: grpc: error while marshaling: proto: Marshal called with nil, which vero reported correctly as that code. The probe uses SayHello with timeout: 5ms instead, close to its answer time. grpc-go sends the deadline as grpc-timeout, nginx forwards it, and grpcbin's copy can expire and answer before the client's own timer has run.

binary passed timed_out failed on DEADLINE_EXCEEDED
319a7a5 (B13, before the fix) 239 57 4
fb78215 and later 168 132 0

Wrong before, right now. Before fb78215, 4 runs in 300 reported vero's own deadline as the server's status, a failed assertion on code. make check had caught the same race under -race on loopback. Here it happens over a real network, through a proxy. The pass and timeout split moves between runs with the machine's load. The zero does not.

A ws socket that nginx closes for idling. On /ws-idle, nginx ends the socket 1 s after echo-server's greeting, with no close frame. Before 100d8fe the step was errored: reading: failed to get reader: failed to read frame header: EOF. Wrong: a proxy dropping an idle socket is the server side ending it, and B11's rule for that is a failure judged on what came. Now it is closed: {code: 1006, reason: "the connection ended without a close frame"}, the code RFC 6455 reserves for this and the one browsers report. The step fails on the added messages count 2 and prints the one message that arrived. Over 20 runs: closed 1006 every time, after 1004 to 1018 ms, failed on messages count 2 only.

A ws server that sends before the plan's first message. echo-server's greeting, above. Right: vero keeps every message in arrival order, and the greeting is messages[0].

Historical teardown — not a shared-namespace cleanup recipe

$ kubectl delete -k deploy/experiments/real-streams
$ kubectl -n testing-platform delete secret streams-proxy-tls
$ kubectl delete namespace testing-platform

The namespace was absent before G05 created it, so it is absent again. Nothing else on the cluster was touched.

This page is docs/content/evidence/real-streams.md in the repository.

verodocs