Evidence
The ws and grpc steps against servers vero did not write
Historical experiment (2026-10-01). Its machines, installed tools, and deployment topology describe that run, not vero prerequisites. For current setup see Getting started and Running the examples.
G05, run on 2026-10-01 between 02:38Z and 02:50Z on host worker, against k3s. Until this run,
the ws step (B11) and the grpc step (B12, B13) had only met fixtures vero wrote itself, in the
same process, over loopback. Here they talk to a gRPC server and a WebSocket server from other
people, through nginx with TLS, the way they would reach a service in a cluster.
Result. Every call shape works as written, and 300 runs at --jobs 1 and at --jobs 8 were
300/300 for every task. The servers were never the problem. The real setup found four places where
vero's report was wrong or unhelpful, and each is fixed in its own commit:
| Found | vero said | Fixed in |
|---|---|---|
a deadline that grpcbin's copy of grpc-timeout answered first |
failed, code = "DEADLINE_EXCEEDED", 4 of 300 runs |
fb78215 B12: now timed_out, 0 of 300 |
nginx's proxy_read_timeout dropping an idle socket |
errored, "failed to read frame header: EOF" |
100d8fe B11: closed.code == 1006, a failure that shows what came |
| a collection window too short for the server under load | "consider holds" on 34 of 40 tasks, "look at the server's state" on 6 | d05b236 E02: names the window, 0 holds hints |
| ws and grpc tasks in the overlap heuristic | "no other task overlapped it on the same host" about 80 tasks on one proxy | d05b236 E02: ws and grpc hosts count |
B11's risk section predicted that a timeout-bounded collection would look like flake under load.
It does, but only once the window gets near the server's pace under load. At 1 s it never came
close (worst case 216 ms). At 10 ms it flaked at --jobs 40 only, which is exactly the pattern the
--jobs hint reads as a missing holds. So the risk showed up in the advice, not in the pass
rates, and the advice is what got fixed.
Historical setup
Manifests in deploy/experiments/real-streams/, applied
with kubectl apply -k into namespace testing-platform, after make -C deploy/experiments/real-streams tls had made a CA and a server certificate for 10.233.1.1 (the
node, $POD) and loaded them as a Secret. The certificate files stay in the git-ignored tls/.
| Piece | Image, as pulled |
|---|---|
| gRPC server | docker.io/moul/grpcbin@sha256:bd8f2ffdd02d0849fad2d1c754eff4402c867e7a3e0552b8992f4590f5687d20 |
| WebSocket server | docker.io/jmalloc/echo-server@sha256:86f2c45aa7e7ebe1be30b21f8cfff25a7ed6e3b059751822d4b35bf244a688d5 |
| proxy | docker.io/library/nginx@sha256:65645c7bb6a0661892a8b03b89d0743208a18dd2f3f17a54ef4b76fb8e2f2a10 (1.27.5, the same image as G02) |
nginx listens with TLS and http2 on on 8443 (NodePort 30562): grpc_pass to grpcbin for
/hello.HelloService/ and /grpcbin.GRPCBin/, proxy_pass with the upgrade headers to
echo-server for /ws and for /ws-idle, which sets proxy_read_timeout 1s. Port 8444 (NodePort
30563) sends every gRPC call to 127.0.0.1:9, where nothing listens. Requests went from worker
to 10.233.1.1:30562, over the veth gateway, by NodePort, with no port-forward. vero trusted the
CA with --ca-cert deploy/experiments/real-streams/tls/ca.pem. The protos are upstream's
hello.proto and grpcbin.proto from github.com/moul/pb at
bca18df4138cc423a9f8513f9c7d5f71f1ea4b35, in
testdata/experiments/real-streams/ with the plans.
vero built from main at each step, go1.26.7. Node port 30560, vero's usual Postgres port, now
belongs to Penpot again, and the testing-platform namespace did not exist when this started.
The run created the namespace and deleted it afterwards. This cleanup was specific to a namespace
created by this experiment; deleting an existing shared namespace is unsafe.
The four call shapes and a ws sequence
streams.yaml has five independent
tasks: SayHello (unary), LotsOfReplies (server stream), LotsOfGreetings (client stream,
three sends), BidiHello (three sends) and a ws step that sends two messages to echo-server with
expect: { count: 3, timeout: 1s }.
$ bin/vero run --ca-cert deploy/experiments/real-streams/tls/ca.pem --repeat 300 --compare-jobs --jobs 8 \
testdata/experiments/real-streams/streams.yaml # 02:38:42Z to 02:38:56Z, exit 0
300 runs at --jobs 1 and 300 at --jobs 8, lifecycle isolated:
task --jobs 1 --jobs 8
unary 300/300 300/300
server-stream 300/300 300/300
client-stream 300/300 300/300
bidi 300/300 300/300
ws-sequence 300/300 300/300
What the real servers do that the fixtures did not:
- echo-server greets first. Its first message is
Request served by echo-674bb44fc8-7mv56, before any echo. vero records it asmessages[0], in arrival order, so the plan asserts on it and reads the echoes at[1]and[2]. A plan written against a fixture that only echoes would fail here on its first index. The step reports the truth, and the plan was the thing to fix. - grpcbin's
LotsOfRepliessends ten replies and the stream's headers come from nginx (server: nginx/1.27.5), soheadersis what the client saw at the proxy, not the upstream's. LotsOfGreetingsanswers once after the half-close (hello a, b, c), andBidiHelloanswers each send in order. Both worked with vero sending and reading at once (B13).
Load, and where the collection window stands
To give B11's risk a fair chance, streams-load.yaml
expands the ws sequence and the bidi call into 40 copies each with a matrix (A07): 80 sockets and
streams through one nginx at --jobs 40.
$ bin/vero run ... --repeat 100 --compare-jobs --jobs 40 testdata/experiments/real-streams/streams-load.yaml
# 02:39:16Z to 02:40:27Z, exit 0
all 80 tasks: 100/100 at --jobs 1 and 100/100 at --jobs 40
When the last message arrived, from the event logs (received[-1].atMs; for ws, time since the
upgrade, and for grpc, since the call started):
| Plan, jobs | Step | n | p50 ms | p99 ms | max ms |
|---|---|---|---|---|---|
| streams, 1 | ws, 3 messages | 300 | 0.39 | 4.40 | 48.88 |
| streams, 8 | ws, 3 messages | 300 | 2.00 | 19.87 | 34.66 |
| streams, 1 | grpc LotsOfReplies, 10 |
300 | 3.59 | 12.00 | 14.76 |
| streams, 8 | grpc LotsOfReplies, 10 |
300 | 12.54 | 51.88 | 69.82 |
| streams, 1 | grpc BidiHello, 3 |
300 | 3.65 | 12.50 | 27.49 |
| streams, 8 | grpc BidiHello, 3 |
300 | 12.53 | 52.13 | 118.36 |
| load, 1 | ws | 4000 | 0.49 | 6.03 | 37.84 |
| load, 40 | ws | 4000 | 8.25 | 88.64 | 215.73 |
| load, 1 | grpc BidiHello, 2 |
4000 | 3.98 | 22.28 | 100.98 |
| load, 40 | grpc BidiHello, 2 |
4000 | 59.86 | 210.60 | 248.92 |
Load moves the ws median from 0.5 ms to 8 ms, 17 times, and the worst case to 216 ms. A 1 s
window is never threatened.
streams-tight.yaml sets the window
to 100 ms, near that p99. It held too, at 100/100 for all 80 tasks at both job counts (02:41:01Z
to 02:42:11Z), because the worst arrival that time was 78.65 ms. The worst case moves from run to
run.
At 10 ms, between the --jobs 1 p99 and the --jobs 40 p50 (the same plan with the one value
changed, run 30 times, not committed):
ws passes at --jobs 1 |
ws passes at --jobs 40 |
bidi | hints | |
|---|---|---|---|---|
before d05b236 |
1193/1200 | 1022/1200 | 1200/1200 both | 34 "consider holds: [target/ws-cNN]", 6 "look at the server's state between runs" |
after d05b236 |
1200/1200 | 1049/1200 | 1200/1200 both | 40 "every failure was step 1's collection window", 0 holds |
Before the fix, vero told the user to add a holds to 34 tasks that share nothing, under a line
claiming "no other task overlapped it on the same host". It said that about 80 tasks calling one
proxy, because the overlap heuristic only counted HTTP exchanges as hosts. After it:
hint: ws-c01 passed 30/30 at --jobs 1 and 25/30 at --jobs 40.
every failure was step 1's collection window: fewer than 3 messages within expect.timeout 10ms.
The latest message that made it in at --jobs 40 arrived at 9.8ms; vero sees nothing after the window closes.
The server's messages come later than the window allows: that is its pace, not state another task
shares, and holds will not change it. Raise expect.timeout above what the server needs, or keep it
and read this as a latency finding.
Did B11's flake risk show up? Yes, and it showed up as wrong advice, not as noise in the pass
rates. A window sized with room for load (1 s against a 216 ms worst case) never flaked in 8 600
collections. A window sized from a quiet run flakes only above --jobs 1, and that pattern is the
same one a missing holds leaves. The pass rates alone cannot tell the two apart. The failure's
shape can (fewer than count, and nothing after the window), and the hint now reads it.
The probes
probes.yaml asserts what vero reports
now.
A gRPC call through nginx to an upstream that is down. nginx answers HTTP 502 with an HTML
page, and grpc-go turns that into UNAVAILABLE with the message unexpected HTTP status code received from server: 502 (Bad Gateway); transport: received unexpected content-type "text/html" and the page's text. vero reports code == "UNAVAILABLE", an assertable code, with
the message shown. Right: the gRPC spec maps HTTP 502 to UNAVAILABLE, and the proxy did
answer, so it isn't B12's "no server answered" (errored), which stays for a refused dial. 300
of 300 runs said the same.
A deadline that nginx passes on. grpcbin.GRPCBin/NoResponseUnary was meant to hang. It
doesn't: grpcbin fails it at once with INTERNAL: grpc: error while marshaling: proto: Marshal called with nil, which vero reported correctly as that code. The probe uses SayHello with
timeout: 5ms instead, close to its answer time. grpc-go sends the deadline as grpc-timeout,
nginx forwards it, and grpcbin's copy can expire and answer before the client's own timer has
run.
| binary | passed | timed_out | failed on DEADLINE_EXCEEDED |
|---|---|---|---|
319a7a5 (B13, before the fix) |
239 | 57 | 4 |
fb78215 and later |
168 | 132 | 0 |
Wrong before, right now. Before fb78215, 4 runs in 300 reported vero's own deadline as the
server's status, a failed assertion on code. make check had caught the same race under
-race on loopback. Here it happens over a real network, through a proxy. The pass and timeout
split moves between runs with the machine's load. The zero does not.
A ws socket that nginx closes for idling. On /ws-idle, nginx ends the socket 1 s after
echo-server's greeting, with no close frame. Before 100d8fe the step was errored: reading: failed to get reader: failed to read frame header: EOF. Wrong: a proxy dropping an idle
socket is the server side ending it, and B11's rule for that is a failure judged on what came.
Now it is closed: {code: 1006, reason: "the connection ended without a close frame"}, the code
RFC 6455 reserves for this and the one browsers report. The step fails on the added messages count 2 and prints the one message that arrived. Over 20 runs: closed 1006 every time, after
1004 to 1018 ms, failed on messages count 2 only.
A ws server that sends before the plan's first message. echo-server's greeting, above.
Right: vero keeps every message in arrival order, and the greeting is messages[0].
Historical teardown — not a shared-namespace cleanup recipe
$ kubectl delete -k deploy/experiments/real-streams
$ kubectl -n testing-platform delete secret streams-proxy-tls
$ kubectl delete namespace testing-platform
The namespace was absent before G05 created it, so it is absent again. Nothing else on the cluster was touched.
This page is docs/content/evidence/real-streams.md in the repository.