• Home
  • Features
  • Pricing
  • Docs
  • Announcements
  • Sign In

kubernetes-sigs / inference-perf / 34513745112
81%

Build:
DEFAULT BRANCH: main
Ran 10 Sep 2026 06:27PM UTC
Jobs 1
Files 118
Run time 1min
Badge
Embed ▾
README BADGES
x

If you need to use a raster PNG badge, change the '.svg' to '.png' in the link

Markdown

Textile

RDoc

HTML

Rst

10 Sep 2026 06:20PM UTC coverage: 78.902% (+0.008%) from 78.894%
34513745112

push

github

web-flow
fix: explain worker process crashes instead of one context-free log line (#747)

Closes #593.

## Problem

A worker dying from an unhandled exception produced exactly one line:

```
ERROR inference_perf.loadgen.load_generator - A worker process died unexpectedly!
```

No worker id, no exit code, no traceback, and no way to separate a
segfault from a bug in the data generator. #662 made the run survive a
dead worker. It did not make the loss explicable.

## What changed

**The catch-all covers the whole worker body.** `Worker.run()` is now a
wrapper around `_run()`. The previous entry point would have left a
failure during numpy seeding or uvloop policy installation uncaught,
which is the same silent-crash class the issue is about.

**The worker reports before it dies.** It logs the exception locally and
hands the parent a `WorkerCrash` (worker id, stage, exception type,
message, traceback, in-flight count). The channel is `mp.SimpleQueue`,
which writes straight to the pipe, so the record is not lost to an
unflushed feeder thread when the process exits immediately after.

**The parent explains what it lost.** `collect_worker_failures()` pairs
each dead worker with its record, at both detection sites (`run_stage`
and `run_session_stage`). Output now reads:

```
ERROR Worker 0 exited with code 1 after an unhandled RuntimeError during stage 0
      with 3 request(s) in flight run-wide: cannot pickle 'generator' object
      Traceback (most recent call last):
      ...
```

The in-flight counter is shared by every worker, so it is labelled
run-wide. A worker that died before pulling its first request reads
`before serving any stage`. A worker killed by a signal never reports,
so the exit code is the only evidence: `-9` is decoded to `SIGKILL` and
the text says plainly that no traceback exists rather than implying one
was lost.

**The policy is #662's, not this PR's.** Losing a worker fails the stage
it happened in, `_respawn_worker` starts a replacement at... (continued)

65 of 82 new or added lines in 1 file covered. (79.27%)

9424 of 11944 relevant lines covered (78.9%)

0.79 hits per line

Uncovered Changes

Lines Coverage ∆ File
17
62.04
1.94% inference_perf/loadgen/load_generator.py
Jobs
ID Job ID Ran Files Coverage
1 34513745112.1 10 Sep 2026 06:27PM UTC 118
78.9
GitHub Action Run
Source Files on build 34513745112
  • Tree
  • List 118
  • Changed 1
  • Source Changed 1
  • Coverage Changed 1
Coverage ∆ File Lines Relevant Covered Missed Hits/Line
  • Back to Repo
  • Github Actions Build #34513745112
  • 330fe86a on github
  • Prev Build on main (#34504974826)
  • Next Build on main (#34620663099)
  • Delete
STATUS · Troubleshooting · Open an Issue · Sales · Support · CAREERS · ENTERPRISE · START FREE TRIAL · SCHEDULE DEMO
ANNOUNCEMENTS · TWITTER · TOS & SLA · Supported CI Services · What's a CI service? · Automated Testing

© 2026 Coveralls, Inc