• Home
  • Features
  • Pricing
  • Docs
  • Announcements
  • Sign In

NVIDIA / nodewright / 36892753378
83%

Build:
DEFAULT BRANCH: main
Ran 01 Oct 2026 04:52PM UTC
Jobs 1
Files 59
Run time 1min
Badge
Embed ▾
README BADGES
x

If you need to use a raster PNG badge, change the '.svg' to '.png' in the link

Markdown

Textile

RDoc

HTML

Rst

01 Oct 2026 04:32PM UTC coverage: 81.981% (-0.02%) from 82.002%
36892753378

push

github

web-flow
fix(operator): count failed attempts while a package's Job retries (#707)

* fix(operator): count a package's failed attempts while its Job retries

Package stage Jobs run restartPolicy Never, so every retry is a fresh
pod whose container RestartCount is always 0. The terminal writes in
JobReconciler already record the Job's status.failed, but the pod watch
wrote that RestartCount, so nodewright_package_restarts_count read 0 for
the whole retry window, exactly when a crash-looping package should be
visible.

For a restartPolicy Never pod the pod watch now records its owning Job's
failed attempts, counting the attempt being reported. The watch sees a
failure before the Job controller counts it (uncountedTerminatedPods,
then finalizer removal, then status.failed), so it adds one while the
pod still carries the job-tracking finalizer or is listed as uncounted.
The Job is resolved through the pod's controller reference and must
match its UID, never the job-name label, which is ambiguous across
reruns of a deterministic Job name. An unresolvable owner still records
erroring and counts this attempt alone.

Interrupt pods (restartPolicy OnFailure) restart in place and keep the
container RestartCount; their Job's status.failed stays 0.

The metric's description in metrics.md, its HELP string, and the Current
Package Status dashboard panel now say it counts failed attempts, and
the failure-nodewright chainsaw test asserts restarts > 0 again.

Fixes #428

Signed-off-by: Alex Yuskauskas <ayuskauskas@nvidia.com>

* test(operator): run the JobReconciler beside the heavy pass in envtest

Nothing in the Controller Suite ran JobReconcile and the heavy pass
together the way production does: the suite's manager runs only the
heavy pass, and the JobReconcile specs drive a fake client.

job_handoff_test.go starts a JobReconciler on its own manager (metrics
disabled, Job cache scoped to its specs' NodeWright names, stopped in
cleanup) and covers the hand-off from both sid... (continued)

34 of 52 new or added lines in 2 files covered. (65.38%)

4 existing lines in 1 file now uncovered.

9500 of 11588 relevant lines covered (81.98%)

7.64 hits per line

Uncovered Changes

Lines Coverage ∆ File
16
73.28
-1.15% operator/internal/controller/job_controller.go
2
86.61
-1.01% operator/internal/controller/pod_controller.go

Coverage Regressions

Lines Coverage ∆ File
4
79.02
0.32% operator/internal/controller/skyhook_controller.go
Jobs
ID Job ID Ran Files Coverage
1 36892753378.1 01 Oct 2026 04:52PM UTC 59
81.98
GitHub Action Run
Source Files on build 36892753378
  • Tree
  • List 59
  • Changed 4
  • Source Changed 3
  • Coverage Changed 3
Coverage ∆ File Lines Relevant Covered Missed Hits/Line
  • Back to Repo
  • Github Actions Build #36892753378
  • 5e754602 on github
  • Prev Build on main (#36890207969)
  • Next Build on main (#36897104756)
  • Delete
STATUS · Troubleshooting · Open an Issue · Sales · Support · CAREERS · ENTERPRISE · START FREE TRIAL · SCHEDULE DEMO
ANNOUNCEMENTS · TWITTER · TOS & SLA · Supported CI Services · What's a CI service? · Automated Testing

© 2026 Coveralls, Inc