• Home
  • Features
  • Pricing
  • Docs
  • Announcements
  • Sign In

kubernetes-sigs / inference-perf / 36454983607
81%

Build:
DEFAULT BRANCH: main
Ran 28 Sep 2026 05:05PM UTC
Jobs 1
Files 126
Run time 1min
Badge
Embed ▾
README BADGES
x

If you need to use a raster PNG badge, change the '.svg' to '.png' in the link

Markdown

Textile

RDoc

HTML

Rst

28 Sep 2026 04:59PM UTC coverage: 81.055% (+0.1%) from 80.907%
36454983607

push

github

web-flow
feat: support the embeddings API (#840)

Fixes #836

Adds an `embeddings` API type for `POST /v1/embeddings`.

## What changed

- **New API type:**
- `api.type: embeddings` sends `model` and `input` (a string or a list
of strings), plus `dimensions` and `encoding_format` when set.
- Input tokens come from the server's `usage.prompt_tokens`, with
client-side tokenization as the fallback.
- Nothing is generated, so TTFT, TPOT and ITL stay `None` rather than 0.
- **Config:**
- `api.embeddings.batch_size` (inputs per request, default 1),
`dimensions`, `encoding_format`.
  - `streaming` and `response_format` are rejected for embeddings.
- **Supported on:**
  - `vllm` and `mock` clients; `mock` and `synthetic` data generators.
- The synthetic generator only needs `input_distribution` for
embeddings.
- **NTPOT fix (separate commit):**
- Requests with no output tokens reported NTPOT as `0.0`, and the
summary averaged it in.
- It's now `None` and left out, as `reportgen/br/v0_2/adapter.py`
already does.
  - Happy to split this into its own PR.
- **Example:** 
  - Config and README in `examples/embeddings/`.
- **License headers (separate commit):** 
- Three files this PR edits already had headers the pre-push check
rejects. Header-only fix.


This builds on the current `apis/` structure rather than waiting for
#798. When #798 lands, the request and response code in
`apis/embeddings.py` should move into the embeddings adapter mostly as
is.

## Testing

- `pdm run validate` and `pdm run test` pass (22 new tests). Coverage
81.04% vs 80.36% on main.
- The NTPOT regression test fails on main and passes with the fix.
- Real vLLM v0.26.0 serving `BAAI/bge-small-en-v1.5` on an RTX 4060
laptop GPU, using `examples/embeddings/config.yml` (400 requests,
~128-token inputs, 8 in flight), each batch size run twice and averaged:

  | batch_size | requests/s | inputs/s | p50 latency | p99 latency |
  |-----------:|-----------:|---------:|------------:|------------:|
  | 1  | 2... (continued)

83 of 85 new or added lines in 9 files covered. (97.65%)

2 existing lines in 2 files now uncovered.

10709 of 13212 relevant lines covered (81.06%)

0.81 hits per line

Uncovered Changes

Lines Coverage ∆ File
2
15.49
0.0% inference_perf/main.py

Coverage Regressions

Lines Coverage ∆ File
1
77.67
3.2% inference_perf/datagen/synthetic/synthetic_datagen.py
1
15.49
0.0% inference_perf/main.py
Jobs
ID Job ID Ran Files Coverage
1 36454983607.1 28 Sep 2026 05:05PM UTC 126
81.06
GitHub Action Run
Source Files on build 36454983607
  • Tree
  • List 126
  • Changed 8
  • Source Changed 8
  • Coverage Changed 6
Coverage ∆ File Lines Relevant Covered Missed Hits/Line
  • Back to Repo
  • Github Actions Build #36454983607
  • 4f253581 on github
  • Prev Build on main (#36409933000)
  • Next Build on main (#36461094743)
  • Delete
STATUS · Troubleshooting · Open an Issue · Sales · Support · CAREERS · ENTERPRISE · START FREE TRIAL · SCHEDULE DEMO
ANNOUNCEMENTS · TWITTER · TOS & SLA · Supported CI Services · What's a CI service? · Automated Testing

© 2026 Coveralls, Inc