Reproducible LLM inference benchmarking across serving backends.
Measures the three numbers that actually matter when you put a model in production — time-to-first-token, steady-state tokens/sec, and peak VRAM — under controlled prompt/generation length distributions, so you can compare backends instead of comparing vibes.
Every serving framework publishes throughput numbers measured under conditions chosen to flatter that framework. Batch-32 with 128-token prompts tells you nothing about your chat workload with 4k-token contexts and a p99 latency SLO. This tool runs your workload shape against multiple backends with identical sampling parameters and reports distributions, not cherry-picked means.
| Backend | Adapter | Notes |
|---|---|---|
| vLLM | nib.backends.vllm |
OpenAI-compatible server or in-process LLM |
| llama.cpp | nib.backends.llamacpp |
via llama-cpp-python, GGUF quants |
| TensorRT-LLM | nib.backends.trtllm |
requires prebuilt engine |
| OpenAI-compatible | nib.backends.oai |
any /v1/completions endpoint |
pip install -e .
# benchmark a local vLLM server against llama.cpp with the same workload
nib run --model meta-llama/Llama-3.1-8B-Instruct \
--backend vllm:http://localhost:8000 \
--backend llamacpp:./models/llama-3.1-8b-q5_k_m.gguf \
--workload chat --prompts 200 --concurrency 8 \
--out results/llama31-8b.json
nib report results/llama31-8b.json --format tableExample output:
backend ttft_p50 ttft_p99 tok/s (agg) peak VRAM
vllm 41 ms 118 ms 2,847 17.2 GB
llamacpp 220 ms 490 ms 612 6.8 GB
chat— lognormal prompt lengths (μ=6.2), short-to-medium generationsrag— long prompts (2k–8k tok), short generations; stresses prefillcodegen— medium prompts, long generations; stresses decode + KV growthcustom— bring a JSONL of{prompt, max_tokens}pairs
- All timing is done client-side at the token-stream level; server-reported metrics are recorded but never used for cross-backend comparison.
- VRAM is sampled via NVML at 50 ms resolution; we report the high-water mark, because that's the number that decides whether the pod OOMs.
- Sampling params are pinned (
temperature=0, fixed seed where supported) so generation length variance comes from the workload, not the sampler.
MIT