Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

neural-inference-bench

Reproducible LLM inference benchmarking across serving backends.

Measures the three numbers that actually matter when you put a model in production — time-to-first-token, steady-state tokens/sec, and peak VRAM — under controlled prompt/generation length distributions, so you can compare backends instead of comparing vibes.

Why

Every serving framework publishes throughput numbers measured under conditions chosen to flatter that framework. Batch-32 with 128-token prompts tells you nothing about your chat workload with 4k-token contexts and a p99 latency SLO. This tool runs your workload shape against multiple backends with identical sampling parameters and reports distributions, not cherry-picked means.

Supported backends

Backend Adapter Notes
vLLM nib.backends.vllm OpenAI-compatible server or in-process LLM
llama.cpp nib.backends.llamacpp via llama-cpp-python, GGUF quants
TensorRT-LLM nib.backends.trtllm requires prebuilt engine
OpenAI-compatible nib.backends.oai any /v1/completions endpoint

Quick start

pip install -e .

# benchmark a local vLLM server against llama.cpp with the same workload
nib run --model meta-llama/Llama-3.1-8B-Instruct \
        --backend vllm:http://localhost:8000 \
        --backend llamacpp:./models/llama-3.1-8b-q5_k_m.gguf \
        --workload chat --prompts 200 --concurrency 8 \
        --out results/llama31-8b.json

nib report results/llama31-8b.json --format table

Example output:

backend      ttft_p50   ttft_p99   tok/s (agg)   peak VRAM
vllm         41 ms      118 ms     2,847         17.2 GB
llamacpp     220 ms     490 ms       612          6.8 GB

Workload shapes

  • chat — lognormal prompt lengths (μ=6.2), short-to-medium generations
  • rag — long prompts (2k–8k tok), short generations; stresses prefill
  • codegen — medium prompts, long generations; stresses decode + KV growth
  • custom — bring a JSONL of {prompt, max_tokens} pairs

Design notes

  • All timing is done client-side at the token-stream level; server-reported metrics are recorded but never used for cross-backend comparison.
  • VRAM is sampled via NVML at 50 ms resolution; we report the high-water mark, because that's the number that decides whether the pod OOMs.
  • Sampling params are pinned (temperature=0, fixed seed where supported) so generation length variance comes from the workload, not the sampler.

License

MIT

About

Reproducible LLM inference benchmarking — TTFT, tokens/sec, and VRAM high-water marks across vLLM, llama.cpp, and TensorRT-LLM

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages