/v1/chat/completions accepts a per-request chat_template (a Jinja template string) that is compiled and rendered server-side. A remote, unauthenticated client therefore controls a template the server executes. The Jinja environment blocks a runaway range() loop (returns 400), but it does not block a string-multiplication expression, so a single request with
{"model":"m","messages":[{"role":"user","content":"hi"}],"chat_template":"{{ \"a\" * 1000000000 }}"}
makes the server allocate a multi-gigabyte string and the process is OOM-killed (observed: Detected 7 oom_kill events … OOM Killed; /health stops answering). One request takes the whole server down for all clients.
That the client template is truly evaluated server-side is confirmed independently: a chat_template of {{ 1/0 }} returns a 400 whose message leaks a ZeroDivisionError from the template. (This also means there is a server-side template-injection surface; HF's apply_chat_template uses a sandboxed Jinja environment, so escalation to RCE is likely blocked, but the memory-exhaustion DoS is not.)
Cause
tensorrt_llm/serve/openai_server.py (~lines 561-562) passes the request field straight into the renderer:
apply_chat_template(..., chat_template=request.chat_template or self.chat_template, chat_template_kwargs=request.chat_template_kwargs or {})
with ChatCompletionRequest.chat_template: Optional[str] (openai_protocol.py ~622). There is no restriction on accepting a client template, and the sandbox does not bound the memory/compute a template may consume (string multiplication, and likely other allocation-heavy expressions, are not limited).
Reproduction
Use repro_trtllm_chat_template_oom.py
trtllm-serve Qwen/Qwen2.5-0.5B-Instruct --host 0.0.0.0 --port 8000 python3 repro_trtllm_chat_template_oom.py --base-url http://localhost:8000 --model-id Qwen/Qwen2.5-0.5B-Instruct
/v1/chat/completionsaccepts a per-requestchat_template(a Jinja template string) that is compiled and rendered server-side. A remote, unauthenticated client therefore controls a template the server executes. The Jinja environment blocks a runawayrange()loop (returns400), but it does not block a string-multiplication expression, so a single request with{"model":"m","messages":[{"role":"user","content":"hi"}],"chat_template":"{{ \"a\" * 1000000000 }}"}makes the server allocate a multi-gigabyte string and the process is OOM-killed (observed:
Detected 7 oom_kill events … OOM Killed;/healthstops answering). One request takes the whole server down for all clients.That the client template is truly evaluated server-side is confirmed independently: a
chat_templateof{{ 1/0 }}returns a400whose message leaks aZeroDivisionErrorfrom the template. (This also means there is a server-side template-injection surface; HF'sapply_chat_templateuses a sandboxed Jinja environment, so escalation to RCE is likely blocked, but the memory-exhaustion DoS is not.)Cause
tensorrt_llm/serve/openai_server.py(~lines 561-562) passes the request field straight into the renderer:with
ChatCompletionRequest.chat_template: Optional[str](openai_protocol.py ~622). There is no restriction on accepting a client template, and the sandbox does not bound the memory/compute a template may consume (string multiplication, and likely other allocation-heavy expressions, are not limited).Reproduction
Use repro_trtllm_chat_template_oom.py