TensorRT-LLM — NVIDIA's Engine, With Model Support Limits
Self-contained guide to TensorRT-LLM — the local inference engine.
- Current: 1.3 (2026-09) documented; PyPI stable is 1.2.1 (2026-04-20) with 1.3.0rc29 in prerelease. NVIDIA-only.
- 1.3 deprecation: the TRITON MoE backend (
TritonFusedMoE) is deprecated and will be removed after a 3-month migration window; CUTLASS replaces it on Hopper forW4A16_MXFP4GPT-OSS (release notes). - 1.2 breaking changes worth knowing: the TensorRT execution backend was removed entirely — PyTorch is now the sole backend, so ahead-of-time engine compilation is gone (
LLM(backend="tensorrt")raisesValueError). Two-model speculative decoding was also removed in favour of one-model paths. 1.2 added beta single-node DGX Spark support with a validated model/precision list (GPT-OSS MXFP4, Llama-3.1/3.3 FP8/NVFP4, Qwen3 family, Nemotron, Phi-4). - The practical problem (measured, Sept 2026): TensorRT-LLM 1.2.1 could not load the newly released Qwen3.8-27B at all, and the 1.3.0rc28 candidate died compiling its FP8 kernels (open bug #14676). On older Llama 3.1 8B it was competitive — 141 tok/s at C=1, 3,219 tok/s at C=50, and the best TTFT (0.54 s) of five engines (winder.ai).
- Best for: NVIDIA-only estates on well-established models, where its CUDA-graph-optimised decode gives the lowest TTFT. Verify your exact model against the latest stable release first.
Quickstart
Assembled from the verified facts on this page and its Sources list. Versions are the ones recorded in research on 2026-10-09.
Install
# NVIDIA-only. Use the official container; TensorRT-LLM 1.3
docker run --rm --gpus all -p 8000:8000 \
nvcr.io/nvidia/tensorrt-llm/pytorch:latest
Run
# Inside the container: build the engine, then serve
trtllm-build --checkpoint_dir ./ckpt/Qwen3-8B \
--output_dir ./engine_out
mpirun -n 1 trtllm-serve ./engine_out --port 8000
Verify
curl http://localhost:8000/v1/models
Expected output: Lowest TTFT of the engines tested when the model is supported (median 2.1s class at concurrency, best-in-class prefill).
Pitfalls: Model support is the gate, and it bit us in testing: TensorRT-LLM could not load a hybrid linear-attention model at all. Confirm your exact model is supported before committing. Also note TRITON MoE is deprecated and 1.2 removed the TensorRT backend entirely.
Sources
- https://github.com/ggml-org/llama.cpp — llama.cpp repo; 130,619 stars, 11,523 commits, active 2026-10-09.
- https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md — authoritative list of llama.cpp backends (CUDA, HIP, Metal, Vulkan, SYCL, MUSA, CANN, ZenDNN, OpenCL, OpenVINO, Hexagon, KleidiAI, WebGPU, BLAS vendors).
- https://github.com/ggml-org/llama.cpp/blob/master/docs/multi-gpu.md — llama.cpp multi-GPU split modes (
none/layer/row), flags and recipes. - https://github.com/ollama/ollama — Ollama repo; 182,473 stars, releases v0.40.1 (2026-10-07).
- https://ollama.com/blog — Ollama blog; MLX engine on Apple Silicon (Mar/Jun 2026), GGUF via llama.cpp (0.30, Jun 2026), $88M raise, 8.9M developers, Anthropic API + Claude Desktop support.
- https://docs.ollama.com/gpu — Ollama hardware support: CUDA compute 5.0+, ROCm v7 on Linux and Windows, Metal, Vulkan defaults and
GGML_VK_VISIBLE_DEVICES. - https://lmstudio.ai/blog/0.4.0 — LM Studio 0.4.0:
llmsterheadless daemon, parallel requests with continuous batching, stateful/v1/chat, unified KV cache. - https://lmstudio.ai/changelog/lmstudio — LM Studio changelog through 0.4.26 (2026-10-08): Splash engine, DFlash/DSpark/MTP drafters, CUDA 13 on Windows ARM, llama.cpp 2.50.0/b11337.
- https://docs.vllm.ai/en/stable/usage/v1_guide/ — vLLM V1 guide: V0 fully deprecated; hardware status (NVIDIA/AMD/Intel/TPU/CPU all functional); plugins.
- https://docs.vllm.ai/en/latest/features/quantization/ — vLLM quantization formats and per-hardware compatibility matrix (AWQ, GPTQ, Marlin, FP8, GGUF, etc.).
- https://vllm.ai/blog/2026-09-22-vllm-metal-v0-28-0 — Announcing vllm-metal: paged varlen Metal kernel, MTP, GGUF/hybrid support, ragged-batch benchmark vs mlx_lm/oMLX/llama.cpp.
- https://www.sglang.io/ — SGLang site: 0.5.21 highlights (Rust prefix-cache core, PD role switching, radix tree), supported hardware NVIDIA/AMD/CPU/TPU/Ascend/XPU.
- https://github.com/sgl-project/sglang — SGLang repo; 36,919 stars, active 2026-10-09, DeepSeek-V4 Rust processor parity work.
- https://github.com/turboderp-org/exllamav3 — ExLlamaV3 repo: EXL3/QTIP quantisation, 2–8 bit cache quant, active ROCm wheels (Oct 2026), "still in development" caveat.
- https://github.com/turboderp-org/exllamav2 — ExLlamaV2 repo: last commit 2026-03-04 — confirms legacy/dormant status.
- https://nvidia.github.io/TensorRT-LLM/latest/release-notes.html — TensorRT-LLM 1.3 release notes: TRITON MoE deprecation; 1.2 removed the TensorRT backend entirely, added DGX Spark beta.
- https://github.com/mlc-ai/mlc-llm — MLC-LLM repo; 23,226 stars, last commit 2026-10-06, TVM refactor work.
- https://github.com/mozilla-ai/llamafile — llamafile repo under Mozilla AI; active 2026-10-08, agent.cpp subtree and "Agentfile" work.
- https://github.com/nomic-ai/gpt4all — GPT4All repo: last commit 2025-05-27 — confirms end-of-life status.
- https://github.com/mudler/localai — LocalAI repo; 49,450 stars, 8,447 commits, 4.11.0 (Oct 2026), 73 backends.
- https://localai.io/blog/what-landed-in-localai-4-8/ — LocalAI 4.8: model variants auto-selection, vllm.cpp (alpha) engine, 3.48× lighter web UI.
- https://localai.io/docs/features/vllm-cpp/index.html — vllm.cpp backend docs: Blackwell-only CUDA images (sm_120a/sm_121a), CUDA 13 required, Vulkan/CPU fallback, KV sizing (144 KiB/token).
- https://github.com/localai-org/vllm.cpp — vllm.cpp repo: C++20 port of vLLM with continuous batching, paged KV, RadixAttention; 6,363 commits, active 2026-10-09.
- https://www.jan.ai/docs/desktop — Jan docs: 0.8.4, Jan Server OpenAI-compatible API, CLI, llama.cpp + experimental MLX engines.
- https://freedom.tech/posts/2026-09-26-koboldcpp-1-122-1/ — KoboldCpp 1.122.1: integrated agent with 9 tools, MCP, AGENTS.md, ubatch handling.
- https://winder.ai/vllm-vs-ollama-vs-sglang-llm-inference-comparison/ — Reproducible H100 benchmark of vLLM 0.30 / SGLang 0.5.20 / TensorRT-LLM 1.2.1 / llama.cpp b11179 / Ollama 0.34.4 on Llama 3.1 8B and Qwen3.8-27B.
- https://www.soothill.io/blog/2026/08/10/sglang-vllm-llamacpp-evox3/ — 14-point AMD Strix Halo (
gfx1151) comparison of SGLang 0.5.17 / vLLM 0.26.0 / llama.cpp b10333; ROCm 7.14 support gaps. - https://ai-tldr.dev/tools/magnitude/ — Magnitude: Rust agent-oriented engine, kernel autotuning, OpenAI+Anthropic API on port 10100, v0.2.0 (2026-09-30), vendor benchmark claims.
- https://www.developersdigest.tech/blog/magnitude-self-tuning-local-inference-engine-2026 — Magnitude launch coverage: Apache-2.0, YC-backed, claims up to 2× faster decode than llama.cpp.
- https://ai-tldr.dev/tools/inco-splash/ — Splash: Apple Silicon engine with precompiled model-specific Metal kernels, DFlash 2 speculative decoding, M3+/macOS 26.4+.
- https://presenc.ai/research/mlx-vs-llama-cpp-throughput-benchmarks-2026 — MLX vs llama.cpp Apple Silicon benchmarks: 10–20% decode advantage above 14B, 4.06× TTFT on M5, crossover past ~40k context.
- https://www.glukhov.org/llm-hosting/comparisons/hosting-llms-ollama-localai-jan-lmstudio-vllm-comparison/ — 14-tool comparison across API maturity, tool calling, GPU support, formats, production readiness.
- https://docs.litellm.ai/docs/proxy_server — LiteLLM Proxy: OpenAI-compatible gateway over 100+ providers including local engines.
- https://particula.tech/blog/ollama-num-ctx-silent-prompt-truncation — Ollama single-slot architectures (mllama, qwen3vl, nemotron_h) forced to one slot since 2026-02-02.
- https://helix.ml/blog/the-ceiling-was-a-state-cache — SGLang mamba state cache silently capping concurrency; 3,833 → 7,883 tok/s after fix.
- https://markaicode.com/benchmarks/gpt4all-production-benchmark-latency/ — GPT4All end-of-life confirmation and CPU-only status.
- https://mortalapps.com/blog/gguf-vs-exl2-vs-mlx-quantization/ — ExLlamaV2 legacy/archived status and EXL3 succession in 2026.
- https://atomic.chat/blog/guides/exl3-vs-gguf — EXL3 vs GGUF quality/speed comparison on RTX 5090.
- https://localai.io/blog/what-landed-in-localai-4-11/ — LocalAI 4.11: diarisation, speaker profiles, model failover chains, Operate → This machine.
- https://developer.apple.com/videos/play/wwdc2026/232/ — WWDC26: MLX-LM and the OpenAI-compatible MLX-LM Server for local agentic AI on Mac.
Date: 2026-10-09