Skip to content

Jan — Open-Source ChatGPT Alternative

Self-contained guide to Jan — the local inference engine.

  • Actively maintained: Jan 0.8.4 (changelog dated 2026-07-21, docs page updated 2026-10-09), 44.9k stars, Apache 2.0. Covers Jan Desktop (macOS/Windows/Linux), Jan Web, Jan Server (a local OpenAI-compatible API server), Jan CLI, Jan Models (e.g. jan-code-4b), Cowork (autonomous task execution), Agents, and integrations with Claude Code / OpenClaw / MCP.
  • Engines: llama.cpp (primary) and MLX (experimental on macOS 14+). Explicitly credits llama.cpp/GGML as its foundation (docs).
  • Best for: users wanting an offline, open-source, self-owned ChatGPT-like app with an API server and agent features. Its local-LLM API maturity is rated Beta and tool calling limited by independent review (glukhov.org).

Quickstart

Assembled from the verified facts on this page and its Sources list. Versions are the ones recorded in research on 2026-10-09.

Install

# Download from jan.ai (open-source ChatGPT alternative)

Run

# GUI app, or the bundled Jan Server for OpenAI-compatible API access
jan serve --port 1337

Verify

curl http://localhost:1337/v1/models

Expected output: Version 0.8.4. Ships llama.cpp plus an experimental MLX engine, a CLI and an OpenAI-compatible API.

Pitfalls: The MLX engine is experimental. On Apple Silicon prefer native MLX unless you specifically want Jan's UI.

Sources

  1. https://github.com/ggml-org/llama.cpp — llama.cpp repo; 130,619 stars, 11,523 commits, active 2026-10-09.
  2. https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md — authoritative list of llama.cpp backends (CUDA, HIP, Metal, Vulkan, SYCL, MUSA, CANN, ZenDNN, OpenCL, OpenVINO, Hexagon, KleidiAI, WebGPU, BLAS vendors).
  3. https://github.com/ggml-org/llama.cpp/blob/master/docs/multi-gpu.md — llama.cpp multi-GPU split modes (none/layer/row), flags and recipes.
  4. https://github.com/ollama/ollama — Ollama repo; 182,473 stars, releases v0.40.1 (2026-10-07).
  5. https://ollama.com/blog — Ollama blog; MLX engine on Apple Silicon (Mar/Jun 2026), GGUF via llama.cpp (0.30, Jun 2026), $88M raise, 8.9M developers, Anthropic API + Claude Desktop support.
  6. https://docs.ollama.com/gpu — Ollama hardware support: CUDA compute 5.0+, ROCm v7 on Linux and Windows, Metal, Vulkan defaults and GGML_VK_VISIBLE_DEVICES.
  7. https://lmstudio.ai/blog/0.4.0 — LM Studio 0.4.0: llmster headless daemon, parallel requests with continuous batching, stateful /v1/chat, unified KV cache.
  8. https://lmstudio.ai/changelog/lmstudio — LM Studio changelog through 0.4.26 (2026-10-08): Splash engine, DFlash/DSpark/MTP drafters, CUDA 13 on Windows ARM, llama.cpp 2.50.0/b11337.
  9. https://docs.vllm.ai/en/stable/usage/v1_guide/ — vLLM V1 guide: V0 fully deprecated; hardware status (NVIDIA/AMD/Intel/TPU/CPU all functional); plugins.
  10. https://docs.vllm.ai/en/latest/features/quantization/ — vLLM quantization formats and per-hardware compatibility matrix (AWQ, GPTQ, Marlin, FP8, GGUF, etc.).
  11. https://vllm.ai/blog/2026-09-22-vllm-metal-v0-28-0 — Announcing vllm-metal: paged varlen Metal kernel, MTP, GGUF/hybrid support, ragged-batch benchmark vs mlx_lm/oMLX/llama.cpp.
  12. https://www.sglang.io/ — SGLang site: 0.5.21 highlights (Rust prefix-cache core, PD role switching, radix tree), supported hardware NVIDIA/AMD/CPU/TPU/Ascend/XPU.
  13. https://github.com/sgl-project/sglang — SGLang repo; 36,919 stars, active 2026-10-09, DeepSeek-V4 Rust processor parity work.
  14. https://github.com/turboderp-org/exllamav3 — ExLlamaV3 repo: EXL3/QTIP quantisation, 2–8 bit cache quant, active ROCm wheels (Oct 2026), "still in development" caveat.
  15. https://github.com/turboderp-org/exllamav2 — ExLlamaV2 repo: last commit 2026-03-04 — confirms legacy/dormant status.
  16. https://nvidia.github.io/TensorRT-LLM/latest/release-notes.html — TensorRT-LLM 1.3 release notes: TRITON MoE deprecation; 1.2 removed the TensorRT backend entirely, added DGX Spark beta.
  17. https://github.com/mlc-ai/mlc-llm — MLC-LLM repo; 23,226 stars, last commit 2026-10-06, TVM refactor work.
  18. https://github.com/mozilla-ai/llamafile — llamafile repo under Mozilla AI; active 2026-10-08, agent.cpp subtree and "Agentfile" work.
  19. https://github.com/nomic-ai/gpt4all — GPT4All repo: last commit 2025-05-27 — confirms end-of-life status.
  20. https://github.com/mudler/localai — LocalAI repo; 49,450 stars, 8,447 commits, 4.11.0 (Oct 2026), 73 backends.
  21. https://localai.io/blog/what-landed-in-localai-4-8/ — LocalAI 4.8: model variants auto-selection, vllm.cpp (alpha) engine, 3.48× lighter web UI.
  22. https://localai.io/docs/features/vllm-cpp/index.html — vllm.cpp backend docs: Blackwell-only CUDA images (sm_120a/sm_121a), CUDA 13 required, Vulkan/CPU fallback, KV sizing (144 KiB/token).
  23. https://github.com/localai-org/vllm.cpp — vllm.cpp repo: C++20 port of vLLM with continuous batching, paged KV, RadixAttention; 6,363 commits, active 2026-10-09.
  24. https://www.jan.ai/docs/desktop — Jan docs: 0.8.4, Jan Server OpenAI-compatible API, CLI, llama.cpp + experimental MLX engines.
  25. https://freedom.tech/posts/2026-09-26-koboldcpp-1-122-1/ — KoboldCpp 1.122.1: integrated agent with 9 tools, MCP, AGENTS.md, ubatch handling.
  26. https://winder.ai/vllm-vs-ollama-vs-sglang-llm-inference-comparison/ — Reproducible H100 benchmark of vLLM 0.30 / SGLang 0.5.20 / TensorRT-LLM 1.2.1 / llama.cpp b11179 / Ollama 0.34.4 on Llama 3.1 8B and Qwen3.8-27B.
  27. https://www.soothill.io/blog/2026/08/10/sglang-vllm-llamacpp-evox3/ — 14-point AMD Strix Halo (gfx1151) comparison of SGLang 0.5.17 / vLLM 0.26.0 / llama.cpp b10333; ROCm 7.14 support gaps.
  28. https://ai-tldr.dev/tools/magnitude/ — Magnitude: Rust agent-oriented engine, kernel autotuning, OpenAI+Anthropic API on port 10100, v0.2.0 (2026-09-30), vendor benchmark claims.
  29. https://www.developersdigest.tech/blog/magnitude-self-tuning-local-inference-engine-2026 — Magnitude launch coverage: Apache-2.0, YC-backed, claims up to 2× faster decode than llama.cpp.
  30. https://ai-tldr.dev/tools/inco-splash/ — Splash: Apple Silicon engine with precompiled model-specific Metal kernels, DFlash 2 speculative decoding, M3+/macOS 26.4+.
  31. https://presenc.ai/research/mlx-vs-llama-cpp-throughput-benchmarks-2026 — MLX vs llama.cpp Apple Silicon benchmarks: 10–20% decode advantage above 14B, 4.06× TTFT on M5, crossover past ~40k context.
  32. https://www.glukhov.org/llm-hosting/comparisons/hosting-llms-ollama-localai-jan-lmstudio-vllm-comparison/ — 14-tool comparison across API maturity, tool calling, GPU support, formats, production readiness.
  33. https://docs.litellm.ai/docs/proxy_server — LiteLLM Proxy: OpenAI-compatible gateway over 100+ providers including local engines.
  34. https://particula.tech/blog/ollama-num-ctx-silent-prompt-truncation — Ollama single-slot architectures (mllama, qwen3vl, nemotron_h) forced to one slot since 2026-02-02.
  35. https://helix.ml/blog/the-ceiling-was-a-state-cache — SGLang mamba state cache silently capping concurrency; 3,833 → 7,883 tok/s after fix.
  36. https://markaicode.com/benchmarks/gpt4all-production-benchmark-latency/ — GPT4All end-of-life confirmation and CPU-only status.
  37. https://mortalapps.com/blog/gguf-vs-exl2-vs-mlx-quantization/ — ExLlamaV2 legacy/archived status and EXL3 succession in 2026.
  38. https://atomic.chat/blog/guides/exl3-vs-gguf — EXL3 vs GGUF quality/speed comparison on RTX 5090.
  39. https://localai.io/blog/what-landed-in-localai-4-11/ — LocalAI 4.11: diarisation, speaker profiles, model failover chains, Operate → This machine.
  40. https://developer.apple.com/videos/play/wwdc2026/232/ — WWDC26: MLX-LM and the OpenAI-compatible MLX-LM Server for local agentic AI on Mac.

Date: 2026-10-09