Evaluating Local Models — lm-evaluation-harness
EleutherAI's lm-evaluation-harness is the standard for reproducible open-model benchmarking (MMLU, HellaSwag, GSM8K, etc.). It evaluates local models directly via transformers or vLLM backend — fully offline.
lm-evaluation-harness (EleutherAI) remains the standard for reproducible open-model benchmarking (MMLU, HellaSwag, GSM8K, etc.). Verified maintained: PyPI lm-eval v0.4.13, last uploaded 2026-08-31 — active, though less frequently than inference runtimes (GitHub, tutorial). It evaluates local models directly via transformers or vLLM backend, so it fits a fully-offline loop. Pair it with task-specific eval sets (Haystack has built-in eval; LangSmith/LangFuse for tracing) — the general leaderboards are table stakes, not a substitute for your own task benchmark.
Sources
- https://vllm.ai — vLLM: high-throughput OpenAI-compatible local serving engine (docs entry point).
- https://ollama.readthedocs.io/en/openai — Ollama's experimental OpenAI-compatible
/v1/...API and its feature/field gaps. - https://github.com/ggml-org/llama.cpp — llama.cpp repo: v0.6.0 (Oct 2026), 130.6k★, commits within hours of Oct 9 2026.
- https://github.com/ggml-org/llama.cpp/releases — llama.cpp release history confirming v0.6.0 on Oct 5 2026.
- https://pypi.org/project/vllm/ — vLLM v0.31.0, last upload 2026-10-05 (maintenance verification).
- https://pypi.org/project/ollama/ — Ollama v0.6.3, last upload 2026-09-29.
- https://pypi.org/project/litellm/ — LiteLLM v1.104.2, 2026-10-08: proxy layer for normalising local endpoints.
- https://modelcontextprotocol.io/specification/2026-07-28 — MCP 2026-07-28 stateless spec: JSON-RPC, tools/resources/prompts, security model.
- https://blog.cloudflare.com/mcp-v2 — Cloudflare on the 2026-07-28 MCP spec and updated SDKs (stateless MCP).
- https://developers.googleblog.com/scaling-ai-agent-infrastructure-with-the-mcp-stateless-updates — Google on MCP stateless updates as foundational agent infra.
- https://www.descope.com/learn/post/mcp — MCP security: local-server privilege and startup-command risks.
- https://pypi.org/project/langchain/ — LangChain v1.4.4, 2026-10-08.
- https://pypi.org/project/llama-index/ — LlamaIndex v0.14.25, 2026-09-21.
- https://pypi.org/project/llama-index-vector-stores-qdrant/ — LlamaIndex↔Qdrant integration v0.10.4, 2026-10-08.
- https://pypi.org/project/crewai/ — CrewAI v1.15.26, 2026-10-08 (multi-agent crews, actively maintained).
- https://pypi.org/project/pyautogen/ — AutoGen v0.10.0, last upload 2025-07-15 (slowing cadence).
- https://pypi.org/project/dspy/ — DSPy v3.4.0, 2026-09-25.
- https://pypi.org/project/haystack-ai/ — Haystack v3.3.0, 2026-10-01.
- https://github.com/sst/opencode — OpenCode (now anomalyco/opencode): open-source coding agent, commits Oct 8 2026, 212k★.
- https://github.com/Aider-AI/aider — Aider: v0.86.2 (Feb 2026), last commit May 2026 — mature but slowed.
- https://platform.claude.com/docs/en/cli-sdks-libraries/libraries/openai-sdk — Anthropic docs: using Claude Code via an OpenAI-compatible SDK/gateway.
- https://github.com/RichardAtCT/claude-code-openai-wrapper — community wrapper to point Claude Code at OpenAI-compatible endpoints.
- https://huggingface.co/Qwen/Qwen3-Embedding-8B — Qwen3-Embedding series (0.6B/4B/8B + rerankers); #1 MTEB multilingual 70.58; instruct adds 1–5%.
- https://huggingface.co/Qwen/Qwen3-VL-Embedding-2B — Qwen3-VL-Embedding/Reranker for multimodal RAG; MMEB-V2 leader.
- https://pypi.org/project/qdrant-client/ — Qdrant client v1.19.1, 2026-09-16.
- https://pypi.org/project/chromadb/ — Chroma v1.5.9, 2026-05-05.
- https://pypi.org/project/faiss-cpu/ — FAISS v1.15.1, 2026-09-16.
- https://pypi.org/project/weaviate-client/ — Weaviate client v4.23.1, 2026-09-07.
- https://pypi.org/project/milvus/ — Milvus v2.3.9, last upload 2024-11-12 (PyPI stale).
- https://pypi.org/project/sentence-transformers/ — sentence-transformers v6.1.0, 2026-09-18 (embedding runtime).
- https://pypi.org/project/unsloth/ — Unsloth v2026.10.3, 2026-10-08 (fast single-GPU LoRA/QLoRA).
- https://github.com/axolotl-ai-cloud/axolotl — Axolotl repo: v0.20.0 (Sep 2026), multi-GPU/FSDP2, QAT, MoE LoRA.
- https://pypi.org/project/axolotl/ — Axolotl v0.20.0, 2026-09-30.
- https://pypi.org/project/mlx-lm/ — MLX-LM v0.32.0, 2026-10-01 (Apple Silicon fine-tuning).
- https://pypi.org/project/mlx/ — MLX v0.32.3, 2026-09-29.
- https://explore.n1n.ai/blog/fine-tune-llm-lora-qlora-guide-2026-2026-04-17 — 2026 VRAM table: QLoRA ~5–8GB for 3–8B, up to 70B at 4-bit.
- https://gigagpu.com/rtx-4090-24gb-for-fine-tuning — RTX 4090 24GB: LoRA to ~14B FP16, QLoRA to ~70B NF4.
- https://www.stack-junkie.com/blog/lora-fine-tuning-consumer-gpu — QLoRA on consumer GPU with Unsloth, export to GGUF for llama.cpp/Ollama.
- https://github.com/EleutherAI/lm-evaluation-harness — lm-evaluation-harness repo (standard open-model benchmarking).
- https://pypi.org/project/lm-eval/ — lm-eval v0.4.13, 2026-08-31 (maintenance verification).
- https://github.com/ggml-org/whisper.cpp — whisper.cpp: active local ASR (commits Oct 2026).
- https://github.com/SYSTRAN/faster-whisper — faster-whisper (CTranslate2) v1.2.1, last upload 2025-10-31.
- https://pypi.org/project/kokoro-onnx/ — kokoro-onnx v0.6.1, 2026-08-19 (active Kokoro TTS runtime).
- https://pypi.org/project/piper-tts/ — Piper TTS v1.8.0, 2026-09-04.
- https://pypi.org/project/mlx-audio/ — mlx-audio v0.5.8, 2026-10-05 (local speech on Apple Silicon).
- https://github.com/comfyanonymous/ComfyUI — ComfyUI (Comfy-Org/ComfyUI): modular diffusion GUI+API, commits Oct 9 2026.
- https://bfl.ai/blog/flux-1-kontext — FLUX.1 Kontext announcement (local/ComfyUI image generation).
- https://bfl.ai/blog/flux-2 — FLUX.2 announcement (Black Forest Labs open weights).
- https://daily.dev/blog/running-llms-locally-ollama-llama-cpp-self-hosted-ai-developers — 2026 overview: Ollama/llama.cpp/LM Studio/LocalAI all expose OpenAI-compatible endpoints.
- https://www.spheron.network/blog/openai-compatible-api-self-hosted — vLLM's OpenAI-compatible server endpoints (
/v1/chat/completionsetc.). - https://www.langchain.com/resources/langchain-vs-llamaindex — LangChain vs LlamaIndex capability comparison for retrieval and agents.