Security — Local vs Cloud, Supply Chain, SafeTensors - Local AI Agent Wiki __md_scope=new URL("../../..",location),__md_hash=e=>[...e].reduce(((e,_)=>(e<<5)-e+_.charCodeAt(0)),0),__md_get=(e,_=localStorage,t=__md_scope)=>JSON.parse(_.getItem(t.pathname+"."+e)),__md_set=(e,_,t=localStorage,a=__md_scope)=>{try{t.setItem(a.pathname+"."+e,JSON.stringify(_))}catch(e){}}
Skip to content

Security — Local vs Cloud, Supply Chain, SafeTensors

Local is private by default, not automatically secure. Inference stays on your machine, but three other data flows leak: tool telemetry, malicious model files, and network exposure.

Local is private by default, not automatically secure. Inference stays on your machine, but three other data flows leak: tool telemetry, malicious model files, and network exposure.

Verified concrete controls (12-point checklist, Aug 2026):

Risk Control
Untrusted model files (supply chain) Download only from Hugging Face or the official Ollama library; verify SHA-256 before loading sensitive workloads
API network exposure Ollama binds to localhost by default — never set OLLAMA_HOST=0.0.0.0; bind all endpoints to 127.0.0.1
Telemetry LM Studio: Settings → Privacy → uncheck "Send anonymous usage data"; GPT4All: disable analytics. Ollama and Jan collect nothing by default (auditable source)
Prompt/log confidentiality Full-disk encryption (macOS FileVault / Linux LUKS); store chat logs in an encrypted folder or disable logging
Lateral risk Dedicated user account for LLM work; audit all extensions/plugins; block outbound from the model process with pf (macOS) or ufw/OpenSnitch (Linux)
Regulated data Air-gap the machine; document approved model versions for audit

Model supply-chain threat model (2026 state of the art): - Pickle deserialization remains the default in many PyTorch checkpoints and is functionally arbitrary code execution at load time. SafeTensors eliminates this vector entirely — refuse non-SafeTensors loads or sandbox pickle loads with no network/filesystem access. - trust_remote_code=True executes tokenizer/config code — disable by default, allowlist only verified publishers. - Weight-level backdoors (BadNets/TrojDiff style) are the hard, active research frontier: detection via Sigstore-style signed artifacts/weight hashes (Meta, Mistral, Alibaba, Anthropic participate for open-weight releases), adversarial behavioural corpora, and weight-statistics tools (ModelScan, NeuralCleanse) — with real false-positive problems. - Dataset poisoning is the least mature corner; CycloneDX 1.6 added AI/ML dataset components for SBOM-style provenance. - Scale of the problem: the Hugging Face Hub took down nine separately reported malicious model uploads in Q1 2026, at least three already pulled into production pipelines. MITRE ATLAS is the de facto tactical reference. - Privacy vs delegated security: cloud providers handle infrastructure security but your prompts transit their servers and are subject to legal process; local gives privacy autonomy at the cost of manual hardening. Local also removes the Art. 28 GDPR processor obligation for the model vendor.

Sources

  • https://huggingface.co/docs/hub/en/local-apps — Hugging Face's official "Use AI Models Locally" docs (llama.cpp, Ollama, Jan, LM Studio one-command flows)
  • https://huggingface.co/learn/llm-course/en/chapter1/1 — Hugging Face LLM Course overview (structure, prerequisites, prerequisites, notebooks)
  • https://huggingface.co/papers — Hugging Face Daily Papers (arXiv ranked by community upvotes)
  • https://www.reddit.com/r/LocalLLaMA/ — r/LocalLLaMA, the local-inference community hub
  • https://subriff.com/guides/best-subreddits-for-ai — r/LocalLLaMA membership/growth stats (844,249 members, ~60 posts/day, Sep 2026)
  • https://prowlo.com/tools/subreddit-stats/localllama — Independent r/LocalLLaMA stats (820,305 members, 9 Sep 2026 crawl)
  • https://dupple.com/learn/ai-news-for-developers — 2026 developer AI news reading list (Simon Willison, Latent Space, The Batch, HN, Daily Papers)
  • https://aiwiki.ai/wiki/open_weight_license_comparison — Open-weight LLM licence comparison, verified July 2026 (per-model table: Apache-2.0/MIT vs Llama 700M MAU vs Gemma/MRL/OpenRAIL)
  • https://opensource.org/ai/open-source-ai-definition — OSI Open Source AI Definition 1.0 (weights + architecture + usage info under OSI terms)
  • https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53 — EU AI Act Article 53 text incl. the 53(2) open-source exemption
  • https://ai-act-service-desk.ec.europa.eu/en/ai-act/faq/how-does-ai-act-apply-general-purpose-ai-models-released-open-source — Official FAQ on the open-source exemption's limits (no systemic-risk models; copyright/training-data duties survive)
  • https://digital-strategy.ec.europa.eu/en/faqs/guidelines-obligations-general-purpose-ai-providers — Commission GPAI guidelines (provider duties, fine-tuning/modification triggers)
  • https://corp-intl.com/news/what-is-the-timeline-for-implementing-the-eu-ai-act — Post-omnibus EU AI Act timeline, 16 Sep 2026 (Digital Omnibus 2026/1744, in force 27 Jul 2026; high-risk pushed to Dec 2027 / Aug 2028)
  • https://www.europarl.europa.eu/legislative-train/package-digital-package/file-digital-omnibus-on-ai — Parliament legislative train on the Digital Omnibus on AI (7 May 2026 trilogue agreement)
  • https://www.promptquorum.com/local-llms/local-llm-security-privacy-checklist — 12-point local LLM security checklist (telemetry defaults per tool, SHA-256 verification, localhost binding, pf/ufw egress blocking, GDPR/HIPAA/APPI/PIPL notes)
  • https://safeguard.sh/resources/blog/model-supply-chain-poisoning-detection-2026 — 2026 model supply-chain threat model (pickle, trust_remote_code, weight backdoors, dataset poisoning, Sigstore signing, MITRE ATLAS, 9 HF takedowns in Q1 2026)
  • https://docs.nvidia.com/cuda//cuda-toolkit-release-notes/index.html — CUDA 13.4 U1 release notes (driver no longer bundled since 13.4 Linux / 13.1 Windows; toolkit→R-branch table; minor-version compatibility)
  • https://docs.nvidia.com/datacenter/tesla/drivers/latest/pdf/NVIDIA_Datacenter_Drivers.pdf — NVIDIA datacenter driver lifecycle (New Feature Branch vs Production Branch)
  • https://rocm.blogs.amd.com/artificial-intelligence/language-models-locally/README.html — AMD's first-party "Practical Guide to Running LLMs on AMD Radeon GPUs" (19 Jun 2026)
  • https://localaimaster.com/blog/amd-rocm-local-llm-setup — ROCm 7.2.x state-of-play, supported/unsupported GPU table, HSA_OVERRIDE_GFX_VERSION, ROCm vs CUDA comparison, ~96 tok/s on 7900 XTX
  • https://rocm.docs.amd.com/projects/ai-ecosystem/en/latest/inference/vllm.html — Official vLLM-on-ROCm setup (prebuilt Docker image, ROCm 7.x)
  • https://d-central.tech/cuda-vs-rocm-local-inference/ — CUDA vs ROCm vs Vulkan backend comparison for local inference
  • https://freedom.tech/posts/2026-10-05-llama-cpp-0-6-0/ — llama.cpp 0.6.0 release notes (5 Oct 2026)
  • https://machinelearning.apple.com/research/exploring-llms-mlx-m5 — Apple ML research: MLX on M5 Neural Accelerators (TTFT up to 3.97×, generation 1.19–1.27×, macOS 26.2+ requirement)
  • https://developer.apple.com/videos/play/wwdc2026/232/ — Apple WWDC26: local agentic AI on the Mac with MLX
  • https://codersera.com/blog/apple-silicon-llms-complete-guide-2026/ — Apple Silicon LLM guide (MLX vs vllm-mlx vs oMLX vs Ollama, unified-memory constraints)
  • https://www.iunera.com/kraken/enterprise-ai/top-20-tools-to-run-llms-locally-in-2026-ollama-anythingllm-open-webui-lm-studio-vllm-and-every-real-alternative-compared/ — 20-tool local LLM comparison (difficulty, open source, enterprise readiness)
  • https://presenc.ai/research/local-llm-vs-cloud-api-cost-2026 — Local vs cloud cost/TCO and breakeven analysis, updated October 2026
  • https://www.promptquorum.com/local-llms/local-llms-vs-cloud-apis — Local vs cloud 8-factor comparison (privacy, cost, speed, quality, regional compliance)
  • https://opentelemetry.io/blog/2026/genai-observability/index.md — OTel GenAI semantic conventions walkthrough (span attributes, metrics, Aspire Dashboard)
  • https://signoz.io/docs/open-webui-monitoring/ — Open WebUI observability with OpenTelemetry
  • https://docs.openwebui.com/features/administration/analytics/ — Open WebUI built-in analytics (usage, token consumption, per-model/per-user)
  • https://github.com/prove-ai/observability-pipeline/blob/main/docs/guides/vllm-guide.md — vLLM + Prometheus GPU monitoring guide
  • https://artificialanalysis.ai/hardware-inference-stack/laptops-workstations — Artificial Analysis local inference benchmark leaderboard
  • https://artificialanalysis.ai/articles/aa-agentperf-local — AA-AgentPerf-Local: open-source local agent benchmark tool (29 Sep 2026)
  • https://simonwillison.net/tags/local-llms/ — Simon Willison's local-LLM blog archive (165+ posts)
  • https://www.turingpost.com/p/tools-for-model-deployment — Turing Post: 2026 open-source model deployment tooling overview
  • https://www.datacamp.com/tutorial/gguf-format-a-complete-guide — GGUF quantization and sizing reference (7B FP16 ~14 GB vs Q4_K_M ~4–5 GB)

Date: 2026-10-09

var target=document.getElementById(location.hash.slice(1));target&&target.name&&(target.checked=target.name.startsWith("__tabbed_"))
{"annotate": null, "base": "../../..", "features": [], "search": "../../../assets/javascripts/workers/search.2c215733.min.js", "tags": null, "translations": {"clipboard.copied": "Copied to clipboard", "clipboard.copy": "Copy to clipboard", "search.result.more.one": "1 more on this page", "search.result.more.other": "# more on this page", "search.result.none": "No matching documents", "search.result.one": "1 matching document", "search.result.other": "# matching documents", "search.result.placeholder": "Type to start searching", "search.result.term.missing": "Missing", "select.version": "Select version"}, "version": null}