
Gemma 4 on a laptop is snappy. The GPU pulls its weight, the answers land in seconds, and the model feels like a real assistant. Move it to a NAS and it slows down enough that a question becomes a small commitment. That gap is the point. When the model lives on shared storage, the whole family or team can call it, and the extra second of latency turns it back into a tool instead of a reflex.
We picked eight apps we run on a NAS to host local LLMs today. Some are the server, some are the front-end, some are the deployment layer. Together they cover the workflow from downloading a model to a family of clients calling a shared endpoint.
What to look for in a NAS-hosted LLM app
- Runs headless. A NAS has no monitor and often no GPU.
- Exposes a stable HTTP API so any client can talk to it.
- Handles model management without repacking a Docker image.
- Streams tokens, because a NAS-hosted model is slower and streaming makes it usable.
- An access control layer, even a simple token, because the endpoint sits on the LAN.
The eight below meet at least four of the five. The rank orders them by how easy they are to get running on a Synology, QNAP, or self-built NAS.
Quick comparison
| App | Best for | Runs on | GPU support | License |
|---|---|---|---|---|
| Ollama | The easiest NAS-hosted server | Linux, macOS, Windows | CUDA, Metal, ROCm | MIT |
| LM Studio | Point-and-click model host | Linux, macOS, Windows | CUDA, Metal, ROCm | Proprietary (free) |
| Open WebUI | Shared web front-end | Docker | Via the backend | MIT |
| LocalAI | OpenAI-compatible drop-in | Docker | CUDA, ROCm, CPU | MIT |
| Text Generation WebUI | Power-user tuning | Linux, macOS, Windows | CUDA, ROCm, CPU | AGPL |
| Jan | Local-first desktop client | Windows, macOS, Linux | CUDA, Metal, CPU | AGPL |
| vLLM | High-throughput inference | Linux | CUDA (server GPUs) | Apache 2 |
| llama.cpp | Bare-metal CPU inference | Any | Any | MIT |
1. Ollama, best for the easiest NAS-hosted server
Ollama is the closest thing to a one-command install for a NAS. Docker image, port 11434, and every client that speaks the Ollama or the OpenAI API talks to it. Model library covers Llama, Gemma, Mistral, Qwen, and a growing list of instruct variants.
Where it falls short: the fine-grained sampling controls are hidden behind flags. Anyone who wants to swap prompt templates by hand ends up in the modelfile syntax.
Pricing:
- Free, MIT.
Platforms: Linux, macOS, Windows (all as Docker or native).
Download: ollama.com
Bottom line: The default recommendation for anyone standing up a NAS LLM for the first time.
2. LM Studio, best for a point-and-click model host
LM Studio is the desktop app that also runs an OpenAI-compatible server. Point it at a NAS-mounted model folder, enable the server, and clients on the LAN can call it. The GUI hides the parts of llama.cpp that scare people off.
Where it falls short: it is a desktop app first, so headless install on a NAS means running it inside a lightweight X server or picking a NAS with a screen for setup.
Pricing:
- Free for personal use.
Platforms: Windows, macOS, Linux.
Download: lmstudio.ai
Bottom line: The right pick when a family member wants to change models and does not want to touch the terminal.
3. Open WebUI, best for a shared web front-end
Open WebUI is the ChatGPT-shaped web interface for a NAS-hosted backend. Point it at Ollama, LM Studio, or LocalAI, add users, and every household member gets their own chats and a shared library of prompts. RAG over local files ships in the box.
Where it falls short: it is a front-end, not a runtime. Ollama or LocalAI still does the model work.
Pricing:
- Free, MIT.
Platforms: Docker on Linux, plus a native NAS install through Portainer.
Download: openwebui.com
Bottom line: The right pairing with Ollama when the goal is a shared, browser-first family chat.
4. LocalAI, best for an OpenAI-compatible drop-in
LocalAI exposes an OpenAI-compatible endpoint (chat, embeddings, images, TTS) against local models. Any app that speaks the OpenAI API talks to it without a change. The backends span llama.cpp, whisper.cpp, and stable-diffusion.cpp.
Where it falls short: the config file is dense and each model needs its own entry. Ollama handles the mapping for you.
Pricing:
- Free, MIT.
Platforms: Docker on Linux, plus native builds for Windows and macOS.
Download: localai.io
Bottom line: The right server when the client is an OpenAI SDK and swapping the base URL is the whole migration.
5. Text Generation WebUI, best for power-user tuning
Text Generation WebUI (oobabooga) is the older, richer front-end that also runs a server. Sampling parameters, LoRA loading, and character-card style prompt templates are all exposed. Anyone who wants to tune a model beyond default sits here.
Where it falls short: the UI has been rebuilt twice and still shows its power-user roots. It is not the tool you hand to a family member.
Pricing:
- Free, AGPL.
Platforms: Windows, macOS, Linux.
Download: github.com/oobabooga/text-generation-webui
Bottom line: The pick for anyone who wants sampling knobs Ollama and LM Studio hide.
6. Jan, best for a local-first desktop client
Jan is the local-first ChatGPT client. It ships with its own runtime for CPU-only NASes, but it also speaks to a remote Ollama or LocalAI server. Chats and prompts stay on disk, so the whole history is a local file.
Where it falls short: as a NAS server, it is not the strongest fit. Use it as the client, and let Ollama or LocalAI do the serving.
Pricing:
- Free, AGPL.
Platforms: Windows, macOS, Linux.
Download: jan.ai
Bottom line: The right desktop companion to a NAS server for anyone who does not want a web tab.
7. vLLM, best for high-throughput inference
vLLM is the production-grade server. Paged attention, tensor parallelism, and continuous batching make it the fastest way to serve a big model to many clients. The catch is it needs a real GPU (Ampere or later).
Where it falls short: it is not a NAS-friendly install. A NAS chassis without a datacenter GPU cannot run it in the way the docs assume.
Pricing:
- Free, Apache 2.
Platforms: Linux (typically Docker with NVIDIA container toolkit).
Download: github.com/vllm-project/vllm
Bottom line: The pick when the “NAS” is a Linux box with a real GPU serving a home lab of clients.
8. llama.cpp, best for bare-metal CPU inference
llama.cpp is the low-level runtime everything else on this list wraps. It runs on CPU-only NAS builds where a real GPU is missing, with quantized models that shrink an 8B parameter model to a few gigabytes.
Where it falls short: there is no UI. It is a server binary and a CLI. Pair it with Open WebUI or Ollama.
Pricing:
- Free, MIT.
Platforms: Any (Linux, macOS, Windows, ARM including Raspberry Pi and Apple Silicon).
Download: github.com/ggerganov/llama.cpp
Bottom line: The right pick when the NAS is CPU-only and every gigabyte of RAM matters.
How to pick the right one
- One-command install on a NAS: Ollama.
- Shared web chat for the household: Ollama plus Open WebUI.
- OpenAI-compatible endpoint any SDK can call: LocalAI.
- Point-and-click model swap: LM Studio.
- Sampling and LoRA control: Text Generation WebUI.
- Local desktop client hitting a NAS: Jan.
- Serious GPU throughput for many clients: vLLM.
- CPU-only NAS: llama.cpp.
FAQ
Which model runs on a NAS with no GPU?
Quantized versions of Gemma, Llama, or Mistral in the 3B-8B range run on a modern NAS CPU at a few tokens a second. That is slow for interactive chat, fine for a one-shot summary or a nightly job.
Do I need a GPU on my NAS?
Only for interactive speed. A CPU-only NAS still runs a small model; response times sit at seconds per sentence instead of tokens per second.
Can I keep the NAS-hosted model private to my LAN?
Yes. Ollama, LM Studio, and LocalAI bind to the LAN interface by default. Add basic auth or a reverse proxy for anything more than a home LAN.
How much disk space do local models need?
A quantized 8B model sits at about 5 GB. A quantized 70B model sits closer to 40 GB. A NAS with a spare 500 GB volume holds a full library.
Can multiple people use the model at the same time?
Ollama, LocalAI, and Text Generation WebUI all serialize by default. vLLM handles concurrent requests natively. For a small family, serialization is fine.
What about privacy?
Every app on this list runs the model locally. No prompt or response is sent to a third party unless the app is configured to.