Run your own private AI, without the GPU headache.
Host open large language models on a server that’s sized right, locked down, and managed. Ollama serves the models; Open WebUI gives your team a ChatGPT-style interface — all on infrastructure you control, with your prompts staying yours.
A ChatGPT-style assistant that never leaves your server.
Ollama runs open models — Llama, Mistral, DeepSeek and more — locally over a simple REST API, and Open WebUI wraps them in a familiar chat interface with users, history and document chat. Together they give a team private AI: no per-token bill, no prompts sent to a third party, and full control over which models run. The hard part isn’t installing them — it’s sizing the machine, storing the models, and keeping the endpoint off the open internet. That’s ours to solve.
Where self-hosted AI usually goes wrong.
- Sizing RAM to the model. Pick a model too big for the box and it crawls or crashes. We match the plan to the models you want to run — starting around 4 GB for smaller models and scaling up for larger ones — so it’s usable, not just installed.
- Model storage adds up fast. Open models are multi-gigabyte each, and teams collect them. We provision and monitor the disk so pulling a new model doesn’t fill the server.
- Keeping the API off the public internet. Ollama’s API has no authentication by default — exposed, it’s an open invitation. We put Open WebUI’s login in front and firewall the raw API so only your people reach it.
- Updates without downtime. Ollama and Open WebUI move quickly. We handle updates and keep the service supervised so it restarts cleanly instead of dying quietly.
What’s included
- One-click Ollama + Open WebUI deploy
- Plan sized to the models you’ll run
- Authenticated web UI; raw API firewalled off
- Model storage provisioned & monitored
- Managed updates with auto-restart
- TLS, custom domain & human support
- Free migration from your current host
What it runs on
Ollama and Open WebUI each want at least 4 GB of RAM; larger models (for example DeepSeek-R1 in our catalog) need 8 GB or more. We recommend a plan based on the models you intend to run — ask us about GPU options if you need faster inference. Compare plans & pricing →
Frequently asked questions
Which models can I run?
Any model Ollama supports — Llama, Mistral, DeepSeek, Gemma and more — within your plan’s memory. Smaller models run comfortably from 4 GB; larger ones need more RAM. We help you pick models that fit.
Do my prompts or data ever leave the server?
No. Inference runs on your own instance, so prompts and responses stay on infrastructure you control — nothing is sent to a third-party AI provider.
Do I need a GPU?
Smaller models run on CPU; larger models are much faster on a GPU. Talk to us about what you want to run and we’ll advise on the right setup.
Can my whole team use it?
Yes — Open WebUI provides multi-user accounts, chat history and document chat, behind a login, at your own domain over HTTPS.
Is the AI endpoint secured?
Yes. The web UI sits behind authentication and the raw Ollama API is firewalled so it isn’t exposed to the internet — the default mistake we specifically prevent.
Related: browse all 74 one-click apps · managed n8n · talk to an engineer
Give your team private AI you actually control.
Deploy Ollama + Open WebUI on a managed, right-sized, locked-down server.
Deploy private AI →