Inference on the box you administer
These notes record the stack built for Integrity Care Resources, a 501(c)(3) that trains people for the trades and for veteran re-entry. The model runs on a Hetzner Ubuntu VPS. Prompts stay on that machine. They are not sent to a public chat API, and this is not a guide to stripping refusals out of a model.
The public internet only needs the nonprofit website. Cloudflare proxies the domain and terminates TLS. Nginx serves static HTML from /var/www/integritycare with root and try_files. The language model is not the homepage. If Ollama or Uvicorn stops, the job-training pages should still return 200.
Inference sits behind that split. FastAPI on 127.0.0.1:8000 is the only application that should call Ollama on 127.0.0.1:11434. The weights we ran were dolphin-mistral:7b. Admin SSH goes over Tailscale to the overlay hostname milleniumfalcon, not to port 22 on the public address. Fail2ban stays on for the VPS. HTML, the Python app, and Nginx config live in three different places: /var/www/integritycare, /root, and /etc/nginx/sites-available.
Two mistakes produced the outages. Pasting the website into a file under sites-available, or pointing location / at proxy_pass http://127.0.0.1:8000 while Uvicorn was down, turned the whole origin into 502 Bad Gateway. proxy_pass belongs on /chat and /api only. Before any tokens are generated, the API picks a system prompt from a rating. G refuses violence, sexual content, illegal instructions, and harm. PG refuses explicit sexual content, graphic violence, and crime how-tos. R still refuses extreme gore, non-consensual content, and real-world crime instructions. Client conversations are not training data, prompt logs are not world-readable, and secrets do not go in git.