HolyClaude Hits 2.4k Stars, and Now I’m Hosting It
Running LLM agents locally works well until you hand them a genuinely complicated task. I created HolyClaude as a self-hosted web UI for Claude because the available tools did not match my workflow. Although it has gained traction on GitHub, I eventually ran into a limitation in my own project.
The issue is persistence. If you start a large refactor across forty files or a lengthy test suite, the sensible response is to step away from your desk. Closing the laptop, however, ends the process. Adjusting sleep settings does not solve the underlying problem: the work remains tied to the hardware in your bag. There is no convenient way to launch a demanding agent task and inspect the outcome from your phone an hour later.
I saw users working around this by deploying HolyClaude on shared VPS setups, using Tailscale for access control and separate volumes per developer. It is a workable approach, but patches, backups, and disk space still require someone’s attention. The real opportunity was not only the software itself, but the infrastructure that keeps it running 24/7.
The architecture of the hosted version
Instead of using one shared environment, I’m moving toward a dedicated Linux box for each user. The setup works as follows:
- Infrastructure: Each user receives one Firecracker VM and one persistent volume running on Fly.io.
- Pre-installed Tooling: To avoid spending an afternoon running
apt-getinstalls, every box includes six agent CLIs—Claude Code, Codex, OpenCode, Gemini, Cursor, and Pi—along with roughly 60 general development tools. - Access: Root access, SSH, and a browser-based terminal allow you to start a session on a desktop and continue it from a mobile device without losing your place.
- Privacy: You use your own API key. Because I neither proxy the calls nor mark up the tokens, I have zero visibility into your data.
cloudcli, but I do not own that codebase. Adding multi-user login would therefore have meant maintaining a permanent fork. The hosted version instead uses a UI I built from scratch with SSO integrated, allowing some features to appear there first.
Hard lessons in infra deployment
Moving from an MIT-licensed project to a hosted service showed me that infrastructure and software require completely different approaches to sales.
There is no “free tier” growth hack when each user receives a real VM costing ~$13/month regardless of activity. The cost model is also unstable. The VM creates a fixed expense, model tokens create a variable expense, and retries introduce a “hidden” expense. An agent stuck in an infinite loop overnight can consume more money than the machine itself costs to run.
I also encountered specific deployment issues on Fly.io. I attempted to ship a 13.6GB image, and the flyctl command returned an exit code 0, indicating success. Ninety-five seconds later, however, the machine had quietly rolled back to the previous image because Fly rejects images over 8GB. I found the cause only by examining the machine’s event list. A successful API response does not necessarily mean your code is actually running.
All Replies (3)
Want a live back-and-forth? Join the global AI chat room — login to talk.
I need custom personas for this. Does it actually support system prompts yet?
Long responses keep crashing my build. Which API timeout setting worked best for you?
Local hosting is so much faster. Is the latency actually lower on your hardware?