HolyClaude Hits 2.4k Stars, and Now I’m Hosting It

HyperNinja Intermediate 8/13/2026 210 views 3 likes 2 min read

Running LLM agents locally works well until you hand them a genuinely complicated task. I created HolyClaude as a self-hosted web UI for Claude because the available tools did not match my workflow. Although it has gained traction on GitHub, I eventually ran into a limitation in my own project.

The issue is persistence. If you start a large refactor across forty files or a lengthy test suite, the sensible response is to step away from your desk. Closing the laptop, however, ends the process. Adjusting sleep settings does not solve the underlying problem: the work remains tied to the hardware in your bag. There is no convenient way to launch a demanding agent task and inspect the outcome from your phone an hour later.

I saw users working around this by deploying HolyClaude on shared VPS setups, using Tailscale for access control and separate volumes per developer. It is a workable approach, but patches, backups, and disk space still require someone’s attention. The real opportunity was not only the software itself, but the infrastructure that keeps it running 24/7.

The architecture of the hosted version

Instead of using one shared environment, I’m moving toward a dedicated Linux box for each user. The setup works as follows:

  • Infrastructure: Each user receives one Firecracker VM and one persistent volume running on Fly.io.
  • Pre-installed Tooling: To avoid spending an afternoon running apt-get installs, every box includes six agent CLIs—Claude Code, Codex, OpenCode, Gemini, Cursor, and Pi—along with roughly 60 general development tools.
  • Access: Root access, SSH, and a browser-based terminal allow you to start a session on a desktop and continue it from a mobile device without losing your place.
  • Privacy: You use your own API key. Because I neither proxy the calls nor mark up the tokens, I have zero visibility into your data.
The UI represented a major architectural change. HolyClaude originally packaged cloudcli, but I do not own that codebase. Adding multi-user login would therefore have meant maintaining a permanent fork. The hosted version instead uses a UI I built from scratch with SSO integrated, allowing some features to appear there first.

Hard lessons in infra deployment

Moving from an MIT-licensed project to a hosted service showed me that infrastructure and software require completely different approaches to sales.

There is no “free tier” growth hack when each user receives a real VM costing ~$13/month regardless of activity. The cost model is also unstable. The VM creates a fixed expense, model tokens create a variable expense, and retries introduce a “hidden” expense. An agent stuck in an infinite loop overnight can consume more money than the machine itself costs to run.

I also encountered specific deployment issues on Fly.io. I attempted to ship a 13.6GB image, and the flyctl command returned an exit code 0, indicating success. Ninety-five seconds later, however, the machine had quietly rolled back to the previous image because Fly rejects images over 8GB. I found the cause only by examining the machine’s event list. A successful API response does not necessarily mean your code is actually running.

AI ProgrammingAI Coding

All Replies (3)

Want a live back-and-forth? Join the global AI chat room — login to talk.

D
Drew36 Advanced 8/13/2026

Local hosting is so much faster. Is the latency actually lower on your hardware?

0 Reply
R
Riley82 Advanced 8/13/2026

I need custom personas for this. Does it actually support system prompts yet?

0 Reply
P
PatFounder Advanced 8/13/2026

Long responses keep crashing my build. Which API timeout setting worked best for you?

0 Reply

Write a Reply

Markdown supported