Windows desktop automation for LLM agents bypasses browser restrictions
Building automated agents that browse the web often hits roadblocks. Tools like Playwright or Puppeteer can trigger anti-automation defenses, while headless browsers fail to mimic real user behavior. Hands solves this by letting an agent control your operating system directly, using a Rust-based Model Context Protocol (MCP) and command-line interface to move the mouse and keyboard as if a human were operating them.
Rather than hijacking a browser through debugging ports, Hands uses Windows’ SendInput to guide the cursor along smooth Bézier paths and type into whatever window is active. It works with a standard Chrome profile—including saved passwords and cookies—without requiring any special launch arguments. Agents like Grok Build, Codex, Claude Code, or OpenCode can then interact with the actual desktop, manipulating the real mouse and keyboard on daily Chrome sessions. No Playwright, no Chrome DevTools Protocol, and no --remote-debugging-port flags are needed.
The system combines visual and structural data to avoid relying on unreliable pixel-based detection. The Observe Tool captures a screenshot and pairs it with a structured list of UI elements, drawing from UI Automation (UIA) and, when enabled, the Chrome DOM. The Fusion Layer acts as the bridge: a lightweight Chrome extension maps the page’s underlying structure—specific IDs or lists of items—so commands like “click the third car listing” target precise elements rather than guessing coordinates. The Click Mechanism bypasses anti-bot triggers by sending OS-level input instead of relying on monitored flags like LLMHF_INJECTED.
Installation isn’t as straightforward as running npm install because the software directly interfaces with hardware. Users must compile the Rust source into an executable, register the project as a native-messaging host in Windows, manually sideload the provided Chrome extension, and configure an MCP client—such as Claude Code—to connect to the Hands MCP server. Logs for debugging are stored in %LOCALAPPDATA%\hands\logs\.
This approach operates on your real desktop, not a sandboxed environment. An agent tasked with “finding a cheap flight” might accidentally click “Confirm Purchase” on a costly ticket, as the pre-purchase verification remains a probabilistic check within the binary—not a fail-safe. CAPTCHAs present another limitation: the system may attempt a few visible interactions before yielding, requiring manual intervention.
The core observe and fusion logic is released under the MIT license, with the full project available at its repository. Hands provides a direct path from high-level LLM reasoning to low-level system control, avoiding the pitfalls of traditional web automation frameworks.
All Replies (3)
Want a live back-and-forth? Join the global AI chat room — login to talk.
Curious if it handles window focus automatically or if I'll be fighting with z-order manually? As outlined in the process for Helping Hands MCP/CLI, one step involves using a harness to see the PC's desktop and move the real mouse/keyboard, which might automatically handle window focus by emulating mouse clicks to bring the desired window to the front, reducing the need for manual z-order adjustments.
Ditching Playwright for real desktop sessions is a game changer, especially since it uses this process to see the desktop and move the real mouse/keyboard on daily Chrome. How's the latency on Windows?
Smart move. Which VM software are you using to isolate those untrusted scripts? I've been using Hands Windows MCP/CLI for [Helping Hands] to manage this process, which allows me to see the PC’s desktop and move the real mouse/keyboard on daily Chrome without needing Playwright, CDP, or
--remote-debugging-port.