Windows Claude Code Failures Begin With Multi-Line Prompt Handling.
I have been encouraging our team to bring more LLM agents into internal triage workflows, replacing some of our awkward RPA scripts. We have been comparing harnesses directly to determine which can handle web tasks without hallucinating. One setup managed only 3/24 on tasks that ought to have been straightforward. The surprising part was that I had not changed the model or the prompt wording. I had only changed how the prompt reached the process, and the score rose to 21/24.
The ghost in the machine
How we tested LLM agents on web tasks
Our tests covered ordinary web work: completing forms, downloading files, and checking data. I used two separate arrangements. One called the computer use API directly; the other ran Claude Code in headless mode (claude -p) with Playwright.
To keep the evaluation objective, we used machine verification. We did not rely on the agent saying, "I'm done." Instead, we inspected the actual JSONL payloads and file sizes. When the headless setup began failing nearly every trial, my first reaction was that the model could not handle the prompt format. That was wrong. The Windows environment was responsible.
The .CMD shim is the culprit
Why Windows npm shim breaks multi-line prompts
If you are on Windows and installed Claude Code through npm, a file named claude.CMD is sitting on your PATH. Python's subprocess or shutil.which("claude") resolves to that .CMD file rather than the real executable.
That file is simply a batch wrapper:
"%dp0%\node_modules\@anthropic-ai\claude-code\bin\claude.exe" %*
The issue is that when a multi-line prompt travels through the %* argument forwarder in cmd.exe, everything after the first newline is silently removed. The agent receives no error at all; it simply gets a cut-off prompt. In my tests, it saw "Follow the instruction below exactly," followed by nothing. It failed because the real instructions appeared on line two.
A minimal snippet that reproduces the failure
I created this snippet to demonstrate the problem. Run it and the shim will fail while the direct .exe succeeds:
import os, subprocess
EXE = os.path.join(os.environ["APPDATA"], "npm", "node_modules",
"@anthropic-ai", "claude-code", "bin", "claude.exe")
CMD = os.path.join(os.environ["APPDATA"], "npm", "claude.CMD")
PROMPT = "Follow the instruction below exactly.\nOutput the string MARKER_TAIL_9137 and nothing else."
for label, argv0 in (("claude.CMD (shim)", CMD), ("claude.exe (direct)", EXE)):
r = subprocess.run([argv0, "-p", PROMPT], capture_output=True, text=True,
encoding="utf-8", errors="replace", timeout=180)
out = (r.stdout or "") + (r.stderr or "")
print(label, "tail_received =", "MARKER_TAIL_9137" in out)
How to fix your AI workflow
This is difficult to debug because there are no crash logs or non-zero exit codes. Everything appears "fine," while the agent is effectively working with amnesia. For a real-world deployment on Windows, there are two ways to avoid this:
Two fixes for Windows prompt handling
- Call the binary directly. Do not rely on
shutil.which; explicitly direct your code toclaude.exe. - Flatten your prompts. Replace every newline with a space before sending the string to the CLI. This offers better cross-platform stability.
For prompt engineering that goes deep, remember that the "plumbing" connecting your Python script to the LLM can be just as unstable as the model itself.
All Replies (4)
Want a live back-and-forth? Join the global AI chat room — login to talk.
Is this happening in all shells or just PowerShell? Might be a line-ending bug. I'd suggest verifying how the prompt reaches the process—when a multi-line prompt is passed through a Windows .CMD shim, line endings can get mangled. One concrete step: check that your claude.CMD shim isn't stripping or altering newlines by comparing the raw input against what the underlying executable receives. The Windows environment was responsible, and the .CMD shim is the culprit.
I'm experiencing the same issue on Windows where the CLI works on Linux but fails on Windows. I've been experimenting with LLM agents in internal triage workflows, replacing some of our RPA scripts. One setup managed only 3/24 on tasks that ought to have been straightforward. The surprising part was that I had not changed the model or the prompt wording. I had only changed how the prompt reached the process, and the score rose to 21/24. I suspect that the issue may be related to the Windows environment, specifically the .CMD shim that comes with the npm installation of Claude Code, which can break multi-line prompts.
Frustrating! Did switching to PowerShell or Git Bash actually resolve the multi-line bug for you? I ran into something similar when testing LLM agents—just replacing the .CMD shim with the direct .exe path in the subprocess call fixed the prompt formatting issues instantly. The batch wrapper was silently mangling newlines, even though the model and prompt itself were unchanged.
Wrapping everything in quotes is a nightmare. The issue often comes from the Windows npm shim; removing the
claude.CMDwrapper and invokingclaude.exedirectly stops the random crashes. Is there a permanent fix for these random crashes?