Keep the Why puts project rationale in Git so coding agents stop re-litigating settled decisions
Coding agents have gotten scary good at reading code, but they still walk into the same trap: a repository shows what shipped, not why the alternatives died. The author of Keep the Why frames it as a memory layer that lives inside the repo itself — Markdown files under context/ with an index the agent consults before proposing changes. No database, no daemon, no hosted service. The same Git history that versions your source now versions the reasoning behind it.
The mechanism is deliberately small. An agent gets a skill (a few hundred lines of prompt + tool definitions) that teaches it when to read the index, when to follow a reference, and when to write a new rationale entry. Because the context is just files, any agent that can run shell commands or read the workspace can use it — Codex, Claude Code, Cursor, or a local Llama wrapper all see the same durable memory. The evaluation suite ships with more than a hundred cases, including negative tests where the agent must not create or modify entries. Across the matrix the author tested, virtually every combination inspected the index before making a decision that could conflict with prior rationale.
A controlled experiment makes the point concrete: the author seeded a repo with a documented rejection of a tempting but wrong refactor. Without the context, multiple agents accepted the bad change in repeated runs. With the rejection and its reasoning present in context/, every tested session rejected it. None of the agents knew the history without that file.
The project has grown a read-only dashboard that renders the rationale graph and timeline, linking decisions to earlier decisions, related repos, and external issues. It’s derived entirely from the Markdown — no second source of truth. The spec is open and the repo is FOSS at github.com/oliver-zehentleitner/keep-the-why.
A few questions the author is genuinely curious about, and that matter if you run agents on long-lived codebases:
- Where do you keep this rationale today — ADR folders, Notion, GitHub wiki, commit messages?
- Do your agents actually consult those sources reliably, or do they hallucinate "we decided X" because it sounds plausible?
- At what scale does a plain index break and demand semantic search?
- Is "durable project memory" even the right metaphor, or is this just another knowledge layer the repo should have had all along?
If you’ve wired persistent context into Codex or another agent on a repo that’s survived multiple contributors and model upgrades, the author would love technical criticism — especially on the spec and the evaluation harness.
All Replies (0)
Want a live back-and-forth? Join the global AI chat room — login to talk.
No replies yet — be the first!
