CyberStrike drops an AGPL harness for AI-driven red-teaming
What makes it different from the usual "ask GPT for a payload" scripts is the harness architecture. You don't just feed it a target and hope. The core loop runs: recon → attack graph generation → tool orchestration → evidence collection → report. Each phase is a pluggable module with a defined schema, so you can swap the LLM backend (local Llama-3-70B, Claude, GPT-4o, whatever) without rewriting your exploit logic.
Key pieces worth knowing
- Attack graph DSL — YAML-based, describes multi-step chains like "enumerate SMB → extract hashes → pass-the-hash → dump LSASS". The LLM expands high-level goals ("get domain admin") into concrete graphs at runtime.
- Tool adapters — First-class wrappers for nmap, bloodhound, crackmapexec, impacket, metasploit modules, and custom binaries. Adapters expose typed inputs/outputs so the planner can chain them reliably.
- Memory layer — SQLite-backed context store persists findings across runs. You can pause a campaign, switch models, resume — the graph state survives.
- Safety rails — Scope enforcement via CIDR/target allowlists, rate limiting per adapter, and a mandatory "dry-run" mode that logs planned actions without executing. The AGPL means any SaaS wrapper must expose these controls.
Getting a local instance running
git clone https://github.com/cyberstrike/cyberstrike.git
cd cyberstrike
pip install -e .[local-llm] # pulls llama-cpp-python, FAISS, etc.
cp config.example.yaml config.yaml
# edit config.yaml — set your target scope, LLM endpoint, adapter paths
cyberstrike init --workspace ./my-campaign
cyberstrike run --goal "achieve domain admin" --dry-runThe dry-run output shows the generated attack graph with confidence scores per node. Once you're comfortable, drop --dry-run and it starts executing against the scope.
Where it shines and where it doesn't
- Strengths: Handles multi-step logic that single-shot prompts butcher. The evidence collector auto-correlates logs, pcaps, and tool output into a timeline — huge for reporting. Local model support means air-gapped environments work.
- Gaps: BloodHound adapter only ingests JSON, doesn't drive the GUI. No built-in C2 framework integration yet (Cobalt Strike, Sliver, Havoc are on the roadmap). LLM hallucination on obscure protocol edges still happens — always verify before firing.
Licensing catch
AGPL-3.0 triggers if you expose CyberStrike as a network service. If you're building a commercial pentest platform on top, you must open-source your modifications. Several vendors have already reached out about dual-licensing; the maintainers seem open but haven't announced anything.
Worth cloning if you run internal red-team exercises and want reproducible, auditable AI assistance. The codebase is clean — typed Python 3.11+, decent test coverage, and the module interfaces are stable enough to build custom adapters without fighting the core.