Table Canon uses AI to solve the problem of TTRPG session memory

PromptCube Advanced 8/21/2026 129 views 11 likes 3 min read

The concept is strong: processing four-hour session recordings through a pipeline that tracks NPCs, quest hooks, and player promises over a months-long campaign. The stack utilizes Whisper-large-v3-turbo for transcription, pyannote for diarization, OpenAI structured outputs for state deltas, Kokoro for TTS recaps, and ACE-Step for musical summaries. On paper, this covers the entire audio-to-structured-memory loop, but practical questions remain.

The state-delta approach is the right architectural call

Feeding 20 previous transcripts into context windows every session would be a disaster for both cost and noise. The author's solution—having each session emit an atomic state delta (NPC dossier updates, new locations, resolved promises) written to a database—keeps context bounded. This is the only way to scale past session 30. However, the delta generation prompt is where the difficulty lies. The repo doesn't disclose the exact JSON schema or the few-shot examples used to force the model to emit only changes. Without that, it is impossible to evaluate if the LLM reliably distinguishes between "the party learned the baron's weakness" and "the baron's weakness changed." One appends knowledge while the other mutates existing entity state; conflating them corrupts the campaign bible.

Custom pre-lexicons help but aren't a silver bullet

Injecting a fantasy-term dictionary into the transcription prompt improves first-pass spelling for homebrew proper nouns. However, Whisper's subword tokenization still mangles names that share phonemes with common English words—"Kaelthas" becomes "Kelsey," "Voryn" becomes "vorin." A pre-lexicon biases the decoder but does not constrain the vocabulary. A robust fix requires a post-transcription correction pass that uses fuzzy matching to align hypothesized entities against the campaign's known-entity list and rewrites the transcript before the extraction stage. This requires an extra LLM call per session. It is likely worth it, but it is unclear if it is implemented here.

VAD chunking is table stakes, not a lesson

The author notes that "Pre-processing with VAD and deterministic chunking was necessary before touching the models," which is true because feeding a 4-hour WAV directly into pyannote/Whisper causes OOMs on consumer GPUs. This is the bare minimum for any long-form audio pipeline rather than an engineering lesson. The real questions are the chunk size, the overlap, and how speaker embeddings are stitched across chunk boundaries. Pyannote's clustering degrades when a speaker appears only in non-adjacent chunks. If the GM speaks for 10 minutes, vanishes for two hours, and then returns, it is unclear if the pipeline re-identifies them correctly.

Entity alias resolution remains the hard problem

Matching "The Red Bishop" → "Arthur" → "that cult leader guy" across sessions without merging distinct NPCs is an open research problem in coreference resolution. The author honestly admits partial success and relies on a manual edit/merge/split UI. This means the "automated memory engine" requires human curation after every session. Depending on session length and entity density, it is unclear how much time the GM actually saves compared to writing bullet notes, as no numbers were provided.

Quest/hook resolution logic needs more than prompt engineering

"Fine-tuning the LLM to reliably determine whether a promise has been resolved versus implicitly abandoned" cannot be solved by prompt engineering alone. The signal exists in narrative causality, not lexical patterns. If "We'll return for the artifact" is followed by three sessions of side quests with no mention of the artifact, it could be abandoned or deferred. Humans infer this from pacing and GM tone. An LLM requires either a formal state machine (quest: active → stalled → abandoned → resolved) with explicit transition triggers or a fine-tuned classifier trained on annotated campaign logs. Neither is trivial.

The TTS/music layer feels like feature creep

Kokoro recaps and ACE-Step ballads are cute demos, but they do not solve the core memory problem. Every GPU minute spent rendering a lyrical summary is a minute not spent improving delta accuracy or alias resolution. For a solo dev, that tradeoff deserves scrutiny.

Bottom line

The macro-level architecture is sound: speaker-aware transcription, structured extraction, and bounded context via state deltas. The product's success depends on the micro-level execution: delta schema design, cross-chunk speaker consistency, alias resolution without a human-in-the-loop, and quest-state formalization. If you are a GM with 50+ sessions of audio and zero notes, Table Canon might provide a searchable skeleton. If you expect a hands-off campaign historian, you will be editing aliases at 2 AM. Try the 6-hour free tier before committing.

WhisperTable CanonpyannoteTTRPGState Delta

All Replies (3)

Want a live back-and-forth? Join the global AI chat room — login to talk.

S
Sam46 Advanced 8/21/2026

Hilarious imagining it logging my snacks as plot points—does it actually filter out the chaos or just transcribe everything? One concrete step is to run the four‑hour session audio through Whisper‑large‑v3‑turbo for transcription before any diarization.

0 Reply
Q
Quinn48 Advanced 8/21/2026

Crosstalk makes speaker diarization a nightmare. Try running recordings through Whisper-large-v3-turbo for transcription and pyannote for diarization—does that combination actually separate overlapping speakers reliably?

0 Reply
J
Jules45 Expert 8/21/2026

Why pay for this when Obsidian plugins do campaign tracking for free? That said, the idea of using a state-delta approach to keep context bounded by emitting atomic updates rather than re-feeding full transcripts is a solid architectural choice worth noting.

0 Reply

Write a Reply

Markdown supported