BL-122 Ideas

Design a live game-running surface for the DM, native to the Claude Code CLI

Came out of the Session 40 retro (docs/session-running-improvements.md) and a follow-up design conversation. Two constraints, both explicit from Jay:

  1. Interaction model is not 1:1 prompt/response. Tonight’s live session ran as an ordinary chat turn-loop, where every DM message was treated as a complete, immediate directive requiring a reply right then. Jay’s own description: “something that me as a DM can type my notes into but doesn’t necessarily mean its a 1:1 prompt/response” – he wants to be able to drop fragments, corrections, and scene-setting notes as they occur to him, with the assistant holding/synthesizing accumulated state rather than reacting to every clause individually. This is the same root cause as most of Session 40’s ~20 corrections (see the retro’s “verify-before-acting” throughline): a chat turn-loop under live-table pressure produces exactly the jump-ahead, talk-past-each-other failures observed that night.

  2. Must run through the Claude Code CLI session interface itself, not a separate web surface. Authenticating against Anthropic via the DM’s own local Claude Code login, not a TxTavern-issued key. This is a direct extension of TxTavern’s core BYOAI pillar (brand-voice.md messaging pillar #2, no built-in chatbot, DM brings their own trusted AI via MCP): the live-running surface should be part of “bring your own AI,” not a parallel in-app AI feature. The DM’s existing MCP connection to TxTavern (already used tonight for Moments, Stages, Articles, etc.) is the data layer; this item is about redesigning the interaction surface on the Claude Code side, not adding new TxTavern app UI.

Design conversation, 2026-09-07, three pieces, recorded in full in docs/session-running-improvements.md’s “Design note” section:

  1. Staged notes, explicit synthesis. Extend the existing live-session-mode hook so an ambiguous DM fragment is appended to a scratch file and briefly acknowledged, not acted on, until an explicit trigger word asks for synthesis. The same “explicit approval word, never inferred” discipline ADR-0102 already proved holds, applied one step earlier.

  2. A multi-session browser dashboard, TxTavern-specific. Verified directly: claude agents --json already lists every active local session with status; claude attach works via a private per-session pseudo-terminal socket the Claude Code daemon manages. The sound way to expose that in a browser is to wrap claude as a real child process with our own pty (Ruby’s PTY.spawn) and stream it to the browser via xterm.js over ActionCable, not reverse-engineer the private daemon socket. Recommended to live inside the Rails app (Devise-authenticated) rather than as a separate standalone tool, given the scope call below.

  3. Referee and thread context isolation. Motivated by a concrete scenario: the players’ ship approached by a threat, an all-NPC pirate ship elsewhere, and no structural reason the pirates’ dialogue couldn’t leak knowledge only the DM-facing session should have. Confirmed not a rare case: party splits, a third party arriving mid-scene, and flash-overs to an unrelated simultaneous thread (all three seen in actual play at this table) are the same underlying need, to spin up a separately-scoped session for a concurrently-running narrative strand, on demand. One Referee session holds full ground truth and is the sole writer to production; thread sessions are scoped-context and narration-only, with cross-thread information moving only by the DM manually relaying a filtered outcome between panes, a hard context wall standing in for a discipline rule that erodes under live pressure. Reuses existing machinery rather than inventing new: the session lore-cache pattern (ADR-0091) for per-thread context scoping, and the existing Calendar System MCP tools for game-clock bookkeeping across threads.

Explicit scope call: TxTavern-specific, not a general cross-project session-orchestration tool, even though the dashboard’s underlying mechanism would generalise.

Open questions this design still needs to resolve:

  • Staged-notes trigger vocabulary needs live-table testing.
  • Rails + ActionCable + PTY.spawn is a recommendation, not yet built or tested end to end.
  • Per-thread lore-cache scoping granularity (per-Stage vs. explicit tag set) is undecided.
  • .claude/agents/archivist.md is stale (still documents the superseded 4-option draft default), a prerequisite fix regardless of build order.
  • No build order committed. Likely candidate: prototype the dashboard’s session-launch/pty plumbing first (useful even before thread isolation exists), then layer thread spin-up on the same plumbing.

Related: ADR-0102 (single-draft approval), ADR-0058 (session-running discipline, out-of-band delivery), docs/session-running-improvements.md (Session 40 retro and this design note).