Thespis Cast: NPCs That Can Be Wrong

Introduction
Most LLM NPCs invent what they “remember.” Thespis Cast — shipped for the Cambridge University × Arcade AI Hackathon (3–4 Oct 2026) by @coderback and @AhmedBokaniUsman — flips that: the game owns ground truth in an append-only ledger, each NPC builds confidence-weighted beliefs that can be false, and a language model may only pick among code-scored actions and speak lines the validator can check. Open MIT code, live demo, measured harness.
- Modality: game NPC minds / agentic dialogue toolkit (not a world-model video stack)
- Open vs closed: MIT open-source core + Crypt Road adapter; optional OpenAI-compatible LLM (Azure works); runs with no model via code choices + template lines
- Where to run today: play The Crypt Road (Watch the 60-second story or drive it yourself), or clone coderback/Thespis and
uvicorn games.crypt_road.app:app - One hard number: on the hosted engine (2026-10-04 02:17 UTC harness), 87% of NPC turns (395/452) were decided by code alone — no model call
Primary artifact: repo created 2026-10-03 with MIT README + results landed 2026-10-04; health endpoint returns {"ok":true}.
What shipped
The Crypt Road is a five-stop race to a relic. Insult rival Kael, duel him, bribe Captain Brenna, frame a lie — NPCs gossip offscreen, update beliefs, and block or clear the gate. Every spoken line in the client opens a why-chain: decision → cited beliefs → ledger event (true or false).
Architecture in one pass:
- Ledger (
thespis/ledger.py) — append-only events; nothing rewrites history - Beliefs — per-NPC claims with source + confidence; truth is derived from the ledger, so testimony can retract a lie
- Drives / utility brain — scores allowed actions from trust and goals; the model may only choose among actions within ~2 points of the best
- State pack + validator (
thespis/expression.py) — model never sees ground truth; replies must cite pack ids, stay 1–160 chars, and name only known characters or they fall back to code - Gateway — any OpenAI-compatible endpoint (primary → backup → fallback, 4s each)
Core (thespis/) imports no game; games/crypt_road/ is one adapter. Tests fail if that boundary breaks.
Numbers that matter for makers
From results.md against the live host with GPT-6 Luna:
MeasureValue
NPC turns with no model call
87% (395 / 452)
Invalid model replies blocked
0 of 80
Model latency p50 / p95
1315 ms / 1585 ms
/act with model p50 / p95
1445 ms / 2173 ms
/act without model p50 / p95
43 ms / 79 ms
Cost per run (every model call priced)
$0.00066
Cost as played (with cache)
$0.00030
Rules routes matching the rules model
8 of 8
That split is the point: drives enforce grudges; the LLM settles near-ties and voices lines you can audit.
Run it tonight
git clone https://github.com/coderback/Thespis.git
cd Thespis
python -m venv .venv && source .venv/bin/activate
pip install -r requirements-dev.txt
pytest
cd client && npm ci && npm run build && cd ..
uvicorn games.crypt_road.app:app --reloadOpen http://localhost:8000. No .env required for a full playable loop. To hear the model, copy .env.example → .env with an OpenAI-compatible base URL, key, and model (see docs/models.md).
Why it fits game-gen
This sits next to ledger-and-permission agent stacks and older belief sims (Talk of the Town, Versu, Generative Agents): the claim is not “NPCs can be wrong” — it is LLM voice under a deny-first action set with cite-or-reject lines. For teams building gossip, factions, or offscreen consequence, Thespis Cast is a same-weekend, click-to-play reference — not a paper stub.
Conclusion
Thespis Cast is a clock-hot NPC toolkit: MIT code, Railway demo, validator-checked dialogue, and an 87% code-alone decision rate on the published harness. Clone the repo or hit the Crypt Road demo and inspect a why-chain before you wire your own ledger.
—Titus
