18trees Character Distillation
Distils a traceable, updatable model of a character's behavior and cognition from a novel or screenplay, with a five-layer evidence contract and offline checks that tie every conclusion back to the source text.
On this page[4]
What it is
18trees-Character-Distillation is a workflow for distilling a character model out of an original work. Give the assistant the path to a novel or screenplay you legally hold, along with the character’s name, and it outputs a character package, by default under workspace/<character-id>/character-os/.
It addresses one problem: most character work produces a prompt, a card, or a dialogue, and whether it is good depends on whether it reads right. This project ties every conclusion back to the source text — what he saw, what information he had, how he weighed it, what action he took, and which passages support or contradict that reading — with each quotation checkable word for word.
It is aimed at creators and developers. A reader meets the character through character-report.md, and a downstream agent follows agent-guide.md to choose snapshots and trace evidence back. There is no chat interface, companionship, voice, or fine-tuning, and web-based chatbots cannot use it: they cannot read a local source text, run the checks, or write the result to disk as a character package.
How it is structured
The project runs on two tracks: the left column is the workflow that turns materials into a package, and the right column is the contracts and tools that constrain and check it.
Inputs. Real materials, task state, and results sit in one private per-character directory: workspace/<character-id>/ holds inputs/, work/, and character-os/, while external materials can be read in place. Neither workspace/ nor .local/ enters Git.
Contract layer. The 15 JSON Schemas in schemas/ split evidence into five layers — source, evidence, experience and behavior, claim, snapshot — and prompts/ holds versioned prompts, each registering its version and hash in registry.json. docs/ and research/ carry the methodology, the SOP, the orchestration convention, and the research comparison.
Workflow. The Host writes the whole TASK.md in one pass, then starts 1–3 coarse-grained primary sessions along work or narrative boundaries. After the sessions produce candidates, two sequential reviews merge and correct them directly, and the final review agent alone writes into character-os/. The Host does not read the source text.
Check layer. scripts/character_os.py offers four commands — validate, report, check-delivery, and check-public — all offline, turning the structural, reference, and time rules into mechanical checks.
Character package. The package contains human entry points (README.md, character-report.md, and agent-guide.md, plus an optional audit-report.md) alongside data files organized by layer; the CLI-generated schema-report.md is a machine data view, not the character report.
Design decisions worth noting
Quotations must match the source verbatim. validate recomputes the source file’s SHA-256, checks the quoted span, and compares the quotation against the source slice character by character. A line written from impression or a misremembered quotation is not a blemish — it is a mechanical error.
Time and what could be known are part of the contract. A claim may not cite evidence that appeared after it, known_from may not predate its supporting evidence, and a snapshot may only use knowledge available before its cutoff — passing off hindsight from training data as canon is rejected.
Counter-evidence, “unknown”, and review each have a fixed place. Supporting evidence and counter-evidence are stored separately, high or strong inferences require two independent scenes, an accepted claim must keep a review record, and psychology with no textual basis stays unknown.
The checks do not depend on a model. validate, report, check-delivery, and check-public all run offline and need no model API key; changing a prompt requires bumping its version and updating the registry hash, and a run record keeps the hashes of its inputs and prompts.
The public boundary is executable. check-public mechanically checks, against an allowlist, which files may go public: real materials, character packages, and .local/ caches must never enter the public candidates, and negative tests cover that boundary.
What problems it solves
- A quotation that does not match the source, an out-of-range span, or a changed source file fails validation instead of waiting for someone to notice.
- Hindsight passed off as canon, collecting only the examples that support you, and treating a single instance as a conclusion are all blocked by the contract.
- Real materials, character packages, and caches have a mechanical public boundary, with tests that keep them out of public candidates.
- The whole check chain needs no model API key, so CI re-runs it on every push and pull request.
- The deliverable has stable entry points: a person meets the character through the report, and a downstream agent gets its reading order and trace-back rules from
agent-guide.md.