Research synthesis that shows what survived.
Turns a folder of interview transcripts into findings, then tries to knock them down — and shows you what survived. Fifteen transcripts, 101,000 words: thirty minutes, $4.53, twenty insights. Every insight carries its receipts, its confidence, its counter-evidence and the critic's outstanding objections. Read the case study for how it was built and measured.
Turns a folder of interview transcripts into findings, then tries to knock them down — and shows you what survived.
Teams and freelancers who synthesise in FigJam or Miro with transcripts in a folder — no research platform, just stickies — and who have been burned by a confident AI summary.
What it does
Synthesis produces eight to fourteen insights from your transcripts. A critic with readable rules checks every one against the exact text it cites, reports what is unsupported, missing or overconfident, and a reviser repairs. The loop stops on a pass, on no progress, or after three rounds — and the report says which. Every insight carries turn-numbered receipts, a computed confidence, its counter-evidence, and any objection the critic still had when the loop stopped.
The loop
Anything that can be checked mechanically is checked before the model is asked. The model is asked only what only a reader can answer. A missing verdict is a failure, never a pass.
Install and run
A Python environment, an Anthropic API key, one command. Rules and thresholds live in a config file in plain language.
The eval
Three conditions on the same five transcripts, three runs each, scored blind against a hand-built ground truth of sixteen themes and twelve traps. The loop halved unsupported claims and cut confidence errors; it also lost eighteen points of coverage in the first eval, most of which a recall rule then recovered. Coverage is still below a single prompt. That result is published, not hidden.
docs/eval1-results.md · docs/eval2-results.md in the repo
The case study — ground truth, what broke, both evals, the trace, the honest ledger.
Pricing
Free. MIT-licensed; run it on your own API key. A hosted tier is not yet decided.