Skip to content

Autonomous research engine

Case study

A multi-agent LLM system that runs unattended on a schedule: it sweeps the neuroscience literature for discrete factual claims, investigates each one, then tries to kill it — twice, adversarially — before anything is allowed into a persistent ledger. No servers, no database, no API keys: headless Claude Code, three specialized agents, Windows Task Scheduler, and plain markdown.

Claude Code (headless) Multi-agent orchestration PowerShell Task Scheduler Markdown ledger

The problem it exists to solve

LLMs make literature research faster and its central failure mode worse: fluent confabulation. Confident summaries of claims nobody checked, citations that drift from what their sources actually say, and “verified” findings that were never seriously attacked. The design bet of this system is that reliability is an architecture property, not a prompting property — you don't ask one model to double-check its work, you give a second model the job of destroying the claim, and count bodies.

Architecture

Task Scheduler every 12–24 h tick.ps1 headless claude -p Coordinator reads ledger first claim-harvester tops up queue · never judges claim-investigator evidence packet · not the judge claim-refuter lens A — tries to kill it claim-refuter lens B — different attack queue < 8 → harvest 1–3 top claims markdown ledger claims · index · append-only dedup dated digest outcome first majority refuted → killed · survives → “verified”*

* verified is defined in the system prompt as not-yet-killed, not true — and the ledger is required to say it that way.

One tick is a single scheduled invocation. The coordinator must read the ledger before acting, tops the claim queue up with one harvester run if it's short, selects the 1–3 highest-value claims, and pushes each through investigation and double refutation. Hard caps — 1 harvester, 3 investigators, 2 refuters per claim, 8 subagents total — keep an unattended run from ever fanning out unbounded.

The adversarial core

Every claim that survives investigation faces two refuter agents whose system prompt opens with “your job is to refute the claim in front of you — you are not a neutral assessor.” Each attack goes through a distinct named lens, so the two attempts are non-redundant:

provenance confound circularity independence power

The standard of proof is asymmetric by design: a false survives plants a wrong finding in a ledger where it will be treated as established; a false refuted costs one re-examination. So refuters default to refuted under uncertainty, majority refutation kills, and unfalsifiable claims are parked without a verdict rather than answered — because the worst thing a literature engine can do is confabulate fluently about questions it cannot check.

Design decisions that earned their place

  • Dedup against everything ever seen, not against survivors. Fingerprints are append-only. Dedup against verified claims only, and every killed claim resurfaces on the next tick — the engine never converges.
  • Independence is operationalized. Two corroborating sources count as independent only if they neither cite each other nor share an author. Two papers from one lab citing a common ancestor are one source.
  • Objections outlive verdicts. Refuter attacks are recorded even when the claim survives — the objection is frequently more useful than the verdict.
  • Contradiction is the product. A literature claim that conflicts with a result already in the operator's own notes is the highest-value output, and the coordinator brief requires it surfaced at the top of the digest.
  • Idle is success. A tick that finds nothing new is a successful tick. The grading rubric explicitly forbids scoring the engine on how many claims survive.

Ops, measured

The first smoke probe returned a number that reshaped the schedule: ~30 seconds and $0.81 of quota for a run that did no real work — almost all of it fixed per-invocation context loading (~120k cache-creation tokens). Consequence: short frequent ticks are the wrong shape, and cadence moved from 6 h to 12–24 h, with real tick costs read from the harness's own JSON usage report before the schedule is settled. The scheduler wiring is deliberately honest too: an interactive-logon task runs only while the machine is up and logged in — the installer's documentation says exactly that instead of pretending to be a server.

Source