Work

System

The Council

Adversarial review. Not two chatbots talking. Not a SaaS product unless that changes.

Models are useful. They are also prone to unsupported confidence, bluffing, shallow agreement, sycophancy, and weak evidence. Asking one model to grade another, in the same chat, usually produces manners — not a verdict.

The Council splits the roles:

  • Claude as defender — argues for the draft or decision.
  • GPT or Codex as critic — attacks it.
  • A deterministic, non-AI referee plus a ledger — scores the exchange and keeps a record that neither model can rewrite.

The referee is the point. Without it, you just have a conversation. With it, you have a review you can inspect.

How the roles sit

Illustrative diagram — not a product screenshot