Tier 3 — do-calculus gates: is the claim true or supportable? No LLM in the path.
Tier 3 · #11
Identifiability Gate
Given a causal DAG and P(Y | do(X)) — is the effect identifiable from observational data? If yes, which criterion (backdoor / front-door / empty-set / instrument) + estimand. Engine: ananke OneLineID (Shpitser–Pearl).
Stage 1 (cheap): self-consistency / disagreement across sampled answers. Stage 2 (symbolic): flagged claims adjudicated against the formal causal oracle — claims the oracle refutes are DENIED, ground truth returned.
Claim: "drug X causes recovery (p<0.01)" · 3 samples agree
Domain Gate: every domain JSON admitted only if structurally valid, internally consistent, and provenance-intact. Flip-Rate Gate: did new causal edges actually improve understanding? The domain's own causal model is the oracle — it tests itself.
new edge: "stress → insomnia" (source cited: PMID 12345)
new edge: "insomnia → stress" AND "stress → insomnia" (both asserted, no source)
The design point
Generate → Verify → Persist → Reuse
The generation side varies (LLM, templates, operators); the knowledge representation varies. What never varies is the gate: a mechanical check of a candidate against ground truth. These gates are why a claim can be structurally verified or refuted without asking a model to grade itself. Abstention is a first-class outcome — "no causal path" is a real answer, not a failure.