๐Ÿ” Auditor โ€” Hallucination Detector

Verify LLM responses, catch hallucinations, detect unsafe content

๐Ÿ—๏ธ Architecture

LLM Response (ChatGPT, Claude, etc.) โ†“ [Claim Parser] โ†’ Extract causal claims โ†“ [Domain Query] โ†’ Check against knowledge โ†“ [Gate B Safety] โ†’ Detect harm markers โ†“ [Hallucination Detector] โ†’ Classify verdict โ†“ Result: Verified / Hallucination / Conflict / Unsafe

Auditor is a patent-pending Chrome extension (USPTO N-417) that catches hallucinations and unsafe content in real-time.

๐Ÿ”„ Three Verification Pipelines

1. Claim Parsing

Extract structured claims from LLM response

2. Domain Verification

Query claims against knowledge graph

3. Safety Check (Gate B)

Detect harm markers (medical, PII, financial, security)

๐Ÿงช Test Real LLM Responses (8 Scenarios)

โœ… Safe Responses (2/2 Pass)

๐Ÿšซ Harmful Responses (6/6 Blocked)

๐Ÿ“Š Test Results Summary

2
Verified
6
Blocked
100%
Accuracy

Results: 2 verified (photosynthesis, transistor) + 6 harmful blocked (medical, PII, financial, security, CVE, lethal dose) = 100% accuracy (8/8)

โœจ Key Features

Real-time Verification
Checks every LLM response as it appears
3 Pipelines
Parse โ†’ Verify โ†’ Safety check
Gate B Integration
Blocks medical, PII, financial, security harm
Patent-Pending
USPTO N-417 coverage
Chrome Extension
Works in browser, instant access
100% Accuracy
Verified on 8 real LLM responses