ProvenanceBench
A faithfulness + justified-abstention benchmark for regulated docs: a correct refusal
is a first-class pass, scored only for the right reason. I ran three systems, including
a real LLM — and showed the default RAG metric returns NaN for a correct refusal.
Offline; every gold label adversarially verified against the corpus.
benchmark
justified abstention
regulated RAG
cite-or-refuse
A documentation Q&A system with one rule: cite a real source, or honestly
refuse — and refusal scores as a first-class eval. Before publishing, an
adversarial review caught it confidently answering questions it had no source
for; I found the root cause and fixed it. 11/11 eval cases pass, 20 offline tests, no API key.
grounded RAG
refusal-as-eval
LLM-as-Judge
promptprint
A Claude Code / Codex companion that measures the one thing nothing else does:
how you ask — and how that's changing. A recurring, local check that ends
with what to fix next — skills to build from the work you keep re-explaining — not a
one-time vanity card. "Nothing leaves your machine" is verifiable in the code.
developer tool
local-first
Claude Code