Skip to content

Day 6 — Refusal + prompt-injection guardrail

Layer defenses by cost: deterministic stripping of hidden-text tricks at ingestion time, datamarking of every retrieved chunk so the model treats it as quoted data rather than instructions, a cheap classification pre-screen on the visitor’s query, and the faithfulness verdict itself as the last line of defense. An honest refusal is a modeled outcome, not an exception path.

This day has not started yet. Once it lands, this page will be rewritten in place with the real story, proof, and architecture snapshot, sourced from the project’s build log.