The protocol inside VibeStack Builder™
What is VERIDEX?
VERIDEX is a forensic evidence-validation protocol: a strict set of rules for how a claim, a plan, or a body of evidence gets assessed — what counts as proof, what must be separated from opinion, and what a verdict is allowed to rest on. In an independent-style comparative assessment of sixteen named validation protocols, VERIDEX placed second — 82.2 out of 100, 1.2 points behind the Federal Reserve’s SR 11-7, the model-validation standard bank examiners enforce. The gap to first place is smaller than the uncertainty on any single score.
How it was measured
Not by us grading ourselves against a curated field. Sixteen protocols were scored on the same eleven-dimension rubric, every score citing the governing text that justifies it — and VERIDEX was scored inside the field, on the same rubric, never held out as the benchmark. The field it stood in:
- SR 11-7 — the Federal Reserve’s model-validation guidance — what bank examiners enforce
- GAGAS — the U.S. Government Auditing Standards (the “Yellow Book”)
- Daubert — the standard U.S. courts use to admit expert evidence
- GRADE & Cochrane RoB 2 — how medical science grades its evidence
- ICD 203 — the U.S. intelligence community’s analytic tradecraft standard
Five dimensions at the top of the scale
Evidentiary rigor
Sources are tiered, cross-checked figures take the weakest contributing tier, and promoter material can never anchor a high score.
Structural completeness
Gates, criteria, and scoring instruments specified in enforceable detail — a procedure, not a philosophy.
Uncertainty calibration
Weak sourcing widens the stated uncertainty instead of hiding in a footnote — the only protocol in the field that propagates source quality into the width of the answer.
Domain portability
It travels — any domain, any subject — while still carrying a hard evidence-admissibility threshold. Nothing else in the field does both.
Failure-mode disclosure
It names its own limits and failure modes in its governing documents — more of them, more plainly, than nearly anything it was scored against.
What no peer protocol has
Substrate-aware debiasing. Every peer protocol was written for human analysts and counters human failure modes. VERIDEX is the only protocol in the field that names and counters the failure modes of the AI executing it — that models inflate scores as evaluators, that an AI assessing AI-generated output accepts plausible claims too easily, that self-evaluation cannot be engineered away with a rubric. The peers could not have done this: they predate the problem.
Disconfirmation you can count. Plenty of protocols encourage looking for contrary evidence. VERIDEX is the only one that requires it as a measurable output property — a report that contains nothing that could discourage the recommended action fails.
Source quality sets the width of the answer. A figure from a single source with undisclosed methodology carries a far wider stated uncertainty than one from an authoritative source with published methodology. Every peer that grades sources stops at a tag; VERIDEX propagates the grade into the number.
Where it’s honest about itself
The same assessment that ranked VERIDEX second also recorded what it lacks, and we’d rather you read it here than discover it. It has no issuing institution — no external body mandates or examines it the way bank examiners enforce SR 11-7, and its governing documents are authored by the entity that operates it. Its consistency between independent operators has not yet been formally measured — though the field’s own record there is humbling: the most-measured medical protocols publish agreement statistics that are famously unflattering. And like every portable protocol in the field, it trades jurisdiction for reach: the protocols with real external enforcement are locked to one domain; everything that travels is self-attested. The assessment found that trade-off unresolved across the entire discipline — no protocol occupies both corners.
A protocol that discloses its own failure modes this plainly is exactly the protocol you’d want auditing your plan. That’s not spin — failure-mode disclosure is one of the five dimensions it maxed.
What this means for your app
Every VibeStack Builder™ blueprint already gets an automatic VERIDEX coverage check before it reaches you — that’s built in, on every plan. Studio members get the deeper instrument: the VERIDEX deep audit, run on demand against your finished blueprint, on your own Anthropic key. A principal-engineer-grade forensic review — stack fitness, data model, security, what production will actually demand — ending in a straight verdict: sound as specified, or here are the material changes, ranked, each with its evidence and its tradeoff. If your plan is sound, it says so and tells you why. If it isn’t, you find out before a single line of code exists — when changes are nearly free.
Then one tap carries the audit into your plan chat, ready for you to send. Your Vibe Agent works through every finding, settles the engineering ones the way any lead programmer would, and comes back to you only on the choices that are genuinely yours — what your app does, what it costs, what it asks of you — in plain language. Your plan is rewritten from what you agree. Audit, converge, improve — before you build.