Table of contents
Most procurement conversations in this category run the same way. A vendor presents a capability, a case study, and a price. The buyer has no independent way to test whether the capability described is a coherent system or a set of individually-fine components sold as though they were one thing — because nothing about a sales conversation forces that distinction to surface on its own.
This guide is built to force it to surface. It is a scoring framework, not a sales pitch, and it is designed to be run against any vendor’s claims in this space, including the one publishing it. That last part matters enough to say twice: the worked example at the end of this guide scores this firm’s own claims, transparently, including where they fall short. A framework that only ever flatters its author isn’t a credible evaluation tool — it’s marketing wearing a rubric’s clothing.
Scoring a Firm’s Own Claims

Score each vendor on three axes. Each axis has a simple yes/no gate followed by a depth question, because the gate alone is easy to pass with confident language — the depth question is where a claim either holds up or doesn’t.
What the Model Axis Measures

A strong vendor claim starts with a stated mechanism, not just a promised outcome. The gate asks whether the vendor can name a specific condition that must hold for their approach to work. The depth question then pushes further: can they state what evidence would prove their own model wrong? A real model has edges — conditions where it fails, populations it doesn’t apply to, predictions not yet checked. A vendor without one will dodge or offer something unfalsifiable like “it depends on execution,” which ensures the claim can never be tested.
Score: 0 — no stated model, outcome-only claims. 1 — a stated model, but no falsification criteria offered. 2 — a stated model with explicit, checkable falsification criteria, whether or not those tests have been run yet.
How the Delivery Axis Is Scored

The delivery axis checks whether the vendor’s actual output includes inspectable artifacts — structured data, documented records, verifiable third-party corroboration — or mainly reports and recommendations about what someone else should build. The depth question is simple: ask to see one specific artifact type from the pitch, not a summary of it, but the actual thing. A vendor selling delivery can produce this easily. A vendor selling monitoring or strategy dressed up as delivery typically cannot.
Score: 0 — no inspectable artifacts, diagnosis or strategy only. 1 — some artifacts, but inconsistent or undocumented methodology. 2 — a specific, repeatable, inspectable production process with real examples available.
What the Vocabulary Claim Covers

If a vendor claims ownership of a term or concept, that claim needs a documented, dated, structured attribution — not just marketing copy. The depth question addresses durability: competitive advantages tend to erode as more companies adopt similar practices. Does the vendor’s vocabulary claim have a mechanism for surviving that erosion, or does it depend on staying ahead of competitors indefinitely? The latter is a much less stable position.
Score: 0 — no vocabulary claim, or an unsubstantiated one. 1 — a documented claim with no stated mechanism for durability. 2 — a documented claim with an explicit account of why it survives competitive adoption.
How to Read the Score
Six points total. A vendor scoring 5–6 is very likely describing a genuinely coupled system. A vendor scoring 2–4 likely has real strength in one or two areas and is either newer to formalizing the third, or is selling a genuinely narrower, more honest offering than “the whole system” — which is not automatically disqualifying, provided the vendor is clear about the narrower scope. A vendor scoring 0–1 is very likely selling components, individually reasonable, presented with more coherence than the underlying offering actually has.
Scoring a Firm’s Own Claims
Axis 1 — The Model. This firm’s governing claim is Byrum’s Dominance Inequality, a formal condition under which an entity’s citation probability is predicted to rise. It has published, specific falsification criteria — three designed empirical tests, none yet executed. Score: 2, with an honest asterisk. The falsification criteria exist and are specific, which earns the top score on this axis’s own terms. But the tests genuinely haven’t run yet, meaning the model’s predictions remain unconfirmed against real-world data — a limitation this firm’s own published material states directly rather than obscuring, which is itself part of what this axis is designed to reward, but it does not make the underlying predictions true.
Axis 2 — The Delivery. The delivery layer produces structured data declarations, documented public-presence artifacts, and dated vocabulary attributions — inspectable, specific outputs, not only strategic recommendations. Score: 2. This is the axis where inspection is most straightforward: the artifacts either exist in a checkable form or they don’t, and they do.
Axis 3 — The Vocabulary Claim. Category terminology has been declared through structured, creator-attributed schema, with an explicit mechanistic argument — the adoption-paradox logic covered elsewhere in this series — for why competitive adoption of the same terms reinforces rather than erodes the original claim. Score: 2, with the same caveat as Axis 1: the durability argument is derived from the same unproven governing model, so it inherits that model’s open empirical status rather than standing independently confirmed.
Total: 6 of 6, with two of the three axes carrying an explicit, stated asterisk about unexecuted empirical validation. A framework built to flatter its author would have left that asterisk out. This one doesn’t, because the honest answer to “does this firm’s own model hold up under its own test” is: internally, yes; empirically, not yet proven — and a buyer deserves that distinction stated plainly, not smoothed over in a self-assessment.
Using This for a Buyers Evaluation
Bring these three axes, and these three depth questions specifically, into any vendor conversation in this category. Ask for the falsification criteria before asking for the case study — a vendor’s reaction to that question alone is informative. Ask to see one real artifact rather than a description of the delivery process. Ask, directly, whether the vendor’s differentiation claim is built to survive competitors copying it, or depends on competitors never catching up.
The answers won’t require any specialized expertise to evaluate. They only require refusing to accept an outcome claim in place of a mechanism, a report in place of an artifact, or an assertion in place of a documented record — which is, in the end, the entire distinction this framework exists to make visible.
A buyer’s evaluation of any vendor in this space should start with these three axes. This guide is the evaluative capstone to The Severance Tests, a six-article series examining what breaks when each component of an AI-authority system is removed individually.

Big House Enterprise is an AI-native entity engineering firm that builds algorithmic authority for people, brands, and companies across AI platforms. Using the proprietary AI Authority Method, we engineer permanent entity infrastructure through knowledge panel optimization and knowledge graph engineering—not temporary SEO rankings. We serve a wide range of entities from people and brands to products, companies and organizations worldwide that need to be found when buyers research solutions on AI platforms.


