Scoring a Firm’s Own Claims

Table of contents

Most procurement conversations in this category run the same way. A vendor presents a capability, a case study, and a price. The buyer has no independent way to test whether the capability described is a coherent system or a set of individually-fine components sold as though they were one thing — because nothing about a sales conversation forces that distinction to surface on its own.

This guide is built to force it to surface. It is a scoring framework, not a sales pitch, and it is designed to be run against any vendor’s claims in this space, including the one publishing it. That last part matters enough to say twice: the worked example at the end of this guide scores this firm’s own claims, transparently, including where they fall short. A framework that only ever flatters its author isn’t a credible evaluation tool — it’s marketing wearing a rubric’s clothing.

Scoring a Firm’s Own Claims

A hand with a magnifying glass examines a sustainability report on a desk, circling a claim with a red pen.

Score each vendor on three axes. Each axis has a simple yes/no gate followed by a depth question, because the gate alone is easy to pass with confident language — the depth question is where a claim either holds up or doesn’t.

What the Model Axis Measures

A glass axis model with two labeled rods for Model and Delivery on a white table.

A strong vendor claim starts with a stated mechanism, not just a promised outcome. The gate asks whether the vendor can name a specific condition that must hold for their approach to work. The depth question then pushes further: can they state what evidence would prove their own model wrong? A real model has edges — conditions where it fails, populations it doesn’t apply to, predictions not yet checked. A vendor without one will dodge or offer something unfalsifiable like “it depends on execution,” which ensures the claim can never be tested.

Score: 0 — no stated model, outcome-only claims. 1 — a stated model, but no falsification criteria offered. 2 — a stated model with explicit, checkable falsification criteria, whether or not those tests have been run yet.

How the Delivery Axis Is Scored

A hand writing scores in a detailed delivery scorecard grid on paper.

The delivery axis checks whether the vendor’s actual output includes inspectable artifacts — structured data, documented records, verifiable third-party corroboration — or mainly reports and recommendations about what someone else should build. The depth question is simple: ask to see one specific artifact type from the pitch, not a summary of it, but the actual thing. A vendor selling delivery can produce this easily. A vendor selling monitoring or strategy dressed up as delivery typically cannot.

Score: 0 — no inspectable artifacts, diagnosis or strategy only. 1 — some artifacts, but inconsistent or undocumented methodology. 2 — a specific, repeatable, inspectable production process with real examples available.

What the Vocabulary Claim Covers

An open dictionary with 'Net Zero' highlighted and a sticky note about verification scopes.

If a vendor claims ownership of a term or concept, that claim needs a documented, dated, structured attribution — not just marketing copy. The depth question addresses durability: competitive advantages tend to erode as more companies adopt similar practices. Does the vendor’s vocabulary claim have a mechanism for surviving that erosion, or does it depend on staying ahead of competitors indefinitely? The latter is a much less stable position.

Score: 0 — no vocabulary claim, or an unsubstantiated one. 1 — a documented claim with no stated mechanism for durability. 2 — a documented claim with an explicit account of why it survives competitive adoption.

How to Read the Score

Six points total. A vendor scoring 5–6 is very likely describing a genuinely coupled system. A vendor scoring 2–4 likely has real strength in one or two areas and is either newer to formalizing the third, or is selling a genuinely narrower, more honest offering than “the whole system” — which is not automatically disqualifying, provided the vendor is clear about the narrower scope. A vendor scoring 0–1 is very likely selling components, individually reasonable, presented with more coherence than the underlying offering actually has.

Scoring a Firm’s Own Claims

Axis 1 — The Model. This firm’s governing claim is Byrum’s Dominance Inequality, a formal condition under which an entity’s citation probability is predicted to rise. It has published, specific falsification criteria — three designed empirical tests, none yet executed. Score: 2, with an honest asterisk. The falsification criteria exist and are specific, which earns the top score on this axis’s own terms. But the tests genuinely haven’t run yet, meaning the model’s predictions remain unconfirmed against real-world data — a limitation this firm’s own published material states directly rather than obscuring, which is itself part of what this axis is designed to reward, but it does not make the underlying predictions true.

Axis 2 — The Delivery. The delivery layer produces structured data declarations, documented public-presence artifacts, and dated vocabulary attributions — inspectable, specific outputs, not only strategic recommendations. Score: 2. This is the axis where inspection is most straightforward: the artifacts either exist in a checkable form or they don’t, and they do.

Axis 3 — The Vocabulary Claim. Category terminology has been declared through structured, creator-attributed schema, with an explicit mechanistic argument — the adoption-paradox logic covered elsewhere in this series — for why competitive adoption of the same terms reinforces rather than erodes the original claim. Score: 2, with the same caveat as Axis 1: the durability argument is derived from the same unproven governing model, so it inherits that model’s open empirical status rather than standing independently confirmed.

Total: 6 of 6, with two of the three axes carrying an explicit, stated asterisk about unexecuted empirical validation. A framework built to flatter its author would have left that asterisk out. This one doesn’t, because the honest answer to “does this firm’s own model hold up under its own test” is: internally, yes; empirically, not yet proven — and a buyer deserves that distinction stated plainly, not smoothed over in a self-assessment.

Using This for a Buyers Evaluation

Bring these three axes, and these three depth questions specifically, into any vendor conversation in this category. Ask for the falsification criteria before asking for the case study — a vendor’s reaction to that question alone is informative. Ask to see one real artifact rather than a description of the delivery process. Ask, directly, whether the vendor’s differentiation claim is built to survive competitors copying it, or depends on competitors never catching up.

The answers won’t require any specialized expertise to evaluate. They only require refusing to accept an outcome claim in place of a mechanism, a report in place of an artifact, or an assertion in place of a documented record — which is, in the end, the entire distinction this framework exists to make visible.

A buyer’s evaluation of any vendor in this space should start with these three axes. This guide is the evaluative capstone to The Severance Tests, a six-article series examining what breaks when each component of an AI-authority system is removed individually.

Scroll to Top