The standard

Scoring methodology

The Wardict Score is a 5–100 viability score. The Warden scores 5 dimensions independently (1–20 each), then sums them. Judgment follows analysis, never the other way around. This is the exact rubric the Warden is held to, published in full. Every verdict is stamped with the methodology version it was scored under, so rulings stay comparable over time.

The 5 dimensions

Market Pull

Is there EVIDENCE people actively want this and will pay for it?

1–4
No evidence of demand, hypothetical problem
5–8
Anecdotal interest or adjacent demand exists, but unproven for THIS specific product
9–12
Moderate evidence: people are searching for solutions, some willingness to pay demonstrated
13–16
Clear demand signals: existing budget, active buyer searches, or proven pain point with ££ attached
17–20
Proven acute pain: people are already paying for inferior alternatives, strong pull

The 1-4 band requires a finding that the problem is not felt, not merely an absence of evidence that it is. A reasoned structural read of existing behaviour IS such a finding: where the behaviour an idea depends on plainly does not exist, or a free alternative already settles the job, say so and score it there. An untested demand claim with nothing against it is different, and belongs in 5-8. Broad market-size statistics never substitute for bottom-up pull: a large TAM with no observed buyer behaviour does not lift this score.

Buyer Clarity

Is there a specific buyer with budget, authority, and urgency to purchase?

1–4
No identifiable buyer, or buyer ≠ user with no clear path to revenue
5–8
Vague segment ("parents", "small businesses") but no evidence of purchase intent or budget
9–12
Identifiable buyer segment with some evidence of budget, but urgency unclear
13–16
Known buyer with demonstrated budget and purchasing behaviour in this category
17–20
Buyer actively seeking solutions, budget allocated, short sales cycle

An unnamed buyer scores low here because clarity is what this dimension measures, but say so as an unknown, not as a finding that nobody will pay. Reserve 1-4 for cases where the buyer is genuinely structurally absent, for example the user cannot pay and no one else has a reason to.

Distribution Feasibility

Can you reach buyers through a REPEATABLE, cost-effective channel?

1–4
No obvious path to buyers, or entirely dependent on virality/luck
5–8
Channels exist in theory but are crowded, expensive, or unproven for this product
9–12
Plausible channel with some evidence of feasibility, but CAC unclear
13–16
Tested distribution channel exists for similar products, reasonable CAC
17–20
Proven repeatable channel with demonstrated unit economics

Judge the channel needed to reach the FIRST buyers, not the channel the ultimate vision would eventually need. An unproven channel is 5-8; 1-4 is for cases where the buyers are structurally unreachable or the only route is virality.

Competitive Edge

What CONCRETE, DEFENSIBLE advantage does this have over existing alternatives?

1–4
Pure commodity, or free/built-in alternatives already exist (Google, Canva, ChatGPT, etc.)
5–8
Minor differentiation that could be easily replicated, or "AI wrapper" over existing tools
9–12
Meaningful differentiation in one area, but not clearly defensible
13–16
Strong advantage from proprietary data, network effects, or deep domain expertise
17–20
Defensible moat: switching costs, regulatory advantage, or compounding network effects

An absence of direct competitors is not evidence either way. It can mean there is no market, that the category is emerging, or that the job is currently done another way. Establish which by looking at substitutes and current behaviour, and score that finding rather than the empty competitor list.

Execution Feasibility

How feasible is building AND SUSTAINING this as a business?

1–4
Requires breakthrough technology or impossible operational constraints
5–8
Technically buildable but significant operational, regulatory, or scaling challenges
9–12
Buildable with standard technology, moderate operational complexity
13–16
Well-understood tech stack and operations, clear path to ship
17–20
Straightforward to build, operate, and scale

Score the business a founder could start, not only the destination they described. Weigh three things: how hard the ultimate vision is, whether the first wedge is buildable and testable with today's technology and capital, and whether a plausible sequence runs from the wedge onward. The wedge carries the most weight, because it is what gets built first, so a vision needing a breakthrough does not by itself force the 1-4 band when the first commercial move is buildable today. Two hard limits apply in the other direction: where the ultimate vision depends on a breakthrough that does not exist, this dimension may not exceed 12 however easy the wedge is; and where no credible wedge exists at all, that IS the execution problem and the score belongs in 1-8.

Verdict thresholds

PARDONED
85–100Strong evidence of viability across most dimensions
PAROLED
70–84Viable with a clear path: strong in some dimensions, acceptable in others
PROBATION
50–69Significant issues but fixable: some strong dimensions, some weak
GUILTY
25–49Fatal structural flaws in multiple dimensions
CONDEMNED
5–24Fails on nearly every dimension. Reserve for truly unviable ideas

Which market a case is tried in

A verdict is only meaningful somewhere. The founder picks the market before the verdict is rendered, or leaves it to the Warden, which reads the filing and nominates one. Whichever happens, the market is stated on the verdict.

UK
Competitors, buyer norms and regulation are those of the United Kingdom. Figures are in GBP.
US
Competitors, buyer norms and regulation are those of the United States. Figures are in USD.
Europe
Competitors, buyer norms and regulation are those of Europe (EU/EEA). Figures are in EUR.
Global
Competitors, buyer norms and regulation are those of the global market. Figures are in USD.

Europe means the EU and EEA, and does not include the United Kingdom, which is a market of its own. Global is reported in US dollars. Where a filing points at one market and is tried in another, the Warden says so in its reasoning rather than quietly moving the case.

How evidence is weighed

Every claim a case rests on sits in one of three states. Two of them are findings. The third is a missing experiment, and it is not scored as though it were a finding.

SUPPORTED
Evidence exists for the claim. It raises the relevant dimension in proportion to how directly and how well the evidence maps to THIS product, not to the category.
CONTRADICTED
Evidence argues the claim is false or unlikely: a failed comparable, economics that do not close, a constraint that cannot be met, a buyer with no reason to act. This is a finding and is penalised heavily. It is also the only state that justifies saying the idea does not work.
UNKNOWN
There is not enough evidence either way. It caps the score, because confidence is what the score measures, and it generates an experiment. It is never written up as though the claim had been disproven.

Ambitious cases

Some filings propose something much larger than a first product. Judging those against the ultimate ambition alone would punish every one of them for being early, so the Warden separates the destination from the first business that could test it: vision, then wedge, then evidence, then expansion.

When it applies
The filing shows at least two structural ambition signals: several new technical layers, a research or regulatory breakthrough, hardware alongside software, infrastructure or platform creation, a required change in established behaviour, new category creation, a long-term ecosystem, genuinely novel technology, replacing an incumbent stack, or a stated destination whose first commercial product is not obvious. An idea is not ambitious for being unusual.
Execution Feasibility
Composed from three separate judgments rather than one: how hard the ultimate vision is, whether the initial wedge is buildable and testable now, and whether a plausible expansion path runs from the wedge toward the vision. The wedge carries the most weight.
What the number means
The score answers how much evidence there is today of a venture-scale business and a credible path to finding it. It does not answer how conventional or how easy the idea is, and it is not a forecast of the vision's ultimate potential. An ambitious, unvalidated filing may sit at 25-40 and the correct reading is low current confidence, not low ceiling.

Checks before a verdict is filed

  1. 1Was the ultimate vision judged where the first product should have been?
  2. 2Was an unknown written up as evidence against the idea?
  3. 3Was a buyer assumed, or ruled out, without evidence?
  4. 4Was execution penalised for work the initial wedge does not require?
  5. 5Was the idea CREDITED for a wedge, an expansion path, a buyer, or a market that the filing did not actually establish?
  6. 6Would this filing have scored the same if it had been written plainly, without the scale of the ambition? If the ambition alone moved a score, remove that movement.
  7. 7Does the proposed experiment actually test the largest uncertainty, and would passing it change the verdict?
  8. 8Is the ambition preserved while the recommended first test is narrower?
  9. 9Is every harsh conclusion carried by evidence rather than by the absence of it, and is every favourable one carried by evidence too?
  10. 10Was market-size data used in place of observed market pull?

Calibration anchors

Worked reference points the Warden uses to gut-check dimension scores, so similar ideas are sentenced consistently.

Score ~80+ (PAROLED/PARDONED)

Most dimensions 14+. Proven pain point, known buyer with budget, demonstrated distribution channel, real competitive advantage.

e.g. “B2B tool replacing a manual process companies already spend £50K+/year on.”

Score ~55-65 (PROBATION)

Mix of strong (12-15) and weak (5-8) dimensions. Real market signal exists but significant gaps in distribution, competition, or buyer clarity.

e.g. “Real pain point, but crowded market and no clear distribution edge.”

Score ~35-45 (GUILTY)

Most dimensions 5-9. Speculative demand, vague buyer, crowded market, or "AI wrapper" over free tools.

e.g. “Consumer app in a category where free alternatives already exist.”

Score ~15-25 (CONDEMNED)

Most dimensions 1-5. No evidence of demand, no buyer, no path to market, no edge.

e.g. “Solution looking for a problem with no evidence anyone cares.”

How the court stays honest

A generic model gives everyone a different opinion with no published standard behind it. The Wardict Score is held to a versioned rubric, reviewed by people, and sharper with every case. That is why a better base model lifts it rather than replacing it.

A published standard

This page is the rubric in full, versioned and citable. Every verdict is stamped with the version it was judged under, so a ruling can always be traced to the exact standard behind it.

A human panel

The Warden is not the last word. Flagged and sampled verdicts are routed to a human panel of founders and investors, and recurring blind spots are encoded into the next version of this rubric. A human always decides what is real and how to encode it.

A growing record

Wardict follows up to learn what actually happened to judged ideas. No foundation model has that record. It is what lets the standard improve as the base models do.

Put My Idea on Trial