Skip to content

ProductEvaluation

Two evaluators.One reproducible number.

Independent OpenAI and Anthropic evaluators score each criterion on your scorecard, with a rationale and cited claims. Deterministic code takes the median, applies weights and gates, and reports confidence and disagreement separately, so no model decides alone.

scoring every criterion independently
2 providers
in the default Fast Product Bet scorecard
9 criteria
synthesis computed in code
Median, then weights
exceptional, strong, hold, weak
4 score bands

How it works

Models give opinions.Code does the arithmetic.

  1. 01

    Choose a scorecard

    Start from Fast Product Bet or define your own criteria, weights, and gates for each project.

  2. 02

    Score independently

    Each evaluator scores every criterion without seeing the other’s answer, citing the claims it relied on.

  3. 03

    Synthesize in code

    The median for each criterion is weighted into a total. Confidence, disagreement, and coverage are computed beside it, never folded in.

  4. 04

    Apply the gates

    Hard gates such as build horizon, legal risk, and acquisition path override the number when they fail.

Decision rules

Bands and gates,written down in advance.

A score maps to a band you can read at a glance: 85 and up is exceptional, 72 strong, 58 hold, and anything lower weak. Gates sit above the number. A failed hard gate overrides any score, and thin evidence caps the recommendation until research closes the gap.

  • Hard gates: build horizon, capability availability, legal and safety risk, acquisition path, exclusions
  • Evidence coverage is a soft gate that triggers targeted research
  • Human overrides are recorded beside the AI recommendation
  • Re-evaluating appends a new run and never rewrites the last
Read the scoring methodology

Capabilities

A score that showsits working.

  • Independent evaluators

    Two providers, two sets of priors. Where they agree you can move faster; where they differ, you know where to look.

  • Weighted scorecards

    Criteria, weights, and gates are data you control per project, starting from a default that covers pain through strategic fit.

  • Rationales with citations

    Every criterion score comes with both rationales and the claims each evaluator cited.

  • Confidence, reported apart

    Confidence never inflates a score. It sits next to it, along with disagreement and evidence coverage.

  • Disagreement surfaced

    Large gaps between evaluators are flagged per criterion instead of being averaged out of sight.

  • Re-evaluation without rewriting

    Evaluate any version again after new evidence or an experiment. Earlier runs stay as they were.

What you getA score you can take apart.

Every number decomposes into criteria, evaluators, rationales, and the claims behind them.

  • Per-criterion scores from both evaluators
  • A weighted total computed in code
  • Confidence, disagreement, and coverage side by side
  • Hard gate results with reasons
  • A score band and a recommendation
  • An immutable record of every run

Let two models argue.Let code keep the score.

Read a complete sample evaluation, then run live evaluations with your own OpenAI and Anthropic keys.