Skip to content

ResourcesSample evaluation

One concept,every number on the page.

A complete evaluation from the LiteSurface pipeline: both evaluators’ scores and rationales, the medians, all six gates, evidence coverage, the cited claims, and an excerpt of the build handoff.

Sample project generated by LiteSurface in fixture mode — illustrative data. The two evaluators here are deterministic stand-ins for OpenAI and Anthropic, and the sources are fictional documents on example.test domains. The medians, weights, gates, coverage, citations, and handoff are the pipeline’s real output from those inputs. What fixture mode changes

01

The result81.6 out of 100: Strong.

Project Device-Native Opportunity Scan, scored on the default Fast Product Bet scorecard. Code computed every figure below from the two evaluators’ stored outputs; no model wrote the final number.

Synthesized score
81.6
Strong band (72 and up)
Recommendation
Validate
From the score band; no gate overrode it
Gates
6 of 6
passed, one of them soft
Evidence coverage
100%
Minimum 35% for a confident call
Confidence
0.63
Reported beside the score, never inside it
Disagreement
0.13
Mean gap between evaluators, 0 to 1

02

The conceptPush Pantry for cooks

Helps home cooks plan meals from what is in the fridge, using push notifications and motion and activity sensing entirely on the device.

Value proposition
Never lose a detail, never upload a word.
Target user
home cooks
Trigger
Finishing a conversation or site visit
Mechanism
On-device pipeline: Push notifications captures, Motion and activity sensing structures, results stay local with optional encrypted sync.
Wedge
Start with the single highest-frequency capture moment for home cooks and win on trust.
Prototype estimate
7 days
Business model
$8/month with a free tier limited by records

Assumptions, ranked by importance and uncertainty

  • desirabilityhome cooks will start a capture within seconds of the moment(importance 5/5, uncertainty 3/5)
  • feasibilityOn-device processing quality is acceptable for domain vocabulary(importance 4/5, uncertainty 4/5)
  • viabilityUsers will pay a monthly subscription for a private capture tool(importance 4/5, uncertainty 3/5)
  • distributionProfessional communities can be reached without paid acquisition(importance 3/5, uncertainty 4/5)

03

Scores, criterion by criterionTwo opinions, one median, a weighted sum.

Each evaluator scored every criterion from 0 to 100 without seeing the other. Code took the median per criterion and multiplied it by the weight. The widest gap is on Why-now / enabling shift (87 vs. 52); it stays visible instead of being averaged away.

Weight, both evaluator scores, median, contribution, and spread per criterion
CriterionWeightOpenAI evaluator*Anthropic evaluator*MedianWeight × medianSpread
User pain / urgency16%80707512.000.10
Why-now / enabling shift11%875269.57.650.35
Differentiation12%91818610.320.10
Prototype feasibility13%90808511.050.10
Distribution plausibility13%89798410.920.10
Monetization quality10%9282878.700.10
Defensibility accrual10%8575808.000.10
Evidence strength10%9688929.200.08
Strategic fitopinion-based5%8070753.750.10
Weighted total100%81.585

* In this sample both evaluator slots are filled by deterministic fixture stand-ins, not live OpenAI or Anthropic models. Spread is the gap between the two scores on a 0 to 1 scale; gaps of 0.25 or more are highlighted.

04

What each evaluator saidEvery score carries a rationale and citations.

Open a criterion to read both rationales and the claims each one cites. Fixture rationales are templated; live evaluators write their own, specific to the concept and the evidence.

  • User pain / urgencyFrequency, intensity, willingness to change behavior.80 · 70 → 75

    OpenAI evaluator80 · confidence 0.65

    user pain: strong signal for "Push Pantry for cooks" based on the supplied context.

    Anthropic evaluator70 · confidence 0.65

    user pain: mixed signal for "Push Pantry for cooks" based on the supplied context.

    Cited claims: [1][2] · evidence coverage 100%

  • Why-now / enabling shiftNew capability, cost curve, regulation, or behavior change.87 · 52 → 69.5

    OpenAI evaluator87 · confidence 0.75

    why now: strong signal for "Push Pantry for cooks" based on the supplied context.

    Anthropic evaluator52 · confidence 0.75

    why now: weak signal for "Push Pantry for cooks" based on the supplied context.

    Cited claims: [2][3] · evidence coverage 100%

  • DifferentiationMeaningful advantage over existing alternatives.91 · 81 → 86

    OpenAI evaluator91 · confidence 0.85

    differentiation: strong signal for "Push Pantry for cooks" based on the supplied context.

    Anthropic evaluator81 · confidence 0.85

    differentiation: strong signal for "Push Pantry for cooks" based on the supplied context.

    Cited claims: [3][4] · evidence coverage 100%

  • Prototype feasibilityCan a credible validation product be built quickly?90 · 80 → 85

    OpenAI evaluator90 · confidence 0.65

    prototype feasibility: strong signal for "Push Pantry for cooks" based on the supplied context.

    Anthropic evaluator80 · confidence 0.65

    prototype feasibility: strong signal for "Push Pantry for cooks" based on the supplied context.

    Cited claims: [4][5] · evidence coverage 100%

  • Distribution plausibilityReachable users through channels the team can actually use.89 · 79 → 84

    OpenAI evaluator89 · confidence 0.55

    distribution: strong signal for "Push Pantry for cooks" based on the supplied context.

    Anthropic evaluator79 · confidence 0.55

    distribution: strong signal for "Push Pantry for cooks" based on the supplied context.

    Cited claims: [5][6] · evidence coverage 100%

  • Monetization qualityPayer clarity, pricing power, margin profile.92 · 82 → 87

    OpenAI evaluator92 · confidence 0.55

    monetization: strong signal for "Push Pantry for cooks" based on the supplied context.

    Anthropic evaluator82 · confidence 0.55

    monetization: strong signal for "Push Pantry for cooks" based on the supplied context.

    Cited claims: [6][7] · evidence coverage 100%

  • Defensibility accrualDoes usage create durable advantage?85 · 75 → 80

    OpenAI evaluator85 · confidence 0.75

    defensibility: strong signal for "Push Pantry for cooks" based on the supplied context.

    Anthropic evaluator75 · confidence 0.75

    defensibility: strong signal for "Push Pantry for cooks" based on the supplied context.

    Cited claims: [7] · evidence coverage 100%

  • Evidence strengthQuality and coverage of supporting evidence.96 · 88 → 92

    OpenAI evaluator96 · confidence 0.85

    evidence strength: strong signal for "Push Pantry for cooks" based on the supplied context.

    Anthropic evaluator88 · confidence 0.85

    evidence strength: strong signal for "Push Pantry for cooks" based on the supplied context.

    Cited claims: [1][2] · evidence coverage 100%

  • Strategic fitAlignment with team strengths and project constraints.80 · 70 → 75

    OpenAI evaluator80 · confidence 0.85

    strategic fit: strong signal for "Push Pantry for cooks" based on the supplied context.

    Anthropic evaluator70 · confidence 0.85

    strategic fit: mixed signal for "Push Pantry for cooks" based on the supplied context.

    Cited claims: [2][3]

05

GatesChecked before the score counts.

A failed hard gate overrides any score. The evidence gate is soft: below the minimum it caps the recommendation and sends the concept back for research.

Gate, how it was checked, severity, and result
GateChecked byResultWhy
Prototype exceeds build horizonRule (code)Pass7 days within the 14-day horizon
Requires unavailable capabilityRule (code)PassAll required capabilities are available
Evidence coverage below minimumsoftRule (code)PassCoverage 100% vs minimum 35%
Conflicts with project exclusionsEvaluatorsPassNo excluded category matched
Legal/safety critical riskEvaluatorsPassEvaluators found no blocking issue
No plausible user acquisition pathEvaluatorsPassEvaluators found no blocking issue

06

Cited evidenceAtomic claims, each with a locator.

Research stored 4 sources and extracted atomic claims from them. These are the claims the scores and the handoff cite, numbered as above.

  1. [1]Models up to 3 billion parameters run on devices released after 2024 with typical latency under 200 milliseconds per request.

    Example Platform Docs, “On-device inference framework” · published 2026-06-10 · paragraph 1 · technical fact · credibility 0.90, freshness 0.85https://example-docs.test/on-device-inference (fictional fixture source)

  2. [2]Analysts expect on-device AI features to become table stakes for premium tiers by 2027.

    Example Market Research, “Personal productivity apps market 2026” · published 2026-03-02 · paragraph 3 · technical projection · credibility 0.60, freshness 0.69https://example-market.test/report (fictional fixture source)

  3. [3]Privacy concerns were cited by 38 percent of surveyed users as a reason to avoid cloud-based note apps.

    Example Market Research, “Personal productivity apps market 2026” · published 2026-03-02 · paragraph 2 · regulatory fact · credibility 0.60, freshness 0.69https://example-market.test/report (fictional fixture source)

  4. [4]NoteWave charges 9.99 dollars per month for unlimited cloud transcription and 99 dollars per year on the annual plan.

    Example Competitor, “NoteWave pricing” · published 2026-08-15 · paragraph 1 · pricing fact · credibility 0.60, freshness 0.98https://example-competitor.test/pricing (fictional fixture source)

  5. [5]The free tier includes 30 minutes of transcription per month.

    Example Competitor, “NoteWave pricing” · published 2026-08-15 · paragraph 1 · pricing fact · credibility 0.60, freshness 0.98https://example-competitor.test/pricing (fictional fixture source)

  6. [6]Several commenters said they would pay 5 dollars per month for reliable offline transcription with good search.

    Example Forum, “Why I quit every voice memo app” · published 2026-07-21 · paragraph 2 · pricing fact · credibility 0.60, freshness 0.93https://example-forum.test/complaints (fictional fixture source)

  7. [7]Every voice memo app I tried uploads recordings to the cloud before transcribing, which I find unacceptable for work conversations.

    Example Forum, “Why I quit every voice memo app” · published 2026-07-21 · paragraph 1 · other fact · credibility 0.60, freshness 0.93https://example-forum.test/complaints (fictional fixture source)

07

Recommendation and next stepsValidate, then test the riskiest assumptions.

Push Pantry for cooks scored in the validate band. The strongest criteria were those with cited evidence; weaker criteria lacked direct claims.

Suggested next action

Run a fake-door test with the target segment this week.

Assumptions to test first

  1. Users start a capture within seconds of the triggering moment
  2. On-device processing quality is acceptable for domain vocabulary
  3. The segment will pay a monthly subscription

Where the evaluators differed

Where evaluators differed, one weighted the enabling capability shift more heavily than the market evidence; the gap narrows with a direct source on adoption.

Evidence still missing

  • A primary source on willingness to pay for the segment
  • Platform documentation confirming background execution limits

Top risks named by the evaluators

  • Capture friction may exceed the value users perceive
  • On-device accuracy on domain vocabulary is unproven

08

Handoff excerpt26 sections, checked before export.

The concept’s implementation handoff was generated section by section, validated for completeness and consistency, and pinned to concept version 1. Two sections are shown in full; sections that do not apply to this concept are marked so rather than padded.

00-product/product-requirements.md

Product requirements

Primary workflow: trigger → capture → process on device → review → retrieve.

Capture

  1. P0: Start a capture within 2 seconds of the trigger moment.
  2. P0: Process content on device; nothing leaves the device by default [1].

Retrieval

  1. P0: Search past captures by keyword and date.
  2. P1: Export a capture as text.

Suggested

  • Onboarding shows the privacy guarantee explicitly.
04-quality/acceptance-criteria.md

Acceptance criteria

  • Capture starts within 2 seconds in 95% of attempts on a mid-range device.
  • No network requests contain capture content (verified by proxy test).
  • Search returns matching items within 500 ms for 1,000 items.
  • Term accuracy ≥ 90% on the domain set.

All 26 sections

  1. Vision
  2. Target user and jobs
  3. Scope and non-goals
  4. Product requirements
  5. Success metrics
  6. Information architecture
  7. User flows
  8. Screens and states
  9. Accessibility
  10. System architecture
  11. Monorepo layout
  12. Domain model
  13. Database schema
  14. API contracts
  15. Async jobs
  16. Security and privacy
  17. AI behavior (not applicable)
  18. Model abstraction (not applicable)
  19. Prompt contracts (not applicable)
  20. Eval plan (not applicable)
  21. Test plan
  22. Acceptance criteria
  23. Launch checklist
  24. Build plan
  25. Implementation agent instructions
  26. Risks and open questions

Completeness checks

  • Target user
  • User job
  • Primary workflow
  • P0 scope
  • Data model
  • Security and privacy considerations
  • Measurable acceptance criteria
  • Implementation phases
  • Unresolved high-risk assumptions
  • Source citations for evidence-backed claims

Consistency note: Success metrics set targets without a measured baseline; record the first cohort before judging retention.

09

What fixture mode changesAnd what it does not.

Stand-ins, not live models

The evaluators are deterministic fixtures in the OpenAI and Anthropic slots, and research read a small fictional corpus instead of the web. That is why rationales repeat a template and the concept text reads generically in places.

The procedure is the real one

Medians, weights, gate checks, coverage, recommendation precedence, citation validation, and handoff checks ran in the same code that runs for live projects. In a live run, each evaluator writes its own rationale against your evidence.

Want the rules without the example? Read the scoring methodology

Now run iton an idea of your own.

Solo and Team are free during early access, and you can open a sample project in your workspace before you connect any keys.

  • Solo and Team are free during early access
  • Explore a sample project first
  • Bring your own OpenAI and Anthropic keys