ResourcesSample evaluation
One concept,every number on the page.
A complete evaluation from the LiteSurface pipeline: both evaluators’ scores and rationales, the medians, all six gates, evidence coverage, the cited claims, and an excerpt of the build handoff.
Sample project generated by LiteSurface in fixture mode — illustrative data. The two evaluators here are deterministic stand-ins for OpenAI and Anthropic, and the sources are fictional documents on example.test domains. The medians, weights, gates, coverage, citations, and handoff are the pipeline’s real output from those inputs. What fixture mode changes
01
The result81.6 out of 100: Strong.
Project Device-Native Opportunity Scan, scored on the default Fast Product Bet scorecard. Code computed every figure below from the two evaluators’ stored outputs; no model wrote the final number.
- Synthesized score
- 81.6
- Strong band (72 and up)
- Recommendation
- Validate
- From the score band; no gate overrode it
- Gates
- 6 of 6
- passed, one of them soft
- Evidence coverage
- 100%
- Minimum 35% for a confident call
- Confidence
- 0.63
- Reported beside the score, never inside it
- Disagreement
- 0.13
- Mean gap between evaluators, 0 to 1
02
The conceptPush Pantry for cooks
Helps home cooks plan meals from what is in the fridge, using push notifications and motion and activity sensing entirely on the device.
- Value proposition
- Never lose a detail, never upload a word.
- Target user
- home cooks
- Trigger
- Finishing a conversation or site visit
- Mechanism
- On-device pipeline: Push notifications captures, Motion and activity sensing structures, results stay local with optional encrypted sync.
- Wedge
- Start with the single highest-frequency capture moment for home cooks and win on trust.
- Prototype estimate
- 7 days
- Business model
- $8/month with a free tier limited by records
Assumptions, ranked by importance and uncertainty
- desirabilityhome cooks will start a capture within seconds of the moment(importance 5/5, uncertainty 3/5)
- feasibilityOn-device processing quality is acceptable for domain vocabulary(importance 4/5, uncertainty 4/5)
- viabilityUsers will pay a monthly subscription for a private capture tool(importance 4/5, uncertainty 3/5)
- distributionProfessional communities can be reached without paid acquisition(importance 3/5, uncertainty 4/5)
03
Scores, criterion by criterionTwo opinions, one median, a weighted sum.
Each evaluator scored every criterion from 0 to 100 without seeing the other. Code took the median per criterion and multiplied it by the weight. The widest gap is on Why-now / enabling shift (87 vs. 52); it stays visible instead of being averaged away.
| Criterion | Weight | OpenAI evaluator* | Anthropic evaluator* | Median | Weight × median | Spread |
|---|---|---|---|---|---|---|
| User pain / urgency | 16% | 80 | 70 | 75 | 12.00 | 0.10 |
| Why-now / enabling shift | 11% | 87 | 52 | 69.5 | 7.65 | 0.35 |
| Differentiation | 12% | 91 | 81 | 86 | 10.32 | 0.10 |
| Prototype feasibility | 13% | 90 | 80 | 85 | 11.05 | 0.10 |
| Distribution plausibility | 13% | 89 | 79 | 84 | 10.92 | 0.10 |
| Monetization quality | 10% | 92 | 82 | 87 | 8.70 | 0.10 |
| Defensibility accrual | 10% | 85 | 75 | 80 | 8.00 | 0.10 |
| Evidence strength | 10% | 96 | 88 | 92 | 9.20 | 0.08 |
| Strategic fitopinion-based | 5% | 80 | 70 | 75 | 3.75 | 0.10 |
| Weighted total | 100% | 81.585 | ||||
* In this sample both evaluator slots are filled by deterministic fixture stand-ins, not live OpenAI or Anthropic models. Spread is the gap between the two scores on a 0 to 1 scale; gaps of 0.25 or more are highlighted.
04
What each evaluator saidEvery score carries a rationale and citations.
Open a criterion to read both rationales and the claims each one cites. Fixture rationales are templated; live evaluators write their own, specific to the concept and the evidence.
User pain / urgencyFrequency, intensity, willingness to change behavior.80 · 70 → 75
Why-now / enabling shiftNew capability, cost curve, regulation, or behavior change.87 · 52 → 69.5
DifferentiationMeaningful advantage over existing alternatives.91 · 81 → 86
Prototype feasibilityCan a credible validation product be built quickly?90 · 80 → 85
OpenAI evaluator90 · confidence 0.65
prototype feasibility: strong signal for "Push Pantry for cooks" based on the supplied context.
Anthropic evaluator80 · confidence 0.65
prototype feasibility: strong signal for "Push Pantry for cooks" based on the supplied context.
Distribution plausibilityReachable users through channels the team can actually use.89 · 79 → 84
Monetization qualityPayer clarity, pricing power, margin profile.92 · 82 → 87
Defensibility accrualDoes usage create durable advantage?85 · 75 → 80
OpenAI evaluator85 · confidence 0.75
defensibility: strong signal for "Push Pantry for cooks" based on the supplied context.
Anthropic evaluator75 · confidence 0.75
defensibility: strong signal for "Push Pantry for cooks" based on the supplied context.
Cited claims: [7] · evidence coverage 100%
Evidence strengthQuality and coverage of supporting evidence.96 · 88 → 92
OpenAI evaluator96 · confidence 0.85
evidence strength: strong signal for "Push Pantry for cooks" based on the supplied context.
Anthropic evaluator88 · confidence 0.85
evidence strength: strong signal for "Push Pantry for cooks" based on the supplied context.
Strategic fitAlignment with team strengths and project constraints.80 · 70 → 75
05
GatesChecked before the score counts.
A failed hard gate overrides any score. The evidence gate is soft: below the minimum it caps the recommendation and sends the concept back for research.
| Gate | Checked by | Result | Why |
|---|---|---|---|
| Prototype exceeds build horizon | Rule (code) | Pass | 7 days within the 14-day horizon |
| Requires unavailable capability | Rule (code) | Pass | All required capabilities are available |
| Evidence coverage below minimumsoft | Rule (code) | Pass | Coverage 100% vs minimum 35% |
| Conflicts with project exclusions | Evaluators | Pass | No excluded category matched |
| Legal/safety critical risk | Evaluators | Pass | Evaluators found no blocking issue |
| No plausible user acquisition path | Evaluators | Pass | Evaluators found no blocking issue |
06
Cited evidenceAtomic claims, each with a locator.
Research stored 4 sources and extracted atomic claims from them. These are the claims the scores and the handoff cite, numbered as above.
[1]Models up to 3 billion parameters run on devices released after 2024 with typical latency under 200 milliseconds per request.
Example Platform Docs, “On-device inference framework” · published 2026-06-10 · paragraph 1 · technical fact · credibility 0.90, freshness 0.85https://example-docs.test/on-device-inference (fictional fixture source)
[2]Analysts expect on-device AI features to become table stakes for premium tiers by 2027.
Example Market Research, “Personal productivity apps market 2026” · published 2026-03-02 · paragraph 3 · technical projection · credibility 0.60, freshness 0.69https://example-market.test/report (fictional fixture source)
[3]Privacy concerns were cited by 38 percent of surveyed users as a reason to avoid cloud-based note apps.
Example Market Research, “Personal productivity apps market 2026” · published 2026-03-02 · paragraph 2 · regulatory fact · credibility 0.60, freshness 0.69https://example-market.test/report (fictional fixture source)
[4]NoteWave charges 9.99 dollars per month for unlimited cloud transcription and 99 dollars per year on the annual plan.
Example Competitor, “NoteWave pricing” · published 2026-08-15 · paragraph 1 · pricing fact · credibility 0.60, freshness 0.98https://example-competitor.test/pricing (fictional fixture source)
[5]The free tier includes 30 minutes of transcription per month.
Example Competitor, “NoteWave pricing” · published 2026-08-15 · paragraph 1 · pricing fact · credibility 0.60, freshness 0.98https://example-competitor.test/pricing (fictional fixture source)
[6]Several commenters said they would pay 5 dollars per month for reliable offline transcription with good search.
Example Forum, “Why I quit every voice memo app” · published 2026-07-21 · paragraph 2 · pricing fact · credibility 0.60, freshness 0.93https://example-forum.test/complaints (fictional fixture source)
[7]Every voice memo app I tried uploads recordings to the cloud before transcribing, which I find unacceptable for work conversations.
Example Forum, “Why I quit every voice memo app” · published 2026-07-21 · paragraph 1 · other fact · credibility 0.60, freshness 0.93https://example-forum.test/complaints (fictional fixture source)
07
Recommendation and next stepsValidate, then test the riskiest assumptions.
Push Pantry for cooks scored in the validate band. The strongest criteria were those with cited evidence; weaker criteria lacked direct claims.
Suggested next action
Run a fake-door test with the target segment this week.
Assumptions to test first
- Users start a capture within seconds of the triggering moment
- On-device processing quality is acceptable for domain vocabulary
- The segment will pay a monthly subscription
Where the evaluators differed
Where evaluators differed, one weighted the enabling capability shift more heavily than the market evidence; the gap narrows with a direct source on adoption.
Evidence still missing
- A primary source on willingness to pay for the segment
- Platform documentation confirming background execution limits
Top risks named by the evaluators
- Capture friction may exceed the value users perceive
- On-device accuracy on domain vocabulary is unproven
08
Handoff excerpt26 sections, checked before export.
The concept’s implementation handoff was generated section by section, validated for completeness and consistency, and pinned to concept version 1. Two sections are shown in full; sections that do not apply to this concept are marked so rather than padded.
Product requirements
Primary workflow: trigger → capture → process on device → review → retrieve.
Capture
- P0: Start a capture within 2 seconds of the trigger moment.
- P0: Process content on device; nothing leaves the device by default [1].
Retrieval
- P0: Search past captures by keyword and date.
- P1: Export a capture as text.
Suggested
- Onboarding shows the privacy guarantee explicitly.
Acceptance criteria
- Capture starts within 2 seconds in 95% of attempts on a mid-range device.
- No network requests contain capture content (verified by proxy test).
- Search returns matching items within 500 ms for 1,000 items.
- Term accuracy ≥ 90% on the domain set.
All 26 sections
- Vision
- Target user and jobs
- Scope and non-goals
- Product requirements
- Success metrics
- Information architecture
- User flows
- Screens and states
- Accessibility
- System architecture
- Monorepo layout
- Domain model
- Database schema
- API contracts
- Async jobs
- Security and privacy
- AI behavior (not applicable)
- Model abstraction (not applicable)
- Prompt contracts (not applicable)
- Eval plan (not applicable)
- Test plan
- Acceptance criteria
- Launch checklist
- Build plan
- Implementation agent instructions
- Risks and open questions
Completeness checks
- Target user
- User job
- Primary workflow
- P0 scope
- Data model
- Security and privacy considerations
- Measurable acceptance criteria
- Implementation phases
- Unresolved high-risk assumptions
- Source citations for evidence-backed claims
Consistency note: Success metrics set targets without a measured baseline; record the first cohort before judging retention.
09
What fixture mode changesAnd what it does not.
Stand-ins, not live models
The procedure is the real one
Want the rules without the example? Read the scoring methodology
Now run iton an idea of your own.
Solo and Team are free during early access, and you can open a sample project in your workspace before you connect any keys.
- Solo and Team are free during early access
- Explore a sample project first
- Bring your own OpenAI and Anthropic keys