Adversarial startup validation
Kill weak startup ideas
before your AI builds them
Describe your idea. CodeKudo attacks it with real complaints, spending signals, competitors and the reasons it fails — then tells you to build it, test it, narrow it, or stop. If it survives, you get the spec in the exact format your coding tool expects.
One spec, 30 correct formats
CLAUDE.md + AGENTS.mdCursor.cursor/rules/*.mdcWindsurf.windsurf/rules/*.mdGitHub Copilot.github/copilot-instructions.mdOpenAI Codex CLIAGENTS.mdGemini CLIGEMINI.md + AGENTS.mdAWS Kiro.kiro/specs/{slug}/requirements.md · design.md · tasks.mdGoogle AntigravityAGENTS.mdZedAGENTS.mdCline / Roo Code.clinerules + AGENTS.mdDevinAGENTS.md + knowledgeAmpAGENTS.mdTraeAGENTS.mdAiderCONVENTIONS.mdLovableKnowledge Base + founding prompt6-part founding promptBase44Single full-stack promptEmergent.shAgent-segmented brief (UI / logic / DB / API / QA)Replit Agentreplit.md + numbered task listv0 (Vercel)Per-screen component prompts + design tokensSoftgenFounding prompt + iterationsCreate.xyzFounding promptDatabuttonPrompt + Python backend briefTempoFounding prompt (React)TrickleFounding promptVybeFounding promptOnSpace.AIFounding promptRorkReact Native + Expo prompt + screen flowFireVibeScreen list + brand tokens + handoff bridgeFlutterFlow AIVisual + AI brief (Flutter)Example card — illustrates the format. Live cards are built from collected signals and every reference resolves to a real source.
Return automation for sub-$50k GMV Shopify stores
Rules-driven return approvals and label generation for small stores that the incumbents price out.
Demand ladder
Complaints are cheap. Money is the signal.
Confidence
Where the data is thin, we say so.
- Demand signal
- High
- Willingness to pay
- Low
- Market size
- No data
- Competitor gap
- Medium
47 signals across 6 sources
Only 3 direct signals — not enough to decide on
No direct data. Anything here would be a guess.
2 sources confirmed
Supporting evidence3
Store operators describe returns as manual work costing roughly 11 minutes per order.
Twelve people said outright they would pay for this rather than keep doing it by hand.
Three products in this niche have Stripe-verified revenue, the largest at $47k MRR.
Falsifying evidence3
Shopify can ship this natively; they have shipped adjacent workflow features twice in 18 months.
Two products serving this exact segment shut down, both citing support load per dollar.
Roughly 70% of the complaints are resolved with a free spreadsheet template people already share.
Most likely cause of death
Shopify adds native return rules and this becomes a feature, not a product. Your defensibility has to be the sub-$50k GMV segment the incumbents refuse to support — not the workflow itself.
We look for reasons to say no
Six competing products all search for supporting evidence and stop the moment they find some. That is confirmation bias with a subscription. Every card here carries a falsifying column and a stated cause of death.
Every claim resolves to a source
Click any [S-4471] and you get the platform, the date, a short excerpt and a link to the original. If a reference doesn't exist in the database, the sentence is dropped before you ever see it.
Specs, not a PDF nobody reads
Pick your tool and get exactly what it expects: AGENTS.md for Claude Code, .cursor/rules for Cursor, a Knowledge Base doc for Lovable, a hard 400-word brief for Bolt.
The chain, unbroken
The tools in this category each own one link and hand you the gap. A complaint corpus with no answer. A perfect PRD with no data behind it. A 200-page report nobody acts on.
- 01
Signal
Six free official APIs, pulled to the edge of every rate limit. Raw text is discarded; the derived signal and its link are kept.
- 02
Opportunity
Clustering, momentum over 30/90/365 days, and a demand ladder that weighs verified revenue 15× a complaint.
- 03
Evidence
Complaint ↔ existing product ↔ Stripe-verified MRR, plus the counter-evidence that argues against building it.
- 04
Decision
A chat that answers from records only. No record, no claim — it will tell you the data doesn't support you.
- 05
Spec
One Universal Core, rendered into whatever your platform reads. EVIDENCE.md is generated by code, never by a model.
What the others left out
We went through every idea-database and idea-validator tool we could find and checked them against the same list. The category looks crowded from the side. Read down the column and it empties out.
| Fundamental | Other tools | CodeKudo |
|---|---|---|
| Real signal data | Standard | |
| Clickable source for every claim | Standard | |
| Stripe-verified revenue | Some, partially | |
| Platform-specific spec output | Some, partially | |
| Evidence traced into the spec | Some, partially | |
| First-100-users plan from the data | Some, partially | |
| Founder-fit score | Some, partially | |
| Outcome feedback loop | Some, partially | |
| Counter-evidence engineonly here | Nobody | |
| Weighted payment evidenceonly here | Nobody | |
| Freshness and momentumonly here | Nobody | |
| Saturation transparencyonly here | Nobody | |
| Live monitoring and alertsonly here | Nobody |
Five tools surveyed. “Some, partially” means at least one does a version of it, not that the category does. Anything markedonly here was found in none of them.
One idea, 30 correct formats
Pasting a 2,000-word PRD into Bolt drowns it. Bolt wants six sections under 400 words. Cursor wants a rules directory, not the deprecated single file. We ship what each one reads.
Bring the idea you can't stop thinking about
We'll show you the strongest case against it. If it still stands, you leave with a spec and the first hundred users' addresses.
cancel any time · credits do not roll over · no free tier