AI Sandbox: From Idea to Validated Demand
How a beginner can validate a product idea with real people - fast, cheap, and without writing a single line of code. The core idea: go talk to users and test real demand through behavior, not your own assumptions.
Where failure starts
The most common beginner mistake: you come up with an idea, fall in love with it, and go build. Three months later you find out you invented the problem, and nobody wants the product.
The right order is the opposite. First you prove the problem is real and actually hurts. Then that people are willing to pay something for a solution: money, time, an email address. Only then do you build.
One rule that changes everything:
Words lie, behavior doesn't. "Cool idea, I'd use that" is not validation - it's politeness. Validation is when someone is already struggling with the problem, already spending money or time on it, and takes a real step toward your solution.
Your job in this phase isn't to convince yourself the idea is good - it's to try to kill it cheaply. If it survives, you dig deeper.
Mini glossary
For the validation phase (this part):
- Hypothesis - a guess you could be wrong about, phrased so it can be tested.
- Validation - testing a hypothesis against real people and their behavior, not your own opinion.
- Problem interview - a conversation where you study someone's life and pain and do NOT pitch your idea.
- The Mom Test - a set of rules for asking questions so that even your mom couldn't lie to you out of politeness.
- Demand signal - proof of interest through behavior (left an email, paid), not words.
- Fake door - a "buy/sign up" button that leads to "launching soon, leave your email." You measure how many click it.
- Concierge MVP - you do the work for your first customers by hand, with no product, and charge for it.
- Pivot - a turn: you change the core hypothesis when the data says the current one is dead.
- RAT (riskiest assumption test) - the most dangerous assumption, the one that kills everything if it's wrong. Test it first.
- MVP (minimum viable product) - the smallest version of the product you can already show a customer.
For deep research (last part):
- TAM / SAM / SOM - market size: total / addressable to you / realistically capturable.
- White space - a gap in the market: the pain exists, but there's no decent solution.
- Core Job / JTBD - the main "job" a customer hires the product to do.
- CAC / LTV - how much it costs to acquire a paying customer / how much they bring in over their lifetime.
- Churn / Retention - customer loss / customer retention.
Step 1. Write your hypothesis (10 minutes)
Not "I'm building an app for X." A testable statement using this template:
I believe that [specific segment] suffers from [specific pain], currently deals with it by [current workaround], and it hurts enough that they'd pay or spend time on [my solution].
Example: "I believe that freelance sole-proprietor bookkeepers lose 3-4 hours a week manually reconciling payments, do it in Excel, and it's annoying enough to pay 20 euros a month for automation."
The narrower the segment, the easier it is to test. "All small businesses" - you can't test that. "Freelance sole-proprietor bookkeepers in Poland" - you can test that in a week.
AI prompt - sharpen your idea into hypotheses:
I'm a beginner building a product. Here's my idea in rough words:
[dictate or describe your idea in your own words]
Help me turn it into 3 testable hypotheses using this template:
"I believe that [segment] suffers from [pain], currently deals with it by [workaround],
and it hurts enough that they'd [action or payment]."
Make the segment as narrow and specific as possible. For each hypothesis,
name the one assumption that's the riskiest -
the one without which everything falls apart.
Step 2. Find the riskiest assumption (RAT)
Every hypothesis rests on several assumptions. One is the most fragile. For a beginner, it's almost never "can I build this." It's "does this pain even exist, and does it hurt enough that people would pay for a fix."
Test first whatever (a) kills the project hardest if it's wrong, and (b) is cheapest to test. Don't write code until you're sure the problem is real.
AI prompt:
Here's my hypothesis: [paste it].
List every assumption it depends on.
Rank them by "how much it kills the project if wrong" multiplied by
"how UNCERTAIN I am that it's true."
Name the top-1 to test first and the cheapest way to test
it within a week, without code.
Step 3. Go talk to users: problem interviews (the heart of it all)
This is the most important skill and the most common mistake. A beginner goes to people and asks "do you like my idea?", gets "yeah, cool", gets excited, builds, and fails. People lie out of politeness.
Rules for the conversation (this is The Mom Test - questions people can't lie to):
- Talk about THEIR life and past, not your idea. Don't pitch at all.
- Ask about specific past behavior, not the future. "When did you last run into this? What did you do?" - not "would you use...?".
- Dig into the numbers behind the pain: how much time it takes, how much money, how often.
- Look for what they're ALREADY doing to solve the problem: workarounds, Excel, other software, hired someone. If they're doing nothing - the pain probably isn't real.
- Compliments aren't data. Data is emotion - "god, this drives me crazy" - and real money already spent.
Signs the problem is real:
- They describe the pain themselves, without you leading them.
- They're already spending money or time on workarounds.
- They get emotional when they talk about it.
How many conversations: 5-8 per segment. Fewer is just luck. More isn't needed at the start.
AI prompt - where to find these people and how to reach out:
My segment: [who exactly, as specific as possible].
1. Where do these people hang out online and offline - specific communities,
subreddits, chats, groups, events. Give links.
2. Write a short, non-salesy message inviting them to a
15-minute chat about their workflow (not about my product).
Goal - they agree to talk without feeling like they're being sold to.
AI prompt - interview script:
Help me put together a 15-minute problem interview guide following
The Mom Test rules. Segment: [who]. Hypothesis: [paste it].
Requirements:
- Questions only about past and current behavior, not about my idea.
- No pitching: I don't mention my solution until the very end.
- 6-8 open-ended questions digging into pain, frequency, current workarounds,
and money and time spent.
- At the end - how to gently ask "who else should I talk to".
AI prompt - break down the interviews (after the conversations):
Here are notes and transcripts from [N] interviews: [paste them].
1. What pain comes up repeatedly across most of them? Back it up with quotes.
2. What are they ALREADY doing to solve it, and what does it cost them?
3. Where did my hypotheses hold up, and where did they fall apart - be honest.
4. Separate real demand signals from polite compliments.
5. Verdict: dig deeper, pivot the hypothesis, or drop it?
Step 4. Test real demand (through behavior, not words)
People said "the problem is real." Good. But they'll say one thing and do another. Now you test whether they're willing to ACT. Still no code.
The ladder of signals, from weakest to strongest. The higher up, the more honest:
- "Interesting, cool" - zero. Throw it out.
- Left an email on a waitlist - weak signal.
- Spent time: filled out a form, came to a demo, let you do the work by hand for them - medium.
- Paid, made a prepayment or deposit - strong. Money doesn't lie.
- Referred a friend, asks "when is this launching already" - very strong.
Cheap ways to test demand without a product:
- Landing page + waitlist. A one-pager with an "I want this" button. Drive a bit of traffic (posts in the same communities, 20-50 euros of ads) and see if people leave their email.
- Fake door. A "buy/sign up" button that leads to "launching soon, leave your email." You measure clicks.
- Concierge. For your first 3-5 customers, do the work by hand, with no automation, and charge for it. If they pay for the manual result - demand is real.
- Pre-order or deposit. The most honest test. If they're willing to put down even a small prepayment - that's a real "yes".
Set the bar in ADVANCE: "success = 10 out of 100 visitors leave an email" or "3 out of 5 pay for the manual version." Otherwise, any result will look good in hindsight.
AI prompt - landing page copy for the test:
Segment: [who]. Pain: [what it is, in the customer's own words from the interviews].
Write copy for a one-page landing page to test demand:
- a headline about the pain, not the features;
- 3 value bullets;
- one clear call to action (leave an email on the waitlist).
No fluff, no buzzwords. Give me 3 headline options for an A/B test.
AI prompt - plan a demand test:
I want to test demand for [idea] without writing code, within 1-2 weeks,
budget up to [amount].
Suggest 3 concrete test methods (landing page, fake door, concierge,
pre-order). For each: exactly what steps I take, what signal
counts as success (with a number), and where to find my first people.
Step 5. Decide: continue, pivot, or stop
Look at the data honestly:
- The problem is confirmed + there's a demand signal (people pay or act) - dig deeper (section below) and start building your MVP.
- The problem exists, but there's no demand or willingness to pay - maybe it's the wrong pain or the wrong segment. Pivot: change the segment or how you frame the pain, and test again.
- No pain, no demand - drop the idea. This isn't a failure, it's three months saved. Take the next hypothesis.
Decide based on the data, not on how much you've fallen in love with the idea.
Dig deeper (after you've tested with real people)
This is the next level. Do it AFTER you've confirmed the problem and demand with real people. This is where AI deep research saves days. But remember: this is desk research - it builds context, it doesn't replace talking to users.
Engines: ChatGPT Deep Research, Gemini Deep Research (depth and sources), Perplexity (fast, pricing and competitors). Run the same prompt through two of them and compare - discrepancies show where the model is hallucinating. And demand sourced links for every number.
Market size (TAM / SAM / SOM):
Role: market analyst. Product context: [brief].
Calculate TAM, SAM, SOM using two methods:
1. Top-down: from the size of the whole industry down to my segment.
2. Bottom-up: number of customers x average deal size x frequency.
For every number - give a source and year; if there's no source, mark it "assumption".
Give a range (pessimistic / base / optimistic) and the 3 main assumptions.
Competitors and their tools:
Role: competitive intelligence analyst. Context: [brief].
1. Find 10-15 players (including "do it in Excel" and "do nothing").
2. For the top 7 - a table: product, segment, key features, price, positioning.
3. 3-5 recurring complaints from reviews (G2, Capterra, Reddit) with links.
4. 2-3 white space opportunities: where the pain exists but there's no solution.
Pull prices from current pricing pages with the date you checked.
Benchmarking:
Role: product benchmarking analyst. Context: [brief] + category: [...].
1. 3-5 best-in-class products in adjacent categories - who to learn from
for onboarding, pricing, retention.
2. Industry benchmark metrics (trial-to-paid, churn, CAC/LTV) with sources.
3. 5 lessons I can apply already at the MVP stage.
Give metrics as ranges with a year. Separate enterprise and SMB.
Advanced level - segmentation by Advanced JTBD. Once the problem is confirmed, segment not by demographics ("startups with 10-50 people") but by jobs ("who needs to get X done, and the current solution is annoying"). Break the Core Job down into sub-jobs and look for underserved nodes - wherever competitors are weak is your entry point. This is generative analysis, not a web-search task: deep research gives you raw material (customer language, current solutions), but you build the segmentation yourself and refine it with live interviews.
Advanced level - work in a project folder, not in the browser. Set up a folder and work through Claude Code or Cursor (IDE - integrated development environment). Put your brief incontext.md, and save each step as a file (interviews.md,demand_test.md,competitors.md). Context accumulates, and that same research later turns into a PRD (product requirements document) and code - without losing context. This is the shift from "AI = chat in a browser" to "AI = working tool".
Checklist: before you write code
- ☐ Written your hypothesis using the template, with a narrow, specific segment?
- ☐ Talked live with 5+ people from the segment?
- ☐ Did they describe the pain themselves, without you leading them?
- ☐ Are they ALREADY spending money or time on workarounds for this problem?
- ☐ Got at least one behavioral demand signal (email, payment, deposit), not just compliments?
- ☐ Set the success bar IN ADVANCE, before the test?
- ☐ Ready to honestly drop or pivot if the data says no?
One last rule: AI speeds up prep and analysis massively - it finds people, writes scripts, synthesizes interviews. But it can't replace the actual conversation with a user or a real demand test. That's the work founders get paid for. Don't delegate it.