A practical, science-grounded framework for the questions enterprise buyers, investors, and regulators are already asking
AI mental health chatbots are moving from pilot to scale faster than the evaluation standards meant to govern them. They are entering a broader mental health app market with more than 10,000 titles and a projected value of $9.45 billion in 2026. Research shows general-purpose AI tools are already being used for mental health support by teens and young adults with no clinical review at all. A 2025 JAMA Network Open study found that 13.1% of U.S. adolescents and young adults use generative AI for mental health advice, rising to 22.2% among 18-to-21-year-olds. This represents roughly 5.4 million people, most without any clinical oversight involved.
Purpose-built AI mental health chatbots and apps are competing for enterprise and health system contracts in a market where buyer standards are tightening fast, and regulators are moving at the same time.
In June 2026, Vermont enacted Act 156, which prohibits corporations and entities from offering mental health services to the public through AI unless those services are provided by a mental health professional or as part of an approved IRB or privacy-board study. It follows Illinois, which passed similar restrictions in 2025, and it lands alongside a Pennsylvania lawsuit against Character.AI for allegedly operating a chatbot that represented itself as a licensed psychiatrist.
The pattern across states is the same: regulators are drawing a hard line between AI that supports a clinician and AI that replaces one. For founders, that means engagement metrics and a good product story no longer clear the bar on their own. Evidence and legal compliance are no longer separate problems, and treating them that way is how a company falls behind one that treats this as part of the same operating challenge.
In May 2026, Common Sense Media released a risk assessment of AI mental health apps in partnership with Stanford Medicine’s Brainstorm Lab. One of the most widely used consumer AI therapy apps scored an “unacceptable” risk rating. Two other apps disappeared from app stores mid-assessment, more than three million users without support or referrals to alternative care.. The report found that purpose-built therapy apps advertising clinical oversight were, in several cases, no safer than general-purpose chatbots like ChatGPT or Gemini.
Companies in this category that build continuous evaluation infrastructure before they need it get a head start because evidence takes time to generate and can’t be assembled after a regulator or buyer asks for it.
What the Research Says About General-Purpose AI in Mental Health Contexts
A 2025 study from Stanford’s Institute for Human-Centered AI, Carnegie Mellon, the University of Minnesota, and UT Austin evaluated five AI chatbots against clinical standards used to assess human therapists.
Results showed that the models responded inappropriately in roughly 20% of interactions, compared to a 7% rate for trained human therapists in the same evaluation. In one test scenario, a user described losing a job and asked for a list of tall bridges in New York. A widely used therapy chatbot responded with bridge heights instead of recognizing the statement as a possible suicide risk. That’s what “design mismatch” looks like in a transcript: a general-purpose model can sound fluent and still miss the one signal a mental health conversation actually depends on.
The Common Sense Media report’s most important finding isn’t just the low risk rating; it’s what caused it. Researchers ran more than 3,100 exchanges across five AI mental health apps, testing responses to 13 clinical and developmental conditions including anxiety, depression, eating disorders, psychosis, and suicidal ideation. The lowest-rated app failed to recognize signs of a genuine psychiatric emergency, responded to OCD symptoms in ways that could reinforce the condition, and allowed a user to exit a suicide crisis pathway with a single denial and no follow-up.
Across both general-purpose and purpose-built tools, researchers documented what they called “missed breadcrumbs”: clear signs of distress that the AI failed to detect because it got sidetracked by tangential details or defaulted to physical health explanations instead of recognizing a mental health condition.
General-Purpose AI and Purpose-Built AI Do Not Have the Same Evidence Requirements
Founders building AI mental health products often get lumped in with general-purpose AI being used off-label for emotional support. That comparison undersells the responsibility. General-purpose AI was never designed for clinical use, so its failures are a design mismatch. A purpose-built mental health product markets itself as clinically informed, and that claim raises the evidence bar considerably.
A general-purpose tool needs disclaimers and referral pathways. A purpose-built clinical AI product needs documented safety testing, evidence of efficacy in the population it serves, and a monitoring system that catches failure before it compounds. Responsible AI in this category cannot be a marketing label. It’s a documented, ongoing evaluation process a buyer or regulator can actually inspect.
Five Things to Evaluate Before Scaling an AI Mental Health Chatbot
Scaling without a structured framework doesn’t just create regulatory exposure. It creates a trust problem that compounds every time the product reaches a new user population. Before scaling, founders should have documented, defensible answers in five categories.
- Clinical safety: Does it recognize a crisis, and what happens next? Has the chatbot been tested against high-severity scenarios, not just typical use cases? Does it reliably recognize crisis language across the many indirect ways people actually disclose distress, rather than only the explicit statements it was trained to catch? Just as important: what happens after it recognizes a crisis? A defensible answer names a specific escalation path, a human reviewer, a crisis hotline handoff, an alert to a clinical team, not a generic disclaimer telling the user to seek help elsewhere.
- Evidence of efficacy: Start with feasibility testing, not just an engagement number. Is there data showing the product improves the outcomes it claims to improve, in a population that resembles the one it will actually serve? A randomized controlled trial is the gold standard for proving an intervention works, but it isn’t the right first step, and treating it as the only credible evidence is how companies delay meaningful evaluation for years. A well-designed feasibility study with a small, well-characterized sample answers a narrower, earlier question: can this be used safely and as intended by the population it’s built for. That’s a necessary step on the way to an RCT, not a lesser substitute for one. An impressive engagement number with no clinical grounding behind it answers neither question.
- User population fit: Was the evaluation run on people like your actual users? Was the evaluation conducted with people who reflect the age, severity, language, and clinical profile of your actual user base? A tool validated on a mild, non-clinical wellness population does not carry the same evidence into a population with elevated clinical need.
- Ongoing monitoring: What happens after launch? Real-time monitoring for drift, a defined escalation protocol, and a process for incorporating user feedback into product changes are not optional add-ons. Without them, a product that was safe at launch can quietly stop being safe as usage scales.
- Regulatory readiness: Can you document it? Can the company answer, in specific and documented terms, what data was collected, how consent was handled, and what an IRB or equivalent oversight body reviewed? Illinois and Vermont have already restricted AI-only therapy tools, as covered in Fit Minded’s earlier blog on AI governance, and more states have similar bills moving through committees. That trend is accelerating, not slowing down.
How Should Founders Build Evaluation Into an AI Mental Health Product?
Fit Minded’s peer-reviewed JMIR AI paper, Rethinking AI Workflows: Guidelines for Scientific Evaluation in Digital Health Companies, lays out a structure for what this looks like in practice. It’s an operating layer built into the product from the start, not a formality on the way to a bigger study. Four elements apply directly here:
- Start with stakeholder needs, not the model. Identify who it actually needs to work for (e.g., patients, clinicians, caregivers), and what their real pain points are before building or evaluating the tool. Skipping this step is how companies end up with a technically capable model nobody trusts enough to use as intended.
- Design the system to check its own work. This means human-in-the-loop oversight built into the workflow itself, not a policy on paper. A structured escalation process, a clinician review panel, or a defined handoff point to a human is what separates a monitored system from one that’s simply being watched after something goes wrong.
- Test in stages before scaling. Beta testing, feasibility testing, and pilot testing each answer a different question: usability, resource and population fit, and initial outcome viability. Skipping a stage doesn’t save time. It moves the risk of failure into a later, more expensive, more public stage.
- Treat evaluation as continuous. Real-time monitoring after launch, regular audits, and publishing findings, even through faster channels like implementation briefs rather than full peer review, keep a system’s safety claims current instead of frozen at launch.
The practical point is not that every company needs a large clinical trial before it can grow. It is that founders need a clear, staged plan for what they will test now, what they will monitor after launch, and what stronger evidence they will need as their product, claims, and user population expand.
What This Means for Founders Building AI Mental Health Products
Scaling without a structured evaluation framework creates regulatory exposure and a trust problem. Engagement numbers and a compelling product story are not enough on their own to move a pilot into a scaled deployment. The Common Sense Media report, the Stanford findings, and the wave of state legislation show why treating them as sufficient is no longer tenable.
The path forward has a sequence: assess stakeholder needs, build human oversight into the workflow itself, test in stages before scaling, document consent processes and any applicable IRB or other oversight requirements, treat monitoring as continuous rather than a launch-day checklist. Founders who build this in early aren’t just avoiding risk. They’re building the kind of evidence trail that buyers, investors, and regulators are increasingly asking to see, an advantage engagement metrics alone can’t buy.
For a direct conversation about what evaluation your specific product needs before it scales, book a discovery call with Fit Minded.
FAQ
What should founders evaluate before scaling an AI mental health chatbot?
Founders should have documented evidence in five areas before scaling: clinical safety testing against high-severity scenarios with a defined escalation path, efficacy data that supports the specific claims they make, evaluation in a population that reflects their actual users, ongoing monitoring infrastructure for drift and safety after launch, and regulatory documentation including IRB review where applicable.
How is a purpose-built AI mental health chatbot different from general-purpose AI?
General-purpose AI tools like ChatGPT were not designed for clinical use, so their limitations in mental health contexts are a design mismatch. A purpose-built AI mental health product markets itself as clinically informed, which raises the evidence bar considerably. Enterprise buyers, investors, and regulators increasingly expect purpose-built tools to meet a clinical evidence standard, not just a general safety disclaimer.
What did the Stanford study find about AI chatbots in mental health contexts?
A 2025 study from Stanford HAI, Carnegie Mellon, the University of Minnesota, and UT Austin found that AI chatbots responded inappropriately in about 20% of test interactions evaluating clinical scenarios, compared to a 7% rate among trained human therapists. Failures included encouraging delusional thinking and failing to recognize suicidal ideation.
What does feasibility testing mean for an AI mental health product, and how is it different from a randomized controlled trial?
Feasibility testing evaluates whether a product can be used safely and as intended by its target population, typically with a small, well-characterized sample before broader recruitment. A randomized controlled trial tests whether the intervention works against a comparison group and is typically a later step. Feasibility testing is a necessary early stage, not a lesser substitute for an RCT.
What new regulations affect AI mental health chatbots?
Vermont’s Act 156, enacted in June 2026, prohibits corporations and entities from offering mental health services to the public through AI unless the services are provided by a mental health professional or are part of an approved IRB or privacy-board study. It follows a similar law in Illinois and joins a growing list of state-level restrictions on AI-only therapy tools.
Why did Common Sense Media rate some AI mental health apps as unacceptable risk?
In its May 2026 assessment, Common Sense Media found that some AI mental health apps, including a consumer app with more than six million users, failed to recognize signs of psychiatric emergencies, responded to certain conditions in ways that could worsen them, and lacked reliable human oversight for escalation. The report found these failures occurred in both general-purpose and purpose-built mental health apps.
Know what your AI product needs before you scale.