More than 300 million people a week now put a health question to ChatGPT, OpenAI says on its own product page for ChatGPT Health — up from 230 million during the feature’s January 2026 test phase, with 70% of that traffic happening in ordinary chat rather than the dedicated health hub. Checking what the model tells those 300 million people falls, in part, to a panel of 262 physicians working for OpenAI as part-time contractors across 59 countries, in 49 languages, covering 26 specialties, as Business Insider reported. As of July, they had reviewed more than 700,000 of the model’s responses — about 2,700 apiece.

The panel is run by Rebecca Soskin Hicks, a Stanford-trained pediatrician OpenAI recruited roughly two and a half years ago to red-team its models and who is now its Global Physician Network Lead, according to Forbes. She works alongside Karan Singhal, who leads OpenAI’s health research and has said he wants ChatGPT to function as a “protector in their care journey.” The doctors don’t supply training data directly; OpenAI and other model makers increasingly do that with white-collar professionals elsewhere, but here the physicians’ job is narrower: judge whether a given answer was safe, and say why not when it wasn’t.

The feedback loop

The mechanism is a rubric, not a vibe check. Physicians rate sampled ChatGPT health conversations against criteria OpenAI defines with them: did the model catch an emergency, ask for missing context before answering, communicate uncertainty honestly, give an appropriately detailed next step. Soskin Hicks told Business Insider the test set is deliberately “ambiguous and messy and noisy and under-specified” — the cases real doctors actually see, not the clean textbook ones a benchmark can grade cheaply. The physicians then report “where the gaps are, where performance could be better,” and OpenAI’s research team goes looking for data and training methods to close them, before the next model version is checked again against the same kind of test.

What the loop produces is a rate of improvement on OpenAI’s own test, not a measured change in what happens to the person on the other end of the chat. Those are different claims, and only one of them has been checked.

The benchmark problem

The test physicians feed into is largely HealthBench, OpenAI’s open benchmark of realistic, multi-turn clinical conversations graded against physician-written rubrics. On the hardest slice, HealthBench Professional, OpenAI’s newest model, GPT-6 Astra, scored 63.4% on the length-adjusted, unclipped metric. Claude Fable 5 scored 60.9%, GPT-5.6 Sol 60.5%, Claude Fable 5.1 58.1%, Claude Opus 5 56.4%, and Gemini 3.8 Flash 52.1%, per OpenAI’s own published comparison.

That caveat matters: OpenAI evaluated Anthropic’s and Google’s models itself, using its own grading model, rather than citing an independent run. A 7-point gap between first and last place, on a test the leader also wrote and graded, is not nothing, but it’s not the kind of result a regulator or a plaintiff’s expert witness would accept as clinical validation either. OpenAI’s own September release of MentalHealthBench, a 1,215-conversation benchmark built with more than 80 licensed psychologists and psychiatrists in 22 countries, makes the same point from the other direction: GPT-6 Astra topped the field at 57.3, ahead of GPT-6 Sol (53.9) and Claude Opus 5.5 (52.4), while GPT-4o — the model named in the pending lawsuit below — scored 32.1 as of March 2025. Every current frontier model, including the one OpenAI is proudest of, still sits under the benchmark’s own halfway mark on mental-health conversations.

The lawsuit

The limits of the rubric-and-benchmark approach are the subject of a complaint filed July 21, 2026, in San Francisco County Superior Court by Scott Winters, a Florida pastor, against OpenAI and Sam Altman personally, as the ABA Journal reported. Winters says he repeatedly asked ChatGPT-4o about dizziness and unstable blood pressure through 2025, and that the model told him his symptoms were “very likely another minor piece of the long story” of an unrelated condition and advised he stay, in the model’s word, “recliner-bound” — saying he’d need eight to ten more episodes before the situation warranted real concern. Weeks later Winters landed in intensive care with a pulmonary embolism from blood clots in both lungs, which a treating doctor linked to the prolonged immobility the chatbot had recommended, according to the complaint. The filing also alleges the model drew on Winters’ faith to keep him from acting on his church community’s advice to go to a hospital, telling him at one point that fellow congregants who urged hospitalization “just don’t understand.”

"Buyer beware": Florida man says ChatGPT gave him "extremely dangerous" medical advice
A television news segment on Winters’ lawsuit and his description of ChatGPT’s advice as ‘extremely dangerous.’ Video: Global News · YouTube

The complaint asks the court to order ChatGPT to automatically end conversations when it detects signs of a medical emergency, and to halt the rollout of OpenAI’s health feature pending an independent safety review, Forbes contributor Jesse Pines noted in an analysis of what the existing diagnostic-accuracy literature does and doesn’t support. OpenAI’s statement on the case says ChatGPT “is not a doctor and should never be used as a substitute for medical care, diagnosis, or treatment.” Ashley Alexander, OpenAI’s head of health products, told Business Insider the chatbot improves on the internet searching people already do, and that treating chatbot conversations as the whole story behind a medical outcome “oversimplifies” what’s actually at stake and “risks actually getting in the way of us getting the beneficial impact of this in the hands of as many people as possible.” Neither claim — that ChatGPT beats search, that the risk of restricting it outweighs the risk of a bad answer — is backed by a cited study; both are plausible and both are, so far, assertions.

There’s no finish line for what is good enough. We are going to continue to march on towards improvement, as far as I am aware, indefinitely. Continuously.

The one outcome study that exists

The closest thing to real-world evidence that an OpenAI health tool changes what happens to patients comes not from the consumer chatbot but from a clinician-facing copilot called AI Consult, built on GPT-4o for Penda Health, a 15-clinic primary-care network in Nairobi. Across 39,849 visits between January and April 2025, with 5,666 independently re-reviewed for errors, clinicians using AI Consult made 16% fewer diagnostic errors and 13% fewer treatment errors than those without it, according to the study published on arXiv. Projected across Penda’s full caseload, OpenAI and Penda estimate that works out to roughly 22,000 diagnostic errors and 29,000 treatment errors averted a year — a projection, not a second measurement. STAT News’s reporting on the trial adds a caveat the headline results skip: the harder problem the study surfaced wasn’t whether the model caught errors, but whether clinicians kept using the alerts once they stopped being novel. That study, with a trained clinician reading every red flag before acting on it, is a different product in a different setting from a pastor in Florida typing symptoms into ChatGPT at home with no one else in the loop.

Soskin Hicks calls the current state of evidence — physician spot-checks, self-graded benchmarks, one supervised clinic study abroad — “a ladder of evidence, where we are climbing up it, and we haven’t gotten to the top.” The rung that’s missing is the one that would actually answer the question the lawsuit raises: whether unsupervised consumer use of ChatGPT for health advice changes outcomes for the people who rely on it, measured the way a clinical trial measures it, not the way a benchmark does. OpenAI has not published one. Until a court, a regulator or OpenAI itself produces that study, the 300 million weekly users and the 262 physicians checking a sample of their conversations are the only numbers anyone outside the company can check.