• Skip to primary navigation
  • Skip to main content
  • Skip to primary sidebar

Techcouver.com

  • News
  • Events
  • Interviews
  • Thought Leadership
  • Jobs
  • About
    • Contact Us

Cortico Launches Benchmark to Test Clinical AI Safety

September 9, 2026 by Techcouver Newsdesk Leave a Comment

Vancouver healthcare technology company Cortico has launched a free, open benchmark designed to measure whether artificial intelligence models can safely support clinical decisions—not merely pass medical exams.

Called MedSafe-Dx, the benchmark evaluates how large language models respond when a patient may require urgent care, when reassurance could be dangerous, and when the available information warrants uncertainty.

Across 11 frontier models evaluated in its launch paper, Cortico found that even the strongest performers faced a significant trade-off between safety and usefulness. The model with the highest safety pass rate, GPT-5.2, passed 97.6% of cases but escalated 71% of routine cases. At the other end, the poorest performer missed 26 of 156 cases classified as urgent.

“High accuracy often masks dangerous overconfidence,” said Cortico CEO and co-founder Clark Van Oyen. “We built MedSafe-Dx to give clinicians and health systems transparent, verifiable safety metrics.”

MedSafe-Dx tests three behaviours: whether a model escalates potentially life-threatening cases, avoids falsely reassuring patients who may be at risk, and expresses appropriate uncertainty when symptoms are ambiguous.

The evaluation presented 250 simulated adult patient cases from the DDXPlus dataset to models developed by OpenAI, Anthropic, Google and DeepSeek. Rather than relying on another AI model to judge the responses, MedSafe-Dx uses deterministic rules, frozen datasets and standardized outputs intended to make results reproducible and auditable.

The benchmark revealed that diagnostic accuracy did not necessarily translate into safer recommendations. Gemini 3 Pro Preview recorded the highest Top-3 diagnostic recall at 87.2%, according to Cortico, but the lowest safety pass rate at 62.4%. Every model evaluated missed at least some cases categorized as requiring escalation.

The findings point to a difficult balance for healthcare AI developers. Models that escalate nearly every questionable case may avoid some dangerous misses, but excessive warnings can burden clinical resources and eventually be ignored—a problem already familiar to providers using electronic health record alerts.

Cortico has since expanded the live leaderboard to 12 models from six AI labs, adding models from Meta and xAI. The company has also made the benchmark’s code and dataset publicly available alongside a medRxiv preprint.

Cortico stresses that MedSafe-Dx is a comparative safety test, not a clinical validation study or evidence that any model is ready for deployment. Its cases are simulated, while escalation labels are derived from dataset severity ratings rather than assessments by practising clinicians.

Still, the company argues that health systems evaluating AI should demand safety testing that examines judgment, confidence and escalation behaviour—not rely solely on exam-style accuracy scores.

Founded in 2015, Cortico provides patient-engagement and healthcare-workflow automation technology to more than 600 clinics and thousands of providers across North America.

Filed Under: News Tagged With: Cortico Health

 
 

Reader Interactions

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Primary Sidebar

Stay Connected

  • Facebook
  • Instagram
  • LinkedIn
  • RSS
  • Twitter

Community Partners

About Us

Techcouver provides real-time reporting and analysis of emerging technology news in Vancouver and throughout British … READ MORE... about About Us

Copyright © 2026 Incubate Ventures | Calgary.tech · CleanEnergy.ca · Decoder.ca · Fintech.ca · Legaltech.ca · Techtalent.ca · | Privacy