Methodology

How we compare AI scribes and healthcare AI assistants

This page records how the figures in our comparison tables are produced: what was measured, where it came from, when it was read, and what it does not prove. Twenty-two products, covering ambient AI medical scribes and the healthcare and general AI assistants used alongside them. BastionGPT appears in both groups because it is both.

Ratings

A rating is the mean of a product’s scores across the review sites where it holds a published rating, read 16 August 2026. Each cell below shows that site’s own score and review count, and links to the page it was read from, so every average can be checked against its parts. The rules are set out beneath the tables.

AI medical scribes

ProductAverageof 5Apple App StoreCapterraG2GetAppGoogleGoogle PlaySoftware AdviceTrustpilotTrust­Radius
Abridge4.204.245listedno rating4.54··3.921·listedno rating·
Ambience Healthcare··listedno rating······
BastionGPT4.84·4.844.844.8455·4.84··
Commure4.414.9155·4.416··3.9212·listedno rating·
DeepCura4.75·4.8154.6134.815··4.815··
DeepScribe4.204394.384.1264.38··4.38·listedno rating
Dragon Copilot4.204.114·4.326··4.813··7.2/1035
Freed AI4.804.85155254··4.4~100···
Heidi Health4.154.72.4Klistedno rating53listedno rating·3.7~350listedno rating3.2490·
Mentalyc4.56·4.5614.854.561··4.5614.570listedno rating
Nabla3.903.98·listedno rating····listedno rating·
Suki AI2.734.165212.5121·3.83321·listedno rating
Tali4.60·4.621·······
Twofold Health4.934.8525153··listedno rating·listedno rating·
Upheal4.433.712listedno rating52··listedno rating·4.674·

Healthcare and general AI assistants

ProductAverageof 5Apple App StoreCapterraG2GetAppGoogleGoogle PlaySoftware AdviceTrustpilotTrust­Radius
BastionGPT4.84·4.844.844.8455·4.84··
ChatGPT3.77·4.43804.62,754····1.63,2809/10640
Claude3.75·4.3464.6395····1.51,6729.2/10236
CompliantChatGPT·········
Doximity GPTnetwork4.8 · 195K·network4.6 · 9··network4.8·network3.0 · 5·
Hathr AI·listedno ratinglistedno ratinglistedno rating··listedno rating··
Microsoft Copilot3.93·4.5274.43584.527··4.5271.63928.1/10187
OpenEvidence4.174.911K·52··4.94.36K·1.926·

listed — a profile exists with no published rating. Not a zero; excluded from the mean. network — the figure belongs to the parent Doximity network app, not to Doximity GPT; shown for provenance, excluded from the mean. A dot means no profile exists on that site.

How the average is calculated

  • Each site counts once, and counts equally. The mean is across sites, not across reviews, so a site holding 3,000 reviews does not outweigh one holding 12.
  • TrustRadius scores are out of 10. They are shown above on their native scale, marked /10, and halved before entering the mean.
  • Consumer complaint pages count, but are not comparable to B2B reviews. The low Trustpilot scores for Claude, ChatGPT, Microsoft Copilot and OpenEvidence are billing and access complaints from consumers, not assessments by the clinicians who buy these tools. They are counted for consistency and flagged here rather than removed.

Sign-in security

A security grade is an MDN HTTP Observatory scan of the host where clinicians sign in, run 26 July 2026 and re-checked 16 August 2026. Every grade links to its own scan, so you can re-run it and see today’s result rather than ours. What the grade does and does not cover is set out beneath the tables.

AI medical scribes

Healthcare and general AI assistants

Why the sign-in host

Marketing sites carry analytics, advertising and chat tags that require loosening exactly the headers Observatory tests, so grading them measures a tag manager rather than a product. The host behind the login is where credentials and patient information are typed. Where a vendor does not publish its application hostname, we resolved the conventional one and confirmed it serves a real sign-in application before scoring it. The host scored is printed under every grade.

What the grade measures

HTTP response headers only: HTTP Strict Transport Security, Content Security Policy, X-Frame-Options, X-Content-Type-Options, Referrer-Policy, cookie flags and subresource integrity. Scores run 0–100, with bonus points above 100.

What it does not measure

  • Not how any vendor stores, encrypts or handles patient data.
  • Not a HIPAA assessment, and no substitute for a BAA or a vendor risk review.
  • Not TLS configuration, authentication design, access control or application code.
  • A weak grade does not mean a weak product, and an A+ proves nothing about the product behind it.

It earns a column for one reason: it is among the very few security signals a buyer can verify independently, in a browser, without taking any vendor’s word for it — including ours. Score reports from BitSight and SecurityScorecard could not be included here because of licensing restrictions.

Pricing

Prices were read from each vendor’s public pricing page on 16 August 2026, and every price in the comparison tables links to the page it was read from.

Five of the ten platforms in the scribe comparison publish no price at all — Abridge, Ambience Healthcare, DeepScribe, Dragon Copilot and Suki AI each route buyers to a sales conversation. Those are marked “Custom”, and any figure beside them is a third-party estimate rather than a quote. Implementation fees, reseller markups, seat minimums and annual prepayment terms sit on top of whatever is eventually quoted and are not visible to us. Prices are per user per month at the entry tier, before annual discounts.

What none of this proves

  • No head-to-head accuracy testing. Nothing here measures note quality, transcription accuracy or clinical safety.
  • Ratings measure reviews, not products. They reflect who chose to write, on which site, at which moment. Sample sizes in this category are small, and ours is one of the smallest.
  • Security grades measure HTTP headers, nothing more.
  • Feature comparisons reflect what vendors publish. Where a vendor documents nothing we record it as not published, which is different from absent. Enterprise contracts routinely include capabilities that never appear on a public page.

If a figure here is wrong, tell us and we will correct it. That includes a price we misread, a review profile we missed, an application hostname we resolved incorrectly, or a capability we recorded as unpublished that is in fact documented.