Methodology

How we compare AI scribes and healthcare AI assistants

This page records how the figures in our comparison tables are produced: what was measured, where it came from, when it was read, and what it does not prove. Twenty-two products, covering ambient AI medical scribes and the healthcare and general AI assistants used alongside them. BastionGPT appears in both groups because it is both.

Ratings

A rating is the mean of a product’s scores across the review sites where it holds a published rating, read 18 July 2026. Each cell below shows that site’s own score and review count, and links to the page it was read from, so every average can be checked against its parts. The rules are set out beneath the tables.

AI medical scribes

ProductAverageof 5Apple App StoreCapterraG2GetAppGoogleGoogle PlaySoftware AdviceTrustpilotTrust­Radius
Abridge4.204.245listedno rating4.54··3.921·listedno rating·
Ambience Healthcare··listedno rating······
BastionGPT4.88·4.84524.8455·4.84··
Commure4.414.9155·4.416··3.9212·listedno rating·
DeepCura4.75·4.8154.6134.815··4.815··
DeepScribe4.204394.384.1264.38··4.38·listedno rating
Dragon Copilot4.204.114·4.326··4.813··7.2/1035
Freed AI4.804.85155254··4.4~100···
Heidi Health4.154.72.4Klistedno rating53listedno rating·3.7~350listedno rating3.2490·
Mentalyc4.56·4.5614.854.561··4.5614.570listedno rating
Nabla3.903.98·listedno rating····listedno rating·
Suki AI2.734.165212.5121·3.83321·listedno rating
Tali4.60·4.621·······
Twofold Health4.934.8525153··listedno rating·listedno rating·
Upheal4.433.712listedno rating52··listedno rating·4.674·

Healthcare and general AI assistants

ProductAverageof 5Apple App StoreCapterraG2GetAppGoogleGoogle PlaySoftware AdviceTrustpilotTrust­Radius
BastionGPT4.88·4.84524.8455·4.84··
ChatGPT3.77·4.43804.62,754····1.63,2809/10640
Claude3.75·4.3464.6395····1.51,6729.2/10236
CompliantChatGPT·········
Doximity GPTnetwork4.8 · 195K·network4.6 · 9··network4.8·network3.0 · 5·
Hathr AI·listedno ratinglistedno ratinglistedno rating··listedno rating··
Microsoft Copilot3.93·4.5274.43584.527··4.5271.63928.1/10187
OpenEvidence4.174.911K·52··4.94.36K·1.926·

listed — a profile exists with no published rating. Not a zero; excluded from the mean. network — the figure belongs to the parent Doximity network app, not to Doximity GPT; shown for provenance, excluded from the mean. A dot means no profile exists on that site.

How the average is calculated

  • Each site counts once, and counts equally. The mean is across sites, not across reviews, so a site holding 3,000 reviews does not outweigh one holding 12. Capterra, GetApp and Software Advice are counted as three separate sites.
  • TrustRadius scores are out of 10. They are shown above on their native scale, marked /10, and halved before entering the mean.
  • Unclaimed profiles count. A vendor not having claimed a page is not a reason to drop an unflattering score. This applies to us as well.
  • Consumer complaint pages count, but are not comparable to B2B reviews. The low Trustpilot scores for Claude, ChatGPT, Microsoft Copilot and OpenEvidence are billing and access complaints from consumers, not assessments by the clinicians who buy these tools. They are counted for consistency and flagged here rather than removed.
  • Tali was read on 26 July 2026, not in the 18 July sweep. Google Play counts marked ~ are approximate.

A rating built on few reviews moves on a single opinion. Several products here rest on a handful of reviews, ours among them. Note also that Capterra, GetApp and Software Advice syndicate a single review pool, so where a product appears on all three the same reviews are represented in each of those columns.

Sign-in security

A security grade is an MDN HTTP Observatory scan of the host where clinicians sign in, run 26 July 2026. Every grade links to its own scan, so you can re-run it and see today’s result rather than ours. What the grade does and does not cover is set out beneath the tables.

AI medical scribes

Healthcare and general AI assistants

BastionGPT Professional signs in at secure.bastiongpt.com; Professional Plus at plus.bastiongpt.com, which scores the same B/75.

Why the sign-in host

Marketing sites carry analytics, advertising and chat tags that require loosening exactly the headers Observatory tests, so grading them measures a tag manager rather than a product. The host behind the login is where credentials and patient information are typed. Where a vendor does not publish its application hostname, we resolved the conventional one and confirmed it serves a real sign-in application before scoring it. The host scored is printed under every grade.

What the grade measures

HTTP response headers only: HTTP Strict Transport Security, Content Security Policy, X-Frame-Options, X-Content-Type-Options, Referrer-Policy, cookie flags and subresource integrity. Scores run 0–100, with bonus points above 100.

What it does not measure

  • Not how any vendor stores, encrypts or handles patient data. Nothing in this scan sees inside an application.
  • Not a HIPAA assessment, and no substitute for a BAA or a vendor risk review.
  • Not TLS configuration, authentication design, access control or application code.
  • A weak grade does not mean a weak product, and an A+ proves nothing about the product behind it.

It earns a column for one reason: it is among the very few security signals a buyer can verify independently, in a browser, without taking any vendor’s word for it — including ours. Treat it as header hygiene, not a verdict.

Pricing

Prices were read from each vendor’s public pricing page in July 2026, and every price in the comparison tables links to the page it was read from.

Five of the ten platforms in the scribe comparison publish no price at all — Abridge, Ambience Healthcare, DeepScribe, Dragon Copilot and Suki AI each route buyers to a sales conversation. Those are marked “Custom”, and any figure beside them is a third-party estimate rather than a quote. Implementation fees, reseller markups, seat minimums and annual prepayment terms sit on top of whatever is eventually quoted and are not visible to us. Prices are per user per month at the entry tier, before annual discounts.

What none of this proves

  • No head-to-head accuracy testing. Nothing here measures note quality, transcription accuracy or clinical safety.
  • Ratings measure reviews, not products. They reflect who chose to write, on which site, at which moment. Sample sizes in this category are small, and ours is one of the smallest.
  • Security grades measure HTTP headers, nothing more.
  • Feature comparisons reflect what vendors publish. Where a vendor documents nothing we record it as not published, which is different from absent. Enterprise contracts routinely include capabilities that never appear on a public page.
  • We are not a neutral party. BastionGPT is one of the products in these tables. The same rules are applied to us — our unclaimed profiles counted, our thin review count shown, our middling security grade published, and competitors that beat us named — but read our comparisons knowing who wrote them, and check the sources we link.

If a figure here is wrong, tell us and we will correct it. That includes a price we misread, a review profile we missed, an application hostname we resolved incorrectly, or a capability we recorded as unpublished that is in fact documented.