Outcome Measure Review Note: Definition, Template & Example

An outcome measure review note is the documentation entry of measurement-based care: it records which standardized instrument was administered, the score, the change from baseline, the clinical interpretation, the discussion with the client, and the treatment decision that follows. Therapists, psychologists, and collaborative care teams write one each time a measure like the PHQ-9 or GAD-7 is given. Most run 75 to 300 words beside a score flowsheet.

Free to use and share. No signup required.
Already have session bullets or a transcript? Generate a structured draft with BastionGPT — you review and sign it.
Who writes it

Therapists, psychologists, counselors, and collaborative care teams

Audience

The clinician and client first; the care team, supervisors, payers, and auditors on request

Typical length

75 to 300 words plus the score flowsheet · 3 to 10 minutes by hand (clinical team estimate)

Format family

Structured flowsheet plus short narrative (also called an MBC note, ROM note, or progress monitoring note)

When it's used

Every administration of a standardized measure: baseline at intake, every 2 to 4 weeks in many programs, and before each treatment plan review

Standards context

Accreditors, some payer codes, and Australian plan rules mandate the measurement; no rule anywhere sets the note's format

What is an outcome measure review note?

An outcome measure review note is the documentation entry for measurement-based care: it records that a standardized instrument was administered, the score it produced, how that score compares with baseline and prior administrations, what the change means clinically, that the result was discussed with the client, and the treatment decision that followed. Measurement-based care itself was defined in a widely cited review as the systematic use of symptom rating scales to drive clinical decision making for the individual patient (Fortney et al., 2017), a practice built through the 1990s patient-focused research of Kenneth Howard, Michael Lambert's work on outcome feedback and not-on-track alerts, and the brief session-by-session scales of Scott Miller and Barry Duncan. The instruments carry precise pedigrees: the PHQ-9 was validated on 6,000 primary care and obstetrics patients (Kroenke, Spitzer & Williams, 2001), the GAD-7 followed in 2006 (Spitzer et al.), and the PCL-5 is public domain and takes 5 to 10 minutes (National Center for PTSD). Clinicians also call the entry an MBC note, a routine outcome monitoring (ROM) note, a feedback-informed treatment note, or a progress monitoring note.

Here is the fact that reframes every "required elements" list you will find online: the instruments are standardized; the note is not. No statute, CPT rule, accreditation standard, or professional college in the US, Canada, or Australia defines the fields, layout, or template of this note. Where requirements exist, they mandate the measurement, never the note format: the Joint Commission requires accredited behavioral health organizations to monitor progress with a standardized instrument (standard CTS.03.01.09, effective January 1, 2018), Medicare's collaborative care codes require validated rating scales and a registry, and Australia's MBS requires an outcome measurement tool when a GP prepares a mental health treatment plan. The gap between evidence and practice is the other defining fact: in a national survey of 504 clinicians, only 13.9% used standardized progress measures at least monthly and 61.5% never used them (Jensen-Doss et al., 2018). The note sits downstream of two neighbors: it supplies the data that a treatment plan review acts on, and it is often written as a labeled section inside the psychotherapy progress note rather than as a separate document.

Who uses outcome measure review notes and when

Any clinician who administers a standardized measure and wants the administration to count writes one: outpatient therapists and psychologists reviewing a PHQ-9 or GAD-7 the client completed through the portal, group practices running symptom scales on an every-2-to-4-week cadence, collaborative care teams whose billing codes require validated rating scales tracked in a registry (CMS MLN909432), Joint Commission accredited organizations meeting the standardized-instrument standard, child and adolescent services scoring the SDQ, and Australian clinicians working under a GP mental health treatment plan or in PHN-commissioned programs that report the K10+ into the national minimum data set. The cadence is the point: a baseline at intake, re-administration every few weeks, and a visible trend by the time the treatment plan review is due. Reach for the psychotherapy progress note when the session narrative is the deliverable and the score is one line inside it, and for the psychological evaluation report when the question calls for clinician-administered testing with interpretation and a written report, the way a full battery write-up like an MMPI-3 or BASC-3 report does. The outcome measure review note earns its place when a brief self-report scale needs to move treatment: score it, compare it, interpret it, discuss it, and change something.

Outcome measure review note structure: what goes in each section

Instrument, date, and administration context. Name the instrument and version, the administration date, and how it was completed: in session, through the portal beforehand, or in the waiting room. Method matters when scores shift with setting. Pitfall: "screen completed" with no instrument name; an unnamed measure supports neither the billing line nor the trend.

Score and scoring integrity. The total score and any subscales, plus who scored it. Brief measures are self-report, but the scoring and documentation are yours. Pitfall: transcribing the number wrong or leaving a half-completed form unscored; a score you cannot reproduce from the responses is a score an auditor cannot credit.

Trend against baseline and prior scores. The line that makes it measurement rather than paperwork: today's score next to baseline and the last two or three administrations, with dates. A flowsheet does this in one row per administration. Pitfall: a score with no comparison; a single number tells nobody whether treatment is working.

Clinical interpretation beyond the number. Severity range, direction and size of the change, and whether the score matches what you see in session. Scores and presentation can disagree; say so when they do. Pitfall: restating the number as its label ("score 11, moderate") with no clinical read; interpretation is the part only a clinician can supply.

The conversation with the client. That you shared the result, how you framed it, and the client's response. Feedback trials keep finding that measurement helps when results reach the clinical conversation, not the chart alone. Pitfall: no evidence the client ever saw the score; measurement-based care without the discussion is data collection.

Item-level flags and risk follow-up. Note any item that changes your plan for the session, and always the risk item: an elevated PHQ-9 item 9 gets a same-day risk conversation, documented with its outcome, linking the suicide risk assessment or safety plan when one follows. Pitfall: filing a positive risk item with no documented follow-up; that single omission turns a routine note into the chart's weakest page.

The treatment decision that follows. What the score changes or confirms: continue, adjust the intervention, change frequency, refer, or carry the target into the next treatment plan review, with one line of rationale linking score to decision. Pitfall: no linkage; a trend that never touches the plan is the pattern payers read as measurement theater.

Signature, date, and billing linkage. Sign and date the entry, and make the billing line match the documentation: one unit per instrument administered, scored, and documented when you report 96127. Pitfall: units billed for instruments the note never names, or a signature that would not survive an authentication check.

Blank template (copy and adapt)

OUTCOME MEASURE REVIEW NOTE

Client: _____________________   Date: ___________
Clinician: __________________   Session #: ______

INSTRUMENT AND ADMINISTRATION
Instrument and version: __________________________
Administered by / method: ________________________
  [ ] In session   [ ] Portal before session   [ ] Waiting room

SCORE
Total score: ________   Subscale scores (if any): ____________
Scored by: __________________

TREND VS BASELINE AND PRIOR SCORES
Baseline (date / score): _________________________
Prior administrations (date / score): ____________
__________________________________________________
Change since baseline: ________  Since last: ______

CLINICAL INTERPRETATION
Severity range and direction of change: __________
__________________________________________________
Consistency with presentation in session: ________
__________________________________________________

ITEM-LEVEL FLAGS AND FOLLOW-UP
Items reviewed / elevated: _______________________
Risk follow-up (if indicated) and outcome: _______
__________________________________________________

DISCUSSION WITH CLIENT
Scores shared and discussed: [ ] Yes  [ ] Declined
Client response: _________________________________
__________________________________________________

TREATMENT DECISION
  [ ] Continue current plan       [ ] Adjust intervention
  [ ] Frequency change            [ ] Referral / consult
  [ ] Goal or target update at next treatment plan review
Rationale linking score to decision: _____________
__________________________________________________
Next administration due: ___________

Clinician signature / credentials / date: ________

Free to use and share, no signup. The PDF includes a one-page cheat sheet with section-by-section pitfalls and a pre-sign checklist; the DOCX is the blank template, ready to adapt.

Sample outcome measure review note

Scenario: an outpatient psychologist reviews PHQ-9 and GAD-7 scores with a 34-year-old client at session 8 of weekly CBT for depression and anxiety, completed through the portal before the visit. Use it as a measurement-based care documentation example you can adapt to any instrument. All details are fictional.

OUTCOME MEASURE REVIEW NOTE
Client: M.T., 34  ·  Date: 07/22/2026  ·  Session: 8 of weekly CBT
Clinician: R. Alvarez, PhD, Licensed Psychologist

Instruments and administration. PHQ-9 and GAD-7, completed by client via secure portal 07/21/2026, evening before session. Both complete; scored automatically, verified by clinician.

Scores and trend.
06/03/2026 (baseline, intake): PHQ-9 16 · GAD-7 15
06/17/2026: PHQ-9 13 · GAD-7 13
07/01/2026: PHQ-9 11 · GAD-7 12
07/22/2026 (today): PHQ-9 8 · GAD-7 11
Change: PHQ-9 down 8 points from baseline; GAD-7 down 4.

Interpretation. Depression scores have moved from the moderately severe range at intake into the mild range and have now crossed below the 10 threshold, past the 5-point change our clinic flowsheet flags as meaningful improvement. The trajectory matches session presentation: brighter affect, resumed morning walks, back to full workdays. Anxiety is improving more slowly and remains in the moderate range, consistent with M.T.'s report that low mood has lifted while worry episodes continue most evenings. Sleep item on the PHQ-9 remains at 2, unchanged across all four administrations.

Item-level review. PHQ-9 item 9 scored 0 at every administration to date, consistent with client's denial of suicidal ideation in session today. No risk follow-up indicated.

Discussion with client. Reviewed the score graph on screen together. M.T.'s reaction: "the mood line finally looks like how I feel." Client expressed frustration that worry has not moved as far; we normalized the differential response and agreed the numbers support a shift in focus rather than a change in course.

Treatment decision. Continue weekly CBT. Beginning next session, shift emphasis from behavioral activation (goals substantially met) to the worry exposure module targeting evening rumination, and add the sleep log given the static sleep item. Both measures to be re-administered in two weeks (due 08/05/2026). Goal 2 target on the treatment plan will be updated at the scheduled treatment plan review on 08/12/2026 to reflect the revised anxiety focus.

Billing. 96127 x 2 units (PHQ-9, GAD-7) reported with today's psychotherapy session per payer contract; administration, scores, and review documented above.

Signature. R. Alvarez, PhD, Licensed Psychologist, signed 07/22/2026, 4:35 PM.

This sample is fictional and for educational purposes. It does not describe a real patient.

↑ Back to the template and downloads

Why this sample works

  • Every element an auditor would look for is present and linked: named instruments, administration method and date, verified scores, a dated trend, interpretation, client discussion, a treatment decision, and a billing line that matches what the note documents.
  • The trend is shown as a four-administration flowsheet with dates, so the 8-point PHQ-9 change is visible at a glance and the differential response between depression and anxiety becomes the clinical story.
  • Interpretation goes beyond the number: severity ranges, correlation with session presentation, and an item-level read (the static sleep item) that feeds directly into the plan.
  • The risk item is reviewed and documented even though it is negative, showing the habit that protects the chart on the day an item 9 comes back positive.
  • The client saw the scores, reacted to them, and shaped the decision, which is the difference between measurement-based care and data collection.
  • The decision is specific and dated: a module shift, a re-administration date, and a hand-off to the scheduled treatment plan review where the goal target changes.

Writing these after every session? BastionGPT drafts complete notes from bullets, dictation, or a transcript.

Generate a note from bullets

Documentation and compliance considerations

Completed instruments and their scores are ordinary clinical record material, and that cuts both ways. Clients can access them: the HIPAA right of access (45 CFR 164.524) reaches the designated record set, and scores cannot be tucked into psychotherapy notes because the psychotherapy-notes definition expressly excludes results of clinical tests and summaries of progress to date (45 CFR 164.501). Write every interpretation as a line the client may read. The same logic drives risk documentation: when an item-level response signals risk, the same-day conversation and its outcome belong in the record, with a suicide risk assessment or safety plan linked when one follows. In Canada no college requires outcome measures at all; Ontario's psychology and behaviour analysis college instead requires registrants to be "familiar with evidence-based tools and techniques" and to justify the decision when they choose not to use them (CPBAO Standards 2024, s. 10.4), while its retention standard keeps individual records at least 10 years after the last contact, or 10 years after a minor client turns 18. Australian records follow the whole-chart rule of 7 years from the last entry, or until age 25 for minors.

On the payer side, sort each rule into its bin, because here the usual line runs backwards: the measurement is the requirement; the note format is the convention. The Joint Commission's standardized-instrument standard binds accredited organizations (accreditation policy), Medicare's collaborative care and behavioral health integration codes build validated rating scales and a registry into the service itself (payer policy), Australia ties an outcome measurement tool to preparing and reviewing the GP mental health treatment plan unless clinically inappropriate (MBS note AN.0.56) and names specific instruments only in the PHN program data set (PMHC MDS v5.0: K10+, or K5 for Aboriginal and Torres Strait Islander clients, and the SDQ, at episode start and end at minimum), and US law reaches only coverage: screening recommendations graded A or B, including depression screening for adults and anxiety screening for adults 64 and younger (USPSTF, June 2023), must be covered without cost sharing by non-grandfathered private plans under 42 USC 300gg-13. For the billing line itself: 96127 carries no physician work value in the fee schedule, is reported per instrument with no time threshold, and its Medicare per-day cap is 3 units with an appealable date-of-service edit; psychotherapy already includes continuing evaluation under NCCI policy, so same-day reporting with therapy codes is payer-specific; and since 2024 marriage and family therapists and mental health counselors enroll with Medicare directly (42 CFR 410.53, 410.54). Medicare's annual depression screening benefit runs under its own code, G0444. When the score triggers a goal change, document it in the treatment plan review so the measurement chain and the medical-necessity chain meet on paper.

Common outcome measure review note errors auditors flag

The audit record here is old, specific, and still the operative precedent. HHS-OIG's review of Medicare Part B mental health services (OEI-03-99-00130) found 42 percent of psychological testing services inappropriate, almost half of them self-administered or self-scored screens such as the Beck Depression Inventory and Geriatric Depression Scale billed as testing when they belong inside the evaluation visit, and about a third of the inappropriate testing records carried no written interpretation of the results. The modern audits repeat the pattern at the treatment-plan layer: one New York provider's psychotherapy claims failed on all 100 sampled beneficiary days, largely on treatment plan and supervision requirements, for an estimated $1,118,789 in overpayments (A-02-21-01006), and another was billed $3.9 million across 23,947 claims with findings that included treatment plan gaps and signatures stamped as digital images (A-02-19-01012). The BastionGPT Clinical Advisory Board sees the same errors most often in outcome measure review note reviews:

  • Self-scored screens billed as psychological testing. A PHQ-9 or Beck inventory reported under the timed testing codes instead of the brief-assessment code. NCCI policy is explicit that the testing codes cannot be reported for self-administered tests, and this exact substitution is what the OIG flagged. Screens ride 96127 or the visit itself.
  • A score graveyard. Flowsheet cells fill up for months while the numbers never touch interpretation, the session narrative, or the plan. A trend nobody acts on documents that measurement changed nothing, which is worse at audit than no measurement at all.
  • No evidence the client saw the result. The feedback literature keeps finding that one-time screening and results delivered outside the clinical encounter do not move outcomes, and a reviewer reads a chart with scores but no discussion the same way. One sentence of client response fixes it.
  • Unit stacking past the cap. Four or five instruments reported on one date when the Medicare per-day edit for 96127 is 3 units, adjudicated per date of service; excess units need an appeal with documentation, not a modifier, and Medicaid programs publish their own separate edit tables.
  • Signatures and score integrity that fail authentication. Unsigned entries, image-stamped signatures like the 109 claims in the On-Site audit, or a recorded total that cannot be reproduced from the completed instrument. The score is evidence; treat its chain of custody accordingly.
How BastionGPT helps

BastionGPT is specifically trained, tuned, and clinically tested on outcome measure review notes.

  • Extract instrument names, dates, and scores from a dictation, session transcript, or pasted portal export, and return them as a dated flowsheet line ready for the chart.
  • Turn a scores flowsheet plus session bullets into the complete note: trend against baseline, interpretation in context, the client discussion, and the treatment decision, in your clinic's format.
  • Check the note before you sign: a score with no comparison or interpretation, no documented client discussion, a positive risk item without follow-up, or a billing line the documentation does not support.

See how clinicians use it day to day on the AI therapy notes page.

Many BastionGPT users report saving more than 90 minutes per day on documentation.

HIPAA-compliant with a signed BAA on every plan. Your data is never used to train models. BastionGPT drafts, you review and sign.

Frequently asked questions

No. No statute, CPT rule, accreditation standard, or professional college in the US, Canada, or Australia defines this note's fields or layout; every "required elements" list you will find online is a vendor synthesis. The requirements that exist mandate the measurement instead: the Joint Commission's standardized-instrument standard for accredited organizations, validated rating scales and a registry inside Medicare's collaborative care codes, and Australia's outcome-tool element in the GP mental health treatment plan. Most working notes run 75 to 300 words beside a score flowsheet and take 3 to 10 minutes, a clinical team estimate.

Six elements cover what payers and auditors expect to find: the instrument name and version, the administration date and method, the score, the comparison with baseline and prior scores, your clinical interpretation, and the treatment decision that follows, plus one line showing the result was discussed with the client. The discussion line is the one most often missing, and the feedback evidence says results that never reach the clinical conversation do not change outcomes.

Medicare's practitioner Medically Unlikely Edit for 96127 is 3 units per date of service in the current quarterly table, with adjudication indicator 3: a per-day edit, so units past the cap are denied and recovery runs through appeal with documentation rather than a modifier. Medicaid programs publish separate MUE tables of their own, and commercial limits are contractual. Bill one unit per standardized instrument administered, scored, and documented in the note.

Neither. The code is reported per standardized instrument, not per minute; the minute thresholds that circulate on billing forums belong to the timed testing codes. It also carries 0.00 physician work RVUs in the Medicare fee schedule, which is the technical way of saying no separately valued interpretation or report exists: the clinician's read of the score is part of the visit it accompanies. A brief score-plus-interpretation line in the note is still good practice, and some payers ask for one.

Payer-specific. NCCI policy treats psychotherapy as including continuing psychiatric evaluation, and many payers bundle brief screens into the session on that logic; others pay separately. Check the current NCCI edits and your contract before building a workflow on the answer. Two related facts age quickly in vendor guidance: marriage and family therapists and mental health counselors have billed Medicare directly since 2024 (42 CFR 410.53 and 410.54), and Medicare's annual depression screening benefit runs under its own code, G0444, not 96127.

Scale, administration, and interpretation. The testing codes 96130 to 96139 describe timed, clinician-administered services that include interpretation and a written report, and NCCI policy bars reporting them for self-administered tests; that work product is a psychological evaluation report. A brief self-report screen like the PHQ-9 is the opposite object: minutes long, self-scored, no separate report, documented in this note. The OIG's most on-point mental health audit found exactly this confusion, self-scored screens billed as testing, to be the largest source of inappropriate testing claims.

No regulation names the physical form; what the rules require is a record sufficient to support the care and the claim, retained on the whole-chart clock: commonly 5 to 10 years by state in the US, at least 10 years after the last contact (or after a minor client turns 18) under Ontario's college standard, and 7 years or until age 25 in Australia. If the completed instrument is the only place the item-level responses exist, keep it or its scanned image: scores are part of the accessible record, and clients can request them under HIPAA's right of access.

At the plan level, yes. The MBS explanatory notes make administering an outcome measurement tool a written element of preparing the GP mental health treatment plan unless clinically inappropriate, with the same tool re-administered at review (note AN.0.56); the choice of tool is the clinician's, with the K10 and DASS-21 named as examples. The psychologist's session items carry no outcome-measure element; their documentation duty is the written report back to the referrer. The named-instrument mandate lives in PHN-commissioned services instead, where the PMHC minimum data set requires the K10+ (or K5) and the SDQ at episode start and end.

Bring the scores however they exist: a portal export, a dictation, or session bullets. BastionGPT drafts the full review note with the trend, the interpretation, the client discussion, and the plan linkage, keeps the flowsheet current across administrations, and checks drafts for the gaps this page lists: a score with no comparison, an undiscussed result, a positive risk item without follow-up, or a billing line the documentation does not support. BastionGPT is HIPAA-compliant with a signed BAA, and data is never used to train models.

Educational content, not legal or billing advice. Sample notes are fictional. Follow your organization's policies and your board, payer, and jurisdiction requirements.