PHQ-9 Documentation: Scoring, Interpretation & Sample Note

The PHQ-9 is a nine-item, public-domain depression severity measure scored 0 to 27, validated by Kroenke, Spitzer, and Williams in 2001. Primary care and behavioral health teams use it for depression screening and measurement-based care. This page covers how to document and interpret PHQ-9 results in the chart, including item 9 follow-up, with a fictional sample note.

Free to use and share. No signup required.
Already have session bullets or a transcript? Generate a structured draft with BastionGPT — you review and sign it.
Who writes it

Patient self-report; any clinician or trained staff can administer and score, no license or qualification level required

Audience

Treating clinicians and care teams, collaborative care registries, quality programs (MIPS, HEDIS), payers and auditors

Typical length

2 to 6 chart lines, plus item-9 follow-up when triggered · patient completion under 3 minutes

Format family

Self-report depression severity measure (9 items, scored 0 to 27)

When it's used

Depression screening, symptom monitoring and treatment-to-target, collaborative care registries, perinatal and adolescent pathways

Standards context

Public domain per the PHQ Screeners terms; no US, Canadian, or Australian authority mandates it by name

What is the PHQ-9?

The Patient Health Questionnaire-9 (PHQ-9) is a nine-item depression severity measure validated by Kurt Kroenke, Robert Spitzer, and Janet Williams in 2001, derived from the clinician-administered PRIME-MD and developed with an educational grant from Pfizer. Patients rate nine symptoms over the past two weeks from 0 to 3, for a total of 0 to 27. It is in the public domain: the official PHQ Screeners terms state that "no permission is required to reproduce, translate, display or distribute" the PHQ family, which covers EHR builds, patient portals, apps, and printed forms at no cost. The family matters for charting: the PHQ-2 is the two-item gateway screen, the PHQ-8 drops the self-harm item for research use (total 0 to 24), and the PHQ-9 Modified for Teens and the broader adolescent PHQ-A are related but not interchangeable forms, so the chart should name the exact version.

The load-bearing distinction for documentation: the PHQ-9 measures symptom severity, it does not make a diagnosis. A total of 10 or higher is a positive screen with pooled sensitivity and specificity near 85% against semistructured diagnostic interviews, and the published bands (minimal, mild, moderate, moderately severe, severe) describe reported symptom burden, not established major depressive disorder; the diagnosis comes from a diagnostic evaluation that weighs duration, impairment, and the bipolar, substance, medical, and grief alternatives. Two published scoring methods also coexist: the dominant continuous total, and a legacy DSM-IV-era symptom-count algorithm that should be labeled provisional if your organization still uses it, because meta-analysis finds it less sensitive than the cutoff of 10. And item 9 is a safety flag, not a suicide risk assessment: a nonzero response starts a direct clinical inquiry, and a zero does not end one.

Who uses PHQ-9 documentation and when

Primary care teams document PHQ-9 results inside screening workflows driven by USPSTF recommendations and quality measures; therapists and psychiatric prescribers trend it for measurement-based care and treatment-to-target; care managers run registries on it in collaborative care; and perinatal and adolescent pathways adapt it with population-specific forms. It pairs naturally with the GAD-7 when anxiety runs alongside depression, and it anchors the depression column of an outcome measure note. Neighbors matter at the boundaries: the licensed BDI-II serves settings that want a longer proprietary self-report, the EPDS owns most perinatal pathways, and Australian Better Access workflows usually reach for the K10 first, so a PHQ-9 entry there should say why a depression-specific, two-week measure fits the clinical question.

How to document PHQ-9 results in the chart

No statute prescribes a PHQ-9 note format. The elements below are the convention that survives review because each one maps to something a future clinician, a quality measure, or an auditor will need. Most results fit in two to six lines; the note grows only when item 9 is positive. Each element carries the pitfall that most often undermines it.

Instrument identity, mode, and language. Name the exact form (PHQ-9, PHQ-2, PHQ-8, PHQ-9 Modified for Teens), the administration mode (paper, portal, tablet, interview, telephone), the language, and any interpreter or staff assistance. Pitfall: "PHQ positive" with no form named. A PHQ-8 total runs 0 to 24, teen variants carry supplemental questions outside the total, and PHQ-A and PHQ-9 Modified for Teens are different instruments, so an unnamed form leaves the number uninterpretable and the quality extract empty.

Total over 27 with the published band. Write the score as a fraction and its severity descriptor: 0 to 4 minimal, 5 to 9 mild, 10 to 14 moderate, 15 to 19 moderately severe, 20 to 27 severe. Pitfall: writing "14/30". The unscored functional item is not part of the total, and a denominator error signals to reviewers that the scorer does not know the instrument.

Item 9, separately, every time. Record the 0 to 3 response even when it is zero. A nonzero response gets a same-visit direct inquiry: passive death wishes versus active ideation, plan, intent, preparatory behavior, history, and protective factors, then a documented risk formulation and a proportionate disposition, with a safety plan where indicated. Pitfall: treating the checkbox as the assessment. The instrument's own manual states that final determination of self-harm risk requires a clinical interview, and item 9 mixes thoughts of death with thoughts of self-harm.

The functional-interference item, unscored. Record the patient's answer (not difficult through extremely difficult) beside the total, never inside it. Pitfall: dropping it. It is the instrument's only direct read on impairment, which is what medical necessity and treatment decisions actually turn on.

Prior score, with change in points and percent. Name the baseline score and date, then give both numbers: points answer the measurement-error question, percent answers the response question. Pitfall: the five-point reflex. A 2026 individual-participant meta-analysis puts the general-practice threshold for change beyond measurement error near six points, response still means a 50% reduction, and remission means below 5; the three constructs deserve three separate sentences.

An interpretation sentence. Say what the result is (a severity screen in clinical context) and what it is not (a stand-alone diagnosis); after the interview, "consistent with" is the defensible verb. Pitfall: promoting a band to a diagnosis. "Moderate depressive symptom severity" is a finding; "moderate depression" is a diagnostic claim the questionnaire cannot carry alone.

Plan linkage and the next measurement. State what continues, changes, or is deferred because of the result, and when the next administration is due, tied to treatment phase rather than habit. Pitfall: an orphaned score. A number with no plan sentence documents administration, not care, and it is the first thing collaborative care reviewers and payer auditors flag.

Blank template (copy and adapt)

PHQ-9 DOCUMENTATION BLOCK
Form: [PHQ-9 / PHQ-2 / PHQ-8 / PHQ-9 Modified for Teens]
Language: [ ]   Mode: [paper / portal / tablet / interview]
Date: [ ]   Completed by: [patient self-report / staff-assisted]
Total: [ ]/27   Severity: [minimal / mild / moderate /
   moderately severe / severe]
Item 9: [0-3]
   If nonzero, same-visit inquiry: [passive vs active ideation;
   plan / intent / preparatory behavior; history; protective factors]
   Risk formulation + disposition: [ ]   Safety plan: [updated / n/a]
Functional interference (unscored): [not difficult ... extremely]
Prior score / date: [ ]   Change: [points] ([percent])
   Construct: [beyond measurement error / response >=50% /
   remission <5 / within error band]
Interpretation: [severity result in clinical context; screen,
   not a stand-alone diagnosis]
Plan linkage: [continued / changed / deferred, and why]
Next measurement: [interval + reason]
Clinician signature / credentials:            Date:

Free to use and share, no signup. The PDF includes a one-page cheat sheet with element-by-element pitfalls and a pre-sign checklist; the DOCX is the blank documentation block, ready to adapt.

Sample PHQ-9 documentation (fictional)

Scenario: a psychotherapy follow-up four weeks into treatment. The total has improved but not to response, and item 9 is newly positive at 1, the combination that gets documented badly most often: the improvement invites a short note, and the item-9 response demands a full one. All details are fictional.

Patient: R.T., 34  ·  Visit: Psychotherapy follow-up, week 4  ·  Clinician: M. Idris, LCSW  ·  Note date: 08/12/2026

Measure: PHQ-9 (English, patient portal, self-completed 08/12/2026): 14/27, moderate depressive symptom severity. Item 9 = 1. Functional interference reported as very difficult. Prior PHQ-9: 20/27 on 07/15/2026 (paper, in office).

Change from baseline: Six-point reduction (30%) over four weeks. The change is at the current general-practice threshold for exceeding measurement error, does not meet the 50% response definition, and is not remission. Trajectory is consistent with early improvement.

Item-9 follow-up (same visit): Reviewed directly today. R.T. describes passing thoughts of being "better off not waking up" on the worst mornings, most recently three days ago. She denies current active suicidal ideation, plan, intent, or preparatory behavior, and has no history of suicide attempt or self-harm. Protective factors include treatment engagement, her partner's support, and stated future plans. Acute risk assessed as low, with reasoning discussed; chronic elevation associated with recurrent depressive symptoms noted. Safety plan reviewed and updated with her today; crisis and after-hours contacts confirmed, including 988. Telephone check-in scheduled within 48 hours.

Interpretation: Today's result is a severity measurement within an established episode of major depressive disorder, single episode, moderate, diagnosed by clinical interview at intake. The score monitors severity; it did not establish the diagnosis.

Plan linkage: Continue weekly CBT with a behavioral activation focus. Prescriber review already scheduled for next week will use today's score and item-9 follow-up; coordination message sent today. Next PHQ-9 in two weeks, per active-treatment cadence. Depression follow-up was provided and documented at today's qualifying encounter for the practice's screening measure.

This sample is fictional and for educational purposes. It does not describe a real patient or record, and the scores are invented for illustration and correspond to no real person.

↑ Back to the template and downloads

Why this sample works

  • It names the exact form, language, mode, and date, so the score is comparable at the next administration and extractable for quality reporting.
  • The total sits over its 27-point denominator with the published band, and the unscored functional item stays separate from the arithmetic.
  • Item 9 gets a documented same-visit inquiry, a risk formulation with reasoning, and a proportionate disposition instead of a checkbox or a reflex emergency referral.
  • Change is stated in points and percent against a named baseline, and each measurement construct (error band, response, remission) is claimed separately.
  • The score links to a plan and a dated next measurement, and the diagnosis stays anchored to the clinical interview rather than the questionnaire.

Writing these after every session? BastionGPT drafts complete notes from bullets, dictation, or a transcript.

Generate a note from bullets

Documentation and compliance considerations

No law in the US, Canada, or Australia mandates the PHQ-9 by name, at any interval, or at any score threshold; the operative rules are PAYER POLICY and CONVENTION, and they run on different clocks. The USPSTF recommendations for depression screening (adults, 2023; adolescents 12 to 18, 2022) are Grade B convention: they name no instrument and specify no interval. MIPS Quality ID #134 is payer policy with teeth: screening with an age-appropriate standardized tool on the encounter date or up to 14 days before, once per performance period, with a documented follow-up plan for a positive screen. In 2026 the MIPS CQM version lets that documentation land up to two calendar days after the encounter while the Medicare Part B claims version requires it the same date, and a suicide-risk assessment alone does not count as the depression follow-up. A documented refusal is an allowed exception; a declined screen entered as a zero is a falsified record. Quality ID #370, depression remission at twelve months, starts a fixed clock at the first qualifying score above 9 with a depression diagnosis and looks for a score below 5 at 12 months plus or minus 60 days; a later, higher score does not restart the clock, so preserve the index date, form, and score. HEDIS runs three more clocks: follow-up within 30 days of a positive screen, a PHQ-9 result in each four-month assessment period for members with depression, and response or remission at 4 to 8 months. Collaborative care billing (99492 through 99494) requires validated rating scales and a registry as a billing element, which the PHQ-9 usually satisfies. For 96127 there is no stable national unit limit: CMS updates its edit files quarterly and commercial daily-limit tables differ, so verify the current plan-specific rule instead of repeating a number. In accredited settings, Joint Commission NPSG 15.01.01 adds the accreditation layer: a positive suicide screen, which an item-9 endorsement can trigger, must lead to an evidence-based assessment covering ideation, plan, intent, behaviors, and risk and protective factors, with the risk level, justification, and mitigation plan documented; item 9 alone does not complete that assessment.

The national postures genuinely diverge, so a cross-border template cannot assume the US pattern. Canada's task force (2025) recommends against routine questionnaire screening of asymptomatic adults, while CANMAT supports targeted case-finding (PHQ-2 at 2 or above, then PHQ-9 at 10 or above) and measurement every 2 to 4 weeks during active treatment with less frequent checks in maintenance; a Canadian chart should show assessment of symptomatic patients rather than universal screening justified by US guidance. Australia's Better Access guidance (March 2026) requires an outcome measure in a mental health treatment plan unless clinically inappropriate but prescribes no tool; the K10 is the incumbent by convention, and the MBS plan-item structure changed on 1 November 2025, so date any billing statement. On measurement itself: official translations may be reproduced freely but are not all independently validated, so record the language and any assistance used; and respect the instrument's error band when interpreting small movements, which the 2026 meta-analytic estimate puts near six points in general practice. If you or a client needs immediate support: call or text 988 (US), 9-8-8 (Canada), or Lifeline 13 11 14 (Australia).

The PHQ-9 and the PHQ instrument family are in the public domain; the official PHQ Screeners site states that no permission is required to reproduce, translate, display, or distribute them. PRIME-MD is a trademark of Pfizer Inc. BastionGPT is not affiliated with, or endorsed by, Pfizer, the instrument authors, or the PHQ Screeners project. This page reproduces no proprietary test items, norms, or scoring materials from any licensed instrument.

↑ Back to the template and downloads

Common PHQ-9 documentation errors reviewers flag

The accountability data here are unusually specific. Validated against an electronic C-SSRS criterion, item 9 carried a positive predictive value of 28.6% (Na et al., 2018: 41.1% of 841 patients were item-9 positive while 13.4% met the risk criterion), and the misses run the other way too: in a large health-system cohort, 39% of suicide attempts and 36% of suicide deaths within 30 days followed an item-9 response of zero (Simon et al.), and across 447,245 VA assessments, 71.6% of subsequent suicides occurred among patients who had answered "not at all" (Louzon et al., 2016). Meanwhile the 2026 individual-participant meta-analysis put the PHQ-9's minimal detectable change near six points, not the five that habit assumes. The BastionGPT Clinical Advisory Board sees the same errors most often in PHQ-9 documentation reviews:

  • Severity bands promoted to diagnoses. "PHQ-9 = 12, moderate depression" reads as a diagnostic statement. The defensible phrase is moderate depressive symptom severity, with the diagnosis carried by the interview; the CMS screening measure itself expects treatment decisions to follow adequate diagnostic evaluation, not the score alone.
  • Item 9 treated as a verdict in either direction. A nonzero response routed straight to the emergency department without an assessment, or a zero response closing a risk inquiry the interview should have kept open. Both patterns ignore the same evidence: the item overflags (positive predictive value 28.6% against a structured criterion) and underdetects (most near-term deaths followed a zero). Document the inquiry, the formulation, and why the disposition is proportionate.
  • The form left ambiguous. "PHQ score 14" without naming PHQ-9 versus PHQ-8 versus PHQ-9 Modified for Teens, the language, or the mode. The denominators differ, the teen supplements sit outside the total, and structured quality extraction fails when the instrument name is missing, which quietly zeroes out measures the organization reports.
  • Change math without its constructs. A five-point drop labeled real change when the 2026 general-practice threshold for exceeding measurement error sits nearer six points, a percent response with no named baseline, or two same-day administrations read as clinical change on an instrument that asks about two weeks. Points, percent, response, and remission each answer a different question; the note should say which one it is answering.
  • Scores invented, prorated, or overwritten. A declined screen entered as zero (which erases the allowed refusal exception and falsifies the record), missing items silently prorated into a "validated" total with no named method, or a patient's corrected answer replacing the original with no audit trail. Document what actually happened: offered, declined, incomplete, or amended, and by whom.
How BastionGPT helps

BastionGPT is specifically trained, tuned, and clinically tested on behavioral health progress notes and screening documentation.

  • Paste the responses, or just the total, item 9, and the prior score, and get a documentation-ready score block: severity band, change in points and percent against the named baseline, and an item-9 follow-up narrative scaffold when the response is nonzero.
  • Cross-check a finished note for the gaps reviewers flag: a total that disagrees with its band, a positive item 9 with no documented follow-up, a missing form name or baseline, or a functional item folded into the total.
  • Summarize a serial score history into response, remission, and measurement-error language for treatment reviews and collaborative care registries.

See how clinicians use it day to day on the AI therapy notes page.

Many BastionGPT users report saving more than 90 minutes per day on documentation.

HIPAA-compliant with a signed BAA on every plan. Your data is never used to train models. BastionGPT drafts, you review and sign.

Frequently asked questions

The total runs 0 to 27, and the published bands are 0 to 4 minimal, 5 to 9 mild, 10 to 14 moderate, 15 to 19 moderately severe, and 20 to 27 severe. Those are severity descriptors for reported symptoms over the past two weeks, not diagnoses. A score of 10 or higher is the conventional positive-screen threshold, with pooled sensitivity and specificity near 85% in the largest meta-analysis. In a low-prevalence clinic, many positives will not turn out to be major depression after interview, which is why chart language should stay at "positive screen, moderate severity" until the diagnostic evaluation speaks.

Three thresholds answer three different questions. Change beyond measurement error: the 2026 individual-participant meta-analysis estimated the minimal detectable change at 5.72 points in general practice, so about six points is the cautious threshold, replacing the older five-point habit. Response: a reduction of at least 50% from a named baseline. Remission: a total below 5, which still says nothing about residual individual symptoms until you review item-level responses. Write the points, the percent, and the construct you are claiming; "improved" with no baseline is the version auditors cannot use.

Record the response, then document a same-visit direct inquiry: passive death wishes versus active ideation, plan, intent, preparatory behavior, past attempts, and protective factors. Close with a risk formulation covering acute and chronic risk with your reasoning, a proportionate disposition, and the follow-up interval. A structured tool can organize the inquiry (the suicide risk assessment page covers it) and a safety plan documents the mitigation. Emergency referral follows the assessment when the assessment supports it; no authority reviewed for this page makes it automatic for every nonzero response.

No. It is one data point with measured limits: in a large health-system cohort, 39% of suicide attempts and 36% of suicide deaths within 30 days followed a zero, and across 447,245 VA assessments, 71.6% of subsequent suicides occurred among patients who had answered "not at all" (Louzon et al., 2016). Record the zero, and let history, behavior, collateral information, and the interview keep their full weight. A zero never overrides contrary clinical evidence, and a note should never cite it as the reason an inquiry stopped.

No authority requires it at every visit. CANMAT describes every 2 to 4 weeks as useful during active treatment, with less frequent measurement in maintenance. MIPS #134 requires screening once per performance period; HEDIS looks for a result in each four-month period for members with depression and checks response or remission at 4 to 8 months; collaborative care programs re-measure on their registry cadence. Pick the cadence from the treatment phase and the program you actually report to, then write it into the plan.

No. The instrument family is in the public domain, and the official PHQ Screeners terms state that no permission is needed to reproduce, translate, display, or distribute it. That covers EHR builds, portals, apps, and paper forms at no cost. Two cautions travel with the freedom: a reworded or restructured version is no longer the validated instrument, and the licensing world next door is different. The BDI-II (Pearson) and the EPDS (Royal College of Psychiatrists) carry reproduction restrictions, so the free-to-copy habit must not migrate to them.

That it was offered, that it was declined, and any reason the patient wants recorded, then the assessment you did instead through conversation and observation. Never enter a zero or a "negative screen" for a form that was not completed. For MIPS #134, a documented patient refusal is an allowed denominator exception, so the honest note is also the one the measure accepts.

They are one family with different jobs, and the chart should name the exact form. The PHQ-2 is the two-item gateway (0 to 6) that triggers a full PHQ-9; the PHQ-8 drops item 9 for research settings (0 to 24, no safety flag); the PHQ-9 Modified for Teens adds youth-specific supplemental questions that sit outside the total; and the broader PHQ-A is a substantially modified adolescent instrument, not a synonym for the teen PHQ-9. Totals from different forms are not interchangeable, and "PHQ positive" without a form name is the ambiguity that starts most downstream errors.

Yes. Paste the responses, or just the total, item 9, and the prior score, and it drafts the documentation block: severity band, change in points and percent against the named baseline, an item-9 follow-up narrative scaffold when the response is nonzero, and the plan-linkage sentence, ready for your review. It can also check a finished note for the classics: a total that disagrees with its band, a positive item 9 with no documented follow-up, or a missing form name. BastionGPT is HIPAA-compliant with a signed BAA on every plan, and your data is never used to train models.

Primary sources

The instrument facts and compliance claims on this page trace to these sources, last verified August 2026:

  1. PHQ Screeners, official screener site and instruction manual: public-domain status, version family, severity bands, the unscored functional item, and the requirement that self-harm risk be determined by clinical interview.
  2. Kroenke, Spitzer & Williams, 2001, Journal of General Internal Medicine, the PHQ-9 validation study: development, original 88%/88% operating characteristics at the cutoff of 10, and reliability.
  3. Negeri et al., 2021, BMJ, updated individual-participant-data meta-analysis (100 studies, 44,503 participants): pooled sensitivity and specificity of 0.85 at the cutoff of 10; Levis et al., 2019, BMJ, the first IPD meta-analysis.
  4. Manea et al., 2015, meta-analysis of the PHQ-9 diagnostic algorithm: the legacy symptom-count method is less sensitive than the total-score cutoff.
  5. Na et al., 2018, Journal of Affective Disorders, item 9 validated against the eC-SSRS: sensitivity 87.6%, specificity 66.1%, positive predictive value 28.6%, negative predictive value 97.2%.
  6. Simon et al., large health-system cohort of PHQ administrations: graded attempt risk by item-9 response and the share of near-term attempts and deaths that followed a zero response; Louzon et al., 2016, Psychiatric Services 67(5), the VA cohort of 447,245 assessments.
  7. 2026 BMJ individual-participant meta-analysis: PHQ-9 minimal detectable change (MDC95) of 5.72 points in general practice and 6.48 in inpatient settings.
  8. CMS Quality Payment Program, 2026 specifications: Quality ID #134 (MIPS CQM), #134 (Medicare Part B claims), and #370 (Depression Remission at Twelve Months).
  9. NCQA, HEDIS depression measures: screening follow-up, PHQ-9 monitoring, and remission-or-response windows.
  10. CMS, behavioral health integration billing guidance (validated rating scales in collaborative care) and the medically unlikely edits program (quarterly updates; no stable national 96127 unit limit).
  11. Joint Commission, standards interpretation FAQs on suicide risk (screening and assessment expectations): what an evidence-based assessment after a positive screen must address.
  12. CANMAT 2023 depression guideline update, Canadian Journal of Psychiatry: two-stage screening pathway, measurement cadence, and response and remission definitions; Canadian Task Force on Preventive Health Care, 2025 depression screening update.
  13. Australian Government Department of Health, Disability and Ageing, Better Access treatment-plan guidance (March 2026): outcome-measure requirement with clinician tool choice.
  14. USPSTF, adult depression screening recommendation (2023) and child and adolescent recommendation (2022); LOINC, 44261-6, the PHQ-9 total-score code used in structured extraction.

Educational content, not legal or billing advice. Sample notes are fictional. Follow your organization's policies and your board, payer, and jurisdiction requirements.