BDI-II Documentation: Score Interpretation & Sample Note

The BDI-II (Beck Depression Inventory, second edition) is a 21-item, licensed depression severity measure published by Beck, Steer, and Brown in 1996 and scored 0 to 63. Behavioral health, medical, and evaluation practices use it for baselines, outcome monitoring, and assessment batteries. This page covers how to document BDI-II results defensibly, with a fictional sample note.

Free to use and share. No signup required.
Already have session bullets or a transcript? Generate a structured draft with BastionGPT — you review and sign it.
Who writes it

Patient self-report on licensed forms or Q-global; scored and interpreted by a qualified (Level B) clinician

Audience

Psychologists and therapists, evaluation and forensic practices, bariatric and transplant programs, research teams, auditors

Typical length

A result line plus interpretation in the note · patient completion about 5 minutes

Format family

Licensed self-report severity measure (21 items scored 0 to 3, total 0 to 63)

When it's used

Depression severity baselines, psychotherapy outcome monitoring, evaluation batteries, bariatric and transplant psychosocial assessments

Standards context

Licensed Pearson instrument (Level B): every administration needs a purchased form or usage; no authority mandates it by name

What is the BDI-II?

The Beck Depression Inventory, second edition (BDI-II) is a 21-item self-report measure of depression severity published by Aaron Beck, Robert Steer, and Gregory Brown in 1996 and distributed by Pearson. Each symptom category contributes 0 to 3 points, for a total of 0 to 63, over a two-week window aligned with the DSM-IV criteria of its era (it was not built against DSM-5-TR). The listed age range for the English edition is 13 to 80, completion takes about five minutes, and purchase requires Pearson's Level B qualification. The manual's severity descriptors are totals through 13 minimal, 14 to 19 mild, 20 to 28 moderate, and 29 to 63 severe. Those are severity bands, not diagnoses and not universal case-finding cutoffs: a 2019 diagnostic meta-analysis found study-optimal thresholds ranging from 10 to 25 depending on the population, which is why the defensible chart phrase is "within the manual's moderate severity range," never "moderate major depressive disorder."

Two facts frame everything else. First, licensing: the instrument is proprietary and test-secure, so every administration consumes a purchased paper form or a Q-global usage, the "unlimited scoring" subscription covers manually entered data only, and photocopying the form or rebuilding it in a survey platform breaches the license. The rights record lists the instrument copyright to Aaron T. Beck (1996, 1987), with Pearson managing distribution and permissions. Second, age: there is no BDI-III (a dated finding as of August 2026; Pearson's own catalog calls the BDI-II the latest edition), so the foundational norms are three decades old, a limitation regulators and high-stakes reviewers increasingly expect reports to address. The free, nine-item PHQ-9 is the standard comparison: totals correlate near .77, but the severity bands do not map one to one, and the BDI-II classifies more patients as severe, so the two are never spliced into one trajectory.

Who uses BDI-II documentation and when

Psychotherapy practices run serial BDI-II totals for outcome monitoring where continuity with the instrument's large treatment literature matters; evaluation practices fold it into batteries for disability, forensic, bariatric, and transplant questions, where the stakes are highest and the instrument never stands alone. Bariatric programs are the emblematic case: the BDI-II is among the most commonly used depression measures in presurgical psychosocial evaluations, yet no national standard mandates it, and transplant policy likewise requires a psychosocial evaluation without naming any instrument. US disability work inverts the usual assumption: Social Security does not require, endorse, or prefer any depression instrument, and since 2017 its mental-disorders rules rate adult functional limitation from functional evidence rather than test scores (intellectual disorder excepted), so a bare total carries weight only through the narrative around it. Medical settings carry the somatic caveat: illness-driven fatigue, sleep, and appetite changes inflate totals (one multiple sclerosis cohort flagged 54.1% for depression on the BDI-II against 30.7% on its somatic-free sibling), which is why the seven-item BDI-FastScreen exists and why perinatal care prefers the EPDS, built without somatic items. Serial-measurement conventions live in the outcome measure note, battery integration in the psychological evaluation report, and session-level care in the psychotherapy progress note.

How to document BDI-II results in the chart

No law prescribes a BDI-II note format, and no US, Canadian, or Australian authority requires the instrument by name. What survives review is a record that treats the total as one measurement with a named method around it: edition, mode, completion, the suicide item, trajectory, integration, and the licensing trail. Each element below carries the pitfall that most often undermines it.

Instrument identity, edition, and mode. Record BDI-II with the language or version, the date, and the administration mode (paper, supervised on-screen, or remote on-screen), with telehealth conditions noted per the publisher's guidance: privacy, interruptions, assistance, technical problems, and whether conditions supported valid interpretation. Pitfall: edition blindness. BDI and BDI-IA scores and cutoffs do not transfer (the 1996 revision changed items and runs a few points higher), so chart the edition at every timepoint and treat any edition change as an instrument transition, not a continuous series.

The result, with the manual's descriptor. Chart the total and the band phrase with attribution: "41-year-old completed the English BDI-II; total 24, within the manual's moderate severity range (Beck, Steer, and Brown, 1996)." Pitfall: converting a band into a diagnosis. The band describes reported symptom burden over two weeks; the diagnosis belongs to the clinical interview and differential, and "severe range" is not "severe major depressive disorder."

The suicide item, handled separately. Any endorsement of the suicide-related item prompts a timely, separately documented risk assessment (a suicide risk assessment with a safety plan where indicated), regardless of the total. Pitfall: arithmetic in either direction. A minimal or mild total cannot neutralize an endorsed suicide item, and a high total does not establish imminent risk without direct assessment; "item 9 positive, safety plan completed" is usually too thin for a high-risk encounter.

Completion and conditions. State whether the administration was complete; if not, record the extent and reason without reproducing item content, and arrange completion or an alternative. Pitfall: prorating. No public publisher rule authorizes averaging or substituting missing responses, Q-global will not report with required responses unresolved, and a partial sum read against the 0 to 63 bands is not a valid total.

Trajectory, with method honesty. Record the prior total and date, the interval, and the absolute and percentage change. If you claim reliable or clinically significant change, name the method and inputs; published conventions genuinely differ (a reliable-change threshold near 5 points under one method, 8.46 under another, 10 in other protocols, and a patient-perceived meaningful difference estimated at 17.5% of baseline). Pitfall: "reliable improvement" from a bare point drop, or a study-specific recovery endpoint (13 or lower in one trial convention) charted as a manual rule. Without a named method, the honest entry is the raw and percentage change.

Integration and limits. Connect the score with the interview, observed functioning, collateral, and medical status, and state the limits that apply: self-report with no validity scale, somatic inflation in medical illness, language and norm limitations, response-style considerations. Pitfall: silent resolution. A score-interview discrepancy is described and investigated, not resolved automatically in favor of either source, and an unusually high total is a hypothesis about reporting style, never proof of exaggeration.

The licensing trail. Keep an organizational record that administrations used purchased forms or Q-global usages under a qualified user, store protocols securely, and keep item content out of the EHR. Pitfall: photocopied forms, survey-platform rebuilds, or test items transcribed into notes; purchase of a kit is not a reproduction permission.

Blank template (copy and adapt)

BDI-II DOCUMENTATION BLOCK
Date: [ ]   Setting: [ ]   Clinician: [ ]
Instrument: BDI-II   Language/version: [ ]   Edition held
   constant across series: [y/n]
Mode: [paper / on-screen / remote on-screen + conditions]
Completion: [complete / incomplete: extent + reason + plan]
Total: [ ]/63   Manual severity range: [minimal / mild /
   moderate / severe]  (Beck, Steer, and Brown, 1996)
Suicide item: [not endorsed / endorsed -> separate risk
   assessment documented at: ]
Trajectory: [prior total + date -> current; absolute + %
   change; method named, or raw change only]
Integration: [interview, function, collateral: consistent /
   discordant + how investigated]
Limits: [somatic confounds, language/norms, response style,
   nonstandard conditions]
Action: [risk, formulation, treatment intensity, referral,
   follow-up interval]
Licensing: [purchased form / Q-global usage; qualified user]
Clinician signature / credentials:            Date:

Free to use and share, no signup. The PDF includes a one-page cheat sheet with element-by-element pitfalls and a pre-sign checklist; the DOCX is the blank documentation block, ready to adapt. Neither reproduces the instrument itself.

Sample BDI-II documentation (fictional)

Scenario: serial outcome monitoring through psychotherapy, including one administration where the suicide item was endorsed and assessed separately, and a change statement that names no threshold it cannot support. All details are fictional.

Patient: J.T., 41  ·  Setting: Outpatient psychotherapy, week 12 review  ·  Clinician: M. Calloway, PhD  ·  Note date: 08/12/2026

Administration: English BDI-II completed today in office on a licensed Q-global on-screen administration; complete, no omitted responses; edition and language constant across the series. Total 16, within the manual's mild severity range (Beck, Steer, and Brown, 1996). Suicide item not endorsed today.

Series: Intake (05/20/2026) total 32, manual severe range; the suicide item was endorsed at a low level at intake and was assessed separately the same day: passive thoughts without intent, plan, or preparation, documented in the intake risk assessment with a collaborative safety plan and 988 contact information provided. Week 6 (07/01/2026) total 24, moderate range, suicide item not endorsed. Today 16, mild range.

Change statement: The total decreased 16 points from intake, a 50% reduction. The direction and pace are consistent with observed functioning: regular work attendance resumed, morning routine re-established, two social commitments kept weekly. No universal reliable-change threshold is assumed; under the trial convention that requires a 10-point improvement with an endpoint of 13 or lower, today's result would meet the change requirement but not that study's recovery endpoint, reported here for transparency, not as a manual rule.

Integration and plan: Score consistent with interview and function; residual symptoms remain clinically significant (sleep, self-critical rumination). Continue weekly CBT with relapse-prevention focus; next BDI-II in four weeks under the same mode and edition; prescriber review shares today's total and series.

This sample is fictional and for educational purposes. It does not describe a real patient or record; the totals, dates, and details are invented to show documentation structure and are not clinical guidance.

↑ Back to the template and downloads

Why this sample works

  • The edition, language, mode, completion status, and licensing basis are all named, and the series explicitly holds them constant, which is what makes the trajectory interpretable.
  • Every total carries the manual's band phrase with attribution, and no band is converted into a diagnosis anywhere in the note.
  • The endorsed suicide item at intake was assessed separately the same day and cross-referenced, and the low total at that visit did not determine the disposition.
  • The change statement gives raw and percentage figures, names the one study convention it cites, and refuses to invent a universal threshold.
  • Integration ties the numbers to observable function, and the plan fixes the next administration's interval and mode, which is what reviewers look for in measurement-based care.

Writing these after every session? BastionGPT drafts complete notes from bullets, dictation, or a transcript.

Generate a note from bullets

Documentation and compliance considerations

United States: Social Security does not require, endorse, or prefer any depression instrument, and its testing policy instructs adjudicators not to rely on test results alone; since the 2017 mental-disorders rewrite, adult functional limitation is rated from functional evidence rather than test scores, intellectual disorder excepted (PAYER POLICY). The same policy expects standardized administration by a qualified specialist, recent and appropriate norms, and a narrative connecting scores to function, which is exactly where a 1996-normed instrument needs explicit justification. In forensic work, Federal Rule of Evidence 702 governs the methodology (LAW), and the convention is to pair the BDI-II with symptom-validity measures: it has no embedded validity scale, and studies of elevated totals as over-reporting indicators found sensitivity too limited for stand-alone use, so a high score is a hypothesis, never proof of exaggeration (CONVENTION). Bariatric and transplant programs use it widely, but no national standard mandates it: ASMBS materials list it among recommended instruments without crowning a gold standard, and OPTN living-donor policy requires a psychosocial evaluation without naming any measure (CONVENTION and PAYER POLICY by program). On billing, the relevant distinction is 96127 versus the 96130 series: a current Medicare contractor policy uses a brief Beck depression questionnaire as its example of testing that should not be billed as psychological testing when it can be performed within the clinical interview, and license cost is not a coding criterion (PAYER POLICY).

Canada and Australia add professional-regulation and program layers. Ontario's College of Psychologists and Behaviour Analysts requires registrants to protect test security and copyright, to avoid outdated or obsolete tests or document a reasoned departure, and to record assessment activity begun but not completed with the reason, direct authority for charting an incomplete BDI-II rather than prorating one (LAW, professional regulation); the Canadian Psychological Association's 2019 position paper found no federal or provincial oversight of the psychological test market, which leaves publisher terms and college standards as the operative controls. The official Canadian French edition differs from the English one (published 1998, listed for ages 18 and older), so a bilingual practice should not silently apply the English 13-to-80 range. In Australia, Better Access guidance expects an outcome measure in a mental health treatment plan unless clinically inappropriate but names the K10 and DASS-21 as examples, not the BDI-II, so the licensed instrument is a discretionary, clinician-funded choice with no MBS item of its own (PAYER POLICY), and registered psychologists remain bound by Board competence and record standards (LAW). If you or a client needs immediate support: call or text 988 (US), 9-8-8 (Canada), or Lifeline 13 11 14 (Australia).

The psychometrics support disciplined use. Internal consistency runs near .90 to .92 with retest reliability of .73 to .96 across 118 studies; a meta-analysis of 62 samples (20,475 participants) supports a two-factor cognitive and somatic-affective structure while the total remains the defensible interpretive unit; and the somatic factor is exactly where medical illness inflates totals, the reason the seven-item BDI-FastScreen exists and the EPDS omits somatic items for perinatal use. The norm question deserves a balanced sentence in high-stakes reports: the BDI-II remains a current Pearson edition with extensive contemporary validation, but its foundational manual and normative frame date to 1996 (a development sample of roughly 500 outpatients, 91% white, mean age 37), and population studies since, including the VA cohort, show meaningful drift, so normative claims should be justified for the population at hand. On rights: the Beck Depression Inventory, second edition, is a proprietary, Level B Pearson assessment; the rights record lists approximately 90 catalogued translations (availability is not local validation) and displays the instrument copyright as 1996 and 1987 by Aaron T. Beck. BastionGPT is not affiliated with Pearson, Beck Institute, or the instrument authors. This page reproduces no items, response options, forms, scoring keys, or proprietary report content.

↑ Back to the template and downloads

Common BDI-II documentation errors reviewers flag

The numbers behind these errors deserve more respect than they get. Study-optimal case-finding cutoffs ranged from 10 to 25 in the 2019 diagnostic meta-analysis; a VA study of 152,260 veterans found scores far above the 1996 normative sample (non-depressed subgroup effect size 1.34) and suggested a local cutoff near 27; one multiple sclerosis cohort flagged 54.1% for depression on the BDI-II against 30.7% on the somatic-free short form; and a 2026 review found only 13% of eligible psychotherapy trials used any reliable-change criterion, spread across at least seven calculation methods. The BastionGPT Clinical Advisory Board sees the same errors most often in BDI-II documentation reviews:

  • Bands read as diagnoses or universal cutoffs. "Severe depressive disorder" charted from a total of 31, or 14 treated as a diagnostic threshold. The manual bands describe reported symptom burden; case-finding cutoffs are population-specific; and the diagnosis belongs to the interview and differential.
  • The suicide item folded into the arithmetic. A mild total used to skip follow-up of an endorsed suicide item, or a severe total charted as imminent risk with no direct assessment. The item is reviewed independently of the total, in both directions, with the risk assessment documented separately.
  • Prorated partials. A total computed from fewer than 21 completed categories and read against the 0 to 63 bands. No public publisher rule authorizes proration; chart the administration as incomplete with the extent and reason, and arrange completion or an alternative.
  • Change verdicts without a method. "Reliable improvement" from a nine-point drop, or a study-specific recovery endpoint presented as a manual rule. Published thresholds differ by method and population; name the method and inputs, or report raw and percentage change only.
  • Instruments and editions spliced. PHQ-9 totals converted to BDI-II totals by formula, BDI-IA history read against BDI-II bands, or an edition switch buried mid-series. Correlation is not equivalence; mark every transition and re-baseline.
  • Licensing shortcuts. Photocopied record forms, the instrument rebuilt in a survey platform, protected item text transcribed into the EHR, or remote administration outside the authorized modes. Purchase is not a reproduction permission, and Canadian regulators make test security a professional obligation, not a courtesy.
How BastionGPT helps

BastionGPT is specifically trained, tuned, and clinically tested on behavioral health progress notes and screening documentation.

  • Give it the facts (edition, language, mode, completion, total, suicide-item status, prior scores and dates, function, limits) and it drafts the full entry: band phrasing with attribution, separate risk-assessment cross-references, an honest change statement, and integration with the interview, ready for your review.
  • Cross-check a finished note for the gaps reviewers flag: a band charted as a diagnosis, a suicide item resolved by arithmetic, a prorated partial, a change verdict with no method, or an edition switch buried mid-series.
  • Draft the outcome summary for a referral or review: the series with modes and intervals, raw and percentage change, functional corroboration, and stated limits, ready to confirm against the record.

See how clinicians use it day to day on the AI therapy notes page.

Many BastionGPT users report saving more than 90 minutes per day on documentation.

HIPAA-compliant with a signed BAA on every plan. Your data is never used to train models. BastionGPT drafts, you review and sign.

Frequently asked questions

Twenty-one symptom categories scored 0 to 3 sum to a total of 0 to 63 over a two-week window. The manual's severity descriptors are totals through 13 minimal, 14 to 19 mild, 20 to 28 moderate, and 29 to 63 severe (Beck, Steer, and Brown, 1996). Those bands describe self-reported symptom burden; they are not diagnoses, treatment mandates, or universal screening cutoffs, and the 2019 diagnostic meta-analysis found optimal case-finding thresholds anywhere from 10 to 25 depending on the population. The defensible chart phrase is "within the manual's moderate severity range," with the diagnosis established separately by clinical interview and differential assessment.

No. Social Security requires no psychological test for depression claims, prefers no instrument, and rates adult functional limitation from functional evidence rather than test scores under its 2017 rules; a BDI-II contributes only through a well-supported narrative tied to function. Bariatric and transplant programs commonly include it in psychosocial evaluations, but no national standard mandates it by name, and payer or program requirements are checked individually rather than assumed. No score band specifies psychotherapy, medication, or level of care; treatment decisions integrate diagnosis, impairment, risk, history, and preference. Where you see a hard number requirement, it is a local program or payer policy, and it should be cited as such in the note.

Assess it directly and document that assessment separately, whatever the total. Any endorsement prompts timely clarification of currency, frequency, intent, plan, preparation, access to means, history, dynamic risks, and protective factors, with consultation, safety planning, and disposition as indicated. The arithmetic works in neither direction: a minimal total cannot neutralize an endorsed item, and a severe total does not establish imminent danger without direct assessment. "Item 9 positive, safety plan completed" is usually too thin for a high-risk encounter; the risk assessment gets its own documentation with reasoning, and the BDI-II note cross-references it.

There is no universal number, and pretending otherwise is the most common charting error. Published conventions include a reliable-change threshold near 5 points under one Jacobson-Truax calculation, 8.46 under another with different inputs, 10 points in other protocols, a trial convention requiring a 10-point improvement plus an endpoint of 13 or lower, and a patient-perceived meaningful difference estimated at 17.5% of baseline (higher, around 32%, in longer-duration treatment-resistant depression). A 2026 review found only 13% of eligible psychotherapy trials used any reliable-change criterion, across at least seven methods. Name the method and inputs when you claim significance; otherwise chart the raw and percentage change and let the functional evidence carry the interpretation.

No. The instrument is proprietary and test-secure: every administration uses a purchased record form or a Q-global usage, the annual "unlimited scoring" subscription covers manually entered data only (on-screen and remote administrations are per-use), and reproduction, translation, platform embedding, and research use are separate permission requests to the publisher. Keep protected item content out of the EHR; chart results and interpretation instead. Canadian colleges make test security and copyright a professional obligation, and Ontario's standards additionally require documenting assessment activity begun but not completed, which is the correct handling for a partial administration. Remote use goes through the authorized modes with the telehealth conditions documented, not through a scanned form in email.

They answer to different economics, not different truths. Totals correlate near .77 overall, responsiveness to change is similar, and neither is more valid because of its price; but severity bands do not map one to one, and the BDI-II classifies more patients as severe. The PHQ-9 is free, shorter, and the default for routine screening and high-volume measurement-based care. The BDI-II earns its licensing cost when continuity with an established BDI-II series or the historical treatment literature matters, when its broader cognitive-affective coverage is clinically useful, or when a protocol or evaluation methodology specifies it. Never convert one total to the other by formula, and never splice the two into one longitudinal graph; migrate with an overlap period and a new baseline.

Document the confound; never rewrite the score. Fatigue, sleep, appetite, and concentration items can be driven by illness, treatment, or depression, and the inflation is measurable: one multiple sclerosis cohort flagged 54.1% for depression on the BDI-II against 30.7% on the somatic-free seven-item BDI-FastScreen, and non-depressed pregnant patients score higher specifically on somatic items. The defensible pattern is to chart the relevant medical conditions and treatments, examine the cognitive-affective and somatic patterns qualitatively, consider a measure designed for the setting (the FastScreen in medical populations, the EPDS in perinatal care), and avoid claiming either that every somatic point is depression or that somatic symptoms never are.

Defensible, if you address them. The severity bands are criterion-referenced and still apply, and the instrument carries an extensive contemporary validation literature; but the foundational manual and normative frame date to 1996, built on a development sample of roughly 500 outpatients that was 91% white with a mean age of 37, and there is no BDI-III (a dated finding as of August 2026). Population drift is documented: a VA study of 152,260 veterans found scores well above the 1996 sample and suggested a local cutoff near 27. High-stakes reports should say why the normative frame remains appropriate for this examinee and question, and disability policy expects recent, appropriate norms, which makes that sentence load-bearing rather than boilerplate.

Yes. Give it the facts (edition and language, mode, completion, total, suicide-item status, prior scores with dates, function, medical confounds, licensing basis) and it drafts the full entry: the band phrase with attribution, the separate risk-assessment cross-reference, an honest change statement with raw and percentage figures, integration with the interview, and stated limits, ready for your review. It can also check a finished note for bands charted as diagnoses, arithmetic around the suicide item, prorated partials, method-free change verdicts, and spliced instruments or editions. BastionGPT is HIPAA-compliant with a signed BAA on every plan, and your data is never used to train models.

Primary sources

The instrument facts and compliance claims on this page trace to these sources, last verified August 2026:

  1. Beck AT, Steer RA, Brown GK, 1996, Manual for the Beck Depression Inventory-II, The Psychological Corporation; the current Pearson product listing (architecture, bands, age range, Level B, pricing model).
  2. ePROVIDE, the BDI-II rights record (instrument copyright 1996 and 1987 by Aaron T. Beck; approximately 90 catalogued translations); Pearson, permissions and licensing policies (purchase is not a reproduction permission).
  3. von Glischinski M and colleagues, 2019, Quality of Life Research, diagnostic meta-analysis (study-optimal cutoffs 10 to 25; pooled sensitivity .86, specificity .78 near 14.5).
  4. Huang C and Chen JH, 2015, Assessment (two-factor meta-analysis, 62 samples, 20,475 participants); Wang YP and Gorenstein C, 2013 (118-study reliability review: alpha near .90, retest .73 to .96).
  5. Kung S and colleagues, 2013, BDI-II versus PHQ-9 comparison (r = .77 overall; discordant severity assignment); Titov N and colleagues, 2011, responsiveness comparison (similar change sensitivity, fair-to-moderate category agreement).
  6. Button KS and colleagues, 2015, Psychological Medicine, patient-perspective MCID (17.5% of baseline; about 32% in treatment-resistant depression); Seggar LB, Lambert MJ, Hansen NB, 2002 (clinically significant change 8.46 under their inputs); Crenshaw and colleagues, 2026, reliable-change usage review (13% of 226 trials; at least seven operationalizations).
  7. 2019, Psychological Services, VA renorming study (152,260 veterans; non-depressed effect size 1.34; suggested local cutoff near 27).
  8. Social Security Administration, POMS DI 24583.050 (no required or preferred instrument; norm recency and integration expectations); Federal Rules of Evidence, Rule 702.
  9. CGS Administrators, LCD L34353 (a brief Beck depression questionnaire performable within the interview is not separately billable psychological testing).
  10. College of Psychologists and Behaviour Analysts of Ontario, Standards of Professional Conduct (test security, outdated-test avoidance, incomplete-service documentation); Canadian Psychological Association, 2019 position paper on test safety.
  11. Australian Department of Health, Better Access treatment-plan information (outcome measure expected; K10 and DASS-21 named); Services Australia, MBS mental health billing rules.
  12. ASMBS, presurgical psychosocial evaluation recommendations; OPTN, living-donor psychosocial evaluation policy (evaluation required; no instrument named).

Educational content, not legal or billing advice. Sample notes are fictional. Follow your organization's policies and your board, payer, and jurisdiction requirements.