The BDI-II (Beck Depression Inventory, second edition) is a 21-item, licensed depression severity measure published by Beck, Steer, and Brown in 1996 and scored 0 to 63. Behavioral health, medical, and evaluation practices use it for baselines, outcome monitoring, and assessment batteries. This page covers how to document BDI-II results defensibly, with a fictional sample note.
Patient self-report on licensed forms or Q-global; scored and interpreted by a qualified (Level B) clinician
Psychologists and therapists, evaluation and forensic practices, bariatric and transplant programs, research teams, auditors
A result line plus interpretation in the note · patient completion about 5 minutes
Licensed self-report severity measure (21 items scored 0 to 3, total 0 to 63)
Depression severity baselines, psychotherapy outcome monitoring, evaluation batteries, bariatric and transplant psychosocial assessments
Licensed Pearson instrument (Level B): every administration needs a purchased form or usage; no authority mandates it by name
The Beck Depression Inventory, second edition (BDI-II) is a 21-item self-report measure of depression severity published by Aaron Beck, Robert Steer, and Gregory Brown in 1996 and distributed by Pearson. Each symptom category contributes 0 to 3 points, for a total of 0 to 63, over a two-week window aligned with the DSM-IV criteria of its era (it was not built against DSM-5-TR). The listed age range for the English edition is 13 to 80, completion takes about five minutes, and purchase requires Pearson's Level B qualification. The manual's severity descriptors are totals through 13 minimal, 14 to 19 mild, 20 to 28 moderate, and 29 to 63 severe. Those are severity bands, not diagnoses and not universal case-finding cutoffs: a 2019 diagnostic meta-analysis found study-optimal thresholds ranging from 10 to 25 depending on the population, which is why the defensible chart phrase is "within the manual's moderate severity range," never "moderate major depressive disorder."
Two facts frame everything else. First, licensing: the instrument is proprietary and test-secure, so every administration consumes a purchased paper form or a Q-global usage, the "unlimited scoring" subscription covers manually entered data only, and photocopying the form or rebuilding it in a survey platform breaches the license. The rights record lists the instrument copyright to Aaron T. Beck (1996, 1987), with Pearson managing distribution and permissions. Second, age: there is no BDI-III (a dated finding as of August 2026; Pearson's own catalog calls the BDI-II the latest edition), so the foundational norms are three decades old, a limitation regulators and high-stakes reviewers increasingly expect reports to address. The free, nine-item PHQ-9 is the standard comparison: totals correlate near .77, but the severity bands do not map one to one, and the BDI-II classifies more patients as severe, so the two are never spliced into one trajectory.
Psychotherapy practices run serial BDI-II totals for outcome monitoring where continuity with the instrument's large treatment literature matters; evaluation practices fold it into batteries for disability, forensic, bariatric, and transplant questions, where the stakes are highest and the instrument never stands alone. Bariatric programs are the emblematic case: the BDI-II is among the most commonly used depression measures in presurgical psychosocial evaluations, yet no national standard mandates it, and transplant policy likewise requires a psychosocial evaluation without naming any instrument. US disability work inverts the usual assumption: Social Security does not require, endorse, or prefer any depression instrument, and since 2017 its mental-disorders rules rate adult functional limitation from functional evidence rather than test scores (intellectual disorder excepted), so a bare total carries weight only through the narrative around it. Medical settings carry the somatic caveat: illness-driven fatigue, sleep, and appetite changes inflate totals (one multiple sclerosis cohort flagged 54.1% for depression on the BDI-II against 30.7% on its somatic-free sibling), which is why the seven-item BDI-FastScreen exists and why perinatal care prefers the EPDS, built without somatic items. Serial-measurement conventions live in the outcome measure note, battery integration in the psychological evaluation report, and session-level care in the psychotherapy progress note.
No law prescribes a BDI-II note format, and no US, Canadian, or Australian authority requires the instrument by name. What survives review is a record that treats the total as one measurement with a named method around it: edition, mode, completion, the suicide item, trajectory, integration, and the licensing trail. Each element below carries the pitfall that most often undermines it.
Instrument identity, edition, and mode. Record BDI-II with the language or version, the date, and the administration mode (paper, supervised on-screen, or remote on-screen), with telehealth conditions noted per the publisher's guidance: privacy, interruptions, assistance, technical problems, and whether conditions supported valid interpretation. Pitfall: edition blindness. BDI and BDI-IA scores and cutoffs do not transfer (the 1996 revision changed items and runs a few points higher), so chart the edition at every timepoint and treat any edition change as an instrument transition, not a continuous series.
The result, with the manual's descriptor. Chart the total and the band phrase with attribution: "41-year-old completed the English BDI-II; total 24, within the manual's moderate severity range (Beck, Steer, and Brown, 1996)." Pitfall: converting a band into a diagnosis. The band describes reported symptom burden over two weeks; the diagnosis belongs to the clinical interview and differential, and "severe range" is not "severe major depressive disorder."
The suicide item, handled separately. Any endorsement of the suicide-related item prompts a timely, separately documented risk assessment (a suicide risk assessment with a safety plan where indicated), regardless of the total. Pitfall: arithmetic in either direction. A minimal or mild total cannot neutralize an endorsed suicide item, and a high total does not establish imminent risk without direct assessment; "item 9 positive, safety plan completed" is usually too thin for a high-risk encounter.
Completion and conditions. State whether the administration was complete; if not, record the extent and reason without reproducing item content, and arrange completion or an alternative. Pitfall: prorating. No public publisher rule authorizes averaging or substituting missing responses, Q-global will not report with required responses unresolved, and a partial sum read against the 0 to 63 bands is not a valid total.
Trajectory, with method honesty. Record the prior total and date, the interval, and the absolute and percentage change. If you claim reliable or clinically significant change, name the method and inputs; published conventions genuinely differ (a reliable-change threshold near 5 points under one method, 8.46 under another, 10 in other protocols, and a patient-perceived meaningful difference estimated at 17.5% of baseline). Pitfall: "reliable improvement" from a bare point drop, or a study-specific recovery endpoint (13 or lower in one trial convention) charted as a manual rule. Without a named method, the honest entry is the raw and percentage change.
Integration and limits. Connect the score with the interview, observed functioning, collateral, and medical status, and state the limits that apply: self-report with no validity scale, somatic inflation in medical illness, language and norm limitations, response-style considerations. Pitfall: silent resolution. A score-interview discrepancy is described and investigated, not resolved automatically in favor of either source, and an unusually high total is a hypothesis about reporting style, never proof of exaggeration.
The licensing trail. Keep an organizational record that administrations used purchased forms or Q-global usages under a qualified user, store protocols securely, and keep item content out of the EHR. Pitfall: photocopied forms, survey-platform rebuilds, or test items transcribed into notes; purchase of a kit is not a reproduction permission.
BDI-II DOCUMENTATION BLOCK Date: [ ] Setting: [ ] Clinician: [ ] Instrument: BDI-II Language/version: [ ] Edition held constant across series: [y/n] Mode: [paper / on-screen / remote on-screen + conditions] Completion: [complete / incomplete: extent + reason + plan] Total: [ ]/63 Manual severity range: [minimal / mild / moderate / severe] (Beck, Steer, and Brown, 1996) Suicide item: [not endorsed / endorsed -> separate risk assessment documented at: ] Trajectory: [prior total + date -> current; absolute + % change; method named, or raw change only] Integration: [interview, function, collateral: consistent / discordant + how investigated] Limits: [somatic confounds, language/norms, response style, nonstandard conditions] Action: [risk, formulation, treatment intensity, referral, follow-up interval] Licensing: [purchased form / Q-global usage; qualified user] Clinician signature / credentials: Date:
Free to use and share, no signup. The PDF includes a one-page cheat sheet with element-by-element pitfalls and a pre-sign checklist; the DOCX is the blank documentation block, ready to adapt. Neither reproduces the instrument itself.
Scenario: serial outcome monitoring through psychotherapy, including one administration where the suicide item was endorsed and assessed separately, and a change statement that names no threshold it cannot support. All details are fictional.
Patient: J.T., 41 · Setting: Outpatient psychotherapy, week 12 review · Clinician: M. Calloway, PhD · Note date: 08/12/2026
Administration: English BDI-II completed today in office on a licensed Q-global on-screen administration; complete, no omitted responses; edition and language constant across the series. Total 16, within the manual's mild severity range (Beck, Steer, and Brown, 1996). Suicide item not endorsed today.
Series: Intake (05/20/2026) total 32, manual severe range; the suicide item was endorsed at a low level at intake and was assessed separately the same day: passive thoughts without intent, plan, or preparation, documented in the intake risk assessment with a collaborative safety plan and 988 contact information provided. Week 6 (07/01/2026) total 24, moderate range, suicide item not endorsed. Today 16, mild range.
Change statement: The total decreased 16 points from intake, a 50% reduction. The direction and pace are consistent with observed functioning: regular work attendance resumed, morning routine re-established, two social commitments kept weekly. No universal reliable-change threshold is assumed; under the trial convention that requires a 10-point improvement with an endpoint of 13 or lower, today's result would meet the change requirement but not that study's recovery endpoint, reported here for transparency, not as a manual rule.
Integration and plan: Score consistent with interview and function; residual symptoms remain clinically significant (sleep, self-critical rumination). Continue weekly CBT with relapse-prevention focus; next BDI-II in four weeks under the same mode and edition; prescriber review shares today's total and series.
This sample is fictional and for educational purposes. It does not describe a real patient or record; the totals, dates, and details are invented to show documentation structure and are not clinical guidance.
Writing these after every session? BastionGPT drafts complete notes from bullets, dictation, or a transcript.
Generate a note from bulletsUnited States: Social Security does not require, endorse, or prefer any depression instrument, and its testing policy instructs adjudicators not to rely on test results alone; since the 2017 mental-disorders rewrite, adult functional limitation is rated from functional evidence rather than test scores, intellectual disorder excepted (PAYER POLICY). The same policy expects standardized administration by a qualified specialist, recent and appropriate norms, and a narrative connecting scores to function, which is exactly where a 1996-normed instrument needs explicit justification. In forensic work, Federal Rule of Evidence 702 governs the methodology (LAW), and the convention is to pair the BDI-II with symptom-validity measures: it has no embedded validity scale, and studies of elevated totals as over-reporting indicators found sensitivity too limited for stand-alone use, so a high score is a hypothesis, never proof of exaggeration (CONVENTION). Bariatric and transplant programs use it widely, but no national standard mandates it: ASMBS materials list it among recommended instruments without crowning a gold standard, and OPTN living-donor policy requires a psychosocial evaluation without naming any measure (CONVENTION and PAYER POLICY by program). On billing, the relevant distinction is 96127 versus the 96130 series: a current Medicare contractor policy uses a brief Beck depression questionnaire as its example of testing that should not be billed as psychological testing when it can be performed within the clinical interview, and license cost is not a coding criterion (PAYER POLICY).
Canada and Australia add professional-regulation and program layers. Ontario's College of Psychologists and Behaviour Analysts requires registrants to protect test security and copyright, to avoid outdated or obsolete tests or document a reasoned departure, and to record assessment activity begun but not completed with the reason, direct authority for charting an incomplete BDI-II rather than prorating one (LAW, professional regulation); the Canadian Psychological Association's 2019 position paper found no federal or provincial oversight of the psychological test market, which leaves publisher terms and college standards as the operative controls. The official Canadian French edition differs from the English one (published 1998, listed for ages 18 and older), so a bilingual practice should not silently apply the English 13-to-80 range. In Australia, Better Access guidance expects an outcome measure in a mental health treatment plan unless clinically inappropriate but names the K10 and DASS-21 as examples, not the BDI-II, so the licensed instrument is a discretionary, clinician-funded choice with no MBS item of its own (PAYER POLICY), and registered psychologists remain bound by Board competence and record standards (LAW). If you or a client needs immediate support: call or text 988 (US), 9-8-8 (Canada), or Lifeline 13 11 14 (Australia).
The psychometrics support disciplined use. Internal consistency runs near .90 to .92 with retest reliability of .73 to .96 across 118 studies; a meta-analysis of 62 samples (20,475 participants) supports a two-factor cognitive and somatic-affective structure while the total remains the defensible interpretive unit; and the somatic factor is exactly where medical illness inflates totals, the reason the seven-item BDI-FastScreen exists and the EPDS omits somatic items for perinatal use. The norm question deserves a balanced sentence in high-stakes reports: the BDI-II remains a current Pearson edition with extensive contemporary validation, but its foundational manual and normative frame date to 1996 (a development sample of roughly 500 outpatients, 91% white, mean age 37), and population studies since, including the VA cohort, show meaningful drift, so normative claims should be justified for the population at hand. On rights: the Beck Depression Inventory, second edition, is a proprietary, Level B Pearson assessment; the rights record lists approximately 90 catalogued translations (availability is not local validation) and displays the instrument copyright as 1996 and 1987 by Aaron T. Beck. BastionGPT is not affiliated with Pearson, Beck Institute, or the instrument authors. This page reproduces no items, response options, forms, scoring keys, or proprietary report content.
The numbers behind these errors deserve more respect than they get. Study-optimal case-finding cutoffs ranged from 10 to 25 in the 2019 diagnostic meta-analysis; a VA study of 152,260 veterans found scores far above the 1996 normative sample (non-depressed subgroup effect size 1.34) and suggested a local cutoff near 27; one multiple sclerosis cohort flagged 54.1% for depression on the BDI-II against 30.7% on the somatic-free short form; and a 2026 review found only 13% of eligible psychotherapy trials used any reliable-change criterion, spread across at least seven calculation methods. The BastionGPT Clinical Advisory Board sees the same errors most often in BDI-II documentation reviews:
BastionGPT is specifically trained, tuned, and clinically tested on behavioral health progress notes and screening documentation.
See how clinicians use it day to day on the AI therapy notes page.
Many BastionGPT users report saving more than 90 minutes per day on documentation.
HIPAA-compliant with a signed BAA on every plan. Your data is never used to train models. BastionGPT drafts, you review and sign.
Twenty-one symptom categories scored 0 to 3 sum to a total of 0 to 63 over a two-week window. The manual's severity descriptors are totals through 13 minimal, 14 to 19 mild, 20 to 28 moderate, and 29 to 63 severe (Beck, Steer, and Brown, 1996). Those bands describe self-reported symptom burden; they are not diagnoses, treatment mandates, or universal screening cutoffs, and the 2019 diagnostic meta-analysis found optimal case-finding thresholds anywhere from 10 to 25 depending on the population. The defensible chart phrase is "within the manual's moderate severity range," with the diagnosis established separately by clinical interview and differential assessment.
No. Social Security requires no psychological test for depression claims, prefers no instrument, and rates adult functional limitation from functional evidence rather than test scores under its 2017 rules; a BDI-II contributes only through a well-supported narrative tied to function. Bariatric and transplant programs commonly include it in psychosocial evaluations, but no national standard mandates it by name, and payer or program requirements are checked individually rather than assumed. No score band specifies psychotherapy, medication, or level of care; treatment decisions integrate diagnosis, impairment, risk, history, and preference. Where you see a hard number requirement, it is a local program or payer policy, and it should be cited as such in the note.
Assess it directly and document that assessment separately, whatever the total. Any endorsement prompts timely clarification of currency, frequency, intent, plan, preparation, access to means, history, dynamic risks, and protective factors, with consultation, safety planning, and disposition as indicated. The arithmetic works in neither direction: a minimal total cannot neutralize an endorsed item, and a severe total does not establish imminent danger without direct assessment. "Item 9 positive, safety plan completed" is usually too thin for a high-risk encounter; the risk assessment gets its own documentation with reasoning, and the BDI-II note cross-references it.
There is no universal number, and pretending otherwise is the most common charting error. Published conventions include a reliable-change threshold near 5 points under one Jacobson-Truax calculation, 8.46 under another with different inputs, 10 points in other protocols, a trial convention requiring a 10-point improvement plus an endpoint of 13 or lower, and a patient-perceived meaningful difference estimated at 17.5% of baseline (higher, around 32%, in longer-duration treatment-resistant depression). A 2026 review found only 13% of eligible psychotherapy trials used any reliable-change criterion, across at least seven methods. Name the method and inputs when you claim significance; otherwise chart the raw and percentage change and let the functional evidence carry the interpretation.
No. The instrument is proprietary and test-secure: every administration uses a purchased record form or a Q-global usage, the annual "unlimited scoring" subscription covers manually entered data only (on-screen and remote administrations are per-use), and reproduction, translation, platform embedding, and research use are separate permission requests to the publisher. Keep protected item content out of the EHR; chart results and interpretation instead. Canadian colleges make test security and copyright a professional obligation, and Ontario's standards additionally require documenting assessment activity begun but not completed, which is the correct handling for a partial administration. Remote use goes through the authorized modes with the telehealth conditions documented, not through a scanned form in email.
They answer to different economics, not different truths. Totals correlate near .77 overall, responsiveness to change is similar, and neither is more valid because of its price; but severity bands do not map one to one, and the BDI-II classifies more patients as severe. The PHQ-9 is free, shorter, and the default for routine screening and high-volume measurement-based care. The BDI-II earns its licensing cost when continuity with an established BDI-II series or the historical treatment literature matters, when its broader cognitive-affective coverage is clinically useful, or when a protocol or evaluation methodology specifies it. Never convert one total to the other by formula, and never splice the two into one longitudinal graph; migrate with an overlap period and a new baseline.
Document the confound; never rewrite the score. Fatigue, sleep, appetite, and concentration items can be driven by illness, treatment, or depression, and the inflation is measurable: one multiple sclerosis cohort flagged 54.1% for depression on the BDI-II against 30.7% on the somatic-free seven-item BDI-FastScreen, and non-depressed pregnant patients score higher specifically on somatic items. The defensible pattern is to chart the relevant medical conditions and treatments, examine the cognitive-affective and somatic patterns qualitatively, consider a measure designed for the setting (the FastScreen in medical populations, the EPDS in perinatal care), and avoid claiming either that every somatic point is depression or that somatic symptoms never are.
Defensible, if you address them. The severity bands are criterion-referenced and still apply, and the instrument carries an extensive contemporary validation literature; but the foundational manual and normative frame date to 1996, built on a development sample of roughly 500 outpatients that was 91% white with a mean age of 37, and there is no BDI-III (a dated finding as of August 2026). Population drift is documented: a VA study of 152,260 veterans found scores well above the 1996 sample and suggested a local cutoff near 27. High-stakes reports should say why the normative frame remains appropriate for this examinee and question, and disability policy expects recent, appropriate norms, which makes that sentence load-bearing rather than boilerplate.
Yes. Give it the facts (edition and language, mode, completion, total, suicide-item status, prior scores with dates, function, medical confounds, licensing basis) and it drafts the full entry: the band phrase with attribution, the separate risk-assessment cross-reference, an honest change statement with raw and percentage figures, integration with the interview, and stated limits, ready for your review. It can also check a finished note for bands charted as diagnoses, arithmetic around the suicide item, prorated partials, method-free change verdicts, and spliced instruments or editions. BastionGPT is HIPAA-compliant with a signed BAA on every plan, and your data is never used to train models.
The instrument facts and compliance claims on this page trace to these sources, last verified August 2026:
Educational content, not legal or billing advice. Sample notes are fictional. Follow your organization's policies and your board, payer, and jurisdiction requirements.