Psychological Testing Scoring Note: Definition, Template & Example

A psychological testing scoring note records the professional evaluation side of psychological or neuropsychological testing: verifying and summarizing scores, integrating them with clinical data, and logging the dated minutes that support CPT 96130 to 96133. Psychologists and neuropsychologists write one per assessment episode, alongside the administration record and the evaluation report. Most run 200 to 400 words plus the score summary and time log.

Free to use and share. No signup required.
Already have session bullets or a transcript? Generate a structured draft with BastionGPT — you review and sign it.
Who writes it

The interpreting psychologist or neuropsychologist; evaluation work cannot be delegated to a technician or trainee

Audience

Medicare and commercial payers, auditors, referring practitioners, regulatory colleges and boards

Typical length

200 to 400 words plus the score summary and time log · 5 to 12 minutes by hand (clinical team estimate)

Format family

Structured worksheet: score summary plus a dated evaluation-services time log (compare: psychological evaluation report)

When it's used

Every assessment episode billed under 96130 to 96133, from first score review through report and feedback

Standards context

No rule names the note; medical-necessity law and MAC payer policy compel its contents, and the discrete document is convention

What is a psychological testing scoring note?

A psychological testing scoring note records the professional evaluation side of psychological and neuropsychological testing: the interpreting clinician's verification of scores, the pulling together of test data with history and collateral, interpretation, planning, report generation, and feedback, together with the dated minutes that make CPT 96130 to 96133 billable. No statute, payer manual, or standards body names the document; the label is functional shorthand, and clinicians also call it a scoring and interpretation note, an evaluation services note, a testing time log, or simply the 96130 note. The code family it serves is young: the AMA CPT Editorial Panel created the evaluation-services codes effective January 1, 2019, retiring the legacy testing codes 96101 to 96103 and 96119 to 96120 and splitting the old per-hour 96101 into evaluation services on one side and administration and scoring on the other (APA Services).

Two boundaries carry the weight. The first is with the administration record: giving tests and mechanical scoring are separately coded work (96136 to 96139) that a trained technician can perform, while everything in this note is the interpreting clinician's own and cannot be delegated. The second is with the psychological evaluation report: the report is the clinical deliverable, and MAC billing articles expect a written report to be generated, but the report rarely proves the minutes. This note is the time-and-activity record that does. One structural rule from APA Services guidance surprises many practices: test selection is bundled into evaluation services, so the evaluation codes are billed together with administration services from the same episode, never on their own. Nothing in federal law names or formats any of it; the obligations flow from medical-necessity documentation rules and each MAC's coverage article, which is why the same note travels under so many names.

Who uses testing scoring notes and when

Psychologists and neuropsychologists write one for every assessment episode billed under the evaluation-services codes: private assessment practices, hospital and academic neuropsychology services, psychoeducational and disability evaluators, and forensic practices whose work must survive hostile reading. The stakes scale with the hours: in the field's professional practice survey of 1,658 neuropsychologists (Sweet and colleagues, 2015, The Clinical Neuropsychologist), a typical evaluation took about six hours of professional time, ranging from half an hour to 25 for complex forensic and diagnostic questions. Use this note whenever those hours span days or the interpretation happens after the patient leaves, which is most of the time. For a single screening questionnaire scored and discussed in session, an outcome measure note is the right record instead; instrument-specific write-up conventions live in guides like our MMPI-3 report guide, and school-context batteries feed the psychoeducational report.

Testing scoring note structure: what goes in each section

Episode header and referral. Patient identifiers, the referring or treating practitioner, the referral question, the ICD-10 linkage, and the episode span from first administration to expected feedback. The evaluation-services claim is an episode claim, so this header is what makes multi-day work cohere. Pitfall: treating each sitting as its own encounter. The base code is reported once per episode with combined time; a header that never defines the episode invites day-by-day billing that auditors unwind.

Instruments and data integrated. Every test interpreted, by exact name, plus the non-test data pulled in: interview, records, informant ratings, prior evaluations. The list must reconcile with the administration record and the report, because reviewers read all three side by side. Pitfall: lists that drift. An instrument that appears in the report but not here, or here but not in the administration record, reads as work that may not have happened.

Score summary and validity. Exact raw and standard scores, percentiles, and validity indices, with anomalies and corrections noted. Preserve exact values and exact test names; this is the section that lets anyone reconcile platform output, note, and report, and it is where verification of computer scoring is proven. Ontario's college makes the stake explicit by barring computer-generated reports as a substitute for the psychologist's professional opinion. Pitfall: pasting the scoring platform's narrative as the summary. An unverified auto-score with a miskeyed age band or wrong norm group corrupts everything downstream.

Evaluation time log. Dated entries naming the activity and the minutes: score verification, integrating data, interpreting results against clinical data, planning, report writing, feedback. Keep evaluation minutes strictly separate from administration and scoring time, which belongs to the technician-side codes. The base code needs 31 documented minutes; add-on hours follow past each full hour. Pitfall: vague time. 'Approximately 1 to 2 hours' is a documented denial risk, and undifferentiated block time cannot be matched to units.

Interpretation and clinical decision-making. The synthesis in your own words: what the verified scores mean against history, presentation, and collateral, the diagnostic conclusion forming, and the treatment planning that follows. This work is the interpreting clinician's alone; Medicare does not pay for evaluation services performed by students or trainees, and none of it can be handed to a technician. Pitfall: a signature on someone else's synthesis. A trainee-drafted interpretation the psychologist merely signs fails the performance requirement, not just style.

Report, feedback, and sign-off. The written report's status and date, the feedback date, the date of service reported for the episode, and signature with credentials. MAC articles expect a written report, and the claim is reported on the date the service concluded, typically feedback, per CMS date-of-service guidance. Referral sources notice speed too: 73 percent in one survey said slow report turnaround hurt patient care (Postal and colleagues, 2017). Pitfall: an open-ended episode. No report status, no feedback date, and no stated conclusion date leave the claim's date of service unsupported and the episode impossible to close.

Blank template (copy and adapt)

PSYCHOLOGICAL TESTING SCORING NOTE
Evaluation services record (CPT 96130 to 96133): score
verification, integration, interpretation, report, and
feedback minutes. Administration time is logged separately.

EPISODE
Patient: __________________  DOB: __________
Referring practitioner: _____________  ICD-10: ______________
Referral question: __________________________________________
Episode span (first administration to feedback): ____________

INSTRUMENTS INTERPRETED AND DATA INTEGRATED
1 ___________________________  3 ___________________________
2 ___________________________  4 ___________________________
Collateral integrated (records, interview, informants): _____
_____________________________________________________________

SCORE SUMMARY AND VALIDITY (exact values, exact test names)
Test / index: _____________ Raw: _____ Std: _____ %ile: _____
Test / index: _____________ Raw: _____ Std: _____ %ile: _____
Validity indices and response style: ________________________
Anomalies, corrections, rescoring: __________________________

EVALUATION TIME LOG (interpreting clinician only)
Date ______  Activity _______________________  Minutes ______
Date ______  Activity _______________________  Minutes ______
Date ______  Activity _______________________  Minutes ______
Total: ______  Units (base code once per episode, 31+ min): _

INTERPRETATION AND CLINICAL DECISION-MAKING
Findings against history and presentation: __________________
_____________________________________________________________
Diagnostic impression and plan: _____________________________

REPORT, FEEDBACK, AND SIGN-OFF
Report completed: __________  Feedback given: __________
Date of service reported (episode conclusion): ______________
Signature / credentials: __________________  Date: __________

Free to use and share, no signup. The PDF includes a one-page cheat sheet with section-by-section pitfalls and a pre-sign checklist; the DOCX is the blank template, ready to adapt.

Sample psychological testing scoring note

Scenario: conclusion of a two-day adult attention and memory battery in an outpatient neuropsychology practice. The neuropsychologist documents her own evaluation services for the episode, from score verification through feedback. All details are fictional.

Psychological Testing Scoring Note. Cedarbrook Neuropsychology Group  ·  Patient: R.T., 34  ·  Episode: 07/21 to 08/04/2026  ·  Clinician: L. Chen, PhD, licensed psychologist (neuropsychology)

Episode and referral: Testing ordered 07/10/2026 by P. Rowan, MD (psychiatry) to clarify adult ADHD versus anxiety-related attention complaints; provisional F41.1. Battery administered by a trained technician on 07/21 and 07/28/2026; technician administration and scoring time (213 minutes) is documented and billed separately under the technician codes. This note records my evaluation services for the episode (96132, 96133).

Instruments and data integrated: WAIS-5 selected indices; CVLT-3; CPT-3; CAARS 2 self-report and observer forms. Non-test data: 07/21/2026 clinical interview, psychiatric and pharmacy records from Dr. Rowan, partner-completed observer ratings, and sleep screening responses.

Score summary and validity: All platform scoring verified against raw responses before interpretation. WAIS-5 Processing Speed Index 92; CVLT-3 Trials 1 to 5 standard score 102, delayed recall within one standard deviation of the mean; CPT-3 detectability T 58. CAARS 2 self-report Inattention T 68; observer form initially T 71, rescored to T 63 after I corrected a miskeyed age band in the scoring platform. Embedded validity indicators within expected ranges on all instruments; response style consistent across sessions.

Evaluation time log (my time only): 07/28/2026, 35 minutes: verified scoring, reviewed validity indicators, began integrating scores with interview and records. 07/30/2026, 58 minutes: interpreted results against history and collateral; drafted report. 08/04/2026, 55 minutes: finalized report (25) and feedback session with R.T. (30). Total evaluation time 148 minutes. Units: 96132 reported once for the episode plus one unit of 96133; a second add-on would have required 151 minutes.

Interpretation and clinical decision-making: Performance-based attention and memory measures are broadly average with low-average processing speed and no decline pattern. Elevated self-reported inattention contrasts with rescored observer ratings in the expected range. Integrated with interview, records, and symptom course, the picture is more consistent with anxiety-related attentional complaints than with a developmental attention disorder. Impression: generalized anxiety disorder (F41.1); adult ADHD not supported. Plan: CBT referral, sleep evaluation, no stimulant trial; full recommendations in the report.

Report, feedback, and sign-off: Comprehensive written report completed and filed 08/04/2026; copy to Dr. Rowan with R.T.'s authorization. Feedback provided 08/04/2026. Date of service for evaluation services reported as 08/04/2026, the date the episode concluded, with work on the three dates above reflected in this record. Signed: L. Chen, PhD, licensed psychologist, 08/04/2026.

This sample is fictional and for educational purposes. It does not describe a real patient, clinician, or practice.

↑ Back to the template and downloads

Why this sample works

  • The scores are verified, not transcribed. The rescored observer form shows the clinician checked the platform, the exact correction is recorded, and every value matches the output it came from.
  • Time is dated, named, and separated. Three dated entries each name the activity, technician administration time is walled off, and 148 minutes is checkable arithmetic rather than an estimate.
  • The unit math shows its work. The base code appears once for the episode, one add-on is earned, and the second add-on is explicitly not claimed, with its 151-minute threshold stated.
  • The date of service is defensible. The episode concludes at feedback, the claim carries that date, and the day-span sits in the record exactly as CMS date-of-service guidance expects.
  • The interpretation is the clinician's own. The synthesis ties verified scores to history and collateral and ends in a signed impression, with nothing delegated and no pasted computer narrative.

Writing these after every session? BastionGPT drafts complete notes from bullets, dictation, or a transcript.

Generate a note from bullets

Documentation and compliance considerations

Sort the obligations by strength before you template anything. LAW: coverage of clinical psychologist services rests on the Social Security Act and 42 CFR 410.71, plus general medical-necessity documentation rules; no federal text names a scoring or interpretation note. State law can go further: Massachusetts Medicaid regulation 130 CMR 411.413 requires, for psychological assessment records, behavioral observations, original responses, a summary of scores, and a comprehensive written report, a rare case of a statute defining the note's contents. PAYER POLICY: MAC coverage articles such as LCD L34646 with Article A57481, and L34520 with A57780, expect the record to support medical necessity, expect a written report to be generated, and set a 31-minute minimum before the base code; Nevada Medicaid adds prior authorization for 96130 to 96133 under announcements effective January 12, 2026. CONVENTION: the titled note itself, start and stop times, and any fixed report length; median reports ran 5, 6, and 8 pages across medical, rehabilitation, and private-practice settings in one study (Donders, 2001), and no regulation sets a length. The work itself is the interpreting clinician's own: Medicare does not pay for evaluation services performed by students or trainees, technicians cannot perform them, and master's-level clinicians are recognized unevenly by payer, so verify before the episode starts. Ontario adds a rule automated scoring makes urgent: the college standard carried forward as 14.8 bars substituting a computer-generated report for the psychologist's professional opinion.

Time and dates are the audit surface. Evaluation work done on a day the patient is not present is billable, and the episode is reported on the date the service concluded, per CMS MLN Matters SE17023, with the day-span reflected in the record; the base code is reported once per episode, and the claim waits until the evaluation is complete, typically the feedback date. Total dated minutes are the requirement; start and stop times are defensive convention, and entries like 'approximately 1 to 2 hours' are a documented denial risk. Expect scrutiny of long episodes: one commercial policy treats 2 to 8 hours as the usual complete evaluation and requires a battery list and rationale beyond 8. A cluster of CO-50 with N115, CO-16, or CO-151 denials on testing claims is the signal to pull your MAC's article and rebut it section by section. The format is a convention; the content is the requirement. Elsewhere the frame changes: Australia has no MBS item comparable to 96130, so psychological services bill as time-tiered attendance items and the record answers the Health Insurance Act's adequate and contemporaneous standard, where Reg 6 prescribes each entry's contents and notes added after the attendance do not count toward consultation time. Canadian obligations run through provincial colleges and privacy statutes rather than a testing code. Keep retention to the longest applicable rule: about 7 years for US adults under the APA guideline with a 6-year federal lookback in practice, Ontario 10 years, British Columbia 16 years from April 1, 2026, Alberta 10 to 11. The claim-side reconciliation is its own artifact, the claim-support billing note, and the clinical deliverable is the psychological evaluation report.

↑ Back to the template and downloads

Common testing scoring note errors auditors flag

The enforcement record around time-based mental health codes is concrete. In a May 2023 federal audit of pandemic-era psychotherapy services (HHS OIG report A-09-21-03021), reviewers examined 216 sampled enrollee-days and found time improperly documented or required information missing for 128 of them, with provider signatures absent for 54; the office estimated $580 million of roughly $1 billion in Medicare psychotherapy payments were improper. That audit is psychotherapy, not testing, and no published improper-payment rate isolates 96130 to 96133; what transfers is the failure modes, time and signatures, which are exactly what this note exists to prove. Australia shows the ceiling: a Professional Services Review determination effective 27 June 2025 ordered a practitioner to repay $300,000 and disqualified her from MBS billing for three years, with failure to keep adequate and contemporaneous records among the central findings on a psychological-strategies item, though the practitioner was a general practitioner rather than a psychologist. The BastionGPT Clinical Advisory Board sees the same errors most often in testing scoring note reviews:

  • Vague or undated time. 'Approximately 1 to 2 hours' and undifferentiated block entries cannot be matched to units and draw CO-151 denials. Each entry needs a date, a named activity, and minutes.
  • A base code every day. Reporting 96130 or 96132 for each sitting instead of once per episode, with combined time on the conclusion date, is the multi-day error auditors unwind first and payers recoup fastest.
  • Delegated interpretation. A technician or trainee drafts the synthesis and the psychologist signs it. Medicare does not pay for evaluation services performed by students or trainees; the performance requirement fails even when the signature is real.
  • Evaluation billed without administration. Test selection is bundled into evaluation services, so the evaluation codes ride with administration services from the same episode. A scoring note with no administration record behind it fails the pairing before medical necessity is even reached.
  • Scores that do not match the platform. Transcription drift and uncorrected auto-scoring errors, or a computer narrative pasted as the interpretation. Preserve exact values and test names, record every correction, and keep the synthesis in your own words.
How BastionGPT helps

BastionGPT is specifically trained, tuned, and clinically tested on psychological testing scoring notes.

  • Extract every score, index, and validity flag from scoring-platform output into a clean summary that preserves exact values and exact test names.
  • Check the note before you sign: dated minutes supporting the units, the base code once per episode, the date of service matching the episode's conclusion, and report and feedback status closed out.
  • Produce a de-identified score summary and time log for consultation, supervision, or teaching, with identifiers stripped.

See how clinicians use it day to day on the AI therapy notes page.

Many BastionGPT users report saving more than 90 minutes per day on documentation.

HIPAA-compliant with a signed BAA on every plan. Your data is never used to train models. BastionGPT drafts, you review and sign.

Frequently asked questions

Yes. Integrating data, interpreting results, and report writing count toward the evaluation-services codes whether or not the patient is present, and no regulation requires the patient in the room for that work. When a service begins on one day and concludes on another, CMS MLN Matters SE17023 directs that the claim carry the date the service concluded, with the record reflecting the span. In practice: log each day's evaluation minutes, report the combined episode on the conclusion date, usually feedback, and use the base code once. Top-ranking billing pages routinely omit this rule, which makes it the most valuable one to get right.

No CMS rule mandates literal start and stop clock times for 96130 to 96133. The codes are time-based, so the record must support the total time behind the units, and dated entries with a named activity and minutes satisfy that. Start-stop notation is a defensive convention that makes audits smoother, while vague entries such as 'approximately 1 to 2 hours' are a documented denial risk. If your MAC's billing article ever adds an explicit start-stop requirement, upgrade the log then; until that day, precise dated totals win.

MAC billing articles set a 31-minute minimum before the base code is assigned, and the base (96130 or 96132) is reported once per assessment episode even when the work spans several days. Add-on hours (96131, 96133) accrue past each full hour of combined evaluation time, so 148 minutes supports the base plus one add-on, and a second add-on would need 151. Expect scrutiny of long episodes: one commercial policy treats 2 to 8 hours as the range of a complete evaluation and asks for a battery list and rationale past 8 hours.

Not for Medicare. Administration and scoring can be delegated to a trained technician under the technician codes, but evaluation services are the interpreting clinician's own work, and Medicare does not pay 96130 or 96132 performed by students or trainees. A technician's role in administration never disqualifies the professional's evaluation claim. Master's-level clinicians such as LPCs and LCSWs are recognized unevenly: some commercial payers allow these codes and Medicare generally does not, so verify each payer in writing before the episode starts.

Three artifacts, three jobs. The administration record proves the giving and mechanical scoring of tests, work a technician can perform. The scoring note proves the interpreting clinician's own evaluation work: score verification, integration, dated minutes, and decisions. The psychological evaluation report or neuropsychological report is the clinical deliverable those minutes produce, and MAC articles expect it to exist. In a time-based-code audit, this note is what survives; the report alone rarely proves the minutes.

Yes, in exact form. Massachusetts Medicaid regulation 130 CMR 411.413 makes behavioral observations, original responses, and a summary of scores part of the required assessment record at the state-law level, and preserving exact values and exact test names is what lets a reviewer reconcile the note, the scoring-platform output, and the report. Verify computer scoring before interpreting anything: Ontario's college standard bars substituting a computer-generated report for the psychologist's professional opinion, a rule written for precisely the moment an auto-scored profile is tempting to paste.

Set retention to the longest rule that touches you. In the US the APA guideline is seven years after last service for adults and three years after a minor reaches majority, with state law able to override and federal auditors using a six-year lookback. Ontario requires ten years, or ten past age eighteen, whichever is later; British Columbia moves to sixteen years from April 1, 2026; Alberta works out to ten or eleven. In Australia, follow state rules and note that the Better Access referral itself must be kept for two years.

No. Australia has no separate scoring-and-interpretation item; psychological services bill as time-tiered attendance items, so US-style evaluation-services claims do not transfer. What Australia regulates hard is the record: the Health Insurance Act requires adequate and contemporaneous records, Reg 6 of the PSR Scheme Regulations prescribes each entry's contents, and MBS notes count only details recorded at the time of attendance toward consultation time, so material added later does not help. A 2025 Professional Services Review determination ordering a $300,000 repayment with a three-year billing disqualification shows the enforcement ceiling.

Yes. Give it the day's work, scores verified, anomalies found, minutes per activity, and it drafts the dated log, the score summary with every value preserved, and the sign-off block for your review. It can also extract exact scores from platform output into a clean summary, check a finished note for the gaps auditors flag (vague time, a missing report status, a base code repeated across days), and produce a de-identified version for consultation or teaching. BastionGPT is HIPAA-compliant with a signed BAA on every plan, and your data is never used to train models.

Educational content, not legal or billing advice. Sample notes are fictional. Follow your organization's policies and your board, payer, and jurisdiction requirements.