The Defensible Assessment Documentation Standard

Version 1.0 · Published 11 September 2026 · Free to read, cite or adopt

The Defensible Assessment Documentation Standard (DADS) sets out six requirements that documentation in psychological assessment must meet when any part of it has been produced with machine assistance. It is written for practising clinicians, for the school and health systems that receive their reports, and for anyone evaluating software that touches a clinical record.

This standard was written by Dr. Chris Barnes, a practising clinical psychologist and the founder of PsychAssist. PsychAssist is one of the systems measured against it. That is a conflict of interest and it is stated here so you can weigh it. The requirements below are written to be applied to any system, including ours.

Why this exists

Psychological reports are already being written with machine assistance. Not occasionally, and not only by early adopters. Drafts are being produced in general purpose chat tools, in dictation software that summarises, in electronic health records that suggest text, and in purpose built platforms. A profession that spent decades establishing what constitutes an adequate evaluation has, within about three years, changed how the resulting document gets written, and has not yet said out loud what that requires.

The vacuum has been filled with assurances. Vendors say their systems are clinician reviewed, HIPAA compliant, and secure. Those are three different claims, none of which addresses whether a sentence in the finished report is true. A platform can be perfectly HIPAA compliant while producing a paragraph that attributes a score to a measure that was never administered. Compliance describes how data is held. It says nothing about whether the argument in the document holds.

What the field lacks is not enthusiasm or caution, both of which are abundant. It lacks a shared description of what an AI assisted report has to be able to do - one specific enough that a clinician can hold a product against it in a demonstration and get a yes or a no. A requirement nobody can be measured against is not a requirement; it is a preference. Everything below is written so that it can be tested in an afternoon, by a clinician, without engineering help.

Two boundaries are worth stating at the outset. This standard describes documentation practice, not clinical practice: it takes no position on which measures belong in a battery or how a differential should be resolved. And it addresses defensibility, not admissibility. Whether evidence is admitted is determined by courts and hearing officers under rules such as Daubert and Frye, and no documentation practice and no vendor can promise it. What good documentation can do is put a clinician in a position to account for every claim they signed.

The six requirements

  1. Provenance

    Every clinical claim must resolve to an identifiable source: a measure, a document, a session note, or a recorded observation.

    A psychological report is an argument. It moves from data to inference to conclusion, and each move is only as sound as the datum underneath it. When a report is drafted with machine assistance, the chain from datum to sentence is the first thing to break, because a language model will produce a fluent sentence whether or not a number stands behind it. Fluency is not evidence.

    Provenance means that for any sentence in the report that makes a clinical claim, a clinician can ask "where did this come from?" and get an answer that is specific: the WISC-V Processing Speed Index of 79, the WAIS-V Working Memory Index in an adult case, the Conners 4 parent-report T-score of 74 on Inattention, the third-grade report card uploaded on 14 March, the observation recorded during the second testing session. Not "the assessment data." Not "the record." The specific datum.

    This requirement is deliberately indifferent to how the sentence was produced. A claim written by a clinician at two in the morning with no source behind it fails the requirement in exactly the same way a generated one does. The standard is about the record, not about the tool.

    What this rules outA generated narrative paragraph that reads plausibly but contains a score, a date, or a history detail that appears nowhere in the case file.

    How to test itOpen a finished report, pick three sentences at random from the Interpretation section, and trace each one to its origin. If any of the three requires you to remember, infer, or ask a colleague, the report has not met the requirement.

  2. Disconfirmation

    A diagnostic conclusion is defensible only when the evidence that diverges from it has been surfaced and considered, not only the evidence that supports it.

    Confirmation bias is the oldest failure mode in clinical reasoning and machine assistance makes it cheaper. A system asked to write the case for ADHD will write the case for ADHD. It will not volunteer that the CPT-3 was unremarkable, that teacher ratings on the BASC-3 fell in the average range while parent ratings were clinically elevated, or that symptom onset as reported does not clearly predate age twelve.

    Divergent evidence is not an inconvenience to be managed. It is the part of the record that makes the conclusion worth anything. A report that presents only convergent findings has not demonstrated that a differential was considered; it has demonstrated that a conclusion was reached and then decorated. Under cross-examination that distinction is the whole case.

    This requirement does not demand that every discrepancy change the conclusion. Most will not. It demands that the discrepancy appear in the record, with the clinician's reasoning about why it does or does not alter the finding. The sentence "teacher ratings did not corroborate parent-reported inattention; this pattern is consistent with the more structured classroom environment and does not preclude the diagnosis" is worth more to a hearing officer than three paragraphs of agreement.

    What this rules outAn Interpretation section in which every cited finding points the same direction, in a case where the raw data did not.

    How to test itTake a completed evaluation and ask the system to list the evidence in the file that argues against your primary diagnosis. If it cannot produce that list, or produces a list you already knew was incomplete, the requirement is not met.

  3. Approval

    No machine-generated content enters a clinical record without an explicit, logged act of clinician approval.

    Approval is not the absence of objection. A draft that appears in the record because nobody deleted it has not been approved; it has been tolerated. The requirement is for a positive act - a clinician looked at this specific content and accepted it - and for that act to leave a trace.

    The distinction matters most at the moments when it is least convenient. On a Friday afternoon, with six reports outstanding, the difference between a system that requires you to accept each section and a system that silently promotes drafts to final is the difference between a record you can defend and a record you will have to explain.

    Logged approval also protects the clinician in the other direction. When a finding is challenged eighteen months later, the log establishes what was reviewed and when. Without it, the clinician's only answer is recollection, and recollection is not a defence.

    What this rules outAutosave-to-final. Bulk 'accept all' with no record of what was accepted. Any workflow in which content reaches a signed document without passing through a human decision.

    How to test itAsk the vendor to show you the approval event for a specific paragraph in a specific report: who approved it, at what time, and what the content was at the moment of approval.

  4. Voice

    The report must read in the clinician's own voice and follow their own standards, because they are the person who signs it.

    This is usually treated as an aesthetic preference. It is not. A signature is an assertion of authorship, and a report that does not sound like its author has a defect at the point where professional responsibility attaches. Parents notice. Attorneys notice. Colleagues who have read your work for a decade notice immediately.

    Voice is also a proxy for something harder to measure. Clinicians differ in how they introduce a diagnosis, how much hedging they consider honest, whether they lead with strengths, and how the same finding is framed in an IEP evaluation report, in a neuropsychological report, and in a primary care summary written for the referring physician. Those choices encode clinical judgement developed over a career. A system that flattens them into a house style has not saved the clinician time; it has substituted its judgement for theirs and left them holding the liability.

    The practical test is whether the clinician has to rewrite. Time spent restoring your own voice to a generated draft is time the system did not save, and it is the most common reason clinicians abandon these tools in the first quarter.

    What this rules outA house style applied to every clinician on the platform. Generated text that requires line-by-line rewriting before the clinician is willing to sign it.

    How to test itHave two clinicians in the same practice, with genuinely different writing styles, generate a report from the same case data. If the two outputs are difficult to tell apart, the requirement is not met.

  5. Auditability

    Every action carries a timestamp, an actor and a role, exportable on demand.

    An audit trail that cannot be exported is not an audit trail; it is a feature of somebody else's software. The requirement has three parts and all three are load-bearing. Timestamp, because sequence is what a challenge turns on. Actor and role, because "the practice" is not a person and a supervising psychologist's approval is not the same event as a psychometrist's data entry. Exportable, because the record may be needed by a party who will never have a login.

    The relevant events are broader than most systems log. Document upload, score entry and correction, draft generation, edit, approval, signature, export, and disclosure to a third party are all points at which a record can be challenged. So is deletion. The record is also the substantiation for a billed service - CPT 90791 for the diagnostic evaluation, CPT 96136 for test administration and scoring - which means an audit trail serves a second purpose the clinician did not ask for but will be glad of.

    Auditability is also what makes the other five requirements verifiable rather than merely asserted. Provenance without a log is a claim. Approval without a log is a claim. The audit trail is the evidence that the rest of the standard was actually followed.

    What this rules outLogs visible only to the vendor. Audit data that requires a support ticket to obtain. Retention windows shorter than the record-retention obligation in the clinician's jurisdiction, which for anyone practising across state lines under PSYPACT means the longest of several.

    How to test itExport the full audit trail for one completed case, yourself, from the interface, in a format you can open. Time how long it takes. If the answer involves contacting anyone, the requirement is not met.

  6. Disclosure

    The use of AI in preparing documentation is stated in the record itself, not buried in a licence agreement.

    Disclosure that lives in a terms-of-service document is disclosure to nobody. The people with a legitimate interest in knowing how a report was prepared - the family, the school team, the referring physician, the hearing officer - will never read the vendor's licence. They will read the report.

    The disclosure does not need to be lengthy or defensive. A short statement of what assistance was used and what the clinician retained responsibility for is sufficient, and it is better placed in the report's methods or procedures section - alongside the battery, the DSM-5-TR criteria applied and any ICD-10 codes assigned - than in a footer. Clinicians who have adopted this consistently report that it defuses the question rather than raising it.

    There is a professional-ethics dimension as well. Informed consent for assessment describes the procedures to be used. Where machine assistance materially shapes how documentation is produced, describing it at consent - and again in the report - is the conservative reading of existing obligations, and it costs the clinician nothing.

    What this rules outAI use disclosed only in the end-user licence agreement, the privacy policy, or the vendor's marketing site. Reports that are silent on how they were produced.

    How to test itRead the last report you signed. If a parent could not determine from the document itself whether machine assistance was used, the requirement is not met.

Applying the standard to psychoeducational evaluations

A psychoeducational evaluation is the setting where these requirements are tested hardest, because it is the setting where the report is most likely to be read by someone looking for a weakness. When a family disagrees with an eligibility determination, the evaluation report becomes the central document in a due process proceeding under IDEA. It is read line by line by an advocate, an attorney, and a hearing officer, none of whom were in the room and all of whom are entitled to understand how the school psychologist got from a set of scores to a recommendation.

Provenance is the first thing examined. If the report states that a student’s WISC-V Working Memory Index fell at the 9th percentile, the protocol is expected to show it, and the WIAT-4 subtest scores supporting a finding of specific learning disorder in reading are expected to reconcile with the discrepancy analysis in the eligibility section. Where a district uses a pattern of strengths and weaknesses model, the entire determination rests on numbers that must resolve to a source. Third grade report cards, prior intervention data, and the referral itself are sources too, and a claim about a student’s classroom history that cannot be traced to a document in the file is a claim that will not survive questioning.

Disconfirmation is where evaluations most often fail, and it is the requirement with the most direct consequences for a family. Parent ratings on the Conners 4 may be markedly elevated while teacher ratings on the BASC-3 sit within normal limits; a CPT-3 may be unremarkable in a student whose inattention is obvious at home. An evaluation report that omits the divergence has not made the case for ADHD, it has curated it, and a hearing officer who discovers the omitted data in the protocol will discount the rest of the document. The report that states the disagreement and reasons through it is the one that holds. The same logic applies to autism spectrum disorder evaluations where observational and rating scale data pull apart, and to cases following traumatic brain injury where premorbid estimates complicate every comparison.

Approval and auditability carry particular weight in school settings because evaluations are frequently team products. A psychometrist administers, a school psychologist interprets, a supervisor signs. If a generated paragraph enters an IEP evaluation report without a recorded act of approval by a named person in a named role, the district cannot establish who is professionally accountable for the sentence being challenged. Records held under FERPA are subject to inspection by parents, which makes an exportable log a practical necessity rather than a theoretical one.

Voice matters here for a reason specific to schools. A school psychologist’s report has to be intelligible to a special education team, a general education teacher, and a parent in the same sitting, while remaining precise enough to support a 504 plan or an eligibility category. Clinicians develop that register over years. A report that arrives in a generic house style will be rewritten before the meeting, or worse, will not be.

Disclosure closes the loop. A parent who learns during a due process hearing that the evaluation report was drafted with machine assistance, and that nobody told them, has been handed an argument that has nothing to do with the quality of the evaluation. A single sentence in the procedures section removes it. Together, the six requirements are what make an evaluation legally defensible: not a guarantee about how a hearing will come out, but a report whose every claim the school psychologist who signed it can stand behind under questioning.

Twenty questions to ask any vendor

Each of these has a right answer and a wrong answer, and most of them can be settled by a demonstration rather than a statement. Ask for the demonstration. If a vendor answers a question about their product with a claim about their compliance posture, that is itself an answer.

  1. ProvenanceShow me a claim in a generated report and trace it to its source - can you do that in the interface, or does it require a support ticket?
  2. ProvenanceWhen the system produces a score in narrative text, is that score read from the entered protocol data, or generated as text? Show me where the value is stored.
  3. ProvenanceIf I upload a school record and the report references it, does the citation point to the specific document and page, or only to the case file as a whole?
  4. DisconfirmationCan the system list, on demand, the evidence in this case that argues against the diagnosis I have selected?
  5. DisconfirmationWhere in the generated draft does divergent evidence appear - is it integrated into the interpretation, or appended where it can be skipped?
  6. DisconfirmationIf parent and teacher ratings disagree, does the draft state the disagreement, or does it report the elevated informant only?
  7. ApprovalCan any machine-generated sentence reach a signed report without a clinician explicitly accepting it? Demonstrate that it cannot.
  8. ApprovalShow me the approval record for one paragraph: who approved it, when, and what the text said at that moment.
  9. ApprovalIf a supervisee drafts and a supervisor signs, are those two distinct logged events with distinct roles?
  10. VoiceHow does the system learn my voice - from my own prior reports, or from a template I select from a list?
  11. VoiceGenerate the same case for two clinicians in my practice with different styles. Are the outputs distinguishable?
  12. VoiceCan I encode my own rules - how I introduce a diagnosis, how I handle discrepancies, how I frame recommendations - or only choose from yours?
  13. AuditabilityExport the complete audit trail for one case, right now, in this call, in a format I can open.
  14. AuditabilityWhich events are logged? Name them. Is document deletion among them?
  15. AuditabilityHow long is audit data retained, and does that period meet the record-retention requirement in my jurisdiction?
  16. DisclosureDoes the generated report contain a statement of AI assistance in the document itself?
  17. DisclosureCan I edit the wording of that disclosure to match my consent form and my jurisdiction's guidance?
  18. DisclosureIs the disclosure on by default, or is it something I have to remember to enable for each report?
  19. GeneralIs my clinical data used to train models that serve other customers? Show me the clause, not the summary.
  20. GeneralIf I leave, what do I take with me - the reports as PDFs, or the structured case data and audit trail as well?

Get the checklist as a PDF.

The twenty questions, one page, generated from this document at build time so it can never drift from what you just read. Take it into a vendor call. Or download it without giving us anything.

How PsychAssist measures against the standard

Stated by the party with the conflict of interest declared at the top of this page. Nothing in this table describes a roadmap item; every row describes behaviour present in customer accounts today. Where the evidence can be seen without an account, the third column links to it.

RequirementHow it is metWhere you can verify it
ProvenanceReport content is generated from entered protocol data and uploaded documents rather than from free text, and each generated section is linked to the case records it drew on.See the worked example, where every claim resolves to a measure, document, or observation
DisconfirmationDivergent findings in the case data are surfaced to the clinician during interpretation rather than being omitted from the draft, and are carried into the narrative where the clinician retains them.See convergent and divergent evidence distinguished in the example report
ApprovalGenerated content is presented as a draft that a clinician must explicitly accept. Nothing reaches a finished report without that acceptance, and the acceptance is recorded.Visible in the platform on any case
VoiceVoice is derived from the clinician's own prior reports and their own configured rules for how diagnoses are introduced, discrepancies handled, and recommendations framed - not from a shared house template.How voice profiles are built
AuditabilityCase actions carry a timestamp, an actor, and a role, and are exportable by the account holder without vendor involvement.Security and record-handling practices
DisclosureAn editable statement of AI assistance is available for inclusion in the report body, so the disclosure travels with the document rather than sitting in a licence agreement.See the disclosure statement in the example report's procedures section

Definitions

Convergent evidence
Findings from independent sources that point toward the same conclusion - for example, elevated parent and teacher ratings on the BASC-3 together with a low WISC-V Working Memory Index in a case of suspected ADHD. Convergence across methods and informants is what raises a hypothesis to a defensible conclusion.
Divergent evidence
Findings in the record that argue against the conclusion being drawn, or that complicate it. A CPT-3 within normal limits in a case of parent-reported inattention is divergent evidence. Its presence in the report is a requirement of this standard, not an optional refinement.
Provenance
The traceable origin of a clinical claim: the specific measure, document, session note, or recorded observation from which it derives, identified precisely enough that a reader can locate it in the case file.
Approval gate
A control that prevents machine-generated content from entering a clinical record until a clinician has performed an explicit, recorded act of acceptance. A gate that can be satisfied by inaction is not an approval gate.
Audit trail
A chronological, exportable record of actions taken on a case, each carrying a timestamp, an actor, and the actor's role. Distinguished from version history, which records what changed but not necessarily who decided it or in what capacity.
Psychoeducational evaluation
An assessment of cognitive, academic, and behavioural functioning conducted to answer questions about learning and school performance, typically informing eligibility determination under IDEA or accommodations under Section 504. Usually combines a cognitive measure such as the WISC-V, an achievement measure such as the WIAT-4, and behavioural rating scales.
Neuropsychological evaluation
An assessment of brain-behaviour relationships across domains including attention, memory, language, visuospatial ability, and executive function, using measures such as the WAIS-V, the D-KEFS, the CPT-3, the BRIEF-A, the TOMM and the PAI, often following traumatic brain injury, in complex differential questions, or where medical aetiology is in play.
Battery
The set of measures selected for a given evaluation. A battery is a clinical decision, driven by the referral question, and the report should make the rationale for its composition legible.
Standardised measure
An instrument administered and scored under fixed procedures, with normative data permitting comparison of an individual's performance to a reference population. Departures from standardised administration must be documented and their effect on interpretation stated.
T-score
A standard score with a mean of 50 and a standard deviation of 10, used by most behavioural rating scales including the Conners 4, the BASC-3, and the BRIEF-A. A T-score of 70 is two standard deviations above the mean and is conventionally treated as clinically significant.
Percentile rank
The percentage of the normative sample scoring at or below a given score. A percentile rank of 16 corresponds to a standard score of roughly 85, one standard deviation below the mean. Percentile rank is not the percentage of items answered correctly and should never be reported to families without that distinction being made.
Due process
The dispute-resolution procedure available to families under IDEA when they disagree with a school district's identification, evaluation, educational placement, or provision of a free appropriate public education. A due-process hearing is the setting in which an evaluation report is most likely to be examined line by line.
Eligibility determination
The decision, made by a team including the family, as to whether a student meets criteria for special education under one of the disability categories defined by IDEA, and whether the disability results in a need for specially designed instruction. The evaluation report is the evidentiary basis for that decision. A DSM-5-TR diagnosis and an ICD-10 code are not the same thing as an IDEA eligibility category, and a report that conflates them creates work for the team.

Frequently asked questions

  • Is it ethical to use AI to write psychological reports?

    The profession's existing obligations already answer most of this. A psychologist is responsible for the content of any document they sign, must base conclusions on adequate data, and must describe their procedures honestly. Machine assistance does not change any of those duties, and it does not create new ones so much as raise the cost of neglecting the old ones. Using AI to draft documentation is ethical where the clinician can trace every claim to a source, has considered evidence that diverges from the conclusion, has explicitly approved what enters the record, and has disclosed the assistance. It is not ethical where a signature is applied to text the clinician cannot account for.

  • Can an AI-assisted report survive a due process hearing?

    A report survives scrutiny on the strength of its evidentiary chain, not on how the sentences were typed. What is examined in a due-process hearing is whether the evaluation was comprehensive, whether the measures were appropriate to the referral question, whether the conclusions follow from the data, and whether the clinician can explain their reasoning under questioning. An AI-assisted report that satisfies the six requirements here is in a stronger position than a hand-typed report that cannot source its claims. An AI-assisted report that cannot is in a considerably weaker one. Note that admissibility is determined by courts and hearing officers, not by vendors; what a documentation practice can do is make a report defensible.

  • Does the clinician still have to review everything?

    Yes. Requirement 3 makes this explicit and it is not negotiable in any jurisdiction we are aware of. The clinician who signs the report is professionally and legally responsible for its contents, and no workflow, disclaimer, or vendor assurance transfers that responsibility. What good tooling changes is not whether you review but how efficiently you can review - a claim you can trace in one click is faster to verify than a claim you have to reconstruct from a stack of protocols.

  • Do I have to tell parents AI was used?

    Requirement 6 says the use should be stated in the record itself. Whether a specific jurisdiction or licensing board compels it is a separate question and the answer varies; several state boards and professional associations have issued guidance since 2024 and more are expected. The conservative position, and the one this standard takes, is that assessment consent describes procedures, machine assistance is part of how the documentation was produced, and disclosing it in the report costs the clinician nothing while removing a line of attack. Clinicians who disclose consistently report that families ask fewer questions about it, not more.

  • What platforms support legally compliant psychoeducational evaluations?

    No platform makes an evaluation legally compliant - a platform supports a clinician who is responsible for compliance. What you should look for is a system that keeps the referral question, the battery rationale, the data, and the decision rules intact and inspectable; that logs who did what and when; that supports the documentation IDEA and Section 504 processes actually require; and that will export your records if you leave. The twenty questions above are written to be taken into a vendor call for exactly this purpose. Be sceptical of any vendor that answers a compliance question with a compliance claim rather than a demonstration.

  • What is the difference between legally defensible and legally admissible?

    Admissibility is a determination made by a court or hearing officer about whether evidence may be considered at all, under rules such as Daubert or Frye and the applicable rules of evidence. No vendor and no documentation practice can guarantee it. Defensibility is a property of the work: whether the evaluation was adequate to the question, whether the conclusions follow from the data, and whether the clinician can account for every claim under questioning. This standard addresses defensibility. Treat any vendor claiming to deliver admissibility as making a claim they are not in a position to make.

  • Does this standard apply to neuropsychological evaluations as well?

    Yes. The six requirements are written to be indifferent to the type of evaluation. A neuropsychological report drawing on the D-KEFS, the CPT-3, the TOMM and the PAI has the same obligations as a psychoeducational report drawing on the WISC-V and the WIAT-4: every claim sourced, divergent findings surfaced, content approved, voice preserved, actions logged, assistance disclosed. If anything the density of data in a neuropsychological battery makes Requirement 1 harder to satisfy by memory and more valuable to satisfy by design.

  • Who wrote this standard and what is their interest in it?

    It was written by Dr. Chris Barnes, a practising clinical psychologist and the founder of PsychAssist. PsychAssist is one of the systems measured against it, which is a conflict of interest and is stated at the top of this page as well as here. The requirements are written so they can be applied to any system, including systems that do not exist yet, and the twenty questions are written so they can be asked of any vendor including this one. If the standard is adopted, stewardship is intended to move to an independent body.

  • Can I use this standard in my own practice policy or district procurement?

    Yes, and that is the intent. The standard is published free to read, cite, or adopt under a Creative Commons Attribution licence. Practices have used the six requirements as a documentation policy, and the twenty questions as a procurement checklist. No permission is required and no attribution beyond naming the standard is expected.

Governance and stewardship

Version 1.0 is maintained by Dr. Chris Barnes. Revisions will be versioned and dated, and superseded versions will remain reachable so that anyone who cited an earlier one can still show what it said. Corrections, disagreements and proposed additions are welcome and will be published with attribution where the author permits. If this standard achieves broad adoption, we intend to move its stewardship to an independent body.

Published under a Creative Commons Attribution 4.0 licence. Cite as: Barnes, C. (2026). The Defensible Assessment Documentation Standard, Version 1.0. Related reading: a complete psychological assessment report, fully sourced and school psychologist AI report writing and IDEA alignment.