Ask ten assessment psychologists what AI does for their reports and you will get ten different answers, because they are describing three different products. One of them is a chat window. One is a drafting tool built for our field. One works on the case rather than the document. They are sold in the same breath, demoed with the same before-and-after, and priced within a few dollars of each other.
The distinction is not marketing. It decides what you are carrying when you sign.
I am a licensed clinical psychologist. I have been doing assessment for about fifteen years, I still see patients, and I still sign reports. What follows is the way my team and I came to categorise the tools in our field after moving a large volume of reports through our own platform, and after several hundred conversations with assessment psychologists about what they are actually using.
The question is not whether it writes well
Every tool in all three categories writes well. That was the easy half of the problem, and it has been solved for two years.
Here is the harder half. Almost every practice is carrying a backlog, and a backlog clears at the speed cases close, not the speed they open. The write-up routinely outlasts the testing that produced it, and it is done by the most expensive person in the practice, transcribing a formulation they have already reached. So the field went shopping for typing speed, and hundreds of tools arrived promising a report in minutes.
Then the reports came back and everyone discovered the same thing: the faster the draft arrived, the more of it you had to check.
The useful question is not "does it write well." It is what does it do with a disagreement in my data? Parent says one thing and teacher says another. The rating scale points one way and the performance measure points another. Your clinical impression from the room does not match the questionnaire that came back in the post.
That moment is the whole job. It is also the moment that separates these three categories completely.
The landscape
There are three kinds of AI in our field
Pick one to see what it is, where it breaks, and what it does with the same piece of contradictory data.
Publicly available assistants · HIPAA-configured builds
What it is
A conversation, not a case file. You paste in what you have and it writes back. It produces the best prose of the three, which is exactly what makes it the most dangerous.
Where it breaks
Nothing persists. No score is tied to the instrument that produced it, and nothing is on record as approved. Fluency that is not standing on a structured case is how a report ends up confident and wrong.
You carry: the data, the reasoning, and the liability.
The same contradiction, every time
Maya, 14, referred for ADHD. On the Conners-4 Inattention scale her mother’s ratings come back at T = 72 and her teacher’s at T = 58. Everything else in the battery points the same way. Here is what this kind of system does with those fourteen points.
“Ratings across home and school settings are consistent with significant attentional difficulties, and Maya continues to present with inattentive symptoms across environments.”
The contradiction is gone. Not resolved, not weighed, gone. The sentence reads well enough that nobody goes looking for the two numbers underneath it, and neither number appears anywhere in the draft.
Purpose-built · template blocks · prompt libraries
What it is
Built for this work, and they solved a real problem. They can read every document you hand them; that was never the limitation.
Where it breaks
The limitation is what they do with a disagreement. A generator is built to keep you happy, so when the teacher and the parent contradict each other it smooths the seam. That gloss holds right up until someone pulls the record and the contradiction it hid is in it.
You carry: the synthesis. You just type less of it.
The same contradiction, every time
Maya, 14, referred for ADHD. On the Conners-4 Inattention scale her mother’s ratings come back at T = 72 and her teacher’s at T = 58. Everything else in the battery points the same way. Here is what this kind of system does with those fourteen points.
“Parent-report ratings on the Conners-4 fell in the Very Elevated range for Inattention, and teacher-report ratings fell in the High Average range. Overall, ratings support attentional concerns across raters.”
Better: both raters are named and the ranges are stated. But the last sentence still resolves the divergence rather than examining it, and the draft never asks the question the gap is actually posing.
Works on the case, not the document
What it is
The unit of work is the case file, not the paragraph. Structured data, a record of what a clinician approved, and reasoning that runs across the whole battery at once rather than section by section.
Where it breaks
It is slower, and it will not hand you a finished report in two minutes. Checking the full record and carrying provenance on every claim takes longer, and it shows its work.
You carry: the judgement. Which was the only part that was ever yours.
The same contradiction, every time
Maya, 14, referred for ADHD. On the Conners-4 Inattention scale her mother’s ratings come back at T = 72 and her teacher’s at T = 58. Everything else in the battery points the same way. Here is what this kind of system does with those fourteen points.
“Divergence flagged, rater discrepancy: Conners-4 Inattention, parent T = 72 (Very Elevated), teacher T = 58 (High Average). 14 points across settings. Candidate accounts to rule in or out before this is written up: structure and scaffolding available in the classroom that are absent at home; homework demands loading the home rater; differing rater thresholds; masking in a structured setting. Two items on the teacher form bear on this and are unscored. Your call: which account does the rest of the file support?”
Nothing is decided for you. The gap is on screen with both T-scores, both instruments and both raters attached, alongside the accounts worth ruling out, because in assessment that contradiction was never noise to tidy. It is the finding.
The case is synthetic and the outputs are illustrative, written to show the difference in handling rather than to reproduce any one product’s wording.
01. General-purpose AI
Publicly available assistants, and the HIPAA-configured builds your group may have stood up.
It writes the best prose of the three, which is exactly what makes it the most dangerous.
What you have is a conversation, not a case file. Nothing persists between sessions in any structured form. No score is tied to the instrument that produced it; the model only knows what you happened to paste and the order you pasted it in. Nothing is on record as approved by anybody, because there is no record, only a thread.
Fluency that is not standing on a structured case is how a report ends up confident and wrong. Run the same paste tomorrow and the emphasis moves. Neither run knows which T-score came from the teacher form and which from the parent form unless you told it, in prose, and remembered to tell it the same way twice.
There is a second problem that clinicians consistently underweight, and it has nothing to do with quality: that thread is a record. It is discoverable. I wrote more about that risk in using Claude and ChatGPT for psychological reports, because the failure case there is not a bad sentence, it is a transcript.
You carry: the data, the reasoning, and the liability.
02. Report generators
Purpose-built for our field: template blocks, prompt libraries, score-to-paragraph mapping.
These solved a real problem and I will not pretend otherwise. They were built by people who understand what a psychoeducational report has to contain, and they removed genuine drudgery from the back end of the process.
They can also read every document you hand them. That was never the limitation.
The limitation is what they do with a disagreement. A generator is built to produce a satisfying document, so when the teacher and the parent contradict each other it smooths the seam and glosses the divergent data. The paragraph reads beautifully. It reads beautifully right up until a records request pulls the file and the contradiction it papered over is sitting in the data that was handed to it.
In assessment, that contradiction was never noise to tidy. It is frequently the finding. A fourteen-point gap between home and school ratings is not an inconvenience on the way to a diagnosis, it is information about where the demands land, where the scaffolding is, and what happens to this child in a structured room versus an unstructured evening.
A document that hides it is worse than a blank page, because it is a blank page you have signed.
You carry: the synthesis. You just type less of it.
03. Clinical reasoning engines
Works on the case, not the document.
The third category treats the case file as the unit of work. Structured data rather than pasted prose. An explicit record of what a clinician approved and when. Reasoning that runs across the whole battery at once rather than section by section, because the patterns worth having do not respect section boundaries.
The test is not whether it writes well. It is whether it surfaces the data that diverges, works that divergence into the clinical reasoning, and helps you find the holes in the autopilot thinking we all deny and all engage in, then shows you the scores that say so.
Two things follow from that, and they are worth being honest about.
The first is that it is slower. We do not compete on the two-minute report. Reading the full record and carrying provenance on every claim takes longer, and it should.
The second is that it does not make the call. It cannot. The clinician renders the diagnosis, and a system that offers to do that for you is selling you something you cannot ethically buy. The engine's job is to make sure that when you make the call, nothing relevant was sitting unexamined in your own file.
You carry: the judgement. Which was the only part that was ever yours.
The two questions I ask before I sign my own name
Strip away the category names and there are two demands I make of anything that touches a report I am going to sign.
One. Show me the evidence that contradicts this diagnosis.
Not a paragraph that sounds careful. I want the T-scores, the instrument, and the session the divergence came from. If a system cannot produce those, it is not reasoning about the case. It is writing around it.
Two. Show me where this claim came from.
Every line in a report I sign has to rest on something I can open: the measure, the document, the session note, the recorded observation. A fluent restatement that points nowhere is precisely the thing that gets found later. You can see what that looks like on a finished document in our worked example of a fully sourced report.
Those two questions will sort any demo you sit through into one of the three categories above, usually inside five minutes.
How to tell which one you are actually looking at
We published six requirements for AI-assisted psychological assessment because nobody else had, and because a standard that only applies to other people's products is a marketing document. This one applies to ours. Put them to any vendor as questions:
- Provenance. Does every clinical claim resolve to an identifiable source, and can you open it from the claim?
- Disconfirmation. Is the evidence that diverges from the conclusion surfaced without being asked for?
- Approval. Can machine-generated content enter the record without an explicit, logged act of clinician approval? It should not be able to.
- Voice. Does the report read in the clinician's own voice and follow their own standards, because they are the person who signs it?
- Auditability. Does every action carry a timestamp, an actor and a role, exportable on demand?
- Disclosure. Is the use of AI stated in the record itself, rather than buried in a licence agreement?
A general-purpose assistant fails one, three and five outright. A report generator usually passes one and four, sometimes three, and struggles with two by design. Nothing in this field passes all six by accident, including us: we have had to build for each one of them deliberately.
What this means for your practice
If you are using a chat window today, you are not doing anything shameful; you are doing what the tooling allowed. But you are personally holding the parts a platform is supposed to hold, and the thread you are holding them in is discoverable.
If you are using a report generator, you have a real tool, and you should keep asking it the second question. Watch what it does the next time two raters disagree. If the draft resolves the disagreement instead of putting it in front of you, you now know exactly what you are checking for on every case.
And if you are evaluating anything that calls itself a reasoning engine, make it prove the first question live, on a case with a contradiction in it. Ask it to argue against your formulation. A system that cannot do that on demand is a typewriter with better marketing.
The report was never the hard part. Deciding what is true about the person in front of you, and being able to show why, is the hard part. That is the part that stays yours.
PsychAssist.ai is an end-to-end platform for assessment psychology, from intake to signed report and every document the report owes after it, built on a clinical reasoning engine. Watch the Summer Release keynote for the live run.
Sources
- The Defensible AI in Psychological Assessment Standard - PsychAssist.ai
- APA Ethical Principles of Psychologists and Code of Conduct - American Psychological Association
- The PsychAssist.ai Summer Release keynote and full transcript - Dr. Chris Barnes