How it works
How the report cards are made
Every grade traces back to a real clinical scenario, a structured rating instrument, and the expertise of our Community Expert & Accountability Panel.
How to read the grades
These report cards are a research and education tool. Here’s what the scores can — and can’t — tell you.
How many raters
Each response is currently rated by a limited pilot sample of expert raters — the exact count is shown on every card.
When it was evaluated
Each report card shows the month (or month range) its ratings were collected, along with the specific model version used.
Important limitations
Grades are preliminary and based on a small pilot sample of raters; as more raters review each response, scores may shift.
Chatbots change constantly — a grade reflects one model version at one point in time.
The same prompt can produce different answers; the response shown is one representative sample raters rated.
Not clinical advice, and not a recommendation to use AI
A higher grade means raters judged a response higher-quality on our instrument — not that a chatbot is safe or appropriate to use with real clients. Nothing here substitutes for professional judgment, supervision, or established clinical protocols.
