A QA score is the grade a customer service conversation receives when it is reviewed against a quality scorecard. The scorecard breaks service quality into weighted criteria, accuracy of the answer, policy compliance, tone, whether the issue was actually resolved, and the reviewer marks each one. The marks roll up into a percentage for the conversation, and conversation scores roll up into QA scores for an agent, a team, or a whole operation. Most programs set a passing threshold around 85 to 90 percent and read the trend line rather than any single review.
Two things decide whether the number means anything. The first is the scorecard: criteria vague enough to argue about produce scores consistent enough to ignore, which is why calibration sessions exist. The second is coverage. In a human operation, QA scoring runs on samples, a few percent of conversations, because reviewer hours are scarce. A sampled score is an inference about the operation. It can be a good inference, but a repeating error in the unsampled majority hides inside it for weeks.
That is also why a perfect score deserves suspicion in a sampled program: 100 QA on five conversations says little about the other thousand. The claim only becomes meaningful when every conversation carries a reviewable record, which is exactly what changes when automation enters the queue. An automated resolution is not scarce to review; it documents itself. Coverage stops being the constraint, and the scorecard becomes the whole game.
Sampled QA score vs full-coverage QA score at a glance
| Dimension | Sampled score | Full-coverage score |
|---|---|---|
| Basis | A few percent of conversations | Every automated resolution |
| What it tells you | An inference about the whole | The state of the whole |
| A repeating error | Hides between samples | Surfaces in the record |
| A 100% score | Statistical noise | A verifiable claim |
Aide, the agentic AI platform for customer experience, is built for the full-coverage column: the Action Trace records what the AI did and why on every conversation, so a QA score reads from complete records instead of a sample, and a finding points at the specific intent to fix rather than a number to file.
Frequently asked questions
- How is a QA score calculated?
- Reviewers mark each conversation against weighted scorecard criteria, accuracy, compliance, tone, resolution. Criteria scores roll into a conversation percentage, and conversation scores aggregate into agent and team QA scores over time.
- What is a good QA score in customer service?
- Most teams set the pass line between 85 and 90 percent, but the honest answer depends on scorecard difficulty and coverage. A high score on a strict, calibrated scorecard with broad coverage means service is good; the same number on a loose scorecard and a thin sample means little.
- Can AI conversations be QA scored?
- Yes, more completely than human ones. Every automated resolution carries its own record, so scoring can cover all of them instead of a sample, and score drops trace back to the specific automation that caused them.