QA evaluation sheets
Definition
QA evaluation sheets
QA evaluation sheets are structured, weighted scorecards a QA analyst uses to grade an agent’s call, chat, or email against fixed criteria — greeting, accuracy, empathy, resolution, and compliance — so scores can feed coaching, calibration, and monthly client reports.
Key takeaways
- A QA evaluation sheet turns a subjective interaction into a comparable, weighted score.
- Most BPOs run one sheet per channel and one per programme, not one master form.
- Contact centres typically grade 4–8 calls per agent per month for statistical validity.
- Fatal errors auto-zero the sheet regardless of other line items.
- Calibration keeps analyst-to-analyst score variance under about 5%.
Most BPO floors run one sheet per channel and one per programme. A voice form for a telco looks different from a chat form for e-commerce, but the bones are the same: header, weighted criteria, comments box, and a final percentage.
The form is the unit of measurement. Calibration sessions, coaching plans, bonus calculations, and even client QBRs all trace back to the score on the sheet. Get the sheet wrong and the rest of the operation drifts.
ICMI’s contact-center resource library treats the scorecard as the single most important QA artefact in a contact centre.
How it works
A scorecard moves through a fixed loop. Most contact centres run 4–8 evaluations per agent per month, enough to be statistically useful without crushing the QA team’s bandwidth or the team leader’s coaching time.
| Step | What happens | Who owns it |
|---|---|---|
| 1. Sample selection | Pull calls randomly or by trigger (escalations, repeat callers, low CSAT) | QA lead |
| 2. Score the form | Mark each line item, add comments, calculate weighted total | QA analyst |
| 3. Calibration | Multiple analysts score the same call to align on standards | QA + ops |
| 4. Agent feedback | Team leader walks the agent through scores and agrees a coaching focus | Team leader |
| 5. Trend reporting | Aggregate scores feed weekly dashboards and client reports | QA manager |
A typical voice scorecard splits into three buckets. Soft skills — greeting, tone, empathy, active listening — usually carries 25–35% of the weight. Process compliance takes another 30–40%. Resolution and accuracy closes out the rest.
Fatal errors auto-zero the score regardless of other line items. Typical fatals include a missed compliance disclosure, a data-protection breach, abusive language, or a rude exchange with the customer.
The COPC Customer Experience Standard, now in Release 8.0, is the benchmark most enterprise BPOs design their forms against.
It pushes operators to grade human agents, chatbots, and self-service against the same outcome metrics, so the scorecard reflects real customer experience, not just script adherence.
Examples
Four operators show how scorecards flex across programmes, channels, and geographies while keeping the same weighted-criteria bones, with voice grading for regulated accounts, chat forms for e-commerce, and dedicated fatals for outbound compliance.
Concentrix (Manila, ongoing). The world’s largest CX provider runs programme-specific scorecards across its Philippine sites for banking, healthcare, and tech clients.
Forms weigh resolution and compliance heavily for regulated accounts, and calibration sessions run weekly to keep analyst-to-analyst variance under 5%.
Teleperformance (global, 2024 deployment). TP’s TAP framework layered speech analytics on top of manual QA sheets across multiple sites.
Auto-scored items (silence, talk-over, sentiment) feed the same sheet humans grade, so the final form blends machine and analyst marks.
Foundever (formerly Sitel Group). Foundever publishes scorecard guidance that treats the form as a coaching artefact first and a compliance tool second. “What good looks like” examples sit beside every line item so agents can self-review before their next call.
A mid-sized Cebu BPO running outbound sales (2025). A 20-line voice scorecard with three fatals: no Do-Not-Call check, no recording disclosure, no payment-card masking.
Analysts grade 6 calls per agent per month, and any fatal triggers an immediate refresher session before the agent returns to queue.
Related terms
- Quality assurance: the broader function that owns the scorecard and the analysts running it.
- Call calibration: the session where analysts grade the same call and reconcile scores.
- Average handle time: a productivity metric that sits alongside QA scores, not inside them.
- First call resolution: an outcome metric many scorecards weight heavily.
- Customer satisfaction score: the customer-side companion to the analyst-side QA score.
- Net promoter score: a loyalty metric clients often pair with QA scores in quarterly reviews.
- Service level agreement: the contract that defines what “passing” QA actually means for a client.
FAQ
What is a QA evaluation sheet used for?
It grades an agent’s interaction against fixed, weighted criteria so you can coach consistently, calibrate across analysts, and report performance to clients. The form turns a subjective call into a comparable number.
How many calls should be evaluated per agent per month?
Most contact centres land between 4 and 8 calls per agent per month. Fewer than four and the sample isn’t statistically useful; more than eight tends to overwhelm the QA team without changing the coaching picture.
Who fills out the QA evaluation sheet?
A dedicated QA analyst, not the agent’s direct team leader. The analyst owns scoring objectivity, the team leader owns the coaching conversation that follows, and calibration sessions keep multiple analysts aligned on what each line item really means.
What’s the difference between QA scores and CSAT?
QA scores are internal: an analyst grading the agent against your standards. CSAT is external: the customer rating their own experience. The two correlate but not perfectly, which is why mature operations track both and reconcile the gap each month.
What counts as a fatal error on a scorecard?
A fatal is anything that auto-zeros the score regardless of other line items. Typical fatals include missed compliance disclosures, data-protection breaches, abusive language, or providing incorrect financial or medical information.
Can QA evaluation sheets work for chat and email too?
Yes. Voice forms grade tone and active listening, chat forms grade typing accuracy and channel etiquette, and email forms grade structure and resolution completeness.
Explore more OA terms and guidance at Outsource Accelerator







Independent




