AI assessment governance: keep training scores out of pay decisions
When scores are tied to people decisions, employees fear mistakes and avoid practice. This guide covers five reasons, Annex III of the AI Act, permitted and prohibited uses, conditions for exceptions and a sample policy text.
Short answer
An AI training score should not be tied directly to pay, bonus, promotion, discipline or dismissal, because it measures behaviour in a simulated conversation, the evaluation model can be wrong, the standard can be incomplete, and the result is affected by factors outside the employee's control. Turning the score into a people decision also destroys practice itself: employees become afraid of mistakes and avoid it.
The right approach is a written policy: what the score is used for, what it is not used for, and — if it is ever considered as one of several sources — under what conditions. The European Union's AI Act classes AI systems that evaluate workers' performance as high-risk, which shows how important such a policy is.
Why a score is unfit for people decisions
- Simulation is not realityThe score measures a conversation with an AI customer. A real customer, real pressure and real outcomes are something else.
- The model can be wrongAI evaluation can misread context, misrecognise speech or misapply a criterion.
- The standard is incompleteA score cannot be better than its standard. If the standard is outdated, correct behaviour scores low.
- No contextThe employee's workload, leads, health, territory — the score knows none of this.
- It distorts behaviourWhen scores are tied to people decisions, employees pick easy profiles, avoid risk and turn practice into an "exam".
The legal context
Point 4 of Annex III to the European Union's AI Act (Regulation 2024/1689) classes as high-risk AI systems intended for decisions affecting work relationships, including those intended "to monitor and evaluate the performance and behaviour of persons in such relationships". Such systems carry requirements for human oversight, transparency and risk management.
Whether the Act applies to your company depends on where it operates. In Azerbaijan, personal data and employment relations are governed by national legislation. Either way, if you are considering using an AI score in a people decision, consult a lawyer. The US NIST AI RMF framework also recommends clearly defining the human role and transparency in managing AI risk.
Policy: permitted uses
- Giving the employee learning feedback.
- Finding the team's recurring skill gaps and planning training.
- Choosing a topic for an individual coaching conversation.
- Checking the quality of profiles and standards.
- A signal about a new hire's readiness for real conversations, read together with the manager's own observation.
Policy: prohibited uses
- Linking pay, bonuses or incentives directly to the practice score.
- Using it as the sole or main basis for dismissal, disciplinary action or promotion.
- Selecting or rejecting job candidates on an AI simulation score alone.
- Showing scores to the team as an open league table.
- Using the score for another purpose without the employee's knowledge.
If a score is ever considered: conditions
Some companies may want to consider practice results as one of several sources — in reviewing a new hire's probation, for example. In that case the minimum conditions are:
- Several sourcesThe score is read only together with real-conversation reviews, the manager's observation and results.
- A human decisionA person makes the decision and writes down the reasoning.
- TransparencyThe employee knows in advance that results will be considered for this purpose.
- A right to challengeThe employee can challenge the result, and the challenge is checked against the transcript.
- Legal reviewThe process has been reviewed by a lawyer.
A sample policy text
The text below is illustrative and should be adapted with your lawyer: "The AI practice system exists so that employees can practise sales conversations in a safe environment. Practice results are used for learning and development. They are not used as a basis for decisions on pay, bonuses, discipline, promotion or dismissal. Employees see their own results and may request a review of a result they believe is wrong. Results can be viewed only by the employee, their direct manager and the person responsible for training."
How it differs from real-call scores
A quality score for real calls is a different situation: there is a real customer and a real outcome, and some companies use it in performance discussions. The same principles — a human decision, evidence, a right to challenge — apply there too; see appealing an agent AI score. A practice score must be handled even more carefully, because its link to real work is more indirect — see training score vs real sales.
Illustrative example: starting without a policy
This is not a real customer case. A retail company starts AI practice, and in the first month management holds "performance conversations" with the three employees with the lowest scores. The next month the number of sessions halves; the rest of the team picks only the easiest profile. The company adopts a written policy: practice scores are not used in people decisions, and results are seen only by the employee and their direct manager. Once the policy is announced, practice on hard profiles picks up again.
Common mistakes
- Starting practice without writing a policy.
- Showing scores to the team as a ranking.
- Saying "only one source" but making the score the main criterion in practice.
- Giving employees no way to challenge a result.
- Calling AI evaluation "objective" and dropping human review.
Limitations
This article sets out general principles and a sample policy; it is not legal advice. Only a lawyer can determine how the AI Act and local employment and personal data laws apply to your company. A written policy is not enough on its own — managers need to actually follow it.
In Vexvon AI Training
Vexvon AI Training is built for a sales team's practice and development: the product page describes it as something that "does not replace your salesperson". Every result comes with an explanation per criterion and references to the employee's messages, the rubric version and weights used are kept, and a conversation can be re-evaluated from the panel — which makes human review possible. Which decisions the score is not used for is set by the company's own policy.
Next step
Take the sample policy text, adapt it with your lawyer and HR, and announce it to the team before practice begins. For correcting a result see correcting a simulation score, and for the manager's role AI feedback and the manager; more articles are in rollout and reliability. To build it together, contact us.