What Should an AI Sales Roleplay Score Actually Tell You?

A rep completes an AI roleplay and scores a 90 against your rubric.

They covered the value proposition, handled an objection, and asked for a next step. The feedback says they demonstrated strong discovery skills.

Then you listen to the conversation.

The rep asked about the buyer’s priorities but never explored the answer. They responded to a pricing concern before understanding what was behind it. The simulated buyer agreed to another meeting anyway.

In that situation, what would an 90 tell you about the rep’s readiness? And would it be accurate?

For enablement leaders evaluating AI roleplay, that’s a question worth answering before scores become part of certification, onboarding, or manager dashboards.

A recent research paper offers a useful starting point. In Sell More, Play Less: Benchmarking LLM Realistic Selling Skill, researchers built 1,805 sales scenarios and a simulated customer trained on more than 8,000 crowdworker-involved conversations. Their evaluation considered both sales progression and the customer’s expressed buying intent. Read the paper.

The researchers tested AI models acting as sellers. They didn’t test whether human reps improved after practicing with AI, so the study doesn’t establish coaching effectiveness or revenue impact.

It does provide useful ideas for evaluating a practice environment. The questions below are practical applications for enablement teams.


Start with the behavior you want the rep to demonstrate.

“Improve discovery” leaves a lot open to interpretation. A more useful objective might be: the rep can uncover how a stated problem affects the buyer’s team and confirm that understanding before recommending a solution.

That gives the exercise something observable to assess.

Imagine a buyer says implementation capacity is their biggest concern. A rep might acknowledge it and continue with the presentation. Another might ask who would own implementation, what else that team has committed to, and what would make the proposed timeline workable.

Your scoring criteria should explain how those responses differ and point to the evidence in the conversation. A manager should be able to understand why the rep received the score and what they should practice next.


Check whether the buyer gives the rep a credible challenge.

The simulated buyer shapes the exercise. If it volunteers every relevant detail, accepts vague answers, or agrees to a meeting after an unresolved concern, you’ll have trouble judging the rep’s performance.

The researchers addressed a related problem: simulated customers could drift into behaving like helpful assistants. Their customer model reduced that role confusion, though it didn’t eliminate it. Customer simulation methodology.

For your own pilot, try a few deliberately weak responses. Skip a follow-up question. Give an unsupported answer. Ignore a stated constraint.

Watch what the buyer does.

A credible response might be to repeat the concern, ask for evidence, or decline the proposed next step. The reaction should fit the scenario. You’re checking whether the conversation gives the rep meaningful consequences for their choices.


Ask what supports the score.

One useful feature of the research was its requirement for customer-side evidence. An objection counted as resolved when the customer’s later responses showed that their concern had eased. Scoring methodology.

You can apply that idea when reviewing a roleplay tool.

If the feedback says “strong objection handling,” ask it to identify the objection, the rep’s response, and the buyer’s reaction. Then have a manager review that explanation.

Consider a buyer who says, “We’ve tried something like this before, and adoption was poor.”

The rep responds with a list of features. The buyer replies, “Okay, send me some information.”

A manager would need more evidence before concluding that the adoption concern had been addressed. The scoring system should be equally careful.

Feedback becomes more useful when a rep can trace it to a specific moment and understand what another response could have accomplished.


See whether the skill holds up in a different conversation.

Repeating an exercise can help a rep become comfortable with a situation. You’ll also want to know whether they can apply the skill when the details change.

Keep the learning objective consistent and vary the buyer’s circumstances.

For an implementation discussion, one buyer might have a clear project owner and an uncertain timeline. Another might have executive support but no available implementation team. A third might reveal the capacity problem only after the rep asks about competing priorities.

Review whether the rep adjusts their questions and recommendations.

You should also define what a good outcome looks like for each scenario. Sometimes the appropriate result is a specific next step. Sometimes it’s recognizing a poor fit, surfacing a dependency, or agreeing that the timing won’t work.

A scoring system that rewards agreement in every situation could encourage behavior you wouldn’t want in a real opportunity.


Connect practice to coaching.

Before using roleplay scores to certify readiness, have managers independently review a sample of the same conversations using a shared rubric.

Compare their assessments with the tool’s feedback. Where do they agree? Where does the tool award credit that managers can’t substantiate? Where do managers disagree with one another?

Those differences can help you refine the rubric, improve the scenario, or identify limits in the tool.

Then follow the practiced behavior into real calls.

If the exercise focused on exploring business impact, managers can look for whether reps ask relevant follow-up questions, confirm what they heard, and use that information later in the conversation.

You’ll need further evaluation to determine how much improvement came from the practice program. But checking for transfer gives you a concrete way to connect roleplay with ongoing coaching.

For a first pilot, choose one skill, build a small set of varied scenarios, and agree on what evidence would demonstrate progress. Review the conversations alongside the scores and give reps a specific behavior to work on.

When a rep earns an 90, their manager should be able to explain what they demonstrated, where they still need help, and what to look for on the next customer call.

Next
Next

Enablement's Forgotten Role in Change Management