Skip across the chapter
Quota Playbook

Sales call scoring questions: 4 vendor rules

Quota Playbook Editorial Team · Published · 10 min read

Write AI call scoring questions an external observer can answer from the transcript.

Key takeaways for how to write ai call scoring questions for sales coaching

  • Write for an external observer because the AI does not know your company-specific terms, abbreviations, or internal processes, according to Outreach, and the model has no knowledge of your company, agents, or internal policies, according to Aircall.
  • Avoid asking about actions taken after the call or information not discussed, a restriction highlighted by Outreach.
  • Use ranged evaluation questions to produce a five-point score from Very poor to Excellent, according to Aircall.
  • Save drafts without adding all questions or with incomplete questions, a feature described by Avoma.
  • Use the AI Scorecard Evaluator tool to review your questions before publishing a scorecard, a tip provided by Aircall.

Write questions for an external observer to ensure accuracy

The AI model does not possess background knowledge of your organization. Aircall explains that since the AI model only receives the call transcript and has no knowledge of your company, agents, or internal policies, well written questions are essential for accurate and consistent scoring Aircall. This limitation means the evaluator cannot infer context from unspoken internal norms. You must provide that context explicitly within the question text itself.

Outreach advises writing for an external observer because the AI doesn't know your company-specific terms, abbreviations, or internal processes Outreach. If you use an acronym like "QBR" or a specific product name, the AI may not recognize it. Define these terms in the question prompt. For example, instead of asking if the rep mentioned "the platform," ask if the rep mentioned "the CRM platform, defined as the software used to manage customer relationships."

Aircall notes that this guide explains how to write clear and actionable custom AI questions for Call Scoring Aircall. Clarity is the primary defense against vague or incorrect evaluations. When you write a question, imagine you are explaining the call to a stranger who has never worked at your company. If the question requires insider knowledge to understand, rewrite it to include that knowledge.

Consider the difference between a vague prompt and a specific one. A vague prompt might ask, "Did the rep handle objections well?" This is subjective and relies on the AI's general understanding of sales. A specific prompt aligned with the external observer rule might ask, "Did the rep acknowledge the customer's concern about pricing by repeating the concern back to them before offering a solution?" This second question provides the exact behavioral cue the AI should look for. It removes the need for the AI to guess what "well" means.

A filled reference table of scoring rules by vendor

The following table summarizes documented rules and limitations for Aircall, Gong, Avoma, and Outreach scorecard features.

Publisher Rule or Feature Documented Detail
Aircall Transcript-only limitation The AI model only receives the call transcript and has no knowledge of your company, agents, or internal policies, making well-written questions essential for accurate and consistent scoring according to Aircall.
Gong Normalization scale Values are normalized, which means they are adjusted to a number within a certain scale; in this case, between 0 and 1 according to Gong.
Avoma Unique scorecard name This name should be unique, as it will be used by the call participants to identify the scorecard to use for that particular meeting according to Avoma.
Outreach External observer rule Write for an external observer because the AI doesn't know your company-specific terms, abbreviations, or internal processes according to Outreach.

Use this table to verify that your scorecard questions align with the specific constraints of your chosen platform before publishing.

See Sales coaching scorecard setup with 6 documented checks.

Avoid asking about actions taken after the call

Focus your scoring questions strictly on the dialogue that occurred during the recorded interaction. According to Outreach, you should avoid asking about actions taken after the call or information not discussed. The documentation explicitly flags questions like "Did the rep send a follow-up email?" as incorrect examples because they reference events outside the audio transcript.

This restriction applies to all question types covered in the guide, including rating scales, yes/no, open-ended, multiple choice, and single choice formats. According to Outreach, the best practices for writing coach card questions aim to produce accurate and consistent AI scoring results. When you include post-call actions, you introduce variables that the AI cannot observe, which undermines the consistency of the evaluation. The goal is to evaluate the sales motion as it happened in real time, not the administrative follow-up that may or may not have occurred later.

Use ranged evaluations for nuanced skill scoring

Ranged evaluation questions produce a five point score from Very poor to Excellent, according to Aircall. This format allows the AI evaluator to assign a specific grade within a defined spectrum rather than a binary judgment. When writing these questions, you must write the question as if the Yes/No criteria fields do not exist, according to Aircall.

The calculation of the final score involves specific mathematical adjustments to handle skipped items and standardize values. If a question is skipped during the evaluation process, the system assigns a specific value to it in the overall score. For example, if the skipped question is a range question from 1 to 5, the value given to it in the overall score is 1, according to Gong. This rule applies specifically to range questions within that numerical boundary.

To understand how individual scores contribute to the final result, you must look at the normalization process. Values are normalized, which means they are adjusted to a number within a certain scale; in this case, between 0 and 1, according to Gong. This step converts raw scores into a consistent format that can be compared across different questions and evaluations.

The normalization formula uses the minimum and maximum values of the range to determine the adjusted number. A specific calculation example shows the process: 7-1 = 6; (3-1)/6 = 1/3; normalized value is 0.33, according to Gong. In this instance, the raw score of 3 is adjusted based on the range limits of 1 and 7 to produce a normalized value of 0.33.

After normalization, the results undergo a final scaling step to align with the display format used in the scorecard. The results are then scaled and rounded to a value between 1 and 5, according to Gong.

Handle short calls and voice mail transcripts

Gong states that AI answers are not available for calls answered by voice mail, or calls with transcripts that have fewer than 100 words, according to Gong. This limitation means your scorecard will lack automated evaluations for these specific interactions. You cannot rely on the tool to grade a conversation that never actually happened between two humans.

Short calls present a different constraint. Gong notes that short calls or calls with limited dialogue may produce fewer outline sections, according to Gong. If the dialogue is minimal, adjust your expectations for the depth of the automated analysis. Use the available outline sections as a starting point, not a complete report. This approach keeps your coaching process realistic and aligned with the tool's documented capabilities.

See Sales call review: 5 notes to write first.

Illustrative example: three scoring questions

A coach drafts 3 questions, each scored from 1 to 5. Question 1 asks whether the rep repeated a pricing concern before a solution. Question 2 asks whether the rep defined a company term in plain words on the call. Question 3 asks whether the rep stayed with what the call actually discussed. The coach does not ask about a follow-up email after the call. The coach then tests the 3 questions before publishing.

Test your questions before publishing the scorecard

According to Aircall, you should use the AI Scorecard Evaluator tool to review your questions before publishing a scorecard. This pre-publish testing step allows you to verify that your wording aligns with the AI model's capabilities before the scorecard becomes active for your team.

When naming your new scorecard, ensure the name is unique. According to Avoma, the name should be unique, as it will be used by the call participants to identify the scorecard to use for that particular meeting. A distinct name prevents confusion when multiple scorecards are available for different types of sales interactions.

You can save your work in progress without completing every field. According to Avoma, drafts can be saved without adding any questions or with incomplete questions of a Scorecard. This draft saving feature lets you build the scorecard incrementally, returning to it later to add specific skill targets or adjust criteria.

Consider the specific context of the call you are evaluating. According to Avoma, you may wish to create scorecards which measure the performance of a participant of a specific type of call, such as Business Review, Training, or Kick-off, targeted at specific skill, such as performing a Discovery or Negotiation, or performing specific tasks, such as sending an agenda or meeting follow up email. Tailoring the scorecard to the call type and specific skill ensures the questions remain relevant to the actual conversation content.

Draft three new scorecard questions using the 'external observer' rule and test them with the AI evaluator tool before publishing.

FAQ: how to write ai call scoring questions for sales coaching

Why does the AI model need clear questions about company context?

The AI model only receives the call transcript and has no knowledge of your company, agents, or internal policies, according to Aircall. Because the model lacks this internal context, well written questions are essential for accurate and consistent scoring, according to Aircall.

What score value is assigned to a skipped range question?

If the skipped question is a range question from 1 to 5, the value given to it in the overall score is 1, according to Gong. This specific value assignment applies when the question is skipped during the evaluation process, according to Gong.

Can I save a scorecard draft without adding all questions?

Drafts can be saved without adding any questions or with incomplete questions of a Scorecard, according to Avoma. This feature allows you to store a scorecard before it is fully populated with evaluation criteria, according to Avoma.

Does AI scoring work on calls answered by voice mail?

Gong AI answers are not available for calls answered by voice mail, according to Gong. The system also excludes calls with transcripts that have fewer than 100 words, according to Gong. These two conditions prevent the generation of AI-based answers for the specified call types, according to Gong.

How are normalized values scaled in the final score?

Values are normalized to between 0 and 1, then scaled and rounded to a value between 1 and 5, according to Gong.

Sources

  1. Sales coaching scorecard setup with 6 documented checks
  2. Sales call review: 5 notes to write first
  3. Log call feedback for every rep: 4 vendor rules to check

More plays