Scoring rubric - Setup
Written By Jozef
Last updated 15 days ago
The scoring rubric is the list of requirements the AI uses to score a candidate. Without it, the AI scores against the job description alone β with it, every answer is measured against criteria you chose, and the result page shows you exactly which ones the candidate met.
This article covers where to find it, how to generate it, how to check its quality, how to test it against real answers before you publish, and what each part of the underlying model actually does.
1. Where it lives
Interview creator βΈ open the interview βΈ Scoring section.
The rubric can only be created or edited while the interview is in draft. Once you set it to active it becomes read-only β that is deliberate: it keeps candidates scored yesterday comparable with candidates scored next month. To change a rubric on a live interview, duplicate the interview, edit the copy, and publish that.
2. What a rubric is made of
A rubric is a flat list of criteria, each tagged with one of three levels. The level is the bar the criterion sets β it is not a description of a weak or a strong candidate.
Two rules follow from this, and they are the ones most rubrics get wrong:
Never put the same topic at two levels. "Basic budget awareness" as weak and "Advanced budget management" as strong is one topic written twice β it inflates the list and confuses the scorer. One topic, one criterion, one level.
The levels do not have to be balanced. Uneven buckets are normal and expected.
How a criterion should be written
Each criterion becomes a chip on the result page with a tick beside it, so it is written as a competency noun phrase, 4β9 words, no leading verb, no full stop:
And it must satisfy three tests:
Verifiable in conversation β provable from what the candidate says. Knowledge, reasoning, attitude. Never a CV fact ("5+ years of Python"), a metric ("98% CSAT"), or a physical skill β the AI cannot hear those.
One thing per chip β no "and"/"or". A compound requirement makes partial impossible to judge; split it into two chips.
Positively phrased β name what the candidate should demonstrate. Never "cannot", "lacks", "no β¦". The assessor assigns the outcome; the chip only names the thing.
3. Generating the rubric
In the Scoring section, pick an interview type and press Generate. It takes roughly 20β40 seconds.
Available interview types:
Add your questions first. The generator reads the interview's actual questions and derives criteria from them, so that every criterion is something the interview can genuinely surface. It knows the live interviewer also asks automatic follow-up probes, so one open question can support several distinct criteria β but a criterion no question could reach is dropped rather than invented.
What else feeds the generation:
job title, short description and long description,
the knowledge base attached to the interview, when there is one,
seniority level β detected from the title and description if you do not set it (entry-level, intermediate, senior, managerial, director, executive). It changes both the vocabulary and how demanding each bar is.
Without questions, the generator falls back to the job description and aims for about 12 criteria. With questions, there is no fixed count β it produces as many as the questions honestly support.
Prefer to write your own? Press From scratch for an empty rubric and add criteria by hand under each tab. Clear scoring removes the rubric entirely and returns the interview to description-only scoring.
The generated rubric is written in the interview's own language, and everything below works the same way in any of the supported languages.
4. Verifying it β the quality check
Press Verify scoring. This is an audit, not a rewrite: it takes about 25 seconds and returns a quality score out of 100 plus one suggestion per problem criterion, which you accept (β) or reject (β) individually. Nothing changes until you accept it.
Issues it flags:
Accepting a suggestion either rewrites the criterion, moves it to another level, or deletes it β the chip on the card tells you which. A score of 100 with no suggestions means the rubric matches the standard; anything below about 80 is worth a second pass.
Re-run Verify scoring after you edit questions. Criteria that were confirmable before may not be any more.
5. Testing it before you publish β score simulation
Interview βΈ Score simulation lets you re-score interviews that already happened using a draft rubric, so you can see the effect on real answers before it touches a real candidate.
Generate or edit the rubric in the simulation panel (it is stored separately β your live rubric is untouched).
Adjust the score weights if you want to test those too.
Run the simulation. Results appear as a scatter chart of original score against simulated score, so outliers are easy to spot.
Happy with it? Copy the criteria into the interview's real rubric. Not happy? Clear the simulation and try again.
Simulation costs 0.15 credits per re-scored result, and the run is queued β the panel shows progress and you can refresh it.
This is the single most useful step before publishing a high-volume role. A rubric that reads well can still score everyone 6.5; the simulation is where you find that out.
6. Question-specific rubrics
You can override the rubric for one question: open the question, switch on Question specific scoring rubric, and fill in its own Strong / Moderate / Weak lists.
A question-specific rubric replaces the interview-level one for that question β it is not added on top of it β and it wins over everything, including a simulation rubric. Use it only where an interview-level criterion genuinely cannot express what that one question needs (a hard technical check, a compliance question with a single correct answer). For everything else, keep the rubric at interview level: it is easier to maintain and produces more consistent scores.
7. How the rubric turns into a score
Two things happen with it, and they are independent.
Per answer. Each answer is scored 1β10 against the rubric. Depth, examples and elaboration move a score within a band; they do not move it between bands. Short but accurate answers are fine by design, and the scorer is instructed to credit what was said rather than how fluently it was said β speech-to-text artefacts, filler words and accents are not penalised.
Per session. At the end, every criterion is marked against the whole transcript as one of:
Each one comes with evidence: a quote from the candidate, an AI assessment note, a link straight to the answer in the transcript, or a line from the CV. That evidence is the part recruiters use most β it turns a number into something you can defend in a hiring conversation.
A large number of not discussed markers is a signal that the rubric and the questions have drifted apart: the criteria are asking for things the interview never gets to. Re-run Verify scoring.
The weights around it
The rubric determines what good looks like. The Scoring section's weight sliders determine how much each signal contributes to the final number. Defaults:
Answer-level components are combined first, then blended with the session score. Same section: the maximum number of attempts (default 3) and the score at which further attempts stop being offered (default 7.5).
8. Showing the rubric to others
On the result page the rubric appears as the Scoring rubric block, with the status chips as filters β click Missing to see only what a candidate did not demonstrate.
In a PDF or export, tick AI scoring rubric. It only renders when AI recruiter assessment is also on, since the rubric is part of that assessment block.
9. Troubleshooting
"Generate" is greyed out, or the whole section is read-only. The interview is not in draft. Duplicate it and edit the copy.
No interview type in the dropdown. You have to pick one before generating. For assessment interviews the only option is process-verification-from-knowledge-base, and it needs a knowledge base attached.
Generation produced very few criteria. Expected when the interview has few questions, or narrow closed ones. The generator will not invent criteria the interview cannot surface. Add or broaden the questions and generate again.
Every candidate scores about the same. Usually the criteria are too easy to satisfy or too vague. Run a score simulation against past results, then move the borderline criteria up a level or make them more specific.
Lots of "not discussed" on results. The rubric no longer matches the questions. Run Verify scoring and act on the No question covers this flags.
Scores don't reflect the criteria I wrote. Check for a question-specific rubric on that question β it silently replaces the interview-level one.
I need to change the rubric on a live interview. Not possible by design. Duplicate the interview, change the copy, publish it, and send new invitations from it. Existing results keep the rubric they were scored with.