OROVA.VN — BIZ AI AGENT
Guides

Interview Questions to Ask Candidates, Built From Your Own Criteria

Orova 35 views
Interview Questions to Ask Candidates, Built From Your Own Criteria

It is Friday afternoon and five people are in a room trying to remember Tuesday. You interviewed six candidates for one opening. Everybody was pleasant. Everybody had a story ready about a difficult project. Now somebody has to say who gets the offer, and the only sentences anyone can produce are "I liked her energy" and "he seemed a bit quiet". The interview questions to ask candidates came off a list somebody pasted into the calendar invite twenty minutes before the first call, so every interviewer asked something slightly different, wrote three lines of notes, and formed a private impression that nobody else can check.

That is not an interviewing skill problem. It is a design problem. A borrowed list of questions produces a conversation, and a conversation produces a feeling. What you needed was a comparison, and a comparison needs the same questions, scored against the same things, for every person who walked through the door.

The fix, in two sentences: build your questions from the weighted criteria you already used to screen resumes, one or two questions per criterion, and write down in advance what a strong answer contains and what a weak one sounds like. Then split those criteria across your interview rounds so no criterion gets tested twice and none gets forgotten.

This article walks through the conversion — criterion to question to scoring note — with a complete worked example for one role, six criteria and the full question set they produce. It also covers how to probe evidence you found in the CV, which follow-up prompts turn a rehearsed story into something you can score, how to take notes that survive a debrief, how to split criteria across rounds, the fairness case for asking everyone the same core questions, and the four question types that waste the slot they occupy.

What are the best interview questions to ask candidates?

The best interview questions to ask candidates come from your own scoring criteria, not a borrowed list. Take each weighted criterion, write one or two questions that force the candidate to describe real work against it, and decide in advance what a strong answer contains and what a weak one sounds like.

That answer sounds obvious written down. It is almost never what happens. The usual sequence is that somebody writes a job description to attract applicants, somebody else screens resumes on gut feel, and then a third person searches for good questions to ask candidates in an interview an hour before the first slot. Three stages, three different definitions of what the role needs, and no way to connect the answer a candidate gave in stage three to the requirement that was written in stage one.

When your questions come from your criteria, three things become true at once. Every answer maps to a number you were already going to record. Every interviewer is measuring the same thing, so their scores can be compared instead of averaged into mush. And when a candidate asks why they did not get the job, you have a specific, statable reason instead of a shrug.

Why generic question lists produce generic interviews

Search for interview questions and you will find the same forty questions on every page. Tell me about yourself. What is your greatest weakness. Where do you see yourself in five years. Describe a time you handled conflict. There is nothing wrong with any of them as English sentences. The problem is what they do to a batch of candidates.

A generic question has no right answer, so it has no wrong answer either. Six candidates answer "tell me about a time you handled conflict" with six competent stories, and every story sounds fine, because anyone who has worked for three years has a conflict story and has told it before. You end up scoring delivery: who was more fluent, who made better eye contact, who had the tidier structure. Delivery is a real skill for a small number of roles and almost irrelevant for the rest, and it correlates with confidence, rehearsal and how much the person resembles you.

There is a second failure that is harder to see. Generic questions are not tied to anything, so nobody knows which of them mattered. If a candidate gives a weak answer to "what is your greatest weakness", what score moves? Nothing moves. The answer goes into an impression, the impression goes into a debrief, and the debrief goes to whoever argues most confidently.

Compare that with a question built from a criterion. If criterion 1 is "has personally owned a renewal number for a book of accounts, weight 9", then the question is about renewals they personally owned, the strong answer contains a book size, a number they were accountable for and what they did when it slipped, and the weak answer describes what the team did. A weak answer moves criterion 1 down. Criterion 1 is worth 9 points of the total. The conversation had a consequence.

The list is not useless, it is just the wrong starting point

Borrowed lists are fine as raw material. Once you know your criteria, skimming a list of interview questions to ask applicants can remind you of a phrasing you would not have thought of, or surface a question that tests a criterion better than your own draft. Use them at step four, after the criteria exist, not at step one instead of criteria. The order matters more than the source.

Where the criteria come from

This whole approach assumes you already have weighted criteria. If you screened resumes properly, you do. The criteria are the requirement lines from the job description, each rewritten to be testable, tagged with a group — must-have, important, preferred, basic, flexible — given a weight from 1 to 10, and paired with the evidence you expect to see. We covered how to build that document in job Posting Template: Write One, Get a Hiring Scorecard Too, and the interview stage is simply the second half of the same file's life.

If you do not have criteria yet, spend the hour before you write questions. You cannot skip to the interview and improvise the standard, because improvising the standard is exactly the thing that made Friday afternoon unpleasant.

One distinction is worth stating clearly, because it decides which criteria belong in an interview at all. Screening criteria are the subset you can verify from a document. Interview criteria are the ones a document cannot show. Most roles have both, and plenty of criteria live in both places at different depths — the CV tells you someone claims to have run onboarding, the interview tells you whether they can describe how.

  • Verified on paper, confirmed in person. Years, titles, tools, scale. The CV asserts it; one question checks the assertion is real and finds out what the person actually did inside that title.
  • Invisible on paper, only testable in person. How someone handles disagreement, whether they ask useful questions, how they explain something complicated to a non-expert, what they do when the information they need does not exist.
  • Neither, and you should admit it. Culture fit, passion, drive. If you cannot say what evidence would settle it, it is not a criterion. It is a preference, and preferences masquerading as criteria are where inconsistency gets in.

Turning one criterion into an interview question

Here is the conversion, five moves, done once per criterion. Give it ten minutes for the first one and about three minutes each after that. Six criteria takes under half an hour, and you reuse it every time you reopen the role.

Move 1: state what the criterion is actually asking

Before you write a question, write the plain-English thing you want to know. Criterion: "has run onboarding for a technical product end to end." What you want to know is whether this person has personally owned the sequence from signed contract to a customer using the product without help — not whether they attended onboarding calls, not whether their company had an onboarding process.

This step catches a lot of bad questions before they exist. Half of all weak interview questions are weak because the interviewer never decided what the question was for.

Move 2: write the question so the candidate must describe real work

Ask about something that happened, with a specific enough frame that a general answer will not fit. The pattern that works most reliably: walk me through the last time you [did the thing]. Start from [a concrete starting point] and tell me what you did, not what the team did.

"The last time" beats "a time". "A time" lets the candidate pick their best-rehearsed story from five years ago. "The last time" gets you something recent, less polished and far more informative. If the last time was genuinely unrepresentative, they will tell you, and that is useful too.

Ask for their part explicitly, in the question. Otherwise you get a fluent account of what a team did, and you spend two follow-ups separating the person from the group.

Move 3: write what a strong answer contains

Before the interview, in the same file, list three or four things a strong answer would include. Not phrasings — contents. For the onboarding criterion: names the steps in order, gives a rough duration, says who else was involved and what those people did, describes at least one thing that went wrong and what they changed afterwards, and can state how they knew the onboarding had succeeded.

This list is what turns your reaction into a score. When you finish the answer, you are not asking yourself "did I like that". You are checking four boxes, and the number of boxes ticked becomes the score.

Move 4: write what a weak answer sounds like

This is the step everyone skips, and it is the one that saves the debrief. A weak answer is rarely silence. It is usually fluent, confident and empty, which is exactly the thing that fools tired interviewers at four in the afternoon.

Write down the specific shape weakness takes for this criterion. For onboarding: describes the company's process rather than their own actions, stays entirely in the present tense and the abstract ("we would typically..."), cannot name a single thing that went wrong, or attributes every decision to a manager. Once that is written on the page in front of you, a smooth empty answer stops being persuasive, because you can see it matching your own description.

Move 5: write two follow-up prompts in advance

Every question needs at least two follow-ups ready, because the first answer is the rehearsed one and the second is where the information is. We will go deeper on follow-ups shortly, but write them at the same time as the question so you are not inventing them under pressure while also taking notes.

Five steps that turn one weighted criterion into a scored interview question: state what you want to know, write the question, list what a strong answer contains, list what a weak answer sounds like, prepare two follow-ups
The conversion. Every criterion you decide to test in an interview goes through the same five moves before the question is allowed on the sheet.

A worked example: six criteria and the question set they produce

Everything below is invented to show the mechanics. The company, the role, the criteria and the numbers are made up. Take the shape, not the specifics.

The role: Customer Success Manager, mid-market accounts, at a 40-person company selling scheduling software to healthcare clinics. The CS team is three people. The person joins to own roughly 45 accounts, take over renewals from the founder, and run onboarding for new clinics.

The screening criteria that survived from the job description, with the weights used to score resumes:

#CriterionGroupWeightTested where
1Has personally owned a renewal or retention number for a book of accountsMust-have9Phone screen and hiring manager
2Has run onboarding for a technical product from contract to active useMust-have8Hiring manager
3Can read product usage data and act on it before the customer complainsImportant7Working session
4Has handled an escalation where the customer was ready to leaveImportant6Hiring manager
5Writes to customers in a way an executive would take seriouslyPreferred5Working session
6Has used a CRM or customer success platform dailyBasic3Phone screen

Six criteria, 38 weight points, three rounds. Now the questions.

Criterion 1: owned a renewal number (must-have, weight 9)

Question. "Walk me through your last renewal cycle. How many accounts were you responsible for, what number were you measured on, and what did you personally do in the eight weeks before the biggest renewal came up?"

A strong answer contains: a book size, the specific metric they were held to and whether they hit it, a description of what they did that was different for the accounts at risk, and at least one renewal they lost with a plain explanation of why.

A weak answer sounds like: "we had a really strong retention rate" with no personal number; account management described as reactive check-ins; the word "we" throughout; no losses, ever, which is either untrue or means the book was not really theirs.

Follow-ups: "Which renewal in that cycle were you least sure about, and what did you do about it?" and "Who decided the discount when one was needed?"

Criterion 2: ran onboarding end to end (must-have, weight 8)

Question. "Take the last customer you onboarded from signature to using the product properly. Tell me the steps in order and what you did at each one."

A strong answer contains: a sequence with rough timings, the handoffs to and from other teams, what data or setup the customer had to supply and how they got it out of them, one thing that went wrong, and how they knew onboarding was finished rather than just concluded.

A weak answer sounds like: a description of the company's documented process in the abstract; "the implementation team handled the technical part" with no visibility into it; onboarding declared complete when the kickoff call happened; no definition of success beyond the customer being live.

Follow-ups: "What was the longest one took and what caused the delay?" and "What would you change about that process if you owned it here?"

Criterion 3: reads usage data and acts on it (important, weight 7)

Question. "Tell me about a time you knew an account was in trouble before the customer told you. What did you see, and what did you do next?"

A strong answer contains: a specific signal — logins dropping, one team using it and another not, a champion leaving, a support ticket pattern — and where they saw it; the action they took; and whether it worked. Bonus: a case where they were wrong about the signal.

A weak answer sounds like: "I always keep in close contact with my customers"; monitoring described as a quarterly business review; the answer moves immediately to relationship language without ever naming a piece of data.

Follow-ups: "What did you look at every Monday morning?" and "Who built that view, and what would you do here if nobody had built one yet?"

Criterion 4: handled a real escalation (important, weight 6)

Question. "Describe the last time a customer told you, or told someone above you, that they were leaving. Start from when you found out."

A strong answer contains: what the customer's actual complaint was under the stated complaint; who they involved internally and how fast; what they promised and whether they could deliver it; the outcome, including if the customer left anyway; and what changed afterwards so it happened less.

A weak answer sounds like: a story where the customer was simply wrong and was talked around; no mention of anyone else being involved; a happy ending with no detail about the mechanism; or a hypothetical answer, because they have not had one.

Follow-ups: "What did you say in the first message you sent them?" and "Looking back, what was the earliest point you could have caught it?"

Criterion 5: writes credibly to customers (preferred, weight 5)

Question, delivered as a short exercise rather than a question. "Here is a situation: a clinic's practice manager emails on a Friday saying the scheduling sync has been dropping appointments since Tuesday and they are considering going back to paper. Take fifteen minutes and write the reply you would send." Then discuss it.

A strong answer contains: an acknowledgement of the specific problem rather than a generic apology, a statement of what is known and not yet known, a concrete next step with a time attached, no blame, and a length a busy person would actually read.

A weak answer sounds like: three paragraphs of apology before any information; promises of a fix with no basis; internal jargon; or a message that reads like a template with the name changed.

Follow-ups: "What would you have done differently if this were the third time this month?" and "What would you send internally at the same time, and to whom?"

Criterion 6: CRM or CS platform in daily use (basic, weight 3)

Question. "Which system did your account data live in, and what did you personally put into it every week?"

A strong answer contains: a named tool, a description of what they recorded and when, and an opinion about what the tool did badly. People who really used a system daily always have a complaint about it.

A weak answer sounds like: a tool name from the skills section with nothing behind it; "the ops team kept it updated"; no view on it at all.

Follow-ups: "What did you do when the data in it was wrong?"

Column chart of weight points across six criteria in the worked example: renewal ownership 9, onboarding 8, usage data 7, escalation 6, customer writing 5, CRM use 3
Weight points per criterion in the invented example, 38 in total. The two must-haves carry 17 of the 38, which is where most of your interview time should go.

Look at what the question set does that a generic list cannot. Every minute of interview time is allocated in proportion to what the role needs: the renewal criterion is worth 9 points and gets a question plus two follow-ups plus a second pass in a different round, while the CRM criterion is worth 3 and gets ninety seconds. Nobody has to decide in the moment what to dig into. And when two interviewers disagree about a candidate, they are disagreeing about a specific criterion with written expectations, which is an argument that can be settled.

Questions that probe evidence you already found in the CV

You screened the resume before you invited the person. That screen produced claims: five years in the title, this tool, that scale, a project with a number attached. Those claims are the cheapest interview material you will ever get, and most interviewers waste them by asking the candidate to summarise their CV out loud.

The rule is simple. Do not ask candidates to repeat what is on the page. Ask about the space around it.

  • Scale claims. The CV says "managed a portfolio of 60 accounts". Ask what the smallest and largest looked like, and how differently they treated them. Someone who really managed 60 accounts has a segmentation opinion. Someone who inherited a spreadsheet does not.
  • Number claims. The CV says "improved retention by a third". Ask what the number was before, how it was calculated, and what else changed in the company that year. A real number survives all three questions. An inherited or borrowed number stops at the second.
  • Title and scope gaps. The title says Manager but the responsibilities read individual contributor. Ask directly how many people reported to them and what they were accountable for at review time. This is not a trap, it is a clarification, and titles genuinely mean different things at different companies.
  • Timeline gaps and short stints. Ask plainly and once: "there is a gap here, what were you doing?" Take the answer at face value and move on. What you are checking is that the CV and the person describe the same working life, not that every month was productive.
  • Tool lists. A skills section with fourteen tools is a keyword list. Pick the two that matter for your role and ask what they used them for last week or last month. Depth on two beats breadth on fourteen.

There is a discipline here worth naming. If you scored resumes against criteria and quoted evidence for each score, your interview prep is already done — you have, per candidate, the exact line in the file that supported each score and the criteria where evidence was thin. Interview time then goes to the thin ones. That is the practical payoff of evidence-linked screening, and it is the same principle that makes any scored decision reviewable: a score without a quotable reason is an opinion with a number on it, which we argued at length in AI Advertising: What It Actually Does for Your Ads.

Follow-up prompts: where the actual information is

The first answer to any behavioural question is the prepared one. That is not dishonest, it is normal — the candidate has told this story before and has smoothed it. The prepared answer tells you what the person thinks is impressive. The follow-up tells you what happened.

You need maybe seven prompts in total, reused across every question. Learn them once.

  1. "What did you personally do?" The highest-yield question in interviewing. Ask it every time an answer arrives in the first person plural. Ask it twice if needed.
  2. "What happened next?" Most prepared stories end at the decision. The interesting part is the consequence — whether it worked, what broke, what they did in month two.
  3. "Who disagreed with you, and what did they say?" Real projects have opposition. An answer with no opposition in it is either a very small project or a story with the difficult parts removed.
  4. "What would you do differently?" Tests whether the person has thought about the work since. A specific regret is a strong signal. "Nothing, it went well" from someone describing a complicated project is a weak one.
  5. "How did you know it worked?" Forces a measurement or an observable outcome. It quietly separates people who ship from people who participate.
  6. "Walk me through the first thing you did on the Monday." Requests concrete detail at a level nobody can fabricate smoothly. If the story is real, the detail is there.
  7. Silence. Not a question. After an answer that felt thin, wait three seconds instead of moving on. A large share of the time the candidate keeps going, and the second half is the useful half.

How hard to push, and when to stop

Two follow-ups per question is a good default; three when the criterion is a must-have and the answer is still not scoreable. Beyond that you are conducting an interrogation, and stressed candidates give worse information, not better. If you cannot score the criterion after three prompts, that is itself the score — write down that the answer did not reach the bar and what was missing, and move on.

Push evenly. A serious failure mode in unstructured interviews is pushing hard on the candidates you already doubt and accepting the first fluent answer from the ones you already like. Same question, same number of follow-ups, everybody. That is the whole point of writing them down beforehand.

Six reusable follow-up prompts for interviews: what did you personally do, what happened next, who disagreed, what would you change, how did you know it worked, walk me through Monday
Six prompts that work on almost any behavioural answer. Use two per question as standard, three when the criterion is a must-have.

Taking notes you can actually score

Most interview notes are useless within 48 hours, and it is usually the format's fault, not the interviewer's. Three lines of prose per candidate, written while listening, reconstructed from memory afterwards. By the debrief they have decayed into a general impression, and general impressions are exactly what you were trying to escape.

Notes that survive have four properties.

Structured by criterion, not chronologically

Your note sheet is not a blank page. It has one block per criterion, in the order you will ask, with the strong-answer contents and weak-answer shape already printed. During the answer you are ticking and annotating a structure that exists, not composing.

Quotes, not judgements, during the interview

Write what the person said, roughly, in their words. "Book of 40, measured on net revenue retention, missed target in Q3 by about 4 points, lost two clinics to a competitor on price" is a note. "Strong on renewals" is a conclusion, and conclusions written in the moment cannot be re-examined later, because the evidence they were based on is gone.

The habit to build: fact into the notes during the interview, judgement into the score in the five minutes after. If a note is a judgement, you will never be able to explain it in the debrief beyond repeating it more firmly.

Scored within five minutes of the candidate leaving

Block five minutes after every interview and use them. Score each criterion on the same 0 to 100 scale you used for resumes, write one line of justification per criterion pointing at what they said, and stop. Scoring later in the day means scoring by memory and by contrast with whoever came next, which is how the last candidate of the day ends up systematically rated against the third instead of against the criteria.

Independent before the debrief

Every interviewer submits scores before anyone sees anyone else's. This single rule removes the loudest source of noise in group hiring decisions: the first person to speak setting the anchor for the room. Once you have independent scores, the debrief has a job worth doing — go straight to the criteria where interviewers disagreed by more than about twenty points and talk only about those. Agreement needs no meeting.

A scoring scale that keeps people honest, for interviews specifically:

ScoreWhat it meansWhat the note looks like
90–100Everything a strong answer should contain, plus depth you did not ask forSpecific numbers, named trade-offs, a failure described without prompting
70–89Most of the strong-answer contents, reached with normal follow-upsReal work described in the first person, one gap in the detail
50–69Relevant experience but thin; needed three prompts to get specificsTeam-level description that only became personal after pushing
25–49Adjacent experience, or an answer that stayed abstract after follow-upsProcess described, actions not; no outcome given
0–24No relevant experience, or the answer contradicted the CVHypothetical answer to a question about the past

Note the rule that carries over from resume screening: a must-have criterion scored below 50 is a fail regardless of the total. A candidate who is delightful, scores 90 on four criteria and 40 on a must-have is not a near miss. They are a different job's candidate.

Splitting criteria across interview rounds

Two or three rounds is standard, and the standard way of running them wastes most of the time. Round one asks about experience. Round two asks about experience again, in slightly more detail, to different people who have not seen round one's notes. Round three is a culture chat where everybody asks whatever occurs to them. The candidate answers the same three questions three times and forms an accurate impression that the company is not organised.

Assign criteria to rounds instead. Every criterion is tested in exactly one round, except must-haves, which can appear twice on purpose, from different angles, because they are worth the redundancy.

RoundLengthCriteria testedDecision it produces
Phone screen25 minutesCriterion 1 (surface level), criterion 6, plus the practical basics: notice period, location, rangeIs this person real, available and roughly in range
Hiring manager60 minutesCriterion 1 (in depth), criterion 2, criterion 4Can they do the core of the job
Working session60 minutesCriterion 3, criterion 5, plus questions from the candidateWhat is it like when they actually work

Three things make this work in practice.

  • Write the assignment down before the first interview. Not in someone's head. In the same file as the criteria, visible to every interviewer, so nobody wonders whether onboarding has been covered.
  • Pass scores forward, not opinions. The hiring manager should see that criterion 1 scored 65 at the phone screen and that the note says the candidate stayed at team level. That tells them where to dig. Sending forward "seemed good" tells them nothing and biases them anyway.
  • Let a later round overwrite an earlier score for the same criterion. The deeper conversation is the better measurement. Keep both numbers in the record so you can see how well your phone screen predicts, which is how you improve the phone screen.

Who asks what, and why it should not be optional

Give each interviewer their criteria and their questions, and ask them not to freelance into someone else's. This feels rigid to people who enjoy interviewing, and it is the single change that most improves the debrief. When four interviewers all ask about experience, you get four overlapping impressions of the same thing and zero coverage of the rest. When each owns two criteria, you get full coverage and four independent measurements.

Leave time at the end of every round for the candidate's questions. It is not a scored criterion for most roles, and it is the part of the process that determines whether your first choice says yes.

Grid showing how six criteria split across three interview rounds: phone screen, hiring manager interview, and working session, with the decision each round produces
Each criterion has exactly one home, except the highest-weight must-have, which is tested twice from different angles on purpose.

Question types to avoid

Some questions are not just neutral, they actively cost you. They occupy a slot that could have measured a criterion, and they generate confident-feeling signal that is not connected to job performance. Four families to cut.

Brainteasers and puzzles

How many windows are there in this city. Why are manhole covers round. Estimate the market for pet food in a country you have never worked in. These were fashionable for years and mostly measure two things: whether the candidate has seen that class of puzzle before, and how comfortable they are performing while being watched.

If your role genuinely requires structured estimation under uncertainty — some analyst, consulting and pricing roles do — then make it a criterion, write it down with a weight, and test it with a problem from your actual business, where a good answer requires knowing what to ask rather than recalling a trick. That is a work sample. A puzzle with a clever answer is not.

Hypotheticals with no right answer

"What would you do if a customer asked for something we do not offer?" invites the candidate to describe an ideal version of themselves, and everyone describes a good one. Hypotheticals test imagination and social calibration. Past behaviour questions test what the person has done.

The exception is narrow and worth keeping: a hypothetical about a situation specific enough that the answer reveals method rather than intention. "Here is the actual usage data for one of our accounts last month. What would you look at first, and what would you do on Monday?" works, because it cannot be answered with a value statement. If your hypothetical could be answered convincingly by someone who has never done the job, drop it.

Questions that measure confidence rather than ability

"Sell me this pen." "Tell me why you are the best candidate for this role." "What is your greatest weakness." All three reward performance skill and rehearsal. The weakness question in particular has been asked so widely that the honest answer has been trained out of the population; you get a prepared strength wearing a disguise, and the only thing you learn is whether the person prepared.

The deeper cost is who these questions favour. Fluency, comfort with self-promotion and a willingness to make bold claims are unevenly distributed across cultures, personalities and levels of interview experience, and in most roles they predict very little about the work. If you need someone who can present under pressure, test presenting under pressure with a real task. Do not use it as a proxy for competence in a role where nobody presents.

Personal questions and everything that drifts near them

Age, marital status, children or plans for them, pregnancy, religion, health and disability, national origin, sexual orientation. In many jurisdictions questions in these areas are restricted or unlawful in hiring, and the specific rules vary by country — check what applies where you are hiring, and if you are unsure, ask someone qualified rather than guessing. Set that aside and there is still a practical reason to avoid them: none of them measure a criterion, so anything they add to your decision is by definition noise.

Most of these questions arrive by accident, through small talk. "Are you planning to stay in the city long term?" is a friendly question with a dangerous shape. If you need to know something operational — whether the person can work the shift pattern, whether they can travel monthly, whether they have the right to work — ask the operational question directly and identically of everyone. "This role needs one week of travel per quarter. Is that workable for you?" is clean. Anything that requires the candidate to disclose a protected characteristic to answer is not.

Why every candidate should get the same core questions

Structure is not bureaucracy for its own sake. Asking every candidate the same core questions, in the same order, scored against the same written expectations, does three separate jobs.

It makes comparison possible at all. If candidate A was asked about renewals and candidate B was asked about relationship building, you do not have two data points, you have two anecdotes. Comparison requires a common measurement. This is the reason that matters most often, and it is purely practical.

It reduces the drift that comes from order and mood. Interviewers are people. The fourth candidate on a Thursday gets a different interviewer from the first candidate on a Tuesday — less patient, more pattern-matched, more likely to accept a fluent answer. A fixed question set with written scoring expectations does not eliminate that, but it narrows the space in which it operates, because the interviewer is checking a list rather than forming an impression from scratch.

It gives you something to stand behind. If a candidate asks why they were not selected, or if a decision is ever questioned internally or externally, "everyone was asked these six questions, here are the scores and the notes behind them" is a defensible position. "We felt someone else was a better fit" is not a reason, it is a summary of a feeling. Rules differ by country and none of this is legal advice, but the general direction is stable: consistent, job-related, documented criteria hold up better than impressions.

Consistency is also the reason to be careful about where AI enters the process. A model can help you draft questions from your criteria, tidy notes, or summarise a transcript. It should not be the thing that decides, and it should not be scoring things you have not defined. We set out where automation genuinely helps in hiring and where it quietly makes things worse in AI in recruitment: what actually works, and the boundary there applies exactly here — the machine can make your standard cheaper to apply, it cannot invent the standard for you.

Structure does not mean a script read aloud

A common objection: fixed questions make interviews robotic, and candidates hate them. That happens when interviewers read questions off a page without listening. Structure applies to the questions asked and the things scored. It does not apply to tone, order of small talk, how you explain the role, or how you respond to what the candidate says. You can be warm and structured at the same time; the follow-up prompts are where the conversation lives, and they are naturally responsive.

It also does not mean identical time. A candidate whose answer to criterion 2 is complete in two minutes frees time for criterion 4. The requirement is that everyone is asked the same core questions and scored on the same expectations, not that a stopwatch runs.

A routine that improves the questions over time

Question sets decay. They encode last year's assumptions, and the only way to find out which questions are doing work is to look back after the round closes. This loop takes about half an hour per role and pays for itself the second time you hire.

  1. After the first three interviews, check the spread per question. If every candidate scores between 60 and 75 on a question, that question is not separating anyone. Either the expectations are too generous or the question is too easy to answer generically. Rewrite it to be more specific.
  2. Find the question everyone failed. If nobody clears the bar on a criterion, the bar is probably wrong, not the market. Check the strong-answer list — you may have described an ideal rather than a requirement.
  3. Compare interviewer scores on the same criterion. Where two interviewers consistently diverge by twenty points or more, the strong-answer definition is ambiguous. Rewrite it together, in specifics, before the next round.
  4. Check phone screen scores against final round scores. If the phone screen barely predicts the deeper interview, it is costing you time and screening out good people. Either fix its questions or shorten it to logistics only.
  5. Note the questions candidates asked you. Recurring candidate questions tell you what the job description left unclear, which is free editing for next round.
  6. After 90 days, look back at the new hire's scores. Which criteria did the interview call correctly, and which did it miss? A criterion that scored high on someone now struggling is a question that measures the wrong thing.
  7. Keep one versioned file per role. Criteria, questions, strong and weak answer definitions, round assignments, scores. Reopening the role in six months should start from this file, not from a blank page and a borrowed list.
Six-item review routine for improving an interview question set after each hiring round
The review loop. Steps 3 and 6 are the ones teams skip, and they are the two that actually change which questions you ask next time.

Doing this by hand, and where a tool helps

For one role with a handful of interviews, this is entirely a paper exercise. A document with six criteria, six questions, twelve strong-answer bullets, six weak-answer descriptions and a round assignment. An hour to build, reusable forever, no software involved. Do not let anyone sell you a system for a problem a shared document solves.

The strain shows up earlier in the pipeline, not at the interview. If 180 people applied, the bottleneck is not writing six questions — it is deciding which twelve people to ask them of, applying the same criteria to file 180 as to file 1 after two hours of reading. That is where consistency breaks first, quietly and gradually, and where the interview stage inherits a shortlist that was assembled by two different standards. The gap between storing applications and judging them is a real one, and we mapped it in what an applicant tracking system actually does.

If you want the criteria half handled for you, that is the specific job Orova Recruit does. You upload the job description as a file or paste the text, and it proposes 8 to 16 criteria sorted into five groups — must-have, important, preferred, basic, flexible — each with a weight from 1 to 10 and a suggested pass threshold. You edit anything you disagree with. It then scores each CV from 0 to 100 per criterion with the supporting text quoted from the file, computes the overall score as a weighted average on the server rather than trusting a number the model adds up, and forces a fail when a must-have scores below 50 regardless of the total.

It then does the step this article is about, from the same criteria: it drafts a question set for each candidate individually rather than reusing one list, and every question carries the scoring line to listen for plus a weight from 1 to 10, both of which you edit or reorder. In the room the questions sit beside a notes panel you type into line by line — each line stamped with the time and the person who wrote it, and an edited line keeps its earlier versions — or you press the mic and let the browser transcribe in any of 56 languages at no quota cost. Afterwards each question is scored 0 to 5 and converted to a weighted total out of 100, with anything neither the notes nor the recording covers scored zero and marked not asked.

It then does the step this article is about. From those criteria it drafts a question set — one shared set per role, and a variant written against an individual scored CV that aims at the claims that particular person has not yet proven. Every question carries what to listen for and a weight of its own, both editable, and cards can be rewritten, reordered by dragging, or polished by AI in a pass that keeps your meaning, your count and your order intact. During the interview the notes panel sits beside the questions so nothing gets asked twice or missed, and afterwards the answers are scored against that same rubric — with anything you never asked recorded as zero and labelled as not covered. The tool does the reading and the drafting; the questions, the expectations and the decision stay yours.

Whichever route you take, keep the editing step and keep the interview human. A criteria list you cannot change is worse than a spreadsheet, because it hides its assumptions instead of making you write them down.

Writing questions per candidate, not just per role

A shared question set per role is the right foundation: it is what makes candidates comparable and keeps a panel honest. But a set that never varies has a blind spot. Two candidates for the same job arrive with different histories, and the thing worth verifying about each of them is different. If you ask both the identical eight questions, you will verify neither properly.

The fix is not to abandon the shared set. It is to keep six shared questions and add two written specifically for the person in front of you.

Where the candidate-specific questions come from

Read the CV and mark three kinds of line.

Claims without proof. "Led the migration to a new CRM." Led how — technically, as project manager, as the person who trained users? A single question here is worth more than three generic ones about teamwork.

Gaps against your criteria. Your fourth criterion is B2B experience and their CV is entirely B2C. Do not silently mark it down; ask. Sometimes the answer is a two-year stint that never made the resume.

Discontinuities. An eight-month gap, three jobs in two years, a sideways move into a lower title. None of these are red flags on their own, and all three have ordinary explanations more often than not. But a candidate who is never given the chance to explain them gets marked down by inference, which is both unfair and unreliable.

How to write the question so the answer is scorable

Three habits separate a question that produces evidence from a question that produces conversation.

Ask for a specific instance, not a policy. "Tell me about the last time a release slipped" beats "how do you handle deadline pressure". People describe their aspirations when asked about policy and their behaviour when asked about instances.

Ask for the boundary of their contribution. "Which part of that was yours and which part was the team's?" This one question does more to separate candidates than any other, and it does it without being adversarial.

Write the scoring line before the interview, not after. One sentence describing a strong answer. If you cannot write that sentence, the question is not ready — you do not yet know what you are testing.

Keeping it fair while making it personal

There is a legitimate worry here: if every candidate gets different questions, are you still comparing like with like? The answer is that you compare on criteria, not on questions. Two candidates can be asked different questions about the same criterion and still receive comparable scores, as long as the criterion and its weight are fixed and the scoring line is written in advance.

Three guardrails keep this defensible. Keep the majority of the set shared, so most of the comparison rests on identical ground. Tie every custom question to an existing criterion rather than inventing a new axis mid-process. And record the custom questions alongside the answers, so six months later you can see exactly what was asked and why — which is what turns a personal question into a documented decision rather than an improvisation.

Frequently asked questions

How many interview questions should I prepare?

One or two per criterion tested in that round, plus two follow-ups each. For a 60-minute round covering three criteria, that is three to six core questions and time to go deep on each. More than eight questions in an hour and you are collecting shallow answers to everything. The most common mistake is a long question list with no follow-ups, which fills the hour with prepared answers.

What are good questions to ask candidates in an interview when the role is entry level?

The same structure, with criteria that do not require a work history. Someone with no professional experience can still describe a project they ran, something they taught themselves and how, a time they had to finish something without supervision, or a piece of work they would now do differently. Keep the "walk me through the last time" shape and drop the assumption of a job. What changes is the criteria, not the method.

Should I send the questions to candidates in advance?

For work samples and exercises, yes — you learn more from prepared work than from panic. For behavioural questions, sending the exact wording removes the value of the follow-ups, but telling candidates the areas you will cover costs you nothing and reduces the advantage held by people who have interviewed recently. A line like "we will spend most of the hour on renewals, onboarding and a difficult escalation" is fair to everyone and does not make the answers less real.

Can I ask different questions to different candidates?

Ask the same core questions to everyone — the ones tied to criteria and scored. Vary the follow-ups freely, because follow-ups depend on what the person just said. Also vary questions that probe CV-specific evidence: each candidate's file is different, and asking about a gap only the person who has one has is not inconsistency. The rule applies to the scored core, not to the whole conversation.

How do I score a candidate who gives a great answer about something I did not ask?

Score it against the criterion it belongs to if it maps to one, and note it separately if it does not. What you must not do is let an impressive off-topic answer raise the score on a criterion it did not address. That is exactly how a strong performer with a must-have gap gets through — the excellence is real, it is just in a place the role does not need.

What if the hiring manager refuses to use a question sheet?

Start with the scoring, not the questions. Ask them to submit a score per criterion with one line of evidence after each interview, and let them ask what they like. Most people find within two candidates that they cannot fill the sheet from an unstructured chat, and they start structuring the chat themselves. Winning that argument in advance is usually slower than letting the empty sheet make the case.

Does structured interviewing remove bias?

It reduces one specific kind: the inconsistency of measuring different people against different things depending on who they remind you of, what order they arrived in, and how tired you were. It does not fix bias built into the criteria themselves. If a criterion effectively requires a background that only one group tends to have, writing it down with a weight of 9 makes you consistent, not fair. Read the criteria list on its own once per round and ask, line by line, whether each one tests the job or a résumé pattern.

What to do this week

Take the role you are interviewing for right now. Open the criteria you screened resumes against — if they do not exist, build them first, because everything here depends on them.

Then do four things. First, pick the criteria that belong in an interview rather than a document, and assign each one to exactly one round. Second, for each of them write the question, three or four things a strong answer contains, the shape a weak answer takes, and two follow-up prompts. Third, build the note sheet from that file so every interviewer is writing into a structure instead of onto a blank page, and put a five-minute scoring block after every interview slot in the calendar. Fourth, tell the panel that scores go in before the debrief starts, and that the debrief only discusses criteria where people disagreed.

That is a bit over an hour of preparation. What you get back is a Friday afternoon where the room can say, out loud and with the evidence in front of it, that candidate three failed a must-have for a stated reason and candidate five cleared every one — instead of five people trying to remember Tuesday.

Let AI read and score resumes against your JD

Orova Recruit turns your job description into weighted criteria and scores every CV with evidence quoted from the file.

Try it free