AI Resume Screening: What Works and Where It Falls Over
AI resume screening is the use of a language model to read every CV in an applicant pool and score each one against criteria you fixed in advance, with the supporting quote attached. It replaces the first-pass read, not the hiring decision — the output is a shortlist with evidence, which a human still has to accept or overturn.
You have sat through three vendor demos this month and all of them opened with the same slide. AI reads the CVs, AI ranks the candidates, you interview the top five, everyone goes home early. Meanwhile the folder for your open role holds 280 files, four of them phone photos of a printed page, and the last screening tool you tried buried a candidate you went on to hire.
Here is the short answer. AI resume screening is genuinely good at two things: reading unstructured documents fast, and applying the same fixed criteria to every one of them with the evidence quoted back to you. It is bad at judgement — potential, culture fit, whether an odd career gap is a warning sign or the most interesting thing about the person. Split the work along that line and it holds. Ask it to decide, and it fails in ways you only notice a month later, when the shortlist turns out to be five versions of the same person.
This article breaks the hiring workflow into stages and marks each one: safe to delegate, delegate with review, or keep human. Then it covers the part nobody demos — how you check that the output is right. Evidence quotes, totals computed by the system rather than written by the model, must-have criteria that override a high score, what happens to files nothing can read, and the audit trail that lets you explain a rejection six months after you made it. Most of what gets written about AI and recruiting swings between "this replaces your hiring team" and "this is dangerous, do not touch it", and neither position helps you decide what to do on Monday. The line is narrower and more boring than either. It ends with the fear everyone has and nobody says out loud in the demo: will this thing throw away good people?
What does AI resume screening actually do well?
AI resume screening reliably does two jobs: it reads unstructured documents fast, and it scores every one of them against criteria you fixed in advance, quoting the evidence it used. It does not judge potential, culture fit, or anything outside the file. Treat it as a reader and a scorer, never as the decision.
The useful way to think about it is to ask what kind of task you are handing over. There is a category of work that is mechanical but not rule-based. Reading a CV is the perfect example. There is no fixed format, no schema, no guarantee the dates are in the same order twice. A regular expression cannot do it. But there is also no real judgement involved in the reading itself — you are looking for whether a document says a specific thing, and if it does, where.
That category is exactly where language models are strong and where older recruiting software was weak. A traditional applicant tracking system parses fields: name, email, employer, dates. It does that reasonably well on clean PDFs and badly on anything else, and it never forms a view about content. If you want the longer version of what those systems do and do not carry, we wrote it up in what an applicant tracking system actually does. The short version is that it is a container, a state machine and a record, and none of those three read the CV.
The second thing AI does well is consistency, which is not a glamorous claim but is the one that changes outcomes. You are not consistent. Nobody is. The fortieth CV of the afternoon gets a different reading than the fourth, and the one that arrives after a bad meeting gets a different reading again. A model applying the same rubric does not get bored, does not get tired, and does not read a candidate more generously because their previous employer is a name you recognise.
Note what both of these have in common: neither requires the system to know anything you did not tell it. Reading is bounded by the document. Scoring is bounded by the criteria. The moment you ask for something outside those two boundaries — will this person stay two years, will they get along with the team, is this the kind of thinker we need — you have left the zone where the technology works and entered the zone where it produces fluent, confident, unfounded output.
The three jobs AI does better than you do
Being specific matters here, because "AI screens your resumes" covers at least three distinct capabilities that fail in different ways.
Reading unstructured documents at speed
CVs arrive as PDFs exported from Word, PDFs exported from Canva with the text baked into a graphic, DOCX files with tables inside tables, plain text pasted into an email, and scans of a printed page. A person handles all of these without thinking about it. Traditional parsing handles about half of them and silently mangles the rest — you have seen the parsed record that lists the candidate's university as their current employer because the two-column layout confused it.
A language model reading the extracted text handles messy layouts far better, because it is working from meaning rather than position on the page. And for the files where text extraction fails entirely — a scan, a photo, a graphic-heavy design CV — a multimodal model can look at the image directly and read it the way you would. That is a real capability gain, not a marketing one. It is the difference between a stack of "could not process" files sitting in a corner and every application actually being read.
Applying one rubric to everyone, at 4pm on a Friday
Give a model a criterion — "has owned a content calendar end to end, not just written to someone else's brief" — and it will apply that same test to CV number 1 and CV number 280 in the same way. It will not decide halfway through that it is being too strict. It will not upgrade someone because the formatting was pleasant.
This is the property that makes the whole thing defensible. Not accuracy in the abstract, but repeatability. If you rerun the same 280 files against the same criteria, you get materially the same answer. If you cannot say that about your current process — and if you read the CVs yourself, you cannot — then the comparison you should be making is not "AI versus perfect" but "AI versus a tired human on a Friday afternoon".
Pulling structured facts out of prose
Name, email, phone number, years of experience, most recent title. These are boring fields and extracting them is boring work, but doing it by hand across a few hundred files is where a real part of your week disappears. A model reading the document pulls them out reliably enough that the exceptions are visible and fixable, which is the standard you should be holding it to. Not zero errors. Errors you can see.
Notice that all three of these are reading tasks. None of them is a decision. That is not an accident — it is the whole finding.
Where AI falls over
The failures are not random. They cluster, and once you know the clusters you can design around them.
Judgement calls it was never given the context for
Ask a model whether a candidate is "a good fit for a fast-moving startup" and it will answer. It will answer confidently, in well-structured prose, with reasons. The answer is worthless, because the model does not know what your company is like on a Tuesday, what the last three people in this role struggled with, what your founder is impossible about, or which of your two hiring managers this person would actually report to.
It is not that the model is guessing badly. It is that it is guessing at all, and the output does not look like a guess. Fluency is the trap. A hedged human answer reads as hedged; a model's answer reads as a finding.
Culture fit is not in the document
This deserves its own line because it is the single most common thing people ask AI screening to do, and it is structurally impossible. Culture fit is a claim about the interaction between a person and a specific group of other people. A CV contains neither side of that interaction. Whatever a tool tells you about culture fit, it derived from proxies: previous employer size, industry, tenure patterns, the vocabulary in the summary paragraph.
Those proxies are exactly where bias lives. A model that learned "people who worked at companies like ours tend to be a good fit" has learned to reproduce your existing hiring pattern, which is the opposite of what most teams say they want. Keep culture questions in the interview, where you can actually ask them.
Unusual career paths read as noise
The three-year gap. The switch from teaching to data analysis. The person who ran their own agency for four years and now wants a salaried job. Human readers can also be bad at these, but a scoring model is bad at them systematically, and systematic is worse — every candidate with that shape gets the same treatment, every round, invisibly.
The specific failure is that the model scores against a pattern of what the criterion usually looks like. "Five years of relevant experience" is easy to verify when someone had one job title for five years, hard when the same five years is spread across freelance work with no employer names. The evidence is there, it just does not present in the expected shape.
It is confidently wrong, in fluent prose
A model asked to score a criterion it cannot verify will not usually say "I cannot verify this". It will produce a plausible score and a plausible-sounding justification, and if you do not force it to quote the document, the justification will sometimes be an invention assembled from context. This is not rare and it is not a bug you can patch — it is the shape of the technology. Every safeguard in the next section exists because of this one property.
Files it cannot read at all
Some files are corrupt. Some are password-protected. Some are 40MB scans at a resolution that produces nothing usable. The dangerous behaviour is a tool that quietly scores these as low rather than flagging them, because a zero looks exactly like a genuinely weak candidate in a sorted list. A file that could not be read should sit in its own status, visible, waiting for you to open it by hand or ask the candidate to resend. It should not consume budget and it should not produce a number.
A stage-by-stage map: delegate, review, or keep human
Here is the whole workflow with a verdict on each step. The verdicts are deliberately conservative — the cost of over-delegating in hiring is a person who never gets a call and never knows why.
| Stage | Verdict | Why |
|---|---|---|
| Writing the job description draft | Delegate with review | Fine as a first draft. The requirements section needs a human pass or you inherit generic lines nobody can score. |
| Turning the JD into scoring criteria | Delegate with review | Good at proposing a structured list. You must read every line, cut the must-haves down, and fix the weights. |
| Extracting text from CV files | Safe to delegate | Mechanical. The only thing you need back is a clear status for files that failed. |
| Pulling name, email, phone, years, last title | Safe to delegate | Boring, high volume, errors are visible against the source file. |
| Scoring each criterion 0 to 100 with evidence | Delegate with review | The core win. Requires quoted evidence and a spot-check routine, not blind trust. |
| Computing the weighted total | Safe to delegate — to the software, not the model | Arithmetic belongs in code. A model asked to add up its own scores will drift. |
| Deciding pass, consider, fail | Delegate with review | Fine as a mechanical band comparison. The bands are yours and the consider band gets read by a person. |
| Ranking the shortlist | Keep human | A sorted list is a suggestion of reading order, not an order of merit. |
| Rejecting a candidate | Keep human | The one irreversible action in the funnel. Never let it happen automatically on a score alone. |
| Drafting outreach and rejection emails | Delegate with review | Good at the words, bad at knowing what you promised this person on a call. |
| Building interview questions from the criteria | Delegate with review | Strong at converting a criterion into probing questions. You pick which ones to actually ask. |
| Judging the interview | Keep human | The context is in the room and the model is not in the room. |
| Assessing culture fit | Keep human | Not in the document. Any AI answer here is a proxy, and proxies are where bias hides. |
| The hire decision | Keep human | Not a close call. |
Read down the "why" column and a pattern appears. Everything marked safe is bounded by an input you supplied. Everything marked keep-human requires information that was never in the file. The middle category — delegate with review — is not a compromise; it is where most of the value sits, and the review is what makes the value bankable.
How to verify what AI recruiting tools tell you
This is the section the demos skip. Every AI screening result you look at should survive five checks. If your current tool cannot support all five, you are not screening with AI, you are gambling with extra steps.
Rule 1: every score carries a quote
A criterion scored 85 with no evidence is a number you cannot argue with and cannot check. A criterion scored 85 with the line "led the content calendar for three product lines, Jan 2023 to Mar 2025" pulled from the CV is a claim you can verify in four seconds by opening the file.
This single requirement removes most of the hallucination risk, because a model forced to produce supporting text from the source has a much harder time inventing the finding. It also changes your review from re-reading the CV to reading the quotes, which is roughly a tenfold difference in time.
When the quote does not support the score — and you will find cases — that is not a disaster, that is the system working. You caught it. Fix the criterion wording and rerun. Criteria that produce bad evidence are almost always criteria that were written vaguely; the same principle applies when you write the JD, which is why it helps to build the posting and the scorecard as one document from the start, the approach in our job Posting Template: Write One, Get a Hiring Scorecard Too.
Rule 2: the system computes the total, not the model
Language models are poor at arithmetic and worse at arithmetic they are asked to perform while also writing prose. If the tool asks the model for an overall score, you get a number that correlates with the per-criterion scores but does not follow from them. It will drift toward whatever the model's general impression was, which is exactly the impressionistic judgement you were trying to remove.
The correct design: the model scores each criterion 0 to 100 with evidence, and the server computes the weighted average from those scores and your weights. Nothing else. Here is a worked example — invented, to show the mechanics — for a content marketer role with eight criteria.
| Criterion | Group | Weight | Score | Contribution |
|---|---|---|---|---|
| Published writing you can read | Must-have | 10 | 85 | 16.7 |
| Owned a content calendar end to end | Must-have | 9 | 70 | 12.4 |
| SEO fundamentals: research and on-page | Important | 8 | 60 | 9.4 |
| Reads their own analytics without help | Important | 7 | 40 | 5.5 |
| B2B SaaS audience experience | Preferred | 5 | 90 | 8.8 |
| Business-level written English | Basic | 6 | 100 | 11.8 |
| Can work core hours with the team | Basic | 4 | 100 | 7.8 |
| Video or podcast production | Flexible | 2 | 0 | 0 |
The weights sum to 51. The weighted total is the sum of weight times score, divided by 51, which comes to 72.4. Every contribution in the last column is weight times score divided by 51, and they add back to 72.4. You can check the whole thing on paper in two minutes, which is the point. A number you can reconstruct is a number you can defend to a hiring manager who disagrees with it.
Look at that chart for a second longer than feels necessary. It tells you something the total does not: this candidate's score is being carried by writing samples and language ability, and dragged down by analytics. That is a completely different interview than the same 72 earned by strong analytics and mediocre writing. The total is a filter. The breakdown is the briefing.
Rule 3: must-have criteria override the total
A weighted average has a structural flaw: strength in several small criteria can bury a fatal weakness in one big one. Take the same candidate and change one number. Suppose the calendar-ownership criterion scored 45 instead of 70. The total drops to about 68 — still above a pass threshold of 65. The candidate passes, on paper, while failing a thing you defined as non-negotiable.
The fix is a hard rule sitting above the arithmetic: any must-have criterion scoring below a floor forces a fail verdict regardless of the total. Fifty out of a hundred is a reasonable floor. What matters more than the exact number is that the rule exists and is applied by the system, every time, without anyone remembering to check.
This also disciplines how you write must-haves. If a rule can knock out an otherwise strong candidate on its own, you will think harder before tagging six things as must-have. Two to four is usually right. Everything else is important, preferred, basic or flexible, and gets its influence through weight rather than veto power.
Rule 4: unreadable is a status, not a zero
When a file cannot be read, the honest output is "unreadable". Not a low score. Not an omission from the list. A visible status that keeps the candidate in the queue and puts the file in front of a person.
Three things should follow. The tool should not call the model on a file it could not extract, because there is nothing to send. It should not charge you for the attempt. And it should keep the file so you can open it, look at it yourself, and either transcribe it or email the candidate for a different format. A candidate whose CV happened to be a scan is not a weaker candidate, and a system that treats them as one is quietly discriminating on file format.
Rule 5: spot-check five files by hand, every round
Not once during evaluation. Every round. Pick five: two that passed comfortably, two that failed, and one from the middle band. Open each CV alongside its scored result and read the evidence quotes against the document.
You are looking for four things. Evidence that does not say what the score claims. Criteria where every candidate got roughly the same score, which means the criterion is not separating anyone and should be rewritten or dropped. Criteria where the model clearly interpreted the wording differently than you meant. And any candidate you would obviously have called who did not pass, which is your most valuable signal of all.
Five files is about fifteen minutes. It is the cheapest quality control in the entire process, and skipping it is how teams end up six months in with a screening setup that has been quietly wrong the whole time.
Will AI reject good candidates?
Yes, sometimes. So do you. The useful question is not whether errors happen but whether they are visible, explainable and correctable — and here the honest comparison is uncomfortable for manual screening, because a human rejection at 4pm on a Friday leaves no record of why at all.
Still, the fear is legitimate and deserves a specific answer rather than reassurance.
The three ways a good candidate actually gets lost
First, the criterion was wrong. If you wrote "5+ years in B2B SaaS marketing" as a must-have, the system will correctly fail a brilliant candidate with four years and a much stronger portfolio. That is not an AI failure. That is your rule, executed exactly as written, and it would have failed the same person if a human applied it. The difference is that a human would have quietly broken the rule, which feels better and is far harder to audit.
Second, the evidence was in an unusual form. The candidate demonstrated the skill, but not in a shape the criterion expected — the work is in a personal project rather than a job, the title was unusual, the achievement is described in a portfolio link rather than in the CV body. The model scored what it could see and what it could see was thin.
Third, the file was bad. Two-column layouts where extraction interleaves the columns, heavy graphic design, scans at low resolution. The content was fine. The reading of it was not.
What to change so it stops happening
For the first cause, apply the ninety-day test to every must-have: could someone genuinely do this job in their first three months without this? If yes, it is not a must-have. Year counts almost never survive that test and are the single most common cause of a good candidate being cut, because years are a proxy for the thing you actually want, and proxies fail on exactly the candidates who are unusual.
For the second, write criteria that name the evidence rather than the credential. "Has shipped work of this type end to end, unsupervised" catches the freelancer, the side-project builder and the salaried employee alike. "Three years at an agency" catches only one of them.
For the third, keep the unreadable pile visible and work it by hand. It is usually a handful of files per round. Fifteen minutes.
The band that saves you
The structural protection is the consider band — a range below your pass threshold where the verdict is neither pass nor fail but "a person reads this one". If your threshold is 65, a consider band from 55 to 65 means every borderline candidate lands in a queue for human eyes instead of a rejection folder.
This is where nearly all of the recoverable mistakes live. A genuinely weak candidate rarely scores 61. A strong candidate with an odd CV frequently does. Widen the band when you are worried about missing people and can afford the reading time; narrow it when volume is crushing you and you accept the trade. Just do not set it to zero, because a hard cut at a single number is where the good-candidate-lost stories actually come from.
The governance question: can you explain the decision six months later?
Set aside efficiency for a moment. There is a second reason to run screening this way, and in some places it is becoming the more important one.
Automated tools used in employment decisions are now explicitly regulated in several jurisdictions. New York City's Local Law 144 requires annual bias audits and candidate notice for automated employment decision tools used on candidates in the city. The EU AI Act classifies AI systems used for recruitment and candidate evaluation as high-risk, with obligations around record keeping, human oversight and transparency. In the United States, the Uniform Guidelines on Employee Selection Procedures (29 CFR Part 1607) have for decades expected selection procedures to be consistent, job-related and documented. None of this is new in spirit. What is new is that "we looked at the CVs and picked the best ones" is a weaker answer than it used to be.
There is also the cautionary case everyone cites, and it is worth citing accurately: Reuters reported in 2018 that Amazon scrapped an experimental recruiting tool after finding it penalised CVs containing the word "women's". The lesson usually drawn is "AI is biased". The more useful lesson is that the system was trained to predict past hiring outcomes rather than scored against explicit criteria — and that the failure was only discoverable because someone could inspect what it had learned. Opacity is the risk. A model that scores against criteria you wrote, and quotes its evidence, is inspectable by design.
So what does a defensible record contain? Concretely: the criteria as they stood when this candidate was scored, the weight on each one, the score on each one, the evidence quoted for each score, the computed total, the threshold and band in force, the verdict, and who ran it and when. If a criterion changed mid-round, that should be visible too, because a candidate scored under the old criteria and one scored under the new ones were not compared on the same basis.
This is the same argument that applies to any AI system that takes consequential action on your behalf: the output is only usable if you can trace it back to a reason. We made that case at length for advertising in AI Advertising: What It Actually Does for Your Ads, and hiring is the higher-stakes version of the same problem, because the thing being decided is a person's next two years rather than a budget line.
Mistakes teams make with AI resume screening
These come up constantly, and each one is fixable in an afternoon.
Letting the tool write the criteria and never reading them
Most AI recruiting platforms will read your JD and propose a criteria list. This is genuinely useful and it is also the highest-leverage place to intervene, because everything downstream is built on it. An auto-generated list typically over-produces must-haves, converts vague adjectives into unscoreable criteria, and weights everything about the same. Ten minutes reading and editing that list changes every score in the round. Skipping those ten minutes means you automated somebody else's opinion of your role.
Treating the score as a ranking instead of a filter
A candidate at 78 is not meaningfully better than one at 74. The precision is false — the underlying per-criterion scores are judgements on a coarse scale, and a four-point gap in a weighted average is noise. Use the score to decide who gets read, then read them. Interviewing strictly in descending score order is a way of outsourcing the decision you said you were keeping.
Screening for things a CV cannot show
"Collaborative." "High ownership." "Comfortable with ambiguity." All real qualities, none of them visible in a document. Put them in the interview rubric where they can be tested with questions, and keep the screening criteria to the subset you can verify from a file. A criterion the document cannot answer produces a score the model invented.
Changing criteria in the middle of a round
You will be tempted after the first thirty CVs. Resist until the round is done, or rescore everyone. Otherwise the first thirty and the last two hundred were measured with different rulers and the comparison between them is meaningless. Note the change, finish the round, apply it next time — or rerun the whole batch, which is cheap when a machine does the reading.
Auto-rejecting on the score alone
The single rule worth being absolute about. A rejection is the one irreversible action in the funnel — the candidate takes another offer, and no amount of later correction gets them back. Let the system sort and flag. Let a person press reject, even if that person is pressing it forty times in a row after skimming forty sets of evidence quotes. Those forty skims are what stands between a wrong criterion and forty wrong rejections.
A routine that keeps this honest
Per round, this is roughly ninety minutes of human work spread across a week, and it is what turns a screening tool into a screening process.
- Before you post. Write the criteria, tag each into a group, set weights from 1 to 10, and cut the must-haves to four or fewer using the ninety-day test. Set the pass threshold and the consider band, and write both numbers down. Twenty minutes.
- After the first twenty applications. Run them, then spot-check five by hand. You are checking the criteria, not the candidates. Fix the wording of any criterion whose evidence quotes look off, and rescore those twenty. Fifteen minutes.
- Mid-round. Check the score distribution. If everyone lands between 60 and 75, your criteria are not separating people and the weights are too flat. If everyone fails, a must-have is unrealistic. Five minutes.
- Work the unreadable pile. Open the files that could not be read, transcribe or request a resend, and score them by hand. Usually under ten files. Fifteen minutes.
- Before you send rejections. Read the evidence for everyone in the consider band and everyone who failed on a single must-have. These two groups hold nearly all your recoverable errors. Twenty minutes.
- After the interviews. For every candidate who passed screening and then clearly failed the interview, find the criterion that should have caught it. That criterion, rewritten, is next round's improvement. Ten minutes.
- After the hire. Save the final criteria, weights and thresholds with the role. Reopening this position in six months should start from a tuned file, not a blank page. Five minutes.
Notice that none of these steps is "check whether the AI is right in general". They are all checks against your own criteria on your own candidates, which is the only evaluation that means anything.
When a tool earns its place, and how far hand work goes
You do not need software to do any of this. A spreadsheet with criteria down the left and candidates across the top works perfectly, and for a role with thirty applicants it is arguably the right answer — you will spend more time configuring a tool than reading thirty CVs.
Hand work runs out at roughly two points. The first is volume: somewhere past a hundred files per role, reading each one against eight criteria and writing down evidence stops being a task you finish and becomes a task you abandon halfway. The second is repetition: hiring the same role three times a year, and rebuilding the criteria from memory each time because last round's version lives in a spreadsheet somebody archived.
What a tool should give you at that point is narrow and checkable. It should take the JD as a file or pasted text and propose criteria you can edit — grouped, weighted, with suggested thresholds. It should accept the file formats candidates actually send, including scans and photos, and read them. It should score each criterion with quoted evidence, compute the total in code from your weights, apply the must-have override, and place each candidate in pass, consider or fail against bands you set. It should mark unreadable files as unreadable without charging you or inventing a score. It should keep the whole record per candidate, with a deleted-items area so a mistaken deletion is recoverable, and it should show who uploaded what.
That list is deliberately unexciting, and it is the specification Orova Recruit is built to. It reads a JD and proposes 8 to 16 criteria across five groups — must-have, important, preferred, basic and flexible — each with a weight from 1 to 10 and a suggested pass threshold, all editable. It takes PDF, DOCX, DOC, TXT, MD, RTF, ODT and images, up to 10MB per file, uploaded singly, in bulk or as a whole folder. It scores each criterion 0 to 100 with evidence pulled from the CV, the server computes the weighted total rather than the model, any must-have under 50 forces a fail, and unreadable files are flagged without an AI call and without consuming quota.
The same restraint carries into the interview, which is where most tools stop. From the approved criteria it drafts a question set for each candidate individually, every question carrying its own scoring line and weight. During the call you type notes line by line or press the mic and let the browser transcribe — 56 selectable languages, and because it uses the browser's own speech engine it costs no quota. The session is recorded, so a disputed answer can be replayed rather than argued from memory. Scoring gives each question 0 to 5 and converts it to a weighted total out of 100, treating the recording as the primary source and your typed notes as the hurried secondary one; a question neither source covers scores zero and is labelled not asked rather than guessed at. With two or more candidates scored, it will compare them and suggest an order — a suggestion only, with the hire-or-pass click, the person who made it and the timestamp recorded separately.
The same restraint is applied to the parts of the loop where AI is most often misused. Interview questions are generated per candidate but always come with a scoring line and a weight, so the interview produces a rubric rather than an impression. Notes can be dictated — the browser listens and types while you keep eye contact — but the judgement still runs off the rubric, and any question your notes never touch is scored zero and labelled as not covered rather than inferred. The candidate comparison stops at a ranking with reasons and a recommendation; the select-or-reject button belongs to a person with the rights to press it, and the offer letter the model drafts leaves salary and start date blank. Signing up is free and comes with 1,000 quota, no card, which is enough to try it on a real role rather than a demo one.
Whatever you use, hold it to the five rules in this article. A tool that cannot show you the evidence behind a score is not saving you work — it is moving the work to the day someone asks you why a particular person never got a call.
Seven places AI earns its keep across the loop — and three where it must not
"AI in recruitment" is usually answered too broadly to act on. Below is the same question answered task by task, with an honest note on how much the machine really contributes.
Seven that work
1. Drafting the job description from bullets. Low risk, because you read every word before it is published. The real gain is not the twenty minutes saved; it is that you stop reusing the two-year-old description because writing a new one felt like work.
2. Turning the description into weighted criteria. Fast, but the value is in your edits. What the model produces is a draft to argue with, and arguing with a concrete draft is always quicker than arguing with a blank page.
3. Reading and scoring applications against those criteria. The clearest use of all, on one condition: every score must carry evidence quoted from the file. Without evidence you have merely swapped a human hunch for a machine hunch.
4. Extracting candidate details into a table. Name, email, phone, years of experience, most recent role. Purely mechanical, nothing to debate, and it removes more typing than anything else on this list.
5. Writing interview questions per candidate. The model reads the CV and points at the claims that are asserted but unproven — which is exactly where the questions belong. You still cut and reorder, but going from ten to five is far easier than going from zero to five.
6. Transcribing notes during the interview. The most undervalued item here relative to what it is worth. It is not clever at all; it just gives the interviewer their eyes and hands back.
7. Drafting candidate emails. Invitations, rejections, scheduling notes. Repetitive text, written once, tuned to your voice, then reused.
Three that should stay human
1. Deciding to reject someone. Not because the scoring is poor, but because accountability cannot be delegated to a model. "The system found you unsuitable" is the answer of an organisation unwilling to own its decisions. The correct shape is: the machine ranks and records the evidence, a person presses the button and signs their name to it.
2. Inferring what the CV does not say. Personality, loyalty, culture fit. Inferring these from the wording of a resume is where bias concentrates, and it is usually systematic bias — by university, by previous employer's name, by writing style. If your tool volunteers judgements of this kind, switch that feature off.
3. Scoring a question that was never asked. Obvious in principle, common in practice: language models fill gaps. The rule needs to be explicit — not in the notes means not asked, and not asked means zero with a label, never a polite average.
A cheap way to test any of this before trusting it
Take twenty applications from a role you filled a few months ago, run them through the tool, and compare its ranking with two things you already know: who you actually interviewed, and who you hired. You are not hoping for a perfect match — a perfect match would mean the tool adds nothing. What matters are the two disagreements. Candidates it ranked high that you skipped: read the evidence and decide whether the machine is wrong or you missed something. Candidates you interviewed that it ranked low: usually a sign your criteria are missing something you value but never wrote down. Both teach you something, and both are legitimate reasons to adopt the tool or walk away from it.
Questions people ask
Can AI screen resumes without bias?
No, and neither can you. What explicit criteria plus quoted evidence gives you is not the absence of bias but its visibility. If a criterion is effectively selecting for a background only one group tends to have, you can read that criterion, see the problem, and change it. Bias inside a human reader's head is not available for inspection. Bias built into a model that predicts past hiring outcomes is not either. Bias in a written criteria list is, which is the entire argument for working this way.
How accurate is AI resume screening?
There is no honest single number, and be suspicious of any vendor who offers one, because accuracy depends entirely on how well your criteria are written. A vague criterion produces unreliable scores no matter how good the model is. A criterion naming specific evidence produces scores you can verify yourself in seconds. Measure it on your own roles with the five-file spot-check rather than trusting a benchmark from a different company hiring different people.
Should candidates be told AI was used?
In several jurisdictions you must — New York City's Local Law 144 requires notice, and the EU AI Act imposes transparency obligations on high-risk employment systems. Even where it is not required, tell them. A one-line note that applications are screened against published criteria, with a human reviewing every rejection, costs you nothing and is a fair description of a process built the way this article describes. If it does not feel comfortable to disclose, that is information about the process, not about the disclosure.
Will AI replace recruiters?
It replaces the reading, not the recruiting. Look back at the stage map: the steps marked safe to delegate are extraction, arithmetic and consistency checks — genuinely the least valuable hours in the week. Everything that decides an outcome stays human. What changes is where a recruiter's day goes: less time opening files, more time on candidate conversations, calibration with hiring managers, and fixing the criteria that keep producing the wrong shortlist.
What about candidates who use AI to write their CVs?
Assume most of them do. This is a reason to score against evidence rather than polish — a well-written CV and a well-evidenced one are different things, and criteria that name specific artefacts ("published work you can read", "owned this end to end") are much harder to satisfy with generated prose than criteria that reward keywords. It also raises the value of the phone screen, where claims meet follow-up questions.
Do I still need an ATS if I have AI screening?
They solve different problems. An ATS is the container and the record; screening is the judgement layer on top. Small teams often start with screening because the bottleneck is reading, not storage — a shared drive plus a scoring tool handles a role with 300 applicants fine. You will want the tracking system when the pain moves from "I cannot read all this" to "I cannot remember where anyone is in the process".
What to do this week
Take the role you have open right now and do four things.
First, write the criteria list on paper before touching any tool. Eight to twelve lines, each one tagged into a group, each with a weight from 1 to 10, and no more than four must-haves after the ninety-day test.
Second, for each line, finish the sentence "I would know this from the CV because it shows ___". Any line where you cannot finish that sentence belongs in the interview rubric, not the screening one.
Third, set a pass threshold and a consider band and write both numbers in the same file. Ten points below the threshold is a sensible starting band.
Fourth, run twenty CVs through whatever process you use, then spot-check five by hand against the evidence. You are testing your criteria, not the candidates.
Do that and you will know, within an hour, exactly what AI is doing for you and exactly which criterion is going to cause a problem. That is a far better position than the one the demo slide puts you in, where everything works until the day someone asks why a particular person never got a call and the only available answer is "the system ranked them low".
Let AI read and score resumes against your JD
Orova Recruit turns your job description into weighted criteria and scores every CV with evidence quoted from the file.
Try it free