Job Description Template That Turns Into a Hiring Scorecard
The ad went live on Monday. By Thursday there are 180 files sitting in a shared folder, and the hiring manager asks the only question that matters: who do we call? You open the first CV, then the tenth, and somewhere around the twentieth you notice something uncomfortable. You have no fixed answer to what "good" looks like for this role. The job description template you started from was built to attract applicants, not to judge them. It has a nice paragraph about the company, a bullet list of responsibilities, and a requirements section that reads like a wish list somebody typed in a hurry.
So the screening turns into taste. The first ten CVs get read carefully. The next fifty get skimmed. The last hundred get sorted by whoever looks familiar. Two weeks later the shortlist arrives at the interview stage and nobody can explain, on paper, why these six and not those six.
Here is the fix, in two sentences. Write the job description so that every requirement line is already a scoring criterion: one testable statement, a group that says how much it matters, a weight, and the evidence you expect to find in a CV. Do that once, and the same document runs the ad, the resume screen, the phone screen and the interview panel, without anybody rewriting it in a spreadsheet at midnight.
This article gives you that template, the conversion rule that turns requirement lines into scoring criteria, a full worked example for a marketing executive plus a shorter one for a customer support lead, and the three ways these templates usually break.
What makes a job description template scoreable?
A scoreable job description template writes every requirement as one testable line, tags it must-have, important, preferred, basic or flexible, gives it a weight from 1 to 10, and names the evidence a CV must show. The job ad and the scorecard then come from the same file.
Most templates you can download today are ad templates. They are optimised for one job: getting a stranger to click apply. That is a real job, and a good ad matters. But an ad template stops the moment the applications arrive. It never tells you how to compare two candidates who both "have strong communication skills".
A scoreable template does one extra thing. It forces the requirements section into a shape a machine or a junior colleague could apply without asking you what you meant. Four properties make that possible.
- One idea per line. "3+ years in B2B content marketing with responsibility for the blog and landing pages" is two ideas glued together. Split it, or you will not know which half a candidate missed.
- Observable in a CV. If you cannot point at a line in a resume and say "that is the proof", the requirement belongs in the interview stage, not the screening stage.
- Graded, not just present or absent. "Has run paid search" is yes or no. "Has personally managed a paid search budget" can be scored on a range: never, helped someone else, owned a small budget, owned a large one.
- Ranked against the other lines. A requirements list where everything is equally important is a list where nothing is important. Weighting is the whole point.
Those four properties are also what any screening tool needs, whether the reader is a person, a keyword filter, or a language model. If your requirements are vague, every downstream step inherits the vagueness. Fix the source document and the rest gets easier for free.
The job description format that produces criteria
Here is the structure. Five blocks. The first three are the ad; the fourth is the ad and the scorecard at the same time; the fifth is admin. Keep them in this order in the file, because the requirements block is the one you will edit most and you want it easy to find.
Block 1: Role in one sentence
Not a paragraph. One sentence that says what the person owns and who they answer to. "You own the content engine for our B2B logistics software: blog, landing pages and lifecycle email, reporting to the Head of Marketing." If you cannot write that sentence, the role is not defined yet and no amount of template will save the hire.
Block 2: What the first 90 days look like
Three to five concrete outcomes with rough timing. This block does double duty: it is the most-read part of the ad by good candidates, and it is where your requirements come from. Requirements that do not trace back to a 90-day outcome are usually decoration. When you later argue about whether a requirement should be must-have or preferred, come back here and ask: can the first 90 days happen without it?
Block 3: How the work actually runs
Team size, tools, meeting rhythm, reporting line, office or remote, the awkward bits. This is where you say the team is four people and the person will write as well as brief, or that the role includes on-call weekends. Candidates who would have quit in month three self-select out here, which is cheaper for everyone.
Block 4: Requirements, written as criteria
The heart of the thing. Every line follows the same pattern:
[Requirement in one testable sentence] — [group] — [weight 1 to 10] — [what counts as evidence in a CV]
You publish only the first part in the ad. The other three columns live in the same file, in a table, and become your scorecard. Nobody outside the hiring team ever sees them, and nobody inside the hiring team can pretend they do not exist.
Block 5: Package and process
Salary range, benefits, the interview stages and roughly how long each takes. Put the number of stages in writing. A candidate who knows there are three steps will not disappear after the second one.
Turning one requirement line into a criterion
This is the part that people skip, so go slowly. You are taking a sentence written for humans and adding three pieces of metadata that make it comparable. Do it for one line and you will do the rest in about a minute each.
Step 1: rewrite the line until it is testable
Take the requirement as first written and ask: if two people read this and looked at the same CV, would they agree? "Good communication skills" fails immediately. Nobody agrees on that. "Has published writing you can read: portfolio, bylined articles, or a public blog" passes, because either the link is in the CV or it is not.
Some rewrites are mechanical. "Experienced in analytics" becomes "Can pull and interpret their own reports in GA4 or an equivalent analytics tool". "Team player" becomes either "Has worked in a team of five or more on a shared roadmap" or, more honestly, gets deleted, because you were not going to measure it anyway.
A test that catches most bad lines: try to write the evidence sentence. If you cannot finish "I would know this is true because the CV shows ___", the line is not a screening criterion. It may still be a real requirement — plenty of important things only surface in an interview — but it does not belong in the resume-scoring block.
Step 2: put it in one of five groups
Five groups is the sweet spot. Three is too coarse, ten and you will argue about the boundaries.
- Must-have. Without this, the person cannot do the job in the first 90 days, and it cannot be learned on the job in that window. Fail this and the candidate is out, regardless of how strong the rest looks.
- Important. Genuinely changes how well the job gets done, but a strong candidate could close the gap in a few weeks.
- Preferred. Makes life easier. Buys you speed, not capability. Never a reason to reject on its own.
- Basic. Administrative or baseline facts: a degree if you truly require one, work authorisation, a location constraint.
- Flexible. Nice extras you would enjoy but will happily trade away. Industry background often sits here, and it should sit here far more often than it does.
The discipline is in the must-have group. Write the list, then cut it in half. If you have seven must-haves you do not have a role, you have a fantasy. Three or four is normal. For a junior role, one or two.
Step 3: give it a weight from 1 to 10
Weight is not the same as group. Group decides whether a miss is fatal; weight decides how much a partial score moves the total. Two must-haves can carry different weights, because one of them is the spine of the job and the other is a hard prerequisite that most applicants clear anyway.
A simple way to set weights without agonising: give the single most important requirement a 9 or 10, give the least important a 2, and place everything else relative to those two anchors. Then check the totals by group. If your preferred and flexible lines together outweigh your must-haves, the scoring will reward well-rounded people over people who can do the job. Sometimes that is what you want. Usually it is not, and you find out only after you interview six pleasant generalists.
Step 4: name the evidence
For each line, write the specific thing you expect to see in the file. "Job title containing content, SEO or growth, at a company selling software to businesses, with dates covering 24 months or more." That sentence is what makes the score checkable later. When somebody disagrees with a score, you do not argue about impressions — you look at whether the evidence is in the document or not.
This is the same principle that makes any automated decision reviewable: a score without a reason attached is just an opinion with a number on it. We wrote about that idea in a different context in why explainable AI matters: every action needs a reason, and it transfers directly to hiring. A screening decision you cannot explain to the candidate, the hiring manager, or yourself six weeks later is a decision you will end up re-litigating.
Step 5: set the pass threshold and the consider band
Once the criteria have weights, the total is a weighted average from 0 to 100. You need two numbers on top of that: the score at or above which a candidate is a pass, and the band just below it where a human should look before rejecting. A pass mark of 70 with a consider band of 60 to 69 is a reasonable starting point for a mid-level role. You will move it after the first batch, and moving it is normal, not a failure.
A worked example: job description examples for a marketing executive
Everything below is an invented example, written to show the mechanics. The company, the role and the numbers are made up; use the shape, not the specifics.
The role: Marketing Executive (Content and Demand) at a 60-person company selling logistics software to mid-sized businesses. The marketing team is four people. The person joins to own content and to take over paid channels from an agency.
The published part of the job description
Role in one sentence. You own our content engine — blog, landing pages and lifecycle email — and take our paid search and paid social in-house from the agency, reporting to the Head of Marketing.
First 90 days. By week 4, you have audited the existing blog and published a rewrite plan. By week 8, you are shipping two publishable articles and one landing page a week and have taken over the paid search account. By week 12, you can show a single dashboard that connects spend, traffic and qualified leads.
How the work runs. Four people in marketing, one designer shared with product. You write, you brief, and you edit — this is a hands-on role, not a manager role. Weekly planning on Monday, one review call with sales on Thursday. Hybrid, three days in the office.
What we need from you.
- Two or more years owning content for a product sold to businesses: blog, landing pages or email, with your name on the output.
- You have personally managed a paid search or paid social budget, not just watched someone else manage one.
- You write to a publishable standard in English, and you can show it: portfolio, bylines, or a public blog.
- You can pull and read your own numbers in GA4 or an equivalent analytics tool without asking an analyst.
- You have worked with an SEO toolset such as Search Console plus one of the common third-party crawlers.
- You have used a marketing automation platform such as HubSpot or an equivalent.
- A degree in any field, or a portfolio that makes the degree question irrelevant.
- You can be in our office three days a week, or you are relocating.
- You have worked in logistics, supply chain or a comparable operational industry.
- You have briefed and managed a freelancer or an agency.
The criteria table it produces
Same ten lines, with the three columns the candidate never sees. Total weight is 49, which is only a denominator — it does not need to add to 100.
| # | Requirement | Group | Weight | Evidence in the CV |
|---|---|---|---|---|
| 1 | 2+ years owning content for a B2B product | Must-have | 9 | Title containing content, marketing, SEO or growth at a company selling to businesses, dates covering 24 months or more |
| 2 | Personally managed a paid search or paid social budget | Important | 7 | Named channel plus a budget figure, or an explicit statement of ownership rather than "supported" |
| 3 | Publishable writing, demonstrable | Must-have | 8 | Portfolio link, bylined article, personal site, or writing samples attached |
| 4 | Can pull and read analytics without help | Important | 6 | GA4, Looker Studio, Amplitude or similar named, tied to a reporting task |
| 5 | SEO toolset experience | Preferred | 5 | Search Console plus one crawler or rank tool named anywhere in the file |
| 6 | Marketing automation platform | Preferred | 4 | HubSpot, Marketo, ActiveCampaign or equivalent named in a role, not only in a skills list |
| 7 | Degree or equivalent portfolio | Basic | 2 | Education section, or a portfolio that satisfies criterion 3 at 80 or above |
| 8 | Can be in the office three days a week | Basic | 3 | Stated city, or a relocation note |
| 9 | Logistics or comparable operational industry | Flexible | 2 | Employer's industry identifiable from the company name or a one-line description |
| 10 | Has briefed or managed freelancers or an agency | Flexible | 3 | Words like briefed, managed, coordinated attached to external suppliers |
Notice what happened to the wish list. "Strong communication skills" became criterion 3 with a piece of evidence attached. "Data-driven mindset" became criterion 4 with named tools. "Industry experience preferred" dropped to flexible with a weight of 2, because on reflection nothing in the first 90 days requires it. That single demotion is often worth more than any other change you make, because industry experience is the requirement that most quietly shrinks a candidate pool.
What three candidates look like through this template
Still the invented example. Pass mark 70, consider band 60 to 69. Every score below is a per-criterion score from 0 to 100; the total is the weighted average, calculated as the sum of (score times weight) divided by 49.
Candidate A scores 85, 70, 90, 60, 80, 30, 100, 100, 0, 50 across the ten criteria. Weighted sum 3,505; divided by 49 that is 71.5. Both must-haves are comfortably above 50. Result: pass. Note that A scores zero on industry and only 30 on marketing automation, and it barely matters, because those lines are worth 2 and 4 points out of 49.
Candidate B looks stronger on almost everything: 40, 95, 85, 90, 90, 80, 100, 100, 100, 90. Weighted sum 3,985, which is 81.3 — the highest total of the three. But that first number is a 40 on a must-have. B ran excellent campaigns agency-side and never owned content for a product. Under this template B is not a fit, and the template says so out loud instead of letting a high average smuggle the gap past you.
Candidate C scores 70, 70, 75, 70, 80, 0, 100, 100, 0, 60. Weighted sum 3,220, which is 65.7 — inside the consider band. C is the person a spreadsheet sorted by total would bury on page three. Someone should read this file properly, because C is a plausible hire who is simply missing the two lowest-weight lines and one preferred tool.
That is the whole argument for the format in one paragraph. The template did not decide anything. It made the decision visible: B is out for a specific, statable reason, C gets a human read instead of an accidental rejection, and A moves forward with a record of why.
A second sample job description: customer support lead
A shorter one, to show the shape holds when the role is not a marketing role. Also invented.
Role in one sentence. You lead a team of six support agents covering email and live chat for our customers in two time zones, reporting to the Head of Operations.
First 90 days. By week 4, you know the top ten reasons customers write in. By week 8, first-response times are inside our stated commitment on every weekday shift. By week 12, you have a written escalation path agreed with engineering.
Requirements, in the same five columns:
| # | Requirement | Group | Weight | Evidence in the CV |
|---|---|---|---|---|
| 1 | Has managed a support team of three or more people directly | Must-have | 10 | Title including lead, supervisor or manager, plus a stated team size or direct reports |
| 2 | Worked in a ticketing or helpdesk system daily | Must-have | 7 | Zendesk, Freshdesk, Intercom, Jira Service Management or similar named inside a role |
| 3 | Has owned a response-time or resolution-time target | Important | 7 | A named metric with a number, or the phrase service level attached to a responsibility |
| 4 | Has written or maintained a help centre or macro library | Important | 5 | Knowledge base, help centre, canned responses or documentation mentioned as a deliverable |
| 5 | Comfortable working across two time zones or shift patterns | Preferred | 4 | Shift work, weekend cover, or a distributed team stated explicitly |
| 6 | Working proficiency in English and one regional language | Basic | 4 | Language section, or a role that clearly required it |
| 7 | Software or SaaS background | Flexible | 2 | Employer identifiable as a software company |
Seven criteria, two must-haves, total weight 39. Shorter than the marketing example and no worse for it. The number of criteria that works in practice sits somewhere between eight and sixteen for most roles; below eight you are not discriminating enough between candidates, above sixteen the low-weight lines stop moving the total and you are just making work.
Where job description templates go wrong
Three failure modes account for most of the damage. They are easy to spot once you know the shape, and all three are fixable in an afternoon.
Failure 1: requirements nobody can verify
"Strong communication skills." "Detail-oriented." "Passionate about our mission." "Ability to thrive in a fast-paced environment." Every one of these is unfalsifiable from a CV. Two reviewers will score them differently, the same reviewer will score them differently on a Monday and a Friday, and a language model asked to score them will produce something confident and arbitrary.
The tell is that you cannot write the evidence sentence. Try it: "I would know this candidate is detail-oriented because the CV shows ___". Nothing goes in the blank except things you are inferring, like typo-free formatting, which is a weak signal at best.
Two ways out. Either convert the line into something observable — "detail-oriented" often really means "has done work where an error would have been expensive and visible", which you can look for — or move it out of the screening block into the interview block, where a structured question and a work sample can actually test it. Both are fine. Leaving it in the screening block is not.
Failure 2: requirements that are measurable but meaningless
The over-correction. Somebody reads advice like the above, and the requirements list fills up with numbers that sound rigorous and predict nothing.
"Minimum 5 years of experience" is the classic. Five years of what? A person who spent five years doing the same quarter repeatedly is not ahead of a person who did two years across three very different problems. Years are a proxy, and a loose one. Where you can, replace the year count with the thing the years were standing in for: has shipped this kind of work end to end, has handled this scale, has operated without supervision.
Other examples in the same family: exact tool names when any comparable tool would do, degree class requirements for roles where nobody has ever checked, and certification requirements copied from a job posting at a much larger company. Each one is precisely measurable and each one filters on the wrong axis.
The test here is different from failure 1. Ask: if a candidate scored zero on this line but 90 on everything else, would I genuinely not want to talk to them? If the honest answer is that you would still take the call, the line is not a must-have, and quite possibly not important either.
Failure 3: requirement inflation
The most expensive one, and the hardest to see from inside. Requirement inflation happens quietly. The hiring manager adds a line. The team lead adds a line because their current headache is on their mind. Somebody copies three lines from a similar job posting at a bigger company, which itself copied them from somewhere else. Nobody removes anything, because removing a requirement feels like lowering the bar.
The result is a job description with eleven must-haves that describes a person who does not exist, or who exists and is being paid substantially more than your range. You post it, you get fewer applicants than expected, the applicants you get score badly against your own criteria, and after three weeks somebody quietly starts interviewing people who fail two of the eleven — which is the same as admitting the eleven were never must-haves.
Three checks catch this before you post:
- The 90-day test. For each must-have, name the specific thing in the first 90 days that becomes impossible without it. No answer means it is not a must-have. Demote it.
- The incumbent test. Would the best person currently doing a similar job on your team pass every must-have on the day they joined? Usually not, and the ones they failed are exactly the ones to demote.
- The count test. More than four must-haves is a warning. More than six is a broken template, not a demanding role.
There is a fourth, softer symptom worth watching: a requirements list where every line is phrased as a minimum. "Minimum 3 years." "At least a bachelor's degree." "No fewer than two enterprise implementations." Minimums are binary gates, and a list of gates is a filter, not a scorecard. A scorecard grades. If your template only knows how to exclude, it will exclude your best unusual candidate and never tell you it did.
The related trap: writing the whole thing with AI and shipping it
Asking a model to draft a job description is genuinely useful for the ad blocks — the role sentence, the 90-day outcomes, the tone. It is much less useful for the requirements block, because a model with no context about your team will produce the median requirements list for that job title, which is the inflated one it read a thousand times in training. You get eleven must-haves and a "5+ years" line for free.
Use it for the first draft and then do the conversion work yourself: rewrite each line to be testable, assign groups, set weights, write the evidence. The same rule that applies to any AI-assisted writing applies here, and we set it out in using AI the right way: the model drafts, you supply the judgement and the facts it cannot know.
Thresholds, the consider band, and what to do when nobody passes
Once the criteria exist, two numbers control the flow of candidates: the pass threshold and the width of the consider band below it.
Set the threshold by what you can process, not by what feels rigorous. If you can run twelve phone screens this week and 180 people applied, you need a threshold that yields roughly twelve to eighteen passes. Start at 70 for a mid-level role, run the first batch, and look at the distribution. Too many passes and the criteria are not discriminating; raise the threshold or raise the weight on the lines that actually separate people. Too few and you are either inflated or your sourcing is wrong.
The consider band exists because a weighted average is a blunt instrument near the boundary. A candidate at 68 and a candidate at 71 are not meaningfully different, and treating the first as a rejection and the second as a pass is false precision. Ten points wide is a sensible default. Everyone in the band gets a human read before a decision, which is a manageable amount of work if the band is narrow and a signal that your threshold is wrong if it is not.
Two rules keep the band honest. First, a must-have failure overrides the total: if a candidate scores below 50 on any must-have criterion, they do not pass, no matter how high the average is. Candidate B in the worked example is exactly that case. Second, when you move a threshold mid-process, rescore the batch you already ran or note the change, because a shortlist assembled under two different thresholds is not a shortlist, it is two half-processes stapled together.
What if nobody passes? Resist the urge to lower the threshold first. Run the three inflation checks above on your must-haves. In most rounds where nobody passes, one must-have is doing all the damage, and it is usually the one somebody added at the last minute.
Keeping the template honest: a short recurring routine
A scoring template decays if you never look back at it. The loop below takes maybe half an hour per round and is the difference between a template that improves and a template that quietly encodes last year's assumptions forever.
- After the first 20 CVs, check the spread. If almost everyone scores between 55 and 65, your criteria are not separating candidates. Usually one or two lines are doing all the work and the rest are noise. Raise the weight on the ones that discriminate.
- Spot-check five scores by hand. Pick two passes, two rejects and one from the consider band. Read the evidence attached to each criterion and ask whether you agree. Disagreement almost always traces to a vague requirement line, not to bad reading.
- Log every rejection reason at the interview stage. When a candidate who passed the screen fails the interview, note which criterion should have caught it. That note is next round's new criterion.
- Log the reverse too. When you interview someone who scored in the consider band and they turn out strong, find out what the template underweighted. This is the loop that fixes inflation, because it is the only one that shows you the people you nearly lost.
- Review the must-have list at the end of every round. Any must-have that no shortlisted candidate actually failed was not doing work. Any must-have you overrode for a candidate you liked was never a must-have.
- Keep one file per role, versioned. When you reopen the role in six months, you want last round's criteria and last round's notes, not a blank template.
Doing it by hand, and when a tool starts to earn its place
For 20 or 30 applications, do this by hand. A table in a spreadsheet with your criteria as columns, one row per candidate, and half a day of reading. It works, it is cheap, and reading CVs yourself for a role you are hiring teaches you things about your own criteria that no summary will.
The manual version breaks somewhere around 80 to 100 files, and it breaks in a specific way. Scoring quality drops after the first hour or two. The criteria stay fixed on paper but drift in your head. The last fifty files get a fundamentally different read from the first fifty, and because the drift is gradual you cannot tell which fifty were scored properly. Splitting the pile across two colleagues does not fix it; it gives you two drifting scales instead of one.
That is the point where software helps, and it helps with exactly one thing: applying the same criteria to file number 180 as to file number 1. Note what it does not solve. It does not tell you whether your criteria are right, it does not decide who to hire, and a bad requirements list processed at speed is just a bad requirements list applied consistently. This is also where the difference between storing candidates and judging them matters — a system of record moves people between stages but rarely tells you who is worth calling, a gap we unpacked in what an applicant tracking system actually does.
If you want the template-to-scorecard conversion done for you, that is the specific job Orova Recruit does: you upload the job description as a file or paste the text, it proposes 8 to 16 criteria sorted into the same five groups with a weight from 1 to 10 and a suggested pass threshold, you edit anything you disagree with, and it scores each CV from 0 to 100 per criterion with the supporting text quoted from the file. The overall score is a weighted average computed on the server rather than a number the model adds up itself, and a must-have scored below 50 forces a fail regardless of the total — the same rule Candidate B ran into above.
Two things extend the template beyond the screening stage. If you have no draft yet, the flow runs the other way: type the role title and a few bullets, press the AI button, and the finished description lands in the editor for you to correct. And because the criteria carry weights, the same weights are reused downstream — interview questions are generated per criterion with what to listen for, and the interview score is the weighted result of those answers on the same 0–100 scale, so a candidate's CV score and interview score can be read next to each other without mental arithmetic.
Whichever route you take, keep the editing step. A tool that proposes criteria and does not let you change them is worse than a spreadsheet, because it hides the assumptions instead of making you write them down.
From scorecard to interview guide: reuse the weights you already set
Most of the effort in a good job description template goes into one thing: deciding what actually matters and how much. That decision usually gets used once, for reading CVs, and then the file closes and the interview starts from scratch with questions invented on the way to the meeting room. It is a waste, and worse, it means the same person is measured with two different rulers within the same week.
Reusing the template downstream takes about twenty minutes. Here is the sequence.
Step 1: decide which criteria need re-testing
Not every criterion belongs in an interview. There are three kinds. The first is what a CV can essentially prove: a degree, a certification, years in a named role. Confirm it, do not spend interview time on it. The second is what a CV asserts but cannot prove: "led a team of five", "strong with tool X". This is exactly what interviews are for. The third is what a CV never touches: how someone handles disagreement, what they do after negative feedback, why they are leaving. Only a conversation reaches it.
A good question set is almost entirely categories two and three. If you notice you are asking a lot of category one, the interview is being used to read the CV out loud.
Step 2: carry the weights across unchanged
This is the small trick that changes everything about comparability. A criterion weighted 9 on the CV scorecard produces a question weighted 9 on the interview rubric. That way a candidate who scored well on paper but falls apart on the heaviest criterion drops exactly as far as they should, rather than being rescued by fluent answers to lightweight questions.
Step 3: write the "what good sounds like" line
Every question should carry one line describing a strong answer. One line is enough: "gives a specific number, separates what they did from what the team did, and names a trade-off they made." That line does three jobs. It lets a new interviewer score close to an experienced one. It speeds up scoring, because the scale is not being reinvented per candidate. And it is the evidence you produce when someone asks why candidate A outranked candidate B.
Step 4: agree the rule for questions you never asked
Interviews always run short and questions always get dropped. The rule should be hard: an unasked question is recorded as not covered and scores zero, rather than receiving a middling score "to be fair". It sounds harsh, but it puts the pressure in the right place — interviewers ask the heavyweight questions first, and the final table reflects what was actually verified rather than what was assumed.
Step 5: revise the template after each hire
When the role is filled, reopen the criteria and mark three things: which criteria separated candidates usefully (high scorers really did perform), which every candidate passed (useless — cut them or lower the weight), and which nobody passed (your requirement may be out of step with the market). Those three marks are the edits for the next version of the job description template. A good template is not one written well the first time; it is one corrected after each round with real data.
Frequently asked questions
How many requirements should a job description have?
Eight to sixteen scoring criteria works for most roles, with two to four must-haves. Below eight, the scores bunch together and you cannot separate candidates. Above sixteen, the extra lines carry so little weight that they cannot change an outcome, and you have added work for nothing. The published ad can show fewer lines than the scorecard holds — combine related criteria into one readable bullet for the ad if the list gets long.
Should the weights add up to 100?
No, and forcing them to is a common waste of an afternoon. Weights are relative, and the total is whatever it is — 49 in the worked example above. The weighted average divides by the total, so the scale takes care of itself. Forcing a sum of 100 means every time you add a criterion you have to re-balance all the others, which people avoid by not adding criteria.
Can I reuse one job specification template across several roles?
Reuse the structure every time. Reuse the criteria only between genuinely similar roles, and even then read every line before you post. The blocks, the five groups, the weight scale and the evidence column are stable across any role. The specific requirements are not, and a template copied from a senior role to a junior one with the year counts edited is the fastest route to requirement inflation.
What if a requirement only shows up in an interview, not a CV?
Keep it, but put it in the interview block rather than the screening block. Plenty of things that matter — how someone handles disagreement, whether they ask good questions, how they think under pressure — are invisible in a document. Screening criteria are only the subset you can verify from a file. Mixing the two is what produces unscoreable lines like "collaborative" sitting in a resume rubric.
Does this work for high-volume hiring?
It works better, not worse. When you are looking at hundreds of applications for a handful of similar openings, consistency is the entire problem, and a fixed set of weighted criteria is the only thing that gives you it. Two adjustments help: keep the criteria list shorter, around eight lines, because long lists are harder to apply consistently at volume, and widen the consider band slightly so borderline candidates get a human read rather than a hard cut.
Will scoring criteria remove bias from hiring?
They reduce one specific kind: the inconsistency that comes from reading the same evidence differently depending on who wrote it, what order it arrived in, and how tired you are. They do not fix bias built into the criteria themselves. If a must-have effectively requires a background only one group of people tends to have, writing it down with a weight of 9 makes it consistent, not fair. The evidence column is your defence — read the criteria as a list and ask, for each one, whether it is testing the job or testing a résumé pattern.
What to do this week
Pick the role you are hiring for right now, or the one you know is coming. Open the job description you were about to post and do four things to it.
First, write the 90-day outcomes if they are not there. Second, take the requirements section and rewrite every line until you can finish the sentence "I would know this from the CV because it shows ___" — delete or relocate the lines that fail. Third, tag each surviving line with a group and a weight from 1 to 10, and force the must-have count down to four or fewer using the 90-day test. Fourth, set a pass threshold and a ten-point consider band, then write both numbers in the file so the next person to touch it knows what you decided.
That is an hour of work. At the end of it you have one document that runs the ad and the screening, a shortlist you can defend line by line, and a version of the role you can improve after the next round instead of rebuilding from scratch.
Let AI read and score resumes against your JD
Orova Recruit turns your job description into weighted criteria and scores every CV with evidence quoted from the file.
Try it free