Gamified LMS: What Actually Lifts Completion Rates
Sixty per cent of your people started the compliance course and thirty per cent finished it. Somebody in the meeting suggested gamification, and the demo you sat through had badges in it, and a streak counter, and a leaderboard with the sales team's names at the top. It looked good. Nobody in the room could say which part of it would raise the number that is actually wrong.
That is the honest position most training teams are in when they start looking at a gamified LMS. The word covers at least three unrelated mechanisms that work through different channels and fail in different ways, and vendors have no incentive to separate them because the bundle sells better than the parts.
So this piece separates them. What the published research found, reported with its confidence intervals rather than as a headline, including the awkward detail that the effects are smallest on the outcome most managers care about. The single design choice that turned out to matter more than badges or points. Where leaderboards quietly lose the bottom half of the room. What timed scoring is really rewarding. And then the unglamorous argument that runs underneath all of it: if your problem is completion, the settings that fix it are not the fun ones, and they are usually already in the product you have.
Does gamification actually improve learning?
Yes, by a small amount, and unevenly. The published meta-analysis found significant small effects on cognitive outcomes (g = .49), motivational outcomes (g = .36) and behavioural outcomes (g = .25). The cognitive effect held up among the most rigorous studies; the motivational and behavioural ones were less stable. Design choices mattered more than whether gamification was present.
That is the honest headline, and it is more encouraging than the cynical position and much less encouraging than the sales deck. Gamification is not a gimmick that does nothing. It is also not a lever you pull to double completion, and anyone promising that is either selling or has not read the literature.
Why completion is low, before you treat it
Worth five minutes before any of the rest, because gamification is a treatment and the diagnosis usually has not been made. Ask the people who did not finish. Not a survey — five conversations, which take an hour and produce more than any dashboard.
In practice the answers cluster into four causes, and only one of them responds to gamification.
They did not know it was mandatory. The most common and the most embarrassing. The course was assigned in a system they check rarely, the email arrived in a busy week, and nothing since has told them it is outstanding. Fixed by making the lesson required and by reminders, not by points.
They could not find the time and nothing forced the issue. The course competes with work that has a deadline, and it does not have one. Fixed by a due date and by a manager who sees the outstanding list, which is an organisational fix rather than a product one.
They started and the material lost them. Too long, too abstract, or written for somebody with more background than they have. This is the one people least want to hear, because the fix is rewriting the course. It is also the one where the drop-off point in your tracking data will tell you exactly where the problem is, if you look at how far each person got rather than at the completion percentage.
They finished the parts they cared about and stopped. Rational behaviour, and arguably the course's fault for bundling material for four audiences into one course. Fixed by splitting it.
Only the third cause is a genuine engagement problem, and it is the one gamification addresses. The other three are addressed by settings, by management, and by editing. That is not a reason to skip the game part; it is a reason to know which of your four groups you are aiming at, because if most of your non-completers are in group one, a leaderboard will change nothing and you will conclude that gamification does not work when what actually happened is that you treated the wrong condition.
Three different mechanisms wearing one word
Before any evidence is useful, the word has to be broken apart, because "we added gamification" describes at least three interventions that share nothing except a marketing category.
Mechanism one: scoring and competition. Points, ranks, leaderboards, timed answers. This works — where it works — through social comparison and the pull of a visible position. It is fast to implement, immediately visible, and the most likely of the three to backfire, because social comparison cuts both ways and the direction depends on where the learner already sits.
Mechanism two: narrative and structured feedback. A story frame, a character, a scenario, a sequence of challenges that build. The research literature calls the story part game fiction, and it is the mechanism people mean when they say a course "felt like a game" rather than "had a score attached". It works through immersion and through making abstract material concrete. It is much more expensive to build, which is why it appears in demos far less often than badges.
Mechanism three: social interaction. Learning with and against other people at the same time: teams, shared rooms, discussion, collaborative challenges. It works through accountability and through the ordinary human fact that doing something alongside other people is different from doing it alone at 4pm on a Friday.
Keep the three separate as you read the rest of this, because the evidence does not treat them equally, and the specific combination that performed best is not the one most platforms ship.
What the evidence says, and how small the effects are
The reference point most worth having is the meta-analysis by Michael Sailer and Lisa Homner, published in Educational Psychology Review in 2020, which pooled the experimental studies on gamification and learning and separated the outcomes into three types.
Cognitive learning outcomes — what people know and can work out afterwards — showed g = .49, with a 95% confidence interval of 0.30 to 0.69, across 19 effect sizes and 1,686 participants. Motivational outcomes came in at g = .36, interval 0.18 to 0.54, across 16 effect sizes and 2,246 participants. Behavioural outcomes — what people actually do — came in lowest at g = .25, interval 0.04 to 0.46, across 9 effect sizes and 951 participants.
Three things worth pulling out of those numbers before anyone builds a business case on them.
They are described as small effects, by the authors. Not "transformative", not "significant" in the everyday sense. Small, real and worth having — the sort of improvement that is worth a modest investment and is not worth reorganising a training programme around.
The order runs the wrong way for most buyers. The strongest effect is on knowledge, the weakest is on behaviour. Organisations almost always buy training to change behaviour: fewer incidents, better calls, correct procedure followed. That is exactly where the evidence is thinnest, with the fewest studies, the smallest sample, and a confidence interval whose lower bound sits close to zero.
Stability differs by outcome. When the authors re-ran the analysis using only studies with high methodological rigour, the cognitive effect stayed stable. The motivational and behavioural effects were less stable. In plain terms: the finding you can lean on hardest is that gamification helps people learn material, and the finding you should hold most loosely is that it changes what they do.
None of this argues against gamifying. It argues against paying a premium for it as though it were the deciding factor, and it argues strongly for measuring the outcome you actually care about rather than the one that improves most easily.
The moderator that mattered more than the mechanism
The more useful half of that paper is not the headline effect sizes. It is the moderator analysis — the part that asks which design choices made gamification work better or worse.
Two moderators came out significant for behavioural learning outcomes: the inclusion of game fiction, and the inclusion of social interaction. And the specific finding worth writing on the wall: including game fiction, and combining competition with collaboration, were particularly effective for fostering behavioural outcomes.
Combining competition with collaboration. Not competition on its own. The design that performed best on the outcome that is hardest to move was the one where people compete as part of a group, or collaborate within a competitive frame — teams against teams, a shared score somebody contributes to, a room where you are both being ranked and working alongside others.
This is a specific, actionable finding and it points away from what most platforms ship. A leaderboard of individuals is pure competition. A badge is pure individual achievement. Neither includes the collaboration half that the analysis identified as particularly effective. If you are choosing between two products and one of them can put a room of people into a shared session while the other can only rank individuals over time, that difference is more supported by the evidence than any feature count.
Game fiction is the harder one to act on, because building a narrative course is genuinely expensive and most compliance material resists storytelling. A realistic middle path is scenario-based questions — a specific situation with a specific character and a decision — rather than abstract recall. It is not a story, but it is a long way from "which of the following is a definition of".
Leaderboards: who they lift and who they lose
The most common gamification feature and the one that needs the most care, because its effect depends on where the learner already stands.
For the top few, a public ranking is genuinely motivating. They can see the gap to first place, it is small, and closing it is a realistic goal that produces effort. For the person in position 40 of 60, the same board says something different: you are behind, the gap is not closable, and the effort required to move up two places is not worth it. That person disengages, and they disengage more thoroughly than they would have without the board, because now the disengagement has a public number attached to it.
This is not an argument against leaderboards. It is an argument against one specific implementation: a permanent, public, individual, all-time ranking. Four design changes fix most of it.
Reset frequently. A board that starts fresh each session gives everybody a real chance today. An all-time board sorts people by how long they have been in the company.
Show a window, not the whole list. Your position plus the two above and two below. The gap is closable, the comparison is with peers rather than with the outlier, and the bottom of the list is not published.
Rank teams as well as individuals. This is the collaboration half from the evidence, and it also means the person in position 40 is contributing to something rather than only losing at something.
Keep the session short. A live round of ten questions is over in ten minutes. Nobody is at the bottom for a month.
One more, less obvious: do not attach anything consequential to leaderboard position. The moment a ranking affects a review, a bonus or a manager's opinion, it stops being a game and becomes a performance measure, and people optimise it accordingly — rushing, guessing, retaking, or getting someone else to play. Quiz scores that count belong in a proper assessment, marked on the server, with the answer key never reaching the learner's device. Games and assessments are different instruments and mixing them ruins both. If you are setting up the assessment side properly, building practice tests that actually diagnose is a different craft from designing a live round.
Timed scoring, and what it is actually rewarding
Most live quiz formats give more points for a faster correct answer. It is worth understanding the shape of that, because it decides what behaviour you are training.
A typical formula splits the points in half: half for being right, half scaled by how much of the question's time limit was left when you answered. Our own implementation is exactly that — a correct answer earns the question's points multiplied by 0.5 plus 0.5 times the fraction of time remaining, rounded, with a wrong or missing answer earning nothing and multiple-choice questions requiring the complete correct set.
Two consequences follow, and both are design decisions rather than accidents. First, being right is worth at least half the points no matter how slowly you get there, which stops the format from purely rewarding reflexes. Second, speed still decides the ranking among people who all answered correctly, which is the whole engine of the format.
So the question to ask is whether speed is a virtue for your material. For recall and recognition — product names, safety signals, policy limits, spotting a phishing email — speed is a genuine part of competence, and rewarding it is honest. For judgement, ethics, prioritisation or anything with a reasonable case on both sides, rewarding speed trains exactly the wrong instinct: answer before you have thought. That material belongs in an untimed assessment with room for a written answer, and putting it in a live round teaches people that hesitation is a scoring penalty.
There is a practical constraint that follows from the same place. A live round can only handle question types that mark instantly — single answer, multiple answer, dropdown. Anything requiring a written response or a scale is skipped when the room plays. That is not a limitation to work around; it is the format telling you which material belongs in it.
Certificates as the end reward
The most underrated motivator in this whole category, and the least game-like.
A certificate works for reasons that have nothing to do with points. It is external to the platform, so it means something outside the course. It is durable, so it is still there next year. And if it carries a serial that anybody can type into a public lookup page, it is verifiable, which makes it worth something to the learner rather than only to the compliance file.
That last property is what separates a certificate from a badge. A badge lives inside the system that issued it and is worth exactly what that system is worth to the person holding it, which for a mandatory internal course is usually nothing. A verifiable certificate is a claim the learner can make to somebody else. If you are choosing where to spend design effort, that difference is larger than anything on a leaderboard, and what makes a certificate worth issuing covers the details that decide whether anybody values it.
The design rule is simple: issue automatically the moment the pass mark is cleared, so there is no administrative gap between finishing and being recognised. A certificate that arrives three weeks later from a person who had to be reminded is not a reward, it is paperwork.
Enforcement is not gamification, and it usually fixes completion faster
Here is the argument this article has been building toward, and it is the one nobody selling a gamified LMS will make for you.
If your problem is that people do not finish courses, gamification is an indirect solution. It works by making the course more appealing, which raises the chance somebody chooses to continue. That is real, and it is small, and it is competing against everything else in that person's afternoon.
The direct solutions are settings, not games, and they are usually already in the platform you have.
Mark the lesson required. Sounds trivial. A large share of non-completion is people not knowing something was mandatory.
Hold learners for a minimum time, counted only while the tab is genuinely in front of them. The second half of that sentence is the entire feature. A timer that runs while the tab sits behind a spreadsheet measures nothing and everybody learns to game it within a week.
Block skipping ahead. Removes the click-to-the-end path that produces a completion record with no learning behind it.
Block downloads. Keeps the material in the tracked environment rather than in a PDF somebody skims once.
Allow one attempt. Changes how seriously the assessment is taken, and stops the quiz being solved by repetition.
None of that is fun. All of it moves the completion number more reliably than a badge, because it addresses the actual mechanism — people not finishing because nothing required them to and nothing noticed that they had not.
The honest version of the argument is that these two things solve different problems. Enforcement raises the number of people who reach the end. Gamification raises the chance that reaching the end left something behind. You want both, and if you have to sequence them, do enforcement first, because it is free, it takes an afternoon, and it gives you a baseline that any later gamification can be measured against.
Measure three different things, or you will fool yourself
Completion is the easiest number to collect and the least informative one. If you only track it, and you turn on enforcement, it will go up, and you will have learned nothing about whether the training worked.
Finished. The completion record. Now includes anybody who was made to sit there. Necessary for the compliance file, close to useless as evidence of learning.
Remembered. A short assessment, ideally a few weeks after the course rather than at the end of it. The gap between end-of-course score and four-weeks-later score is the single most informative number in corporate training, and almost nobody collects it because it requires sending a second thing.
Applied. The behaviour the training existed to change: incident counts, error rates, call scores, audit findings. Slow, noisy, confounded by everything else happening in the business — and the only measure that answers the question the budget was approved against. This is also, per the meta-analysis, where gamification's effect is weakest and least stable, which is a reason to measure it rather than a reason to skip it.
The per-learner tracking that makes the first two possible is ordinary LMS plumbing: how far each person got, how many seconds they stayed, how often they came back. Together those three tell you the difference between somebody who worked through the material and somebody who left the tab open, and no amount of gamification tells you that.
How to test this on one course before changing the system
Six weeks, one course, no procurement.
Week one: measure what you have. Completion rate, average time in the material, and a short knowledge check at the end. This is your baseline and you cannot skip it, because everything after this is a comparison.
Week two: turn on enforcement only. Required lessons, a minimum time counted properly, no skipping ahead, one attempt. Change nothing else. Measure the same three things.
Weeks three and four: add one live session. Take the material people find driest, build a ten-question round from it in the three question types that mark instantly, and run it with a room of people at once. This is the mechanism the evidence supports most for behaviour — competition combined with being in a group — and it costs an hour to build.
Week five: knowledge check again, on the same questions. Not the ones used in the live round. This is your retention signal.
Week six: compare. If enforcement moved completion and the live round moved retention, you have learned which lever does what in your organisation, which is more useful than any vendor's case study. If neither moved anything, your problem is the content, and no format fixes content.
Do this before you evaluate platforms, not after. It converts the shortlist conversation from "which has more features" into "which does the two things we now know work here", and it is the cheapest procurement research available. The wider platform question is a separate one, and what a corporate LMS is actually for is the frame for it.
What it costs to run, which is not what it costs to buy
The licence line is the small part. Three running costs decide whether a gamified programme survives its first year, and none of them appear in a quote.
Somebody has to run the room. A live session needs a host: opening it, reading the question out, keeping the pace, handling the person whose phone will not scan. That is a person's half hour, repeated every time. It is not much, and it is a real recurring commitment that has to belong to somebody by name or it quietly stops happening after the third month.
Questions wear out. A live round is only interesting once per group. Run the same ten questions with the same team next quarter and you are testing memory of the quiz rather than of the material. Budget for writing new rounds, which is where most of the ongoing effort actually goes — the format is cheap, the content for it is not.
Novelty decays. The first live session is genuinely popular. The fourth is a normal Tuesday. Every measurement you take in month one is inflated by the novelty and every comparison you make against it later will look like a decline. The way to avoid fooling yourself is to take your baseline before the first session and your comparison after the third, not after the first.
There is one cost that people expect and that does not usually appear: learner-side setup. If joining a room means scanning a code and typing a name, there is no account provisioning, no password resets and no support queue. If it requires an app and a login, that is a per-session administrative cost you will pay in every session forever, and it is the single largest predictor of whether a room of forty people actually starts on time.
What to ask a vendor, and what to ignore
Four questions that separate the products, and three demo features that tell you nothing.
Ask: can a group play at the same time, in the same room? Simultaneous play is the collaboration-plus-competition combination that the evidence favours. Asynchronous individual scoring is not the same thing, however many points it awards.
Ask: where is the marking done? If a live quiz marks in the browser, the answers are in the browser, and somebody will find them. Server-side marking, with the key never sent to the player, is the only version that survives a determined room.
Ask: what do learners need in order to join? An account, an app and a password before a ten-minute quiz is enough friction to kill the session in a room of forty people. Scanning a code and typing a name is not.
Ask: which enforcement settings exist, and does the timer count only focused time? This is the completion question, asked precisely. Vendors who have thought about it will know immediately what you mean.
Ignore: badge galleries. Cheap to build, individually earned, and the mechanism least supported by the moderator analysis.
Ignore: streak counters. They reward daily habit, which is the wrong shape for a course somebody takes once. They also punish annual leave.
Ignore: point stores and avatars. Long to build, quick to lose novelty, and no evidence in the literature that they carry any of the effect.
The question of what the whole platform must do sits underneath this, and if you are earlier in the process than feature comparison, the plain guide to employee training software and the tiers of authoring tools are the two pieces that come before this one.
Where Orova Training fits, and where it does not
Plainly, because the whole argument above is that vague claims are how this category sells.
The ranked quiz mode is real and it is the shape the evidence favours. You open a room from the presenting screen; learners join by scanning a code, with no account and no app and no cap on numbers; they wait in a lobby; the question appears for everybody at once; and the board reshuffles after each one. Scoring is done on the server, and the correct answer never reaches a player's device — even at the reveal, a player is told only whether their own answer was right. Only the three instantly-markable question types play in a live round; anything written or scaled is skipped, and the opening screen says so.
Ranking between questions shows the top five rather than the whole list, which is the windowed leaderboard argued for above. The ordinary take-it-alone quiz is still there for material that does not belong in a room.
The enforcement half exists too: mark a lesson required, hold learners for a set number of seconds that only count while the tab is actually in front of them, block skipping ahead, block downloads, allow one attempt each. Progress is tracked per learner — how far they got, how many seconds they stayed, how often they came back. Certificates issue automatically once the pass mark is cleared, each with its own serial that anybody can type into the public lookup page.
What is not there, said once so nobody has to infer it: no badge gallery, no streak counters, no point store, no avatars, no experience levels. If those are what you mean by gamification, this is not that product. The bet the design makes is that a shared room and a certificate somebody can verify do more than a cabinet of icons, and the moderator analysis is the reason for the bet rather than a justification found afterwards.
One more limit worth knowing, since it comes up straight after this decision: courses can be exported as SCORM 1.2 packages, but a live ranked session is not part of a package. A package runs offline inside somebody else's system, and a shared room needs a server. If your material has to be handed to a client's platform, what SCORM compliance actually requires covers what does and does not travel.
Common questions
Does gamification increase course completion?
Indirectly and modestly. The meta-analysis found small effects, strongest on knowledge and weakest on behaviour. Completion specifically responds faster to enforcement settings — required lessons, a properly counted minimum time, no skipping ahead — because those address the mechanism directly rather than by making the course more appealing.
Are leaderboards a good idea?
They help the people near the top and can actively lose the people near the bottom. Reset them often, show a window around the learner's position rather than the whole list, rank teams as well as individuals, and keep sessions short. Never attach a consequence to a position, or it stops being a game.
What does the research actually say?
Sailer and Homner's 2020 meta-analysis in Educational Psychology Review reported significant small effects: g = .49 on cognitive outcomes, .36 on motivational, .25 on behavioural. The cognitive effect held up among high-rigour studies; the other two were less stable. Game fiction and social interaction were significant moderators for behavioural outcomes, and combining competition with collaboration was particularly effective.
Are badges worth building?
They are the cheapest feature to ship and the one with the least support in the moderator analysis, since they are purely individual and carry no collaboration. A verifiable certificate does the same job better, because it means something outside the platform that issued it.
Should quiz scoring reward speed?
For recall and recognition, yes — speed is part of competence there. For judgement or ethics, no, because rewarding speed trains answering before thinking. Formats that split the points, half for correctness and half scaled by remaining time, keep the ranking interesting without making a slow correct answer worthless.
Can gamification replace good content?
No, and the six-week test above is designed to expose that. If neither enforcement nor a live round moves any of your numbers, the problem is the material. A format change makes dull content briefly more tolerable and leaves it dull.
How do we stop learners cheating a live quiz?
Mark on the server and never send the answer key to the player's device. That is the whole defence, and it is architectural rather than procedural. Anything marked in the browser can be read in the browser.
Does any of this apply to compliance training?
The enforcement half applies directly and is usually what compliance teams need. The gamification half applies to retention, which compliance training rarely measures and probably should — a completion record proves attendance, not understanding. The quiz side of that is where most of the retention evidence gets collected.
What to do with this
Open the settings of the course with your worst completion rate and check four boxes: is it marked required, is there a minimum time, does that timer count only focused time, and is skipping ahead blocked. In most organisations at least two of those are off, and turning them on costs nothing and takes an afternoon.
Then take the driest twenty minutes of that course and build a ten-question live round from it. Run it with a room of people rather than sending it out. That single change contains both of the moderators the evidence identified — a group, and a competition — and it costs an hour.
Measure three things afterwards, not one: how many finished, how much they remembered four weeks later, and whether anything changed in the work. The first will move because you enforced it. Whether the second moves is the only real evidence you will ever get about the gamification part, and it is the number that tells you whether to do it again.
Everything else in this article exists to stop you buying the expensive version of a small effect. That is a less exciting promise than the demo, and it is the one supported by the numbers.
The live room and the enforcement switches, both included
Orova Training has the ranked quiz mode built in: you open a room from the presenting screen, learners join by scanning a code with no account and no app, and scoring happens on the server, so the answer key never reaches a player's device. The unglamorous half is there too, from required lessons to blocked skipping. You pay for usage rather than for tiers.
See Orova Training