Skip to content
What You Are Measuring

Assessment predicts the work to the extent that it resembles the work

Fifty notes on skills assessment: what the evidence actually supports, how to design a process from the job rather than from a catalogue, and the gap between what organisations measure and what determines success.

The one finding to build on

Selection has been studied for a century and the most robust result is also the simplest: assessments that resemble the work predict performance better than assessments that do not.

Once the criteria described in “Assessment predicts the work to the extent that it resembles the work” are fixed, the practical challenge is running them consistently across assessors. A coordinator can use long-term time tracking software to plan review time, compare the effort spent at each stage and spot avoidable delays, without confusing administrative efficiency with the quality of the evidence collected from candidates.

For an independent benchmark, compare the process with U.S. OPM assessment resources; the useful test is whether the local method remains job-related, proportionate and explainable.

Everything else follows from it. A work sample outperforms an interview because you are watching the thing you want to predict rather than inferring it. A structured interview outperforms a conversation because it asks about what somebody actually did rather than how they present. And unstructured conversation predicts weakly while feeling highly informative to the person conducting it, which is the largest gap between confidence and accuracy in the field.

The ordering between methods is stable: work samples, structured interviews and job knowledge tests near the top; unstructured interviews, personality questionnaires used alone and years of experience near the bottom.

A note about the numbers

You will see precise validity coefficients quoted confidently, frequently in vendor material. Treat them with caution. A major re-examination in recent years found that statistical corrections applied to the original meta-analyses had inflated those figures, sometimes considerably.

The conclusion is not that selection methods do not work. It is that the ranking between them is more reliable than the numbers attached, and that everything predicts less strongly than the older figures suggested. This collection quotes the ordering and not the coefficients, deliberately.

The gap that produces bad hiring

Ask what makes somebody good at a role and you get one list. Look at what the assessment measures and you get another. The distance between them is where most expensive hiring mistakes live.

What usually gets measured: credentials, because they are on the application. Years of experience, for the same reason. Interview performance, which measures the ability to perform in interviews. And availability to attend four rounds, which nobody intends to select on and everybody does.

What usually matters: whether the person can do the core task, which is frequently not tested directly. Judgement in ambiguous situations, which is what most roles actually consist of. Working with the specific people and systems in place. And learning speed, which matters more than any current knowledge in a role that changes.

Closing the gap is not complicated. Start from a written job analysis, assign each requirement to exactly one method that can assess it, and remove anything that fails that test. Most processes lose two stages and improve.

Two kinds of error, only one of them visible

Hiring somebody who cannot do the job is expensive and obvious. It is what most processes are built to prevent.

Rejecting somebody who could is invisible, unmeasured and usually more common. Nobody follows up on the candidates they turned down, so nothing ever corrects the impression that the process is working.

Every additional hurdle reduces the first error and increases the second. That trade-off is the central design decision in any assessment process and almost nobody makes it explicitly. Instead, stages accumulate: each was added for a reason, none is ever removed, and a role that was two interviews becomes five rounds and a project.

What length actually selects for

Add up what your process asks of a candidate: application, tests, exercises, interviews, travel. Above about four or five hours in total, you are selecting on availability rather than on capability.

That filter is not random. It removes people in jobs with no flexibility, people with caring responsibilities, people who cannot afford unpaid leave, and disabled candidates for whom each appointment costs more. None of those correlates with capability and all of them correlate with the diversity outcomes organisations say they want.

Strong candidates with options also leave long processes, which means length selects against exactly the people you were competing for.

Fairness and prediction point the same way

This is the useful part, and it surprises people: the things that make a process defensible are the same things that make it predict better.

A written job analysis justifies your requirements and tells you what to assess. Structure removes the space where individual bias operates and improves validity. Contemporaneous records with evidence defend a decision and enable calibration. Adjustments offered by design widen the pool and reduce measurement error. And measuring your own outcomes demonstrates that you looked and tells you whether any of it works.

None of that is compliance bolted onto a process. It is the process, and organisations that treat the two as separate concerns do both badly.

What a job analysis actually takes

Everything above depends on knowing what the role requires, and most organisations design an assessment from the job advertisement, which was written from the last job advertisement.

The fix is an afternoon. Talk to two or three people who do the job well, separately. Ask what they actually did last week, in order, with rough proportions of time. Ask what goes wrong when somebody is struggling in the role. Then ask for two specific stories: a time somebody handled a situation notably well, and a time somebody handled one badly.

Those critical incidents are the raw material for your structured interview questions, and they are automatically relevant in a way that generic competency questions are not.

Then apply one test to every requirement on the resulting list: could a competent person learn this in the first month? If yes, it is an onboarding item rather than a selection criterion. Most requirement lists shrink by half under that question, and the shrinkage widens the candidate pool without lowering the bar.

The step almost nobody takes

Record assessment scores. Revisit them when you know how the hire worked out.

Small numbers will not prove anything statistically, and they will show you whether your top scorers are conspicuously failing or whether the people you scored poorly and hired anyway turned out fine. If there is no relationship at all between your scores and later performance, that is worth knowing before running the same process for another year.

The same applies to fairness. Pass rates by group, measured stage by stage rather than only at the end, because a process can look balanced overall while one stage removes a group sharply and a later stage compensates. The obstacle here is not the data or the arithmetic: it is the willingness to look.

Calibration, which costs an hour

Take three real responses — a strong one, a weak one, and one that divided opinion. Have your assessors score them independently, without discussion. Then compare.

The disagreement will be wider than anybody expects, particularly on the borderline case. That is the finding, and it means your criteria are too vague to apply rather than that one assessor is wrong.

Without this, you are not running one assessment. You are running as many assessments as you have assessors, and which one a candidate gets is a matter of scheduling.

What assessment cannot do

It cannot fix a role nobody wants, a manager people leave, or a job nobody can describe. It cannot predict retention, which depends far more on the first ninety days than on anything detectable at selection. It cannot remove judgement — structure constrains and informs it rather than replacing it. And it cannot reach people who never applied, which is why a high rejection rate should prompt a look at sourcing before a look at standards.

For each thing you hope a better process will fix, ask what would change if the assessment were perfect. Where the answer requires a different salary, a different manager or a clearer role, the assessment was never the constraint.

Where to start

Designing from scratch: the job analysis, then which method predicts what.

Your process is too long: how long an assessment should take, then the audit.

Hires are not working out: what assessment cannot do, then the audit.

Worried about fairness: adverse impact, then experience proxies.

Assessors disagree: training assessors, then calibration.

Candidates complain: candidate communication, then take-home tasks.

Foundations

6 notes

Assessment narrows uncertainty about capability. It does not remove it, and processes designed as though it does fail predictably.

  1. What an Assessment Is Actually For
  2. Prediction, and What the Evidence Supports
  3. The Gap Between What Is Measured and What Matters
  4. Job Analysis: Knowing What You Are Selecting For
  5. Reliability and Validity, Without the Statistics
  6. The Question to Ask Before Designing Anything

Methods

8 notes

The ordering between methods is more reliable than the numbers attached to them, and it is stable across analyses.

  1. Work Samples
  2. Structured Interviews
  3. Unstructured Interviews, and Why They Persist
  4. Cognitive Ability Tests
  5. Personality Questionnaires
  6. Situational Judgement Tests
  7. Portfolios and Past Work
  8. Which Method Predicts What

Each requirement assigned to exactly one method that can assess it. Anything unassigned comes out.

  1. Starting From the Job, Not From the Catalogue
  2. Designing a Work Sample That Fits
  3. Writing Structured Interview Questions
  4. Scoring Rubrics That Interviewers Can Apply
  5. How Long an Assessment Should Take
  6. Take-Home Tasks and Unpaid Labour
  7. Sequencing: What to Assess and When

Fairness

7 notes

Every additional hurdle selects on slack as well as on capability, and the people it removes have less of the first rather than less of the second.

  1. Adverse Impact, Measured
  2. Accessibility and Adjustments
  3. Language, and What You Are Actually Testing
  4. Experience Proxies and What They Exclude
  5. Assessment Anxiety and Its Effects
  6. Candidates Who Cannot Take Time Off
  7. Auditing Your Own Process

Running it

6 notes

Calibration is the highest-return hour available to most organisations, and almost nobody spends it.

  1. Training Assessors
  2. Calibration Between Assessors
  3. Recording Decisions
  4. Candidate Communication
  5. Giving Feedback
  6. When an Assessment Goes Wrong

Internal candidates frequently receive less assessment than external ones, which is the wrong way round.

  1. Internal Promotion and Its Different Problems
  2. Assessing Existing Staff
  3. Apprenticeships and Entry-Level Routes
  4. Certification and Licensing
  5. Skills Gaps and What Assessment Reveals
  6. Reassessment Over Time

Reference

5 notes

The end state as a checklist, twelve common failures, and where to start.

  1. What a Working Process Looks Like
  2. What Assessment Cannot Do
  3. Common Failures, Listed
  4. Costs, Honestly
  5. Glossary and Where to Start

Tool guides

3 comparisons

Practical comparisons for teams running assessment and hiring workflows.

Independent guidance on skills assessment, selection design and fair hiring practice. External tools are included for practical comparison; evidence from the job remains the basis for decisions.