Every hire you make is a bet on future performance. The core question for hiring managers is, how can you be more confident that the person you hire today will be a strong performer tomorrow?
Relying on resumes and gut feelings from unstructured interviews has been shown to be a weak predictor of who will succeed. Research indicates these traditional methods often fall short.
So, how do you move from guesswork to a more evidence-based approach?
The Core Idea: Connecting the Test to the Job
Think like a professional sports scout. A scout doesn’t just clock a player's 40-yard dash time in an empty gym. Those are just metrics. What the scout really cares about is how that athlete performs under pressure, in a real game.
That’s the essence of criterion-related validity. It addresses a fundamental question for any hiring tool: Does this test actually tell me who will be good at the job?
Criterion-related validity is the statistical evidence that your hiring assessment is meaningfully linked to real-world job success. It helps shift hiring from a game of chance toward a process of prediction.
In this framework, it's all about connecting two key pieces:
- The predictor: This is your assessment—the skills test, work sample, or structured interview you use to evaluate candidates.
- The criterion: This is the outcome you want to predict, the "in-game performance" that matters. Usually, this means job performance, but it could also be employee retention, team collaboration, or sales numbers.

A Principle for Fair and Effective Hiring
This isn't a fleeting trend. For decades, the U.S. government has supported this principle through the Uniform Guidelines on Employee Selection Procedures, which outline standards for creating fair, defensible, and merit-based hiring processes. The goal is to ensure you’re evaluating candidates on their ability to do the job, not on factors that have little bearing on performance.
This is why assessment platforms that focus on creating role-specific challenges, such as Cohesyve, can be effective. They are often built to strengthen the link between the assessment (the predictor) and on-the-job success (the criterion).
By understanding and prioritizing criterion-related validity, you're better equipped to build a team that can consistently meet performance goals. It's a key part of making predictions you can act on.
A Key Standard for Predictive Hiring
Let's say a doctor suggests a new diagnostic test. You wouldn't be impressed by its technology alone. You’d ask a simple question: how well does it actually predict my health outcome? That same logic is at the heart of criterion-related validity—an important concept for showing that a hiring method works.
In hiring, your assessment is the "predictor," and the outcome you care about—great job performance—is the "criterion." Criterion-related validity is the statistical evidence that connects them. It's what takes hiring out of the realm of gut feelings and helps anchor it in objective, measurable evidence.
A Journey from Battlefields to Boardrooms
The idea was developed over a century ago out of necessity. During World War I, psychologists faced a significant task: figure out, and fast, which recruits had the aptitude to be successful pilots.
They developed some of the first cognitive and psychomotor tests (the predictors) and then tracked who actually succeeded in training and combat (the criterion). By finding the correlation between test scores and real-world pilot performance, they discovered which tests could reliably spot a future ace. This was criterion-related validity in a practical application.
By focusing on the statistical link between an assessment and a real-world outcome, you can transform your hiring process. It stops being a search for a "good-looking resume" and becomes a focused mission to find a "high-performing employee."
That principle became a cornerstone of industrial-organizational psychology. It’s the scientific method that allows us to say with more confidence that a specific hiring tool can give you an edge in finding talent.
The Proven Power of Predictive Assessments
The proof is in the data. In talent acquisition, criterion-related validity is the benchmark for showing that an assessment predicts on-the-job success. A landmark meta-analysis covering 85 years of research revealed that general mental ability tests have an average validity coefficient of 0.51 for predicting job performance. For a closer look at how these work, our guide on cognitive ability tests for hiring is a helpful resource.
What’s even more telling? Structured interviews can reach a coefficient of 0.58. These figures mean that well-designed hiring tools can account for a significant portion of the difference between an average and a great employee.
Why This Matters for Modern Hiring
In today's market, hiring mistakes can be costly. A bad hire can affect more than just a salary—it can drain team morale, hurt productivity, and impact the bottom line. Understanding how to use predictive analytics in HR provides a framework for gaining insight into a candidate's potential performance.
When you build your hiring process on criterion-related validity, you're giving your team a clearer, more defensible, and effective way to build a workforce. It’s about making smarter decisions on people by using evidence to see the future more clearly.
Predictive Versus Concurrent Validity
When we discuss criterion-related validity, the conversation often comes down to timing. The goal is to show your assessment can predict who will succeed on the job. But when you collect that job performance data is what splits this concept into two distinct approaches: predictive and concurrent validity.
Imagine you need to hire salespeople who can consistently meet their quotas. You have an assessment you believe can spot top talent, but now you have to demonstrate its effectiveness. This is where you face a choice.
The Crystal Ball: Predictive Validity
Predictive validity is often considered the gold standard. It’s about playing the long game to show that your assessment can forecast future success. It requires patience, but the results can be powerful.
Here’s the plan:
- Test Your Candidates: You give your sales assessment to a new group of applicants.
- Hire "Blind": This is a key step. You hire people from that applicant pool using your existing methods, without looking at their new assessment scores. This is important to prevent the scores from influencing who you hire, which would skew the results.
- Wait and Watch: Now, you wait. After a set period—perhaps six months or a year—you gather performance data for those new hires. This is your "criterion," and for a sales team, it would be things like revenue numbers, client retention rates, or how often they hit their targets.
- Connect the Dots: Finally, you analyze those initial assessment scores and compare them to the on-the-job performance data you collected.
If the people who scored highest on your assessment are now your top sellers, you have strong evidence of predictive validity—your test successfully predicted who would thrive.
The Snapshot: Concurrent Validity
Realistically, sometimes you can’t wait six months to see if a test works. You need answers sooner. This is where concurrent validity comes in, giving you a quick, in-the-moment check.
Here’s how you'd use it with that same sales team:
- Test Your Current Team: You administer the same assessment to your existing employees—your top performers, your solid middle-of-the-pack folks, and your struggling reps.
- Grab Current Performance Data: At the same time, you pull their most recent performance metrics. Since they're already on the job, this data is readily available—think last quarter's sales numbers or their latest performance review scores.
- Find the Pattern: You then analyze the data to see if there’s a clear connection between their assessment scores and their job performance.
If your top-performing veterans also score highest on the assessment, you’ve got a strong signal. This is good evidence of concurrent validity, showing that your test accurately reflects the skills that separate your best employees from the rest, right now.
The Real Difference: Predictive validity answers the question, "Can this test tell me who will become a great hire?" Concurrent validity answers, "Does this test distinguish my best employees from the rest today?"
Which Path Is Right for You?
So, which approach should you take? Neither is inherently "better." The right choice depends on your company's situation—your timeline, your resources, and what you’re trying to achieve. While we're on the topic of different validity types, you can see how this compares by reviewing some examples of content validity.
Go for Predictive Validity if:
- You're building a highly defensible, long-term hiring strategy for a core role.
- You have the time and buy-in to conduct a formal study over several months.
- Your goal is the highest level of proof that your assessment works for new candidates.
Go for Concurrent Validity if:
- You need to move fast and get insights quickly, like in a startup or high-growth environment.
- You want to pilot a new assessment and get a "good enough" signal before rolling it out widely.
- You have a clear distinction between high and low performers on your current team to use as a benchmark.
Ultimately, both paths lead away from guesswork and toward a hiring process built on evidence. Whether you’re looking into the future or just need a clear picture of the present, knowing how to use these two tools can help you build a team that delivers.
Cohesyve
See what candidates can do before you interview them
Cohesyve turns a job description into a role-specific assessment with a scoring rubric. Each candidate gets a different version, so questions cannot be shared. Ten candidates free, no card.
How to Measure and Interpret Validity
Alright, we've covered the "what" and "why" of criterion-related validity. Now for the "how." How do we actually figure out if an assessment truly predicts who will succeed in a role? It's less about complex math and more about finding a single, powerful number that tells the story.
We call this number the validity coefficient. Think of it as a confidence score for your hiring process. It's the data that shows just how well your assessments are predicting future top performers.
Understanding the Validity Coefficient
A common way to get this number is with a statistical tool called Pearson's correlation coefficient (often shown as an r). This simply measures how two things are related—in our case, scores on an assessment and actual job performance.
The result is always a number between -1.0 and +1.0.
- A score of +1.0 is a perfect positive match. Imagine every single person who aced the test also became a top performer. This is a statistical ideal you'll likely never see in the real world.
- A score of -1.0 means you have a perfect inverse relationship. High test scores are linked to poor job performance. If you ever see this, your assessment isn't just ineffective—it's actively working against you.
- A score of 0 means there’s absolutely no connection. Your test is no better at predicting success than flipping a coin.
It's simple: the closer your coefficient is to +1.0, the more predictive power your assessment holds.
This chart breaks down how you collect performance data for the two main types of validity studies—one looking forward (predictive) and one looking at the present (concurrent).

While predictive studies give you more robust, long-term evidence, a concurrent study is a great way to get a quick signal on whether your assessment is on the right track by using your current team.
What Does a Good Score Look Like in Hiring?
So, what's a realistic number to aim for? We're dealing with people, not machines, so perfection isn't the goal. Human performance is too nuanced for that. But decades of research have given us some clear benchmarks.
According to standards in industrial-organizational psychology, here’s a practical guide for what those numbers mean for hiring:
- Below .15: Generally not considered useful for making hiring decisions.
- .15 to .34: Useful. This shows a meaningful relationship that can improve your selection process.
- .35 and above: Very useful. You've found a strong predictor of job performance.
Of course, to measure this accurately, you need a consistent way to define "good performance." Exploring different performance evaluation examples and templates can help you standardize how you collect that all-important criterion data.
Why Validity Is the Engine Behind Smart Hiring
This isn't just an academic exercise; it's the engine driving more successful companies. The evidence is there. In a 2003 review, the U.S. Office of Personnel Management looked at 300 federal hiring tools and found their validity coefficients averaged 0.36 for predicting performance—a very useful score.
A 2021 Gartner report studying 200 large companies found that top-tier assessment platforms were achieving coefficients as high as r=0.65 to r=0.75. The impact of these high-validity tools included explaining a large portion of the variance in employee performance.
The lesson is clear. Higher validity is directly linked to better hires, less turnover, and a healthier bottom line. This is precisely why investing in high-quality pre-employment skills testing is so critical.
When an assessment is designed specifically for the role—measuring the skills that actually matter on the job—it will naturally have a much stronger connection to performance. Proving your test works isn't just about compliance; it's about making a data-driven investment in the future of your team.
Common Roadblocks to Accurate Validation
So, you’ve decided to put your hiring process to the test. That's a huge step forward. You’re ready to move from gut feelings to hard evidence, proving that your assessments actually predict who will succeed on the job.
But even the best-laid plans can hit a few snags. Think of a validation study like a sensitive scientific experiment—if the conditions aren't right, the results can get skewed. Knowing what to watch out for ahead of time is the best way to protect your results and ensure you’re making fair, accurate decisions.
Let's walk through the two biggest culprits that can throw your validation efforts off track.
Criterion Contamination
Picture this: a manager is getting ready to write a six-month performance review for a new hire. Out of curiosity, she pulls up the employee's original pre-hire assessment report and sees he scored in the 98th percentile.
Whether she realizes it or not, that little piece of information is almost guaranteed to color her review.
We call this criterion contamination, and it’s a classic case of the data "cheating." It’s what happens when the person rating job performance (the criterion) already knows the person's test score (the predictor).
This knowledge can easily create a self-fulfilling prophecy. A manager who sees a stellar test score might subconsciously look for confirming evidence of success, while downplaying mistakes. On the flip side, knowing an employee barely scraped by on their assessment could lead that same manager to magnify every small error. The evaluation is no longer objective.
The Fix: The solution here is simple in theory but requires real discipline in practice: keep the predictor scores and performance reviews completely separate. The managers rating performance should be "blind" to the original assessment results. This is the only way to ensure their feedback is based on what they actually see on the job, not on what a test score told them to expect.
Range Restriction
The other major roadblock is a statistical gremlin called range restriction. This one is a little trickier. It pops up because you only have performance data for the people you hired.
Think about it. Your total applicant pool likely had a massive range of assessment scores, from top-tier to bottom-of-the-barrel. But you only hired the people who scored well, right? So when you try to connect their test scores to their job performance, you’re only looking at a small, high-scoring slice of the original group.
This inevitably shrinks the correlation between scores and performance, making your assessment appear less effective than it really is.
- The Problem: You have no idea how the lower-scoring candidates would have performed. If you had hired them and they performed poorly, it would have created a much stronger, more obvious link between test scores and job success.
- The Impact: Because you’re only analyzing a narrow "range" of top talent, the statistical connection looks weaker.
This is a very common issue, especially with predictive studies where the whole point is to hire the best. The good news is that I-O psychologists have developed statistical corrections for range restriction. These formulas help you estimate the "true" validity coefficient, as if you had performance data from your entire original applicant pool.
Side-Stepping Risk with Better Systems
Both of these challenges—contamination and restriction—point to a bigger truth: human bias and messy data are the enemies of good validation. This is exactly where more objective, role-specific assessment systems, like those from Cohesyve, can give you a leg up.
When you use a platform that gives every candidate unique, real-world tasks to solve, you’re already building a more objective foundation. You're measuring concrete skills—how someone actually navigates a tricky client call or debugs a piece of code—instead of relying on abstract questions or a manager’s subjective opinion.
This creates much cleaner performance data from the very start, which helps neutralize the risk of contamination. And while range restriction will always be a statistical factor to account for, starting with more reliable and objective data makes your entire validation process stronger, clearer, and far more defensible.
How to Increase Validity with Role-Specific Assessments
We've reviewed the theory and the stats. Now for the practical question: how do you actually increase your assessment’s criterion-related validity?
A powerful method is grounded in a simple truth: make the test look like the job.
The closer your hiring assessment is to the real work someone will be doing, the stronger the connection to performance will likely be. This is where we're seeing an industry-wide shift away from generic brain teasers and toward highly relevant challenges that reflect the day-to-day realities of a role.
From Puzzles to Problems
For decades, hiring sometimes felt like a game of abstract puzzles. Candidates were asked to solve logic riddles or answer personality quizzes, with the hope of finding a signal about their intelligence or temperament.
But does acing a pattern-recognition puzzle tell you if a Senior Software Engineer can hunt down a bug in a legacy codebase? Not directly. That disconnect is why those methods can produce weaker validity scores.
That’s why many talent teams now focus on content validation—a term for ensuring your assessment is built from the same material the job is.

Instead of a generic puzzle, a better predictor for an engineering role might be a realistic coding challenge. For a marketing manager, it could be a case study asking them to allocate a launch budget. For a sales rep, it might be a simulated role-play where they have to navigate a tough client objection.
When an assessment feels like a "day in the life," you're not just guessing at abstract skills. You are watching a candidate perform the kinds of tasks you’ll be paying them to do. That direct link is what drives validity coefficients up.
The Power of Job-Specific Simulation
This is where modern assessment platforms like Cohesyve can be useful. By analyzing a job description, these systems can generate unique challenges tailored to the specific demands of the position, moving far beyond static question banks.
They create realistic simulations that give you a true-to-life preview of a candidate's abilities.
- For Technical Roles: Imagine a candidate placed in a real-world coding environment, tackling problems that scale in difficulty based on their skill. You see where their expertise shines—and where it stops.
- For Client-Facing Roles: You can put candidates into interactive scenarios that simulate tough customer calls or partner negotiations, giving you a front-row seat to their communication and problem-solving skills.
- For Strategic Roles: A candidate might get a business brief and be asked to outline their 90-day plan, revealing not just what they know, but how they think.
By grounding your assessments in the daily reality of the job, you create a stronger signal. The link between how a candidate performs in the simulation and how they'll perform on the job becomes clearer and more reliable.
This is a direct path to boosting the criterion-related validity of your hiring process. It lets you stop guessing and start measuring what actually matters—giving you the confidence to hire people who have shown they can hit the ground running.
Putting Validity into Practice: Your Questions Answered
We’ve covered the theory behind validity, but what does this look like in the thick of hiring? It's one thing to understand the concepts; it's another to apply them confidently.
Here are answers to common questions from hiring managers and recruiters who are ready to make criterion-related validity a practical advantage.
What Is a Good Validity Coefficient in Hiring?
This is a frequent question. Everyone wants to know, "What's a good score?" While a perfect +1.0 correlation is a statistical ideal you'll never find in the real world, decades of research have given us some solid goalposts.
In the world of hiring, a validity coefficient of 0.35 or higher is considered very useful. It signals that your assessment is a strong predictor of who will succeed in the role.
Don't discount smaller results, though. Coefficients between 0.15 and 0.34 are still useful and can bring real, measurable value to your hiring process. Anything below 0.15, however, is likely too weak to base important decisions on.
How Long Does a Predictive Validity Study Take?
There’s no getting around it: a proper predictive study requires patience. You’re playing the long game. Typically, you should plan for a study to last anywhere from 6 to 12 months.
Why so long? You need to give your new hires enough time to truly settle in, learn the ropes, and show what they're capable of. That's when you get reliable performance data (the "criterion") to compare against their initial test scores. Rushing this step is the fastest way to get misleading results.
Can You Validate Assessments for Soft Skills?
Absolutely. This is a common point of confusion. It might feel difficult to quantify skills like leadership or communication, but it's entirely possible to establish strong criterion validity for them.
The key is to define a clear, observable criterion that reflects what that soft skill looks like in action.
For instance, you could use a role-play simulation to gauge a candidate's negotiation skills (that’s your predictor). Six months later, you could check their actual client retention rates, customer satisfaction scores, or even 360-degree feedback from their team (that's your criterion). You then see how well the first predicted the second.
Why Is This So Important for Fairness and Legal Compliance?
Think of criterion-related validity as the foundation of a hiring process that’s not just effective, but also fair and defensible. It gives you the evidence to show you’re hiring people based on their genuine ability to do the job—not on a whim, a gut feeling, or factors that introduce bias.
This commitment to merit-based hiring aligns well with guidelines from bodies like the EEOC. When you can demonstrate a clear, statistical link between your assessment and on-the-job performance, you build a process that is more resistant to legal challenges. It’s why well-designed, role-specific assessments are so valuable; they create that provable connection from day one.
Ready to build a hiring process with proven predictive power? Cohesyve uses AI to create dynamic, role-specific assessments that measure what truly matters. See how you can identify top performers with confidence.
