How to assess · For hiring teams

How to Assess A/B Testing Skills When Hiring

The test formats that actually work for A/B Testing, what a strong answer looks like, sample questions and a scoring rubric you can use as-is.

The short answer

Assess A/B Testing with a task, not a conversation: design a test from a proposal, interpret a flawed result, ai-scored assessment (e.g. cohesyve) or programme conversation. Score it against written criteria you fix before you see any submissions, and weight the criteria that the role actually depends on.

  • Defines the hypothesis, primary metric and minimum detectable effect before the test starts
  • Calculates the sample size and duration and does not stop early on a green result
  • Checks randomisation actually worked: sample ratio, pre-period balance
  • Chooses guardrail metrics and knows what would make them stop a winning test

Paste a job description; Cohesyve generates a role-specific assessment and rubric. Ten candidates free, no card.

A/B testing looks simple — split traffic, compare metrics — and most organisations do it badly. Tests stopped the moment they turn green, metrics chosen after the fact, variants that were never actually randomised, and a winner declared from a sample that could not have detected the effect. The skill is not running tests; it is designing them so the answer means something and interpreting them without fooling yourself. This page covers how to assess A/B testing for product analyst, growth and experimentation roles: design, power, randomisation, interpretation and the discipline to report honestly.

Why A/B Testing is worth testing

Bad experimentation is worse than no experimentation, because it produces confident wrong answers. A team that peeks and stops early ships changes that do nothing; a team that tests ten metrics and reports the one that moved rewrites the product on noise. Testing shows whether a candidate understands power, peeking and selection, and that determines whether the experimentation programme creates knowledge or theatre.

What strong A/B Testing looks like

  • Defines the hypothesis, primary metric and minimum detectable effect before the test starts
  • Calculates the sample size and duration and does not stop early on a green result
  • Checks randomisation actually worked: sample ratio, pre-period balance
  • Chooses guardrail metrics and knows what would make them stop a winning test
  • Interprets results with uncertainty and does not cherry-pick secondary metrics
  • Understands common pitfalls: novelty effects, interference, seasonality, multiple variants
  • Reports null results as useful and keeps a record of what was learned

Ways to assess A/B Testing

Design a test from a proposal

Describe a product change and the current metrics. Ask for the hypothesis, primary and guardrail metrics, sample size reasoning, duration, and what could invalidate the result. Forty-five minutes.

Pros

Tests design discipline and pitfall awareness.

Cons

Design-based; pair with interpretation.

Best for Any experimentation role.

Interpret a flawed result

Provide a test readout with a sample ratio mismatch, a test stopped at day three, and a secondary metric presented as the win. Ask what can be concluded.

Pros

The real failure modes; scoreable.

Cons

Needs a crafted readout.

Best for Mid and senior roles.

AI-scored assessment (e.g. Cohesyve)

Generate an experimentation scenario from the job description — a design, a readout critique, a programme question — with a rubric. Each candidate receives a different variant; reasoning is scored in writing.

Pros

Asynchronous and consistent; unique per candidate; experimentation judgement is written reasoning.

Cons

No computation needed for most roles.

Best for Screening a pool.

Programme conversation

Ask how they would set up experimentation for a team that has never done it: tooling, process, culture.

Pros

Reveals maturity and leadership.

Cons

Talk-based.

Best for Senior and lead roles.

Cohesyve

Run a A/B Testing assessment on your next opening

Cohesyve generates a unique A/B Testing task per candidate from your job description, with the scoring rubric attached. Questions are different for every applicant, so they cannot be shared or looked up.

What to test

Design

Whether the test can answer the question.

Choose a primary metric and justify itEstimate sample size for a given minimum detectable effectDecide the duration and what to do about weekly seasonality

Validity

Whether the result can be trusted.

Diagnose a sample ratio mismatchExplain the effect of stopping earlyIdentify interference between variants

Interpretation

Whether conclusions match the evidence.

Read a result with a confidence interval spanning zeroDecide what to do with a winning primary and a losing guardrailHandle a secondary metric that moved when the primary did not

Programme

Whether experimentation creates knowledge.

Write up a null result usefullyPrioritise a backlog of test ideasDesign a review process that prevents cherry-picking

Sample A/B Testing questions

Why should you not stop a test the moment it shows significance?

Entry

Look for Peeking inflates false positives; fixed horizon or sequential methods.

What is a sample ratio mismatch, and what does it tell you?

Mid

Look for Split deviates from the intended ratio; randomisation or logging is broken; results are not trustworthy.

The primary metric is flat but a secondary metric is up 12%. What do you conclude?

Mid

Look for Likely noise from multiple comparisons; treat as a hypothesis for a new test; do not ship on it.

How do you choose the minimum detectable effect?

Mid

Look for The smallest change worth acting on, given cost; drives sample size; a trade-off with duration.

Two teams run tests on the same page simultaneously. What is the risk and how do you manage it?

Senior

Look for Interaction effects; mutual exclusion, layered experiments, or accept and test for interaction.

Red flags

  • Stops tests when they turn green
  • Chooses metrics after seeing results
  • Has never checked a sample ratio
  • Cannot explain power
  • Reports only winning tests

Scoring rubric

CriterionWeightWhat strong looks like
Design discipline30%Hypothesis, metric, power and duration set up front.
Validity checks25%Randomisation and stopping rules are verified.
Interpretation25%Conclusions respect uncertainty and avoid cherry-picking.
Programme thinking20%Experimentation produces durable learning.

Mistakes hiring teams make

  • Testing tool knowledge instead of design judgement
  • Not including a peeking or cherry-picking scenario
  • Accepting a design with no power reasoning
  • Ignoring guardrails
  • Assuming analytics experience includes experimentation discipline

Roles that need A/B Testing

Product AnalystGrowth AnalystData ScientistExperimentation ManagerProduct ManagerMarketing Analyst

Common questions

Do I need an experimentation platform to assess A/B testing?

No. Design and interpretation are the skills, and both can be tested on paper with a proposal and a readout.

What is the best single A/B testing question?

Show a flat primary metric and a big secondary win and ask what to do. Multiple-comparison awareness and discipline show immediately.

Should product managers be assessed on A/B testing?

A light version, yes: what makes a result trustworthy, and why stopping early is a problem. They make the decisions that experiments inform.

How long should an A/B testing assessment take?

Forty-five minutes for a design exercise; thirty for a readout critique.

Cohesyve · Skill assessments for hiring

Test A/B Testing before the first interview

Generate a role-specific A/B Testing assessment from your job description and see who can do the work before you spend interview time on them.

1,500+

assessments completed

50%

faster time-to-hire

90%

completion rate

5 min

from JD to assessment

No credit card · 10 free candidates · Plans sized to your hiring volume

See Cohesyve in action

Free 30-min walkthrough

See it on your role