Interview questions

Data Scientist Interview Questions

Hire top Data Scientists with these 25 interview questions covering ML, statistics, Python, and problem-solving with expert evaluation guidance.

What's here

10 Data Scientist interview questions, grouped into 3 areas: machine learning questions, statistics & experimentation questions and applied problem-solving questions. Each one comes with why it is worth asking and what a strong answer contains, so the same question can be scored the same way by different interviewers.

  • Include a take-home or live coding exercise focused on a realistic data science problem. Abstract questions alone don't predict job performance.
  • Ask candidates to explain their approach at multiple levels of detail. The ability to zoom in and out demonstrates true understanding.
  • Evaluate communication skills carefully. Data Scientists who can't explain their work to non-technical stakeholders have limited impact.

Hiring a Data Scientist requires evaluating deep technical skills in machine learning and statistics alongside the ability to solve real business problems. The best candidates combine mathematical rigor with practical engineering skills and clear communication. Use these questions to assess candidates across the full data science skill set.

Machine Learning Questions

These questions test depth of ML knowledge and practical experience.

  1. 1

    Explain the bias-variance tradeoff. How does it influence your model selection process?

    Why ask it

    Fundamental ML concept that reveals depth of understanding.

    What to look for

    Clear explanation of overfitting vs. underfitting, how model complexity affects each, and practical strategies (cross-validation, regularization, ensemble methods) to manage the tradeoff.
  2. 2

    Walk me through how you would build a recommendation system from scratch. What approaches would you consider?

    Why ask it

    Tests breadth of ML knowledge and system design thinking.

    What to look for

    Should discuss collaborative filtering, content-based filtering, and hybrid approaches. Bonus for mentioning cold start problems, evaluation metrics (precision, recall, nDCG), and scalability considerations.
  3. 3

    How do you handle class imbalance in a classification problem?

    Why ask it

    Tests practical ML problem-solving for common real-world challenges.

    What to look for

    Multiple strategies: oversampling (SMOTE), undersampling, class weights, ensemble methods, appropriate metrics (F1, AUC-ROC over accuracy), and understanding of when each approach is appropriate.
  4. 4

    Explain how gradient descent works. What variants exist and when would you use each?

    Why ask it

    Tests understanding of optimization fundamentals.

    What to look for

    Clear explanation of the algorithm, understanding of learning rate importance, knowledge of SGD, mini-batch, Adam, and their tradeoffs. Bonus for discussing learning rate scheduling and convergence issues.

Cohesyve

Screen Data Scientist candidates before you ask any of these

Cohesyve builds a Data Scientist assessment from your job description so interview time goes to people who have already shown they can do the work.

Statistics & Experimentation Questions

These questions evaluate statistical rigor and experimental design skills.

  1. 5

    Design an A/B test to evaluate whether a new checkout flow increases conversion rates. Walk me through every step.

    Why ask it

    Tests end-to-end experimentation skills.

    What to look for

    Hypothesis formulation, sample size calculation, randomization, metric selection (primary and guardrail), test duration, statistical significance testing, and practical considerations (novelty effect, seasonality).
  2. 6

    What is the difference between Type I and Type II errors? Give a real-world example where each would be particularly costly.

    Why ask it

    Tests statistical literacy and practical understanding.

    What to look for

    Clear definitions with real examples. Should understand the relationship to significance level and power, and how to adjust based on the cost of each error type.
  3. 7

    When would you use Bayesian methods instead of frequentist methods?

    Why ask it

    Tests breadth of statistical knowledge.

    What to look for

    Understanding of prior information, small sample sizes, sequential testing, and interpretability advantages. Should be able to discuss practical applications rather than just theory.

Applied Problem-Solving Questions

These questions assess how candidates approach real-world data science challenges.

  1. 8

    You built a model with 95% accuracy on test data, but it performs poorly in production. What could be going wrong?

    Why ask it

    Tests ability to diagnose real-world ML problems.

    What to look for

    Should identify: data drift, train/test distribution mismatch, data leakage, feature pipeline differences, and feedback loops. Demonstrates experience shipping models to production.
  2. 9

    How do you decide whether a business problem needs ML or if simpler analytics would suffice?

    Why ask it

    Tests pragmatism and business judgment.

    What to look for

    Preference for simplicity when possible, understanding of when ML adds value (complex patterns, scale, automation), and ability to articulate ROI of ML investment.
  3. 10

    Tell me about a model you deployed that had a measurable business impact. How did you measure the impact?

    Why ask it

    Evaluates end-to-end data science impact.

    What to look for

    Should describe the business problem, model choice, deployment process, and quantified business outcome. Look for awareness of online vs. offline metrics and attribution challenges.

Running the interview well

  • 1Include a take-home or live coding exercise focused on a realistic data science problem. Abstract questions alone don't predict job performance.
  • 2Ask candidates to explain their approach at multiple levels of detail. The ability to zoom in and out demonstrates true understanding.
  • 3Evaluate communication skills carefully. Data Scientists who can't explain their work to non-technical stakeholders have limited impact.
  • 4Look for intellectual honesty. The best Data Scientists know the limits of their models and are transparent about uncertainty.
  • 5Assess engineering maturity. Can they write production-quality code, not just notebook experiments?

Common questions

How long should a Data Science interview process take?

A typical Data Science interview process includes 3-5 rounds: phone screen, technical screen (coding/SQL), take-home or live ML exercise, system design, and behavioral/culture fit. The entire process usually takes 2-4 weeks. Avoid excessive rounds that lose top candidates.

Should I test for coding skills or statistical skills?

Both are essential. A Data Scientist who can't code can't ship models. One who doesn't understand statistics will build unreliable models. Balance your interviews to cover Python/SQL proficiency, statistical reasoning, and ML knowledge.

How do I evaluate a Data Scientist candidate without deep technical knowledge myself?

Use structured scoring rubrics, involve technical team members in the interview, focus on communication clarity, and consider using standardized assessment tools like Cohesyve to objectively evaluate technical competencies.

Cohesyve · Skill assessments for hiring

Assess Data Scientist candidates before you interview them

Cohesyve turns a job description into a role-specific assessment with a scoring rubric. Each candidate gets a different version, so questions cannot be shared between applicants.

1,500+

assessments completed

50%

faster time-to-hire

90%

completion rate

5 min

from JD to assessment

No credit card · 10 free candidates · Plans sized to your hiring volume

For candidates

Preparing for a Data Scientist role yourself? Practise on the same AI job simulations companies use — 5 free assessments a month, no card required.

See Cohesyve in action

Free 30-min walkthrough

See it on your role