# Data Scientist interview questions and answers

Source: https://digitalcvmaker.com/interview-questions/data-scientists
Last updated: 2026-10-01

Data scientist interviews go beyond reporting. They test whether you can frame a business problem as a modelling task, choose and validate the right algorithm, avoid leakage and overfitting, run sound experiments, and get a model working in production. Expect statistics, machine learning theory, Python coding and a case study. Practise the 20 questions below out loud, compare with the sample answers and prepare for the follow-ups.

Data scientist hiring usually runs as an online Python, SQL and statistics test or take-home modelling task, one or two technical rounds on machine learning and your projects, a business case round, and a final managerial and HR discussion.

## Role and technical questions

### 1. Explain the bias-variance trade-off with an example from a model you built.

Round: Role · Level: Fresher

**What they’re checking:** Whether you understand underfitting and overfitting at a conceptual level and can connect it to choices like model complexity and regularisation.

**Sample answer:**

> Bias is error from a model being too simple to capture the real pattern, so it underfits and does badly on both training and test data. Variance is error from a model being too sensitive to the training data, so it fits noise and does well on training but poorly on new data. In my college project predicting house prices, a linear regression on three features had high bias, with similar but poor error on train and test. A deep decision tree had almost zero training error but much worse test error, which is high variance. A random forest with limited depth balanced the two and gave the best validation error.

**Likely follow-ups:**

- How do learning curves help diagnose this?
- Does adding more data reduce bias or variance?

### 2. You are building a fraud detection model where only 0.5 percent of transactions are fraud. How do you handle the imbalance?

Round: Role · Level: All levels

**What they’re checking:** Whether you know that accuracy is misleading on imbalanced data and can choose suitable metrics, sampling and threshold strategies.

**Sample answer:**

> First, I would not use accuracy, because predicting no fraud for everything already gives 99.5 percent. I would use precision, recall and the area under the precision-recall curve. I keep the class ratio in train and test splits with stratification, or split by time for fraud. To help the model learn, I try class weights in the loss, or resample only the training data, such as undersampling the majority or SMOTE, never the test set. Tree-based models like gradient boosting often handle imbalance well with weights. Finally I tune the decision threshold with the fraud team, based on the cost of a missed fraud versus a blocked genuine customer.

**Likely follow-ups:**

- Why should resampling never touch the validation set?
- How would you explain the chosen threshold to the business?

### 3. For a loan default model, would you optimise for precision or recall?

Round: Role · Level: Fresher

**What they’re checking:** Whether you link evaluation metrics to the real business cost of each type of error, rather than quoting definitions only.

**Sample answer:**

> Precision is the share of predicted defaulters who actually default, and recall is the share of actual defaulters the model catches. For loans, it depends on the cost of each error. Missing a defaulter, a false negative, means the lender loses principal, which is usually expensive. Rejecting a good customer, a false positive, means lost interest income and an unhappy customer. In most lending settings recall matters more, but not at any cost, because rejecting too many good borrowers kills growth. So I would agree a cost for each error with the credit team and pick the threshold that minimises total expected cost, then report both metrics.

**Likely follow-ups:**

- What is the F1 score and when is it useful?
- How does ROC AUC differ from precision-recall AUC?

### 4. How do random forests and gradient boosting differ, and when would you choose each?

Round: Role · Level: All levels

**What they’re checking:** Whether you understand how the two main ensemble methods work, not just their library names, and can reason about tuning and overfitting.

**Sample answer:**

> A random forest builds many deep decision trees independently, each on a bootstrap sample of rows and a random subset of features, and averages their predictions. That mainly reduces variance, works well with little tuning and is hard to overfit badly. Gradient boosting builds shallow trees one after another, each fitting the errors, or gradients, left by the previous trees, which mainly reduces bias. Libraries like XGBoost and LightGBM usually give better accuracy on tabular data, but need careful tuning of learning rate, depth and number of trees, with early stopping. I use a random forest as a quick, solid baseline, and boosting when I need the extra performance.

**Likely follow-ups:**

- What does the learning rate control in boosting?
- How do you get feature importance from these models?

### 5. What is data leakage? Give an example and explain how you prevent it.

Round: Role · Level: Experienced

**What they’re checking:** Whether you can spot the subtle mistakes that make models look excellent offline and fail in production, a sign of real project experience.

**Sample answer:**

> Leakage is when the model uses information during training that will not be available at prediction time, so validation scores are unrealistically good. In a churn project I reviewed, a feature called days since last complaint resolution was hugely predictive, but it was updated after customers had already asked to cancel, so it was effectively recording the outcome. Another common case is scaling or imputing using the whole dataset before splitting, which leaks test statistics into training. I prevent it by building features only from data available before a cut-off date, putting all preprocessing inside a pipeline fitted on training folds only, and splitting by time when the data has time order.

**Likely follow-ups:**

- How can target encoding cause leakage?
- What would you check if a model scores suspiciously high?

### 6. What is the difference between L1 and L2 regularisation?

Round: Role · Level: Fresher

**What they’re checking:** Whether you understand how regularisation controls overfitting and why L1 can be used for feature selection.

**Sample answer:**

> Both add a penalty on the size of model coefficients to the loss function, which discourages overly complex models. L1, used in Lasso, adds the sum of absolute values of the coefficients. Its shape pushes some coefficients exactly to zero, so it also performs feature selection, which is useful when there are many weak or irrelevant features. L2, used in Ridge, adds the sum of squared coefficients. It shrinks all coefficients smoothly towards zero but rarely makes them exactly zero, and it handles correlated features better by spreading weight across them. Elastic Net combines both. The strength of the penalty is a hyperparameter I tune with cross-validation.

**Likely follow-ups:**

- Why should you scale features before regularised regression?
- How does dropout relate to regularisation in neural networks?

### 7. How do you decide the sample size and duration of an online experiment before it starts?

Round: Role · Level: All levels

**What they’re checking:** Whether you understand statistical power and the dangers of stopping experiments early, which separate rigorous data scientists from casual testers.

**Sample answer:**

> I need four inputs: the baseline rate of the primary metric, the minimum detectable effect that would matter to the business, the significance level, usually 5 percent, and the power, usually 80 percent. With these I calculate the sample size per group using a power calculation. Smaller effects or noisier metrics need much larger samples. Then I divide by expected daily traffic to get duration, and round up to full weeks to cover weekly patterns. I write this down before launch and do not stop the test early when results look good, because repeated peeking inflates false positives. If early stopping is needed, I use a sequential testing method designed for it.

**Likely follow-ups:**

- What is a sample ratio mismatch?
- How would you handle a metric like revenue with heavy outliers?

### 8. Your model performed well at launch but its accuracy has dropped over six months. What do you do?

Round: Role · Level: Experienced

**What they’re checking:** Whether you understand model monitoring, data drift and concept drift, and have a practical process for retraining and validation.

**Sample answer:**

> First I check for pipeline problems: a changed upstream column, missing values from a broken source, or a feature now computed differently. Then I compare the distributions of input features and predictions now against training, using measures like population stability index. If inputs have shifted, that is data drift, for example a new customer segment after a marketing campaign. If inputs look similar but the relationship to the outcome changed, that is concept drift, such as fraud patterns changing. I retrain on recent data, validate on the latest period, and compare with the current model before replacing it. Going forward I set up monitoring with alerts on drift and on delayed actual performance.

**Likely follow-ups:**

- How do you monitor performance when labels arrive months later?
- How often should a model be retrained?

### 9. Which features would you engineer for a customer churn model at a telecom company?

Round: Role · Level: All levels

**What they’re checking:** Whether you can translate domain knowledge into useful, leak-free features, which usually matters more than the algorithm.

**Sample answer:**

> I would define churn first, for example no recharge for 30 days after a plan expires for prepaid users. Then features from the period before the prediction date: tenure, plan type and price, recharge amount and frequency, and trends such as data usage this month versus the last three months, because declining usage often signals churn. Network experience matters: call drops and slow data sessions in the user’s area. Customer service features like complaints in the last 30 days, and whether a port-out request code was generated, if allowed. I would also include whether the user has a second SIM active, if that data exists. Every feature must be computed only from data before the cut-off.

**Likely follow-ups:**

- How would you choose the prediction window?
- How would the business use the churn scores?

### 10. Why can you not use normal k-fold cross-validation for time series, and what do you use instead?

Round: Role · Level: Experienced

**What they’re checking:** Whether you respect time order in validation and understand how random splits leak future information into training.

**Sample answer:**

> Normal k-fold shuffles data into random folds, so the model can train on future observations and be tested on the past. For time series that leaks information, because patterns like trends and seasonality from the future help predict the past, and validation scores look better than real performance. Instead I use forward-chaining or rolling-origin validation: train on months one to six and test on month seven, then train on one to seven and test on eight, and so on. scikit-learn’s TimeSeriesSplit does this. I also leave a gap between training and test periods when features use lagged values, and keep a final hold-out period that is never touched during tuning.

**Likely follow-ups:**

- What is the difference between expanding and sliding windows?
- How would you validate a demand forecast across many stores?

### 11. How do you explain a complex model’s predictions to business users and regulators?

Round: Role · Level: All levels

**What they’re checking:** Whether you can make models understandable and trustworthy, which is required in lending, insurance and other regulated settings.

**Sample answer:**

> At the global level I show which features matter most and in which direction, using permutation importance or SHAP summary plots, and partial dependence plots for key features. At the individual level, SHAP values show why one customer got a particular score, for example high utilisation of existing credit lines pushed the risk up. For business users I convert these into plain language reason codes, such as the top three reasons for each decision. Where regulators or policy require it, I compare against a simpler model like logistic regression or a scorecard, and sometimes use that if the accuracy loss is small, because explainability can matter more than a small gain.

**Likely follow-ups:**

- What are the limitations of SHAP values?
- How do you check a model for unfair bias?

## Behavioural questions

### 12. Tell me about a model you built that did not create the business value you expected.

Round: Behavioural · Level: Experienced

**What they’re checking:** Whether you measure impact beyond model metrics and understand that adoption and action matter as much as accuracy.

**Sample answer:**

> At an e-commerce company, I built a model predicting which customers would buy again within 30 days, with good AUC on validation. Marketing used it to send discount coupons to high-probability buyers. After a month, an experiment showed almost no extra revenue, because those customers would have bought anyway. We were giving discounts to people who did not need them. I learned to ask what action the model drives before building it. We rebuilt it as an uplift model, targeting customers whose purchase probability rose most with a coupon, and tested it against a control group. That version showed a clear gain in incremental orders.

**Likely follow-ups:**

- How does an uplift model work?
- How did you explain the first result to marketing?

### 13. Describe a time a stakeholder expected far more accuracy than the data could support.

Round: Behavioural · Level: All levels

**What they’re checking:** Whether you can manage expectations with evidence, explain uncertainty clearly and still deliver something useful.

**Sample answer:**

> A regional sales director wanted a model to forecast each store’s weekly sales to within 2 percent, so he could set targets. Many stores had only a year of history and sales were heavily affected by local festivals. I built a quick baseline and showed him that even the best model had around 10 to 15 percent error for smaller stores, and explained why using charts of their own sales swings. I suggested forecasting at the district level, where error was much lower, and giving store forecasts as ranges. He agreed, and the ranges actually helped his managers plan staffing better than a single number.

**Likely follow-ups:**

- How did you calculate the ranges?
- What if he had insisted on point forecasts?

### 14. Tell me about working with engineers to put one of your models into production.

Round: Behavioural · Level: All levels

**What they’re checking:** Whether you can work beyond notebooks, collaborating on deployment, latency, monitoring and code quality with engineering teams.

**Sample answer:**

> I built a model to rank support tickets by urgency for an insurance company. My first version was a Jupyter notebook with heavy text preprocessing that took several seconds per ticket. The engineering team needed under 200 milliseconds. I worked with a backend engineer to refactor the code into a Python package with tests, simplified the text features and saved the fitted pipeline as one artefact. They wrapped it in an API with logging. We agreed together on what to monitor: latency, input volume and the share of tickets marked urgent. Shipping it taught me to think about production limits from day one, not after modelling.

**Likely follow-ups:**

- Which tools did you use to track experiments?
- How did you handle model versioning?

### 15. Describe a time you found a serious problem in your analysis after presenting it.

Round: Behavioural · Level: Fresher

**What they’re checking:** Integrity, and whether you correct mistakes quickly and openly, and have since built checks into your workflow.

**Sample answer:**

> During my internship, I presented a clustering of app users into five segments. Two days later, while extending the work, I noticed I had not removed test accounts created by the QA team, and they formed most of one segment. I told my manager the same day, reran the analysis without them and sent a short corrected summary to everyone who attended the presentation. The remaining four segments held up. It was embarrassing, but my manager appreciated the quick correction. Since then I always start with data profiling and a list of exclusions, like test users, staff accounts and duplicates, before any modelling.

**Likely follow-ups:**

- How did you validate the clusters?
- What checks do you now run before presenting?

### 16. How have you handled very messy or incomplete data in a modelling project?

Round: Behavioural · Level: Fresher

**What they’re checking:** Whether you approach data quality systematically and make sensible, documented decisions rather than dropping rows blindly.

**Sample answer:**

> For my final-year project on predicting crop yields, I used district-level government data with missing years, inconsistent district names after boundary changes, and rainfall in different units across sources. I first profiled every column and documented issues. I standardised district names with a mapping table, converted units, and checked rainfall against a second source for a sample. For missing values, I used the district’s own historical median where gaps were short and added a flag column for imputed values, and dropped districts with too little history. I wrote all these steps into a reproducible script so anyone could rerun them and see each decision.

**Likely follow-ups:**

- When would you drop rows instead of imputing?
- Did the imputation flag turn out to be useful?

### 17. Tell me about a time you chose a simpler model over a more accurate one.

Round: Behavioural · Level: Experienced

**What they’re checking:** Whether you weigh accuracy against cost, explainability and maintenance, which shows maturity in applied data science.

**Sample answer:**

> For a lead scoring model at a B2B software company, a gradient boosting model with about two hundred features beat a logistic regression with twelve features by a small margin in AUC. But the sales team needed to understand why a lead scored high, the boosting model depended on several features from an unreliable third-party source, and retraining it needed a heavier pipeline. I tested both in a small pilot with sales reps and found conversion was essentially the same. We chose logistic regression, shared the scoring factors with the team, and they trusted and used it more. I documented the boosting model so we could revisit it later.

**Likely follow-ups:**

- How did you design the pilot?
- When would you have chosen the complex model?

## HR round questions

### 18. Why do you want to move from data analytics into a data scientist role?

Round: HR · Level: All levels

**What they’re checking:** Whether the move is backed by real modelling work and skills, not just a preference for a better title.

**Sample answer:**

> For three years as an analyst, I answered what happened through dashboards and SQL. Increasingly I was asked what will happen and what we should do, and I found those questions more interesting. I started building forecasting and churn models in Python on my own time, and one of them, a simple demand forecast for our top products, is now used by the planning team. I have also completed coursework in statistics and machine learning. My analytics background helps, because I know the business data and how stakeholders use numbers. Now I want a role where modelling and experiments are the main work, with mentors who have shipped models at scale.

**Likely follow-ups:**

- What did you learn from deploying the forecast?
- Which part of machine learning do you still find hard?

### 19. Our data science team works closely with business teams in Mumbai. Would you relocate?

Round: HR · Level: All levels

**What they’re checking:** Whether you are open to the location and working style the role needs, and whether any constraints will affect your joining.

**Sample answer:**

> Yes, I am open to relocating to Mumbai. I understand the value of sitting near the business teams. Most of my best project ideas came from informal conversations with operations people, not from formal requirement documents. I would need about a month after my notice period to move, since I have to find a place and shift my family. If the role allows a couple of remote days later on, that would be welcome, but I am happy to be in the office full time while I learn the business and build relationships with stakeholders.

**Likely follow-ups:**

- Is there any reason you might not be able to move?
- How do you work with stakeholders remotely?

### 20. What are your salary expectations as a data scientist, and what is your notice period?

Round: HR · Level: Experienced

**What they’re checking:** Whether your expectation is reasonable for your experience and the role, and whether your joining date fits their needs.

**Sample answer:**

> I currently earn ₹16 lakh CTC as a data scientist with three years of experience, including two models in production. For this role, which includes ownership of the pricing models, I am looking for ₹20 to 22 lakh fixed. I am open to discussing the variable part and any stock component. My notice period is 90 days, but my company has a policy of early release if the handover is complete, so I expect I could join in about 60 days. I would make sure my current models are well documented for the person taking over.

**Likely follow-ups:**

- Would you buy out your notice period?
- What matters to you besides salary?

## How to prepare for a data scientist interview

- Prepare to walk through one end-to-end project: problem framing, data, features, validation method, model choice, business impact and what you would do differently.
- Revise probability, hypothesis testing, confidence intervals and power, because statistics questions catch many candidates who focus only on algorithms.
- Practise writing clean Python with pandas and scikit-learn pipelines without notes, as many companies run live coding or take-home tasks.
- Be ready to explain how you avoid leakage and validate properly, since interviewers trust candidates who distrust their own high scores.
- Keep a GitHub repository or notebook with clear structure and a README, but expect to be asked about decisions, not just to show code.

## Questions about data scientist interviews

### How is a data scientist interview different from a data analyst interview?

Data analyst interviews focus on SQL, Excel, dashboards and explaining past performance. Data scientist interviews add machine learning theory, model validation, statistics for experiments, Python coding and often a modelling case or take-home task. Interviewers also ask how you deployed or monitored models and how a model changed a business decision.

### What should a fresher study for data scientist interviews?

Cover statistics and probability, linear and logistic regression, decision trees and ensembles, evaluation metrics, cross-validation and regularisation. Practise SQL and pandas. Build two projects on real, messy data with a clear question and a documented approach. Be ready to explain every choice in your projects, because interviewers will dig into them more than into textbook definitions.

### Do data scientist roles require deep learning?

Many roles in finance, retail and operations mainly use classical machine learning on tabular data, where gradient boosting and regression are standard. Deep learning matters more for roles involving text, images, speech or recommendations, and for teams working with large language models. Read the job description carefully and prepare deep learning topics in depth only if the role needs them.
