20 questions · sample answers

DevOps Engineer interview questions and answers

DevOps engineer interviews test whether you can build and run reliable delivery and infrastructure: CI/CD pipelines, Linux, Docker, Kubernetes, Terraform, cloud networking, monitoring, secrets and incident response. Expect hands-on troubleshooting scenarios rather than definitions alone, plus questions about outages and on-call. Practise the 20 questions below out loud, compare your answers with the samples and prepare for the follow-ups interviewers usually ask.

DevOps hiring usually runs as a screening call, a hands-on round with a live Linux, Docker or Kubernetes troubleshooting task or a take-home pipeline exercise, a system design round on cloud architecture and reliability, and a final managerial and HR discussion.

Last updated

DevOps Engineer interview questions and answers
Round
Level

20 of 20 questions shown

Role and technical questions

Walk me through a CI/CD pipeline you built, stage by stage.

Role

What they’re checking: Whether you have actually built a pipeline end to end and understand why each stage exists, including quality gates and rollback.

Sample answer

For a Node.js API at a fintech startup, a push to a feature branch triggered GitHub Actions to install dependencies, run linting and unit tests, and run a dependency vulnerability scan. On merge to main, it built a Docker image tagged with the commit SHA, scanned the image, and pushed it to Amazon ECR. Then it deployed to a staging namespace in our EKS cluster using Helm, ran smoke and API tests, and waited for a manual approval. Production deployment used the same image and Helm chart with production values, a rolling update and readiness probes. If health checks failed, Helm rolled back to the previous release automatically.

Likely follow-ups
  • Why build the image once and promote it rather than rebuilding?
  • How did you keep pipeline run time down?

What is the difference between a Docker image and a container, and how do you keep images small?

Role Fresher

What they’re checking: Whether you understand container basics and good image-building practice, which affects build time, security and deployment speed.

Sample answer

An image is a read-only template made of layers, built from a Dockerfile, that contains the application and its dependencies. A container is a running instance of an image with its own writable layer, process space and network. Many containers can run from one image. To keep images small, I use multi-stage builds, so compilers and build tools stay in the build stage and only the final artefact is copied into a slim runtime image such as an alpine, slim or distroless base. I add a .dockerignore file, combine related RUN steps, clean package caches in the same layer, and order instructions so dependency layers are cached between builds.

Likely follow-ups
  • Why should containers not run as root?
  • What is the difference between CMD and ENTRYPOINT?

A Kubernetes pod is stuck in CrashLoopBackOff. How do you debug it?

Role Experienced

What they’re checking: Your practical Kubernetes troubleshooting method, using the right commands in the right order rather than restarting things at random.

Sample answer

CrashLoopBackOff means the container starts and exits repeatedly. I run kubectl describe pod to check events, the last state, the exit code and the reason. Exit code 137 with OOMKilled means the memory limit is too low or there is a leak. Then kubectl logs with the previous flag to see the output from the crashed container, since the current one may have no logs yet. Common causes I check are a missing environment variable, ConfigMap or Secret, a wrong command or entrypoint, the app failing to reach its database, and a liveness probe that kills the container before it finishes starting. If needed, I run the image locally or with a debug container.

Likely follow-ups
  • When would you use a startup probe?
  • How do you debug a pod stuck in Pending?

Compare rolling, blue-green and canary deployments. When would you use each?

Role

What they’re checking: Whether you understand release strategies and their trade-offs in risk, cost and rollback speed, and can match them to the service.

Sample answer

A rolling deployment replaces instances gradually, a few at a time. It is the Kubernetes default and needs no extra capacity, but old and new versions run together and rollback is another rollout. Blue-green runs a full new environment alongside the old one, then switches traffic at the load balancer. Rollback is instant, but it doubles capacity during the release, and database changes must work for both versions. Canary sends a small share of traffic, say five percent, to the new version, watches error rates and latency, then increases gradually. I use rolling for routine services, blue-green for critical systems needing instant rollback, and canary for high-traffic user-facing changes.

Likely follow-ups
  • How do you handle database migrations in blue-green deployments?
  • Which tools have you used for canary releases?

What is Terraform state, where should it be stored, and why does locking matter?

Role

What they’re checking: Whether you have used infrastructure as code in a team and understand the risks of state corruption and leaked secrets.

Sample answer

Terraform state is a file that maps resources in my code to real resources in the cloud, with their IDs and attributes. Terraform uses it to work out what to create, change or destroy. In a team it must be stored remotely, for example in an S3 bucket with versioning and encryption, or in Terraform Cloud, never in Git, because it can contain secrets like database passwords. Locking prevents two people or pipelines from running apply at the same time and corrupting the state. With S3 this has commonly been done with a DynamoDB table. I also split state by environment and component, so one mistake cannot affect everything.

Likely follow-ups
  • How do you bring an existing resource under Terraform?
  • What happens when someone changes a resource manually in the console?

How do you manage secrets such as database passwords and API keys across environments?

Role

What they’re checking: Whether you follow secure practice for secrets and know the tools, instead of storing credentials in code or plain environment files.

Sample answer

Secrets never go into Git, Docker images or pipeline logs. I store them in a secrets manager such as AWS Secrets Manager, HashiCorp Vault or Azure Key Vault, with access controlled by IAM roles per service and environment. In Kubernetes I use the External Secrets Operator to sync them into Kubernetes Secrets, with encryption at rest enabled for etcd. Pipelines use short-lived credentials, for example OIDC federation from GitHub Actions to AWS instead of long-lived access keys. I enable rotation for database passwords, and I run secret scanning on repositories to catch anything committed by mistake. If a secret leaks, I rotate it first and investigate afterwards.

Likely follow-ups
  • Why are Kubernetes Secrets not secure by default?
  • How would you rotate a database password without downtime?

A Linux server shows its disk is 100 percent full. How do you find and fix the problem?

Role Fresher

What they’re checking: Basic Linux troubleshooting under pressure, including less obvious causes like deleted files held open and inode exhaustion.

Sample answer

I run df -h to see which filesystem is full, and df -i to check whether inodes are exhausted rather than space. Then du -sh on the top-level directories, drilling down with sort to find what is large. Common culprits are application or system logs in /var/log, old Docker images and volumes, and core dumps. If df shows full but du does not add up, a process may be holding a deleted file open, which I find with lsof and the deleted filter, then restart that process. To fix it, I clear or compress safe files, then set up logrotate and monitoring alerts at around 80 percent, so it does not happen again.

Likely follow-ups
  • How do you safely clean up Docker disk usage?
  • What would you do if the root filesystem is full and you cannot log in normally?

What would you monitor for a customer-facing web service, and how do you avoid alert fatigue?

Role Experienced

What they’re checking: Whether you design monitoring around user impact and service level objectives, not just CPU graphs, and keep on-call alerts actionable.

Sample answer

I start with the four golden signals for the service: latency, traffic, errors and saturation. Then I define service level objectives with the product team, for example 99.9 percent of checkout requests succeeding and 95 percent completing under 500 milliseconds over 30 days. Alerts that page someone are based on how fast we are burning the error budget, not on raw CPU, because high CPU with happy users is not an emergency. Infrastructure metrics and logs go to dashboards for investigation. I review every page after on-call weeks: alerts that were not actionable get tuned or removed. Each alert has a runbook link explaining what to check first.

Likely follow-ups
  • What is an error budget used for?
  • Which tools have you used for metrics, logs and tracing?

What is the difference between git merge and git rebase, and which branching model do you prefer?

Role Fresher

What they’re checking: Whether you understand version control well enough to set up sensible workflows and avoid rewriting shared history.

Sample answer

Merge combines two branches by creating a merge commit, keeping the full history as it happened. Rebase replays my branch’s commits on top of another branch, giving a straight, cleaner history, but it rewrites commit hashes. So I never rebase a branch that others are already working on. I use rebase to update my own feature branch before opening a pull request, and merge or squash merge into main. For most teams I prefer trunk-based development: short-lived feature branches, small pull requests, required checks and reviews, and feature flags for unfinished work. Long-lived branches like in GitFlow tend to cause painful merges and slower releases.

Likely follow-ups
  • How do you recover a commit after a bad reset?
  • What branch protection rules do you set on main?

How does the Horizontal Pod Autoscaler work, and why might it fail to scale?

Role

What they’re checking: Whether you understand autoscaling mechanics in Kubernetes, including its dependencies and the link between pod and node scaling.

Sample answer

The Horizontal Pod Autoscaler periodically compares a metric, usually average CPU or memory utilisation as a percentage of the pods’ resource requests, against a target, and adjusts the number of replicas within a minimum and maximum. It can also use custom metrics, such as queue length or requests per second. Common reasons it fails: metrics-server is not installed, pods have no CPU requests set so utilisation cannot be calculated, the maximum replicas is already reached, or new pods stay Pending because there is no node capacity. For the last case, a cluster autoscaler or Karpenter adds nodes. I also tune scale-down stabilisation so traffic spikes do not cause flapping.

Likely follow-ups
  • When would you use the Vertical Pod Autoscaler instead?
  • How would you scale on a message queue backlog?

Design a highly available setup on AWS for a web application with a database.

Role Experienced

What they’re checking: Whether you can design for failure across availability zones, covering compute, data, networking, backups and recovery targets.

Sample answer

I would use a VPC across at least two, preferably three, availability zones, with public subnets for the load balancer and private subnets for the application and database. An Application Load Balancer sends traffic to containers on ECS or EKS, or to an EC2 Auto Scaling group, spread across zones with health checks. The database would be RDS with Multi-AZ for automatic failover, automated backups and point-in-time recovery, plus read replicas if reads are heavy. Static files go to S3 behind CloudFront. Sessions live in ElastiCache or a database, not on instances. I would agree recovery time and recovery point objectives, and test restores and failover regularly.

Likely follow-ups
  • What changes if you need to survive a full region outage?
  • How would you control the cost of this setup?

Behavioural questions

Walk me through a production outage you handled while on call.

Behavioural Experienced

What they’re checking: How you act during an incident: restoring service first, communicating clearly and following up with a blameless review and real fixes.

Sample answer

One night our payments API started returning errors for about a third of requests. I acknowledged the page, opened an incident channel and posted updates every fifteen minutes for support. Dashboards showed errors only from pods on two new nodes. Those nodes had come up with a wrong security group from a recent Terraform change, so they could not reach the database. I cordoned and drained them, which restored service within about twenty-five minutes, then fixed the Terraform module. In the blameless review we added a pipeline check for security group changes and a synthetic database connectivity test on new nodes before they accept traffic.

Likely follow-ups
  • How do you decide when to roll back versus fix forward?
  • What goes into a good postmortem?

Tell me about a time you made a mistake that affected infrastructure. What happened next?

Behavioural

What they’re checking: Honesty, quick recovery and whether you improved the system so others cannot make the same mistake, rather than just being more careful.

Sample answer

Early in my role, I ran terraform apply against what I thought was the staging workspace and deleted a production queue that matched a renamed resource. A background job started failing. I told my lead immediately, recreated the queue from code within about fifteen minutes, and replayed the failed jobs from logs. Nothing was lost for customers. Then I fixed the process, not just my habits: we moved applies into the pipeline only, added a plan review step showing any destroy actions in red, and enabled deletion protection on stateful resources. I also wrote a short guide for new engineers on how workspaces were set up.

Likely follow-ups
  • How did your team respond?
  • What other guardrails would you add today?

How did you convince developers to adopt a new deployment process or tool?

Behavioural

What they’re checking: Whether you drive change by making developers’ work easier and involving them, rather than enforcing tools from the platform team alone.

Sample answer

Our developers were deploying by SSHing into servers and running scripts, and they resisted a move to containers and pipelines because it seemed like extra work. I picked one team with frequent deployment problems, worked with them for two weeks and built a pipeline template that deployed their service with one merge. Their deploy time dropped from about forty minutes to under ten, and they stopped needing server access. They presented it at our engineering meeting. Then I turned the setup into a reusable template with documentation, and held office hours twice a week. Within a few months most services had moved.

Likely follow-ups
  • What did you do with teams that still refused?
  • How did you measure the improvement?

Tell me about a time you reduced cloud costs significantly.

Behavioural Experienced

What they’re checking: Whether you understand cloud cost drivers and can cut waste without harming reliability, which most companies now expect from DevOps teams.

Sample answer

Our monthly AWS bill had grown steadily and nobody owned it. I enabled cost allocation tags, then built a simple breakdown by team and service. Three things stood out: development environments running all night and at weekends, oversized RDS instances, and old EBS snapshots. I scheduled non-production clusters to scale down outside working hours, right-sized databases based on two weeks of metrics, set a snapshot retention policy, and moved steady workloads to Savings Plans. The bill dropped by roughly a third over two months, with no incidents. We also added a monthly cost review so each team saw its own spend.

Likely follow-ups
  • How do you use spot instances safely?
  • How do you stop costs creeping back?

Describe how you learned a new tool quickly to solve a problem at work or in a project.

Behavioural Fresher

What they’re checking: Whether you can teach yourself new tools in a structured way, since the DevOps toolset changes constantly.

Sample answer

In my internship, our team needed to deploy a small Python service on Kubernetes, and I had only used Docker. I gave myself a week. I installed kind on my laptop, followed the official tutorials for deployments, services and ConfigMaps, and then wrote manifests for our actual service rather than a sample app. I kept notes of every error and how I fixed it, such as image pull errors from a private registry. By the end of the week the service ran on our test cluster, and my notes became the team’s getting-started guide. My lead then asked me to convert the manifests into a Helm chart.

Likely follow-ups
  • What was the hardest Kubernetes concept to understand?
  • How do you keep up with new tools?

Tell me about a time you had to balance urgent requests from several teams.

Behavioural Fresher

What they’re checking: Whether you can prioritise interruptions against planned work and communicate clearly, which is constant in platform and DevOps teams.

Sample answer

In my first DevOps role, I supported four product teams, and requests came through chat at all hours: new environments, permission changes, pipeline failures. I was switching tasks constantly and missing my own deadlines. I proposed a simple system to my lead: all requests through a ticket form, with production issues marked urgent and handled immediately, and everything else reviewed twice a day. I also wrote short self-service guides for the five most common requests, like rerunning a failed pipeline. Request volume dropped noticeably, teams knew when to expect a response, and I finished our planned monitoring work on time.

Likely follow-ups
  • How did teams react to the ticket form?
  • How do you decide what counts as urgent?

HR round questions

This role includes an on-call rotation. How do you feel about that?

HR

What they’re checking: Whether you understand on-call responsibilities and expectations, and whether you will handle the rotation reliably and sustainably over months.

Sample answer

I am comfortable with it. I have been on a one-week-in-four rotation for the last two years. I take it seriously: laptop and charger with me, reachable within the agreed response time, and no long travel during my week. I would like to understand how often I would be on call, how pages typically come in each week, whether there is a secondary on-call, and how on-call time is compensated or given back as time off. I also care about improving the rotation itself, for example by fixing noisy alerts, because a calmer on-call week means better decisions when a real incident happens.

Likely follow-ups
  • What is the worst on-call week you have had?
  • How do you hand over at the end of a rotation?

Why do you want to join our platform team specifically?

HR

What they’re checking: Whether you have researched their infrastructure and engineering culture and can link your experience to their challenges.

Sample answer

I read your engineering blog about moving from virtual machines to Kubernetes and building an internal developer platform so product teams can deploy without filing tickets. That is the work I enjoy most. In my current role I built pipeline templates and Terraform modules that teams use themselves, and I want to do that at a larger scale with more services. I also noticed you run a blameless postmortem culture and publish learnings internally, which matches how I like to work. And your stack of AWS, EKS, Terraform and GitHub Actions overlaps closely with mine, so I could contribute quickly.

Likely follow-ups
  • What would you change about our setup from what you read?
  • How do you measure developer experience?

What salary are you expecting for this DevOps role, and when can you join?

HR Experienced

What they’re checking: Whether your expectation is reasonable and clearly explained, and whether your notice period works for the team.

Sample answer

My current CTC is ₹14 lakh with four years in DevOps, mostly on AWS and Kubernetes, including on-call for production systems. This role adds ownership of the Kubernetes platform and the observability stack, so I am looking for ₹18 to 19 lakh fixed, plus the on-call allowance. I am open to discussing the structure. My notice period is 60 days. I have already documented most of my current work, so I could ask for a shorter period if needed, but I would want to finish handing over the on-call runbooks properly before leaving.

Likely follow-ups
  • Do you have any certifications we should know about?
  • Would you consider a lower fixed with a joining bonus?

Practise these questions
Answer them aloud against a timer, then compare with the sample answers.

Start practice →

How to prepare for a devops engineer interview

  • Practise on a real Linux terminal and a local Kubernetes cluster such as kind or minikube, because many interviews include live troubleshooting tasks.
  • Be ready to draw your current pipeline and infrastructure on a whiteboard, explaining why each tool was chosen and what you would change.
  • Prepare one detailed incident story with timeline, actions, communication and follow-up fixes, as almost every DevOps interview asks for one.
  • Revise networking basics such as DNS, TCP, load balancers, subnets and security groups, since many production problems turn out to be networking problems.
  • Keep a public repository with Terraform modules, a Dockerfile and a pipeline file you wrote, with a short README explaining the design.
One place for your job

Everything for DevOps engineers

FAQ

Questions about devops engineer interviews

Expect Linux and shell scripting, Git, Docker, Kubernetes, CI/CD tools, Terraform or other infrastructure as code, cloud services on AWS, Azure or GCP, networking, monitoring and security basics. Many interviews include a hands-on troubleshooting task and a design question about a reliable deployment. Behavioural questions focus on incidents, on-call and working with developers.

Yes, though many people move into DevOps from system administration, support or development. As a fresher, build strong Linux and networking basics, learn Git, Docker and one cloud, and create a project that deploys an app through a pipeline to a cloud environment with infrastructure as code and monitoring. Show that project in interviews and explain what broke and how you fixed it.

Certifications like AWS Solutions Architect Associate, Certified Kubernetes Administrator or Terraform Associate can help with shortlisting, and some service companies need certified staff for client contracts. But interviewers test hands-on skill. A certification without practical experience will be exposed quickly in a troubleshooting round, so pair it with real projects.

Give the interviewer a link to your work

A personal website with your resume, projects and certificates — live in about five minutes.

● Live in 5 minutes · free to start · no auto-renew

Recruiters Google you before the interview

Get a page that shows up: your experience, projects and contact details at your own link. Live in minutes.

Start free
Chat on WhatsApp