How to Vet a Senior Engineer in a 45-Minute Interview

Last updated October 2026

To vet a senior engineer in a 45-minute interview, spend 20 minutes on a code review of a flawed pull request they did not write, 15 minutes on one system they shipped and what broke in production, and 10 minutes on their questions. Score every answer against a written rubric. Drop the from-scratch coding puzzle.

Why has the standard way to vet a senior engineer stopped working?

Because the exercise at the center of it tests the one thing that is no longer scarce. The 2025 DORA research, published by Google across nearly 5,000 technology professionals, put AI adoption among software development professionals at 90 percent, with a median of two hours a day spent working with AI tools. In the same survey only 24 percent reported trusting AI output a lot or a great deal, and 30 percent reported little or no trust.

That gap is the job now. The 2025 Stack Overflow Developer Survey, with roughly 49,000 respondents, found 84 percent using or planning to use AI tools and 51 percent of professional developers using them daily, while the single most cited frustration, named by 66 percent, was AI solutions that are almost right but not quite. Just over 75 percent said that when they do not trust an AI answer they still want to ask another person. Senior engineers are that person. A candidate who can produce a working rate limiter from a blank file has demonstrated a commodity. A candidate who can look at 250 lines of plausible, passing, almost-right code and name the one line that will corrupt data in six weeks has demonstrated the thing you are hiring.

Most published advice on the senior technical screen has not caught up. The current top-ranking guides still allocate half the session to a greenfield coding task and none of it to code the candidate did not write, which is how a senior engineer spends most of a real week.

Which interview signals actually predict performance?

There is good evidence on this, and it is more recent than most hiring playbooks assume. Sackett, Zhang, Berry and Lievens re-estimated the validity of selection methods after correcting a long-standing statistical overcorrection, published in the Journal of Applied Psychology in 2022 and summarized by the Society for Industrial and Organizational Psychology. The ranking changed the conventional wisdom.

Selection method Revised validity (r) What it looks like in a 45-minute screen
Structured interview .42 Same questions, same order, scored against anchors
Job knowledge test .40 Reviewing real code from your own stack
Empirically keyed biodata .38 Verified history of comparable scope
Work sample test .33 The take-home or live coding task
Cognitive ability test .31 The algorithm puzzle

Read the first row and the last row together. Structure is doing more work than the exercise. A structured conversation run by an average interviewer beats an unstructured one run by your best engineer, because the structure is what makes two candidates comparable. So the design question is not which clever problem to pose. It is whether you will ask every candidate the same things in the same order and write down a score before you discuss them with anyone.

How do you vet a senior engineer in 45 minutes?

One flawed pull request, one incident story, and their questions. Here is the allocation.

Minutes What you run What you are scoring Fail signal
0 to 3 Frame the session and state that AI tools are allowed if narrated Nothing None
3 to 23 Review of a 250-line pull request with five planted defects Defect detection, severity ranking, reasoning Comments on formatting before correctness
23 to 38 One system they shipped: what broke in production and what changed after Scope owned, failure literacy, follow-through No incident they can describe in detail
38 to 45 Their questions, answered honestly What they probe for No questions, or only compensation

The review block runs first on purpose. It is the highest-information segment, so it gets the protected time, and a candidate who is going to be strong there usually relaxes into the rest of the conversation. The incident block is deliberately narrow: one system, one failure, what they changed. Breadth questions invite rehearsed answers, and a rehearsed answer about a system that never broke tells you nothing.

What goes in the planted pull request?

Take a real merged diff from your own history, revert the fixes, and keep it to something a reader can hold in their head. Five defects, in these classes:

  1. A data-model defect. A nullable column read as non-null, a denormalized field written in two places, a missing unique constraint. This is the one that must be found, because it is the defect class that gets more expensive every week it ships.
  2. A security defect. Something from the ordinary list rather than something exotic: an unparameterized query, a missing authorization check on an object lookup, a secret read from a committed file. We keep a longer inventory in the nine security flaws we find most often in AI-generated code.
  3. An error-masking construct. A broad catch that logs and continues, a retry with no backoff, a default value that hides a failed call. Candidates who have run production systems spot these in seconds.
  4. A dependency problem. A package that does not do what the code assumes, a version pinned to something yanked, or an import that does not resolve at all.
  5. A red herring. Code that is ugly, repetitive and completely correct. This is a calibration test. A senior engineer notes it and moves on; a candidate who ranks it alongside the authorization hole has told you how their code reviews will go.

Use your own code rather than a public exercise. Public problems leak, and a model can retrieve a canonical answer to anything that has been posted. Your merged diff from eighteen months ago has no canonical answer, which is the point. If you want the sharper version of why tooling will not do this screening for you, we have written up why AI code review tools do not catch architecture problems: they read the diff rather than the system, and so does an unprepared candidate.

How do you score it without arguing about it afterwards?

Write the rubric before you run the first interview, score during the session, and submit the score before any debrief. Three dimensions, four anchors each.

  1. Detection and ranking. 4: found the data-model and security defects, ranked them above the cosmetic issues, named the blast radius of each. 3: found both, ranking muddled. 2: found one. 1: found neither, or ranked the red herring as critical.
  2. Reasoning. 4: explained the failure mode and the conditions that trigger it, proposed a fix and named its cost. 3: correct diagnosis, vague remedy. 2: pattern-matched to a rule without explaining why it applies here. 1: asserted a problem that is not one.
  3. Calibration and communication. 4: said plainly what they were unsure about, asked for the context they were missing, framed comments as a reviewer would to a colleague. 1: confident on everything, including the parts they had wrong.

Two hard rules make the scores decisive. Missing the data-model defect is a no at senior level regardless of the rest. Treating the red herring as a blocker is a no on its own, because an engineer who cannot rank severity will burn a team’s review capacity on preferences. Everything else is a discussion.

What about AI use and identity during the interview?

Allow the AI and require narration. The task you set is adjudication of code, which is exactly the task that stays human when the assistant is on, so a candidate who uses a model to summarize the diff and then reasons about what it got wrong is demonstrating the skill rather than evading the test. The exercise is deliberately resistant to a copy-paste answer: there is no correct output to produce, only a ranked set of judgments to defend out loud.

Identity is the separate problem, and it is real. Gartner’s survey work, released on 31 July 2025 and covering roughly 3,000 job candidates, found 6 percent admitted to interview fraud by posing as someone else or having someone pose as them, and the firm projects that by 2028 one in four candidate profiles worldwide will be fake. Jamie Kohn, senior research director in Gartner’s HR practice, framed the stakes directly: candidate fraud creates cybersecurity risks that can be far more serious than making a bad hire. Keep the response proportionate. Camera on, one question about a specific decision in their own commit history, and formal verification at offer stage rather than in a 45-minute screen.

Book a Code Review

A screen built this way is cheap to run and it is yours to keep: one flawed diff from your own repository, a three-dimension rubric, and the discipline to score before you discuss. Build it once and it outlives whichever framework your stack is on.

Where this gets expensive is volume. Vetting at scale is a different problem from vetting well once, which is why our own bench runs to 15,643 vetted developers, averages more than ten years of engineering experience, and holds 96 percent team retention, with 56 percent of engineers arriving through internal referral. If you would rather buy the output of that process than run it, Delivery Pods start at $15,000 a month on 30-day notice, and the trade-offs against hiring directly are laid out in our pillar on delivery pods versus staff augmentation versus project outsourcing. What happens in the first week either way is covered in how to structure an engineering pod that ships in week one.

If you want a second opinion on the code before you hire anyone to maintain it, book a code review and a senior engineer will read your repository and reply within one business day.