Dor Meiri

GitHubLinkedIn
>

AI writes the code. What do we interview for?

_
With Daniel Sinai
October 9, 20267 min read

Over shawarma a few weeks ago, Daniel Sinai and I realized we were both fed up (pun very much intended) with the same moment in our coding interviews. A strong candidate struggles to type out code they'd never normally write by hand, then pauses to ask the question we'd heard more times than we could count: "Can I use AI to solve this?"

It was a fair question. Every developer at Port writes code with AI daily, yet our interview asked candidates to work without it. Were we really testing the skills developers need today?

This post is about how we answered that question and rebuilt our R&D coding interview from scratch.

The interview we loved

Our old interview looked simple on paper: write a recursive function that filters files.

That's easy to say. Once you start writing it, the decisions pile up. How do you model the filter parameter? A string? A list of extensions? A predicate function? Something composable, so that "large and not hidden" doesn't need its own flag? Every choice affects the next one, and a "simple" task quickly turns into a real design problem.

We didn't care about syntax. We didn't care if you remembered the exact name of the OS or standard library function. We didn't even run the code. We wanted to see whether you have a passion for solving problems with code, and how you make decisions about a problem that looks trivial but has hidden depth.

It worked really well, and we can vouch for it from both sides of the table: one of us went through it as a candidate before ever running it. It felt collaborative and interesting, and the interviewer left a strong impression with their questions. That's a big deal in an interview, because the candidate is evaluating you too.

Then AI became the default

Gradually, candidates started saying things like:

  • "I'm a bit rusty writing code, so I might be a little slow."
  • "I haven't written code in a text editor for a year and a half."

These weren't weak candidates. They were reflecting how engineering work had changed. Writing code by hand in an empty file had stopped being part of the daily job.

So we asked the obvious question: what if we just let candidates use AI? We want to see how people really work, and this is how people really work.

We tried it ourselves first. The AI solved the entire interview in one shot.

The design decisions were the whole point, and the AI made all of them, and made them well. With AI, writing good code for a small, self-contained problem is just too easy. The interview no longer told us anything.

What are we actually looking for?

Before designing something new, we had to step back and define what engineering looks like today. We ended up with three phases:

  1. Knowledge crunching (a term borrowed from Eric Evans' Domain-Driven Design). Gather context: the idea, the system, the trade-offs, and what "done" means. Write down the facts and assumptions the implementation needs. Too much detail and the spec turns into the implementation, too rigid for anyone to discuss. Too little and the task drifts somewhere unexpected, with poor trade-off decisions along the way.
  2. Implementation. Turn the spec into working code that scales, can be extended, and stays easy to maintain.
  3. Judgement and verification. How do you test it? What are the next steps? How does it get to production? What are the risks, and how do you reduce them?

AI has sped up phase 2 dramatically. Phases 1 and 3 still depend almost entirely on the engineer's understanding.

That changed which signals we looked for. Typing fast? Writing code from an empty page? Knowing clever TypeScript tricks?

Nope.

What we care about now:

  1. Understanding. Can you take in a piece of information and actually understand it?
  2. Judgement. Can you review code and architecture and tell good from bad?
  3. Control. Do you drive the AI agent, or does it drive you? That includes which tools you reach for, how you set them up, and how you use them.

Each one maps to a phase. Understanding is knowledge crunching, control is how you get through implementation, and judgement is verification.

Greenfield or brownfield?

We considered two formats:

  1. Greenfield: here's a problem, here's an AI, go build it from nothing.
  2. Brownfield: here's an existing product with some bugs and missing features. Work on it.

We chose brownfield, for two reasons.

First, we think it's fairer. The scope of each task is limited, but there's still plenty of room to show how deeply you think it through. Everyone starts from the same codebase, and the depth comes from you.

Second, it's the job. Port is a mature product. Most of the time you aren't writing something new in isolation. You're building a feature that has to fit with the rest of the system. An interview should look like that.

Getting it wrong, then less wrong

We built a small product related to our real one, small enough to understand in a short time. We added a few missing features to implement and some bugs to fix. We didn't put much effort into the environment the interview ran on, or into the AI agent setup.

Then we ran an internal test flight, and the feedback was not good.

It wasn't clear what we were looking for. The timing was stressful. The product was too complicated, with too many paths and rabbit holes. The README was too long, and our internal "candidate" didn't know where to start or what mattered most.

So we cut. We simplified the README a lot and narrowed it to three main tasks with a clear starting point. We also invested properly in the environment and the AI agent, so candidates could focus on the tasks instead of fighting their setup.

With those changes, it started to feel like a real interview: we could watch how someone approached a problem, made decisions, and communicated.

Honestly, it didn't go well with the first real candidate, mostly because of technical problems with the setup. That stung, but it showed us exactly what to fix.

The second interview went much better. The candidate was excited throughout, gave us positive feedback afterwards, and it was the first time we felt we were testing the skills we actually cared about. After a few more candidates and some refinements, we decided to move away from the classic coding interview and make the new format our standard.

Making it boring to run

An interview that only its creators can run doesn't scale. Before adding more interviewers to the pool, we wanted it simple to deploy and maintain.

Today, an interviewer tags Port AI in Slack, an environment spins up, and it's destroyed automatically afterwards. The candidate gets in with a single click. Nobody has to babysit infrastructure in the middle of an interview.

Under the hood, Port tracks the state of each interview environment, self-service workflows deploy it with AWS CloudFormation, and automations handle its lifecycle. If you want something similar, Port has a guide on managing ephemeral environments.

Infrastructure was only half of it. The other half was making sure every interviewer looks for the same things. We wrote a cheat sheet that lists good and bad signals for each task. A good signal: the candidate gets evidence of the issue before jumping to a solution, then verifies that their change actually fixes it. A bad one: they pick the AI's recommended solution without considering other paths and trade-offs. The goal is that a new interviewer knows what to watch for, and the evaluation doesn't depend on who happens to run the session.

Where we landed

This is now the main first coding interview we run for R&D candidates, and we're very happy with it. It's fair, it reflects how we actually work, and it shows us what AI can't do for you: understanding the problem, judging the code, and staying in control of the agent.

So, back to that question: "Can I use AI to solve this?"

Yes. Please do. That's the point. The AI will write the code. We want to see what you do with it.