Python interview questions: the ones a backend round actually asks

Most Python interview lists are written for people learning the language. A backend round isn't that. It assumes you can write Python and goes looking for the places where Python's convenience quietly costs you something in production.

PracticeDepth team10 min read

The questions below have one thing in common: each has an easy answer that's correct and a second layer that decides the round. The easy answer is the doorway, and interviewers use these particular questions precisely because the doorway is so well known that it tells them nothing on its own.

If you'd rather work through the topic than read about it, the Python track covers the same ground one idea at a time.

Question 1: what does the GIL actually stop you doing?

Everyone has heard that Python has a global interpreter lock and that threads are therefore useless. Half of that is right, and the half that's wrong is the half that matters when you're deciding how to make a slow endpoint faster.

You have a slow endpoint. Would threads help?

Junior answer

Probably not, because of the GIL. Python can only really run one thread at a time, so threading doesn't give you a speedup. You'd need multiprocessing.

Senior answer

It depends what it's slow on. The lock is held while interpreting bytecode, so two threads never execute Python code at the same time, and a CPU-bound handler gets nothing from threads. But the lock is released around blocking I/O, so if the endpoint is waiting on a database, an HTTP call or the disk, threads overlap that waiting perfectly well. So: I'd measure where the time goes first. I/O-bound, threads or asyncio. CPU-bound, move the work to processes, or into a library like NumPy that drops the lock in its C code, or out of the request entirely onto a queue.

The natural follow-up is what multiprocessing costs you, and the honest answer is: real memory per process, and everything you send between them has to be picklable, which rules out more objects than people expect.

Worth knowing rather than claiming: recent CPython versions ship a build with the lock removed, and it's a genuinely large change to the runtime. Saying it exists is a good signal. Saying your service would get faster by switching to it is a claim you'd have to defend, and single-threaded code on that build isn't automatically quicker.

Question 2: why does this function remember the last call?

The most-asked Python gotcha, and still worth asking, because the mechanism behind it explains a whole category of surprises.

def add_item(item, basket=[]):
    basket.append(item)
    return basket

add_item("apple")   # ['apple']
add_item("pear")    # ['apple', 'pear']  <- the same list
The default is created once, when the function is defined, not each time it's called.

A def statement runs once. Its defaults are evaluated at that moment and stored on the function object, so every call that doesn't pass the argument shares one list. The fix is basket=None and if basket is None: basket = [] inside the body.

Where else does that same rule bite?

Junior answer

Mostly with lists and dictionaries as defaults. As long as you use None instead, you're fine.

Senior answer

Anywhere an expression is evaluated once at definition time and then reused. Default arguments are the famous one, but class attributes behave the same way: a list assigned in the class body is shared by every instance, which is why dataclasses make you use field(default_factory=list) instead of a bare list. The same thing bites with decorators that build state at import time. The rule worth carrying is that class and def bodies run once, at import, and anything mutable created there is shared for the life of the process.

Question 3: async, threads or processes?

The question a backend round almost always reaches, because it's the one where a wrong mental model shows up directly in production behaviour.

@app.get("/report")
async def report():
    data = requests.get("https://slow-api.example.com").json()  # blocking
    return summarize(data)
An async endpoint making a blocking call: every other request on that worker waits, however many of them there are.

An event loop runs on one thread. A coroutine that blocks doesn't yield, so nothing else on that loop runs until it returns, including requests that have nothing to do with it. It's the same failure as blocking the event loop in Node.js, and the fix is the same shape: use an async client so the wait is awaited, or push the blocking call off the loop with asyncio.to_thread or a run-in-executor call.

  • asyncio for lots of concurrent I/O in one process: thousands of open connections, where a thread each would be wasteful. It costs you a whole parallel ecosystem of libraries, because one blocking call in the wrong place undoes it.
  • Threads for I/O when the code is already synchronous. Simpler, and fine at the scale most services actually run at.
  • Processes for CPU-bound work, with real memory cost and pickling at the boundary.
  • A task queue when the work doesn't need to finish inside the request at all. Frequently the right answer, and the one candidates forget to offer.

Question 4: this endpoint loads a million rows into memory

Generators are a first-year topic that turns into a senior one the moment the question is about a real dataset instead of a toy.

rows = [transform(r) for r in cursor.fetchall()]   # whole table in memory
rows = (transform(r) for r in cursor)              # one row at a time
Two characters of difference, and an entirely different memory profile.

When would you not use a generator here?

Junior answer

Generators are more memory efficient, so I'd use one wherever I'm iterating over something large. There's not much reason not to.

Senior answer

When you need the data more than once, or need its length, or need random access, because a generator is consumed and gone. And when laziness moves work somewhere it can't fail safely: if the generator is still pulling rows while the request is being serialised, the database transaction and connection have to stay open that whole time, and an exception now surfaces halfway through a response you've already started writing. For a streaming export that's the right trade. For something the caller retries, I'd rather page explicitly and know each chunk is finished.

A good follow-up to expect: what happens if you call the generator function and never iterate it. Nothing does, the body hasn't run yet, which is also why a bug inside it can show up nowhere near the line that created it.

Question 5: the ORM query that gets slower as the data grows

This one isn't really a Python question, which is exactly why it separates backend candidates from people who know the language.

for order in Order.objects.all():        # 1 query
    print(order.customer.name)           # 1 more query per order
The N+1: the loop looks like attribute access and is actually a round trip each time.

Lazy loading is the ORM's best feature and its sharpest edge. Attribute access issues a query, so a loop over a hundred orders is a hundred and one queries, each fast in isolation, and the endpoint gets slower in production as the table grows while staying fine on your laptop's seeded data. The fix is to tell the ORM what you're going to need: a join for a single related row, a second batched query for a collection. In Django that's select_related and prefetch_related; in SQLAlchemy, joinedload and selectinload.

What makes the answer senior is saying how you'd catch it rather than how you'd fix it, because you'll write this bug again. Log query counts per request in development, assert on the count in a test for the endpoints that matter, and watch the count in production rather than only the latency. Database interview questions for backend engineers goes further into what the database itself is doing with these.

Two smaller ones that come first

  • is versus ==. == asks whether two objects are equal, is asks whether they're the same object. They agree often enough by accident to be dangerous: small integers and short strings are cached by the interpreter, so a is b can be True for 256 and False for 257 with nothing else changed. Use is for None, True and False, and == for values.
  • Type hints. They're not enforced at runtime. The interpreter ignores them, a type checker reads them before the code ever runs, and passing a string where the annotation says int fails nowhere until something downstream breaks. Pydantic and FastAPI are the exception people are thinking of when they disagree: those validate at runtime, deliberately, because the data came from outside.

The live task: write a retry decorator

If there's a keyboard moment in a Python backend round, it's usually a decorator, and retry is the most common one. It's short, it uses closures, and it has an opinionated second half that most candidates skip.

  1. Say what you're building before typing: a wrapper that calls the function, and on a failure waits and tries again up to a limit.
  2. Reach for a closure over the arguments, and functools.wraps on the inner function, so the wrapped function keeps its name and docstring instead of becoming a nameless wrapper in every traceback.
  3. Forward *args and **kwargs unchanged, so the decorated function stays interchangeable with the original.
  4. Back off between attempts rather than retrying immediately, because a service that just failed under load is exactly the one you shouldn't hit three more times in a row.
  5. Catch the exceptions you can retry, not every exception. Retrying a timeout is sensible, retrying a 400 is a bug that hides a bug.
  6. Say the limit out loud: retries are only safe if the operation is idempotent. This is the sentence that separates the answer from the implementation.
import functools, time

def retry(attempts=3, delay=0.5, exceptions=(TimeoutError,)):
    def decorator(fn):
        @functools.wraps(fn)
        def wrapper(*args, **kwargs):
            for attempt in range(attempts):
                try:
                    return fn(*args, **kwargs)
                except exceptions:
                    if attempt == attempts - 1:
                        raise
                    time.sleep(delay * 2 ** attempt)
        return wrapper
    return decorator
Three nested functions, because the decorator takes arguments. Re-raising on the last attempt is what keeps the failure visible.

Expect to be asked what happens when every caller retries a struggling service at once, and whether you'd add jitter. And expect the version of this question that's really about trade-offs: would you write this, or take the dependency that already does it properly.

Question 6: what keeps an object alive?

The closing question in a lot of senior Python rounds, and the one people are least ready for, because Python's memory management is usually invisible until a long-running worker starts growing.

CPython counts references. When the count of things pointing at an object drops to zero, it's freed immediately, which is why the memory behaviour is mostly predictable. Cycles are the exception: two objects referring to each other keep each other's counts above zero forever, so a separate cycle collector runs periodically to find groups that only reference each other. That's why a parent holding children that hold the parent is a real leak shape, and why a cache that never evicts is a leak with no cycle in it at all.

Where would you look if a worker's memory grows all day?

Junior answer

Probably a memory leak somewhere. I'd restart the workers periodically and look for anything obviously holding data, like a big list that keeps growing.

Senior answer

First I'd separate growth from a leak: memory that plateaus is usually fragmentation or a pool, while a straight line up is something accumulating. Then I'd look at the usual suspects in order: module-level caches and anything decorated with lru_cache on a method, which keeps every instance it was ever called on alive; long-lived containers keyed by request data; and objects with references from a closure or a registry nobody remembers registering. tracemalloc gives you allocations by line between two snapshots, which answers where it came from. Restarting workers is a real mitigation, and I'd still want the snapshot before I shipped the restart as the fix.

weakref is the tidy answer when something needs to observe an object without keeping it alive: caches keyed by object, parent pointers, and registries of live instances.

Common questions

Do I need to know Django or FastAPI specifics?

Know the one on the job description well enough to explain what it does for you: how a request becomes a handler, where validation happens, and how it talks to the database. Framework trivia is rarely the question. Framework behaviour under load usually is.

Is asyncio expected for a backend Python role?

You're expected to know what it's for and when it's the wrong tool, which is most of the value of the question. Deep asyncio internals only come up when the job is already built on it.

How much algorithm practice do I need for a Python interview?

It depends entirely on the company, and it's worth asking the recruiter. Product companies with an algorithm stage expect it. For most backend roles the round looks like this article: the language's behaviour, the data layer, and what happens in production. The technical interview prep guide covers the shape the rest of the round takes.

Will I be asked about packaging and dependencies?

Often, as a short practical question rather than a deep one: how you pin dependencies, what a lock file is for, and how the environment in CI matches production. Naming the tool you actually use and why is enough.

Does this overlap with a general backend round?

Heavily. Questions 3, 5 and 6 are backend questions that happen to be asked in Python, which is why those answers travel with you if you switch languages. The Python-specific half is questions 1, 2 and 4, plus the decorator.

PythonBackendTopic guide

Keep reading

All posts